All-media fusion method and system based on complementary fusion

By performing feature extraction and pseudo-query vector generation on multimedia data, combined with implicit interaction module and vector database, the problem of cross-modal content fusion is solved, efficient matching and real-time display is achieved, and it is suitable for a variety of high-concurrency scenarios.

CN120123970AActive Publication Date: 2025-06-10SHANGHAI SHENGTONG ZHIMING TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510170015.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-10
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The prior art is difficult to achieve deep fusion and efficient matching of cross-modal content, especially in the need of real-time and high concurrency environments, and the demand for multimedia convergence is difficult to meet.

Method used

By collecting different types of media data from multiple media sources, dividing them into media units, and performing feature extraction separately to generate basic feature vectors. Use the pseudo-query module to generate a pseudo-query vector, combine it with the implicit interactive module to output the fused media representation vector, and store it in the vector database to establish an index to support fast retrieval and matching.

Benefits of technology

Cross-modal retrieval and complementarity of different media types are realized, suitable for online teaching, audio and video entertainment, live conference broadcasts, telemedicine and other scenarios, significantly improving the speed of convergence response and supporting large-scale distributed expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123970A_ABST
    Figure CN120123970A_ABST
Patent Text Reader

Abstract

The invention provides an all-media fusion method and system based on complementary fusion, and relates to the field of fusion communication. The method comprises the following steps: acquiring different types of media data from a plurality of media sources, and dividing the media data into a plurality of processable media units according to a preset segmentation strategy; feature extraction is carried out on different types of media data, and corresponding basic feature vectors are generated; generating an associated pseudo query vector by using a pseudo query module; inputting the basic feature vector and the pseudo query vector into an implicit interaction module, and outputting a fused media representation vector; storing the media representation vector and the corresponding media metadata into a vector database, and establishing an index for the media representation vector in the vector database; and in response to a received fusion request for the target media unit, obtaining media resources matched with the target media unit by retrieving the vector database so as to perform cross-media display of the target media unit.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of converged communication, and in particular, to an all-media convergence method and system based on complementary convergence. Background Art

[0002] Multi-modal data such as text, images, audio, video, and live streams are increasingly widely used in various industries. Different modal data have significant differences in structure, semantics, and temporal distribution. Traditional single-modal or simple splicing processing methods often struggle to achieve deep fusion and efficient matching of cross-modal content. Especially when multiple modal data need to be combined and displayed (such as video and text subtitle alignment, live stream and image / audio mixing, etc.), without an effective complementary mechanism for multi-modal features, problems such as low retrieval and matching efficiency, large data redundancy, and inaccurate alignment of cross-modal information are likely to occur, making it difficult to meet the requirements for multimedia fusion in real-time and high-concurrency environments.

[0003] For example, the Chinese patent application with the publication number CN113033647A discloses a multi-modal feature fusion method. Its main idea is: first, extract the features of each modality of the multimedia resource separately, combine the features of each modality in the modality dimension to form multi-channel features, and then generate fusion features through multi-channel convolution processing to achieve information complementarity between different modal features. This solution provides multiple convolution operations based on the convolution kernel in the D direction of the dimension and supports technical means such as feature aggregation and stretching processing, hoping to improve the effect of multi-modal feature fusion to a certain extent. This solution emphasizes that through convolution processing between channels, the influence of a relatively large single feature value can be effectively reduced, and the expression ability of the fusion features can be improved, and it is applicable to feature fusion scenarios of multimedia resources such as video image frame data, audio data, and text data.

[0004] However, the above solutions mainly focus on using convolutional operations for relatively static "same - dimension concatenation" and channel convolution of multimodal features, without deeply considering how to reduce the computational overhead of online large - scale comparison in cross - media retrieval and real - time interaction scenarios, nor proposing an incremental fusion strategy for media types with strong temporal continuity such as live streams. Although it has a certain complementary effect between channels in multimodal feature combination, when the system needs to dynamically supplement, synchronize in real time, or present on multiple devices the potential demands and associations of different modalities in an online or high - concurrency environment, it lacks a more perfect implicit interaction and offline perception mechanism, and is difficult to balance real - time performance and fusion accuracy. In addition, the processing method of only performing channel convolution on same - dimension features is also difficult to further reduce network load or improve cross - media retrieval efficiency when complementary queries need to be performed on each media unit and online fusion scenarios are realized. Therefore, there is still a need for an all - media fusion technology solution that can better adapt to offline pre - processing of multimodal data, fast online matching, and can perform low - latency, multi - device synchronized display for real - time data such as live streams, in order to overcome the limitations of existing solutions in cross - modal association accuracy and real - time hybrid presentation. Summary of the Invention

[0005] In view of the deficiencies of the prior art, embodiments of the present disclosure provide an all - media fusion method and system based on complementary fusion.

[0006] In a first aspect, embodiments of the present disclosure provide an all - media fusion method based on complementary fusion, including:

[0007] Collect different types of media data from multiple media sources and perform preliminary formatting processing on the media data; wherein, the media data includes: text, images, audio, video, and live stream data;

[0008] For the formatted media data, divide it into media units according to a preset segmentation strategy; wherein, a media unit refers to an independent media segment within a time or space range, and each media unit includes at least one of the following: a segment of text, a frame or a segment of video, an image or a group of images, and a segment of audio;

[0009] Extract features for different types of media data respectively to generate basic feature vectors corresponding to each media unit;

[0010] Use a pseudo - query module to generate pseudo - query vectors associated with each media unit based on the basic feature vectors;

[0011] Input the basic feature vectors and the pseudo - query vectors into an implicit interaction module to output a fused media representation vector; store the media representation vector and its corresponding media metadata in a vector database, and establish an index for the media representation vector in the vector database;

[0012] When a fusion request for a target media unit is received, retrieve media resources matching the target media unit from the vector database, and synchronously synthesize the target media unit and the matching media resources according to the timestamp information of the matching media resources for cross-media display of the target media unit.

[0013] As an alternative implementation, generating the pseudo query vectors associated with each media unit includes:

[0014] Performing an attention mechanism operation on the basic feature vectors to generate a potential demand representation;

[0015] Generating pseudo query vectors based on the potential demand representation and optimizing the pseudo query vectors using reconstruction loss or contrastive loss.

[0016] As an alternative implementation, outputting the fused media representation vectors includes:

[0017] Concatenating the basic feature vectors and the pseudo query vectors to form an input sequence;

[0018] Processing the input sequence through a multi-head self-attention network to generate fused media representation vectors;

[0019] Outputting the fused media representation vectors for indexing and retrieval.

[0020] As an alternative implementation, storing the media representation vectors and their corresponding media metadata in the vector database and establishing an index for the media representation vectors in the vector database includes:

[0021] Establishing a mapping relationship between the media representation vectors and the media metadata, where the media metadata includes media type, timestamp, and frame index;

[0022] Storing the media representation vectors and their media metadata in the vector database;

[0023] Establishing an index for the media representation vectors based on the approximate nearest neighbor search algorithm or the hash index algorithm.

[0024] As an alternative implementation, the cross-media display includes:

[0025] Receiving a fusion request for the target media unit;

[0026] Retrieving candidate media units in the vector database whose media representation vectors are more similar to the media representation vector of the target media unit than a preset threshold;

[0027] Stitch and fuse the target media unit and the candidate media unit according to the set fusion strategy;

[0028] Output the fusion result to the front-end for display.

[0029] As an alternative implementation, the cross-media display further includes:

[0030] Slice the live stream data according to time periods, and generate the basic feature vectors and pseudo query vectors for each time period;

[0031] Incrementally write the fused media representation vectors of the live stream data into the vector database, and use the latest media representation vectors for matching during online retrieval.

[0032] As an alternative implementation, the feature extraction includes:

[0033] For text, images, audio, and video, respectively, use convolutional neural networks, Vision Transformers, pre-trained language models, or speech recognition models to extract the basic feature vectors from the original media data.

[0034] As an alternative implementation, the cross-media display further includes: embedding the matched media resources into the rendering window corresponding to the target media unit.

[0035] In a second aspect, the embodiments of the present disclosure further provide an all-media fusion system based on complementary fusion, including: an acquisition module, a first processing module, a feature extraction module, a second processing module, a third processing module, and a display module;

[0036] The acquisition module is configured to collect different types of media data from multiple media sources and perform preliminary formatting processing on the media data; wherein, the media data includes: text, images, audio, video, and live stream data;

[0037] The first processing module is configured to divide the formatted media data into media units according to a preset segmentation strategy; wherein, a media unit refers to a media segment that is relatively independent in terms of time or space, and each media unit includes at least one of the following: a segment of text, a frame or a segment of video, a single image or a group of images, and a segment of audio;

[0038] The feature extraction module is configured to perform feature extraction on different types of media data respectively to generate basic feature vectors corresponding to each media unit;

[0039] The second processing module is configured to use a pseudo query module to generate pseudo query vectors associated with each media unit based on the basic feature vectors;

[0040] The third processing module is configured to input the basic feature vector and the pseudo query vector into an implicit interaction module, and output a fused media representation vector; store the media representation vector and its corresponding media metadata in a vector database, and create an index for the media representation vector in the vector database;

[0041] The display module is configured to, when receiving a fusion request for a target media unit, retrieve a media resource that matches the target media unit from the vector database, and synchronously synthesize the target media unit and the matched media resource according to the timestamp information of the matched media resource, so as to perform cross-media display of the target media unit.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows: Different media (text, image, audio, video, live stream) can utilize the same vector database to achieve cross-modal retrieval and complementarity, and are suitable for applications in scenarios such as online teaching, audio-visual entertainment, conference live broadcast, and telemedicine. By using approximate nearest neighbor search in the vector database and timestamp synchronization, only dot product calculation or similarity calculation is required in the online stage to complete cross-media matching and fusion, greatly improving the fusion response speed. As the media data increases, the system only needs to perform feature extraction and implicit interaction on the newly added media units and then write them into the vector database, and the index is updated, supporting large-scale distributed expansion. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is a flowchart of an all-media fusion method based on complementary fusion provided by an embodiment of the present disclosure;

[0044] Figure 2 is a flowchart of a method for creating an index in a vector database provided by an embodiment of the present disclosure;

[0045] Figure 3 is a schematic diagram of an all-media fusion system based on complementary fusion provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0047] The present invention provides an all-media fusion method based on complementary fusion. Refer to Figure 1 As shown, it is a flowchart of an all-media fusion method based on complementary fusion provided by an embodiment of the present disclosure. The method includes steps S101 to S106, where:

[0048] S101: Collect different types of media data from multiple media sources and perform preliminary formatting on the media data; wherein, the media data includes: text, images, audio, video, and live stream data;

[0049] S102: For the media data after formatting, divide it into several processable media units according to a preset segmentation strategy; wherein, a media unit refers to a relatively independent media segment within a time or space range, and each media unit includes at least one of the following: a segment of text, one or more frames or segments of video, one or a group of images, and a segment of audio;

[0050] S103: Extract features for different types of media data respectively to generate basic feature vectors corresponding to each media unit;

[0051] S104: Use a pseudo-query module to generate pseudo-query vectors associated with each media unit based on the basic feature vectors;

[0052] S105: Input the basic feature vectors and the pseudo-query vectors into an implicit interaction module to output a fused media representation vector; store the media representation vector and its corresponding media metadata in a vector database, and establish an index for the media representation vector in the vector database;

[0053] S106: When a fusion request for a target media unit is received, retrieve media resources matching the target media unit from the vector database, and synchronously synthesize the target media unit with the matching media resources according to the timestamp information of the matching media resources for cross-media display of the target media unit.

[0054] Regarding the above S101:

[0055] In the scenario of all-media fusion, the media data from different sources are diverse in type and inconsistent in format. If directly subjected to subsequent analysis, it is prone to difficulties in recognition or processing. By uniformly collecting media data from different sources and performing preliminary formatting, it can lay a foundation for subsequent feature extraction and cross-media matching.

[0056] In specific implementation, different types of media data can be obtained from multiple media sources, including text databases, image storage services, audio recording and storage servers, video streaming platforms, and live servers, etc.

[0057] For example, for text data, it can be obtained from existing document libraries, news sources, or documents uploaded by users through an API interface; for image data, it can be read from materials uploaded from image storage or photographic devices; for audio and video data, it can be loaded from streaming media platforms or local files; for live stream data, streaming media segments can be captured through a real-time acquisition port.

[0058] In specific implementation, the preliminary formatting process of the collected original media data according to their respective media types may include:

[0059] For text, remove redundant whitespace, unify the encoding format (such as UTF-8), and perform necessary cleaning;

[0060] For images, perform proportional scaling in terms of length and width resolution or unify the image encoding format (such as JPEG / PNG);

[0061] For audio, unify the sampling rate, bit rate, or audio encoding format (such as MP3 / WAV);

[0062] For video, perform preliminary transcoding based on frame rate, resolution, or encoding format (such as H.264 / H.265);

[0063] For live streams, split the real-time stream according to a predetermined time period (such as 5 seconds or 10 seconds) and store it as a temporary buffer file for subsequent processing.

[0064] In this way, through the above processing, a basic media file or data stream for subsequent analysis and feature extraction is generated.

[0065] Regarding the above S102:

[0066] After completing the unified formatting, the entire segment or the whole media data still needs to be further divided into smaller "media units" so that the system can perform more refined feature extraction and matching on segments of different types and different time / space ranges in subsequent steps.

[0067] In specific implementation, the formatted media data is divided into several processable media units according to a preset segmentation strategy.

[0068] Among them, for text data, it can be divided based on paragraph or sentence-level granularity, with each paragraph of text corresponding to a media unit; for image data, each image or group of images is regarded as a media unit; for audio or video data, a time segmentation strategy can be adopted, such as every 10 seconds or each key frame interval as a media unit; for live stream data, the same time segmentation strategy can also be used to split the live stream into multiple relatively independent media units.

[0069] It should be noted that in the present disclosure, a media unit refers to a relatively independent media segment within a time or space range, and may include one or more of the following types: a piece of text, a frame or a segment of video, a single image or a set of images, and a segment of audio. Through this classification method, in subsequent steps, the "media unit" can be used as the processing object to simplify the management and analysis of multimodal data.

[0070] Regarding the above S103:

[0071] The key to multimodal fusion lies in converting different types of media into comparable and retrievable vector representations. Without a unified feature vector representation, it is difficult to achieve effective comparison and complementarity among media such as cross-text, images, audio, and video.

[0072] In specific implementation, for each media unit obtained by partitioning, a feature extraction algorithm adapted to its media type needs to be used respectively to obtain the basic feature vector. For example:

[0073] For text feature extraction, a pre-trained language model (such as BERT, RoBERTa) or a Transformer-based semantic encoder can be used to perform vectorization operations on text units to generate semantic expression vectors.

[0074] For image feature extraction, CNN (such as ResNet, VGG) or Vision Transformer can be used to perform convolutional / self-attention analysis on images to extract visual feature vectors.

[0075] For audio feature extraction, Mel-Spectrogram or MFCC feature analysis can be performed on audio media units (including live stream segments), and an RNN or Transformer structure can be combined to obtain audio feature vectors.

[0076] For video feature extraction, spatio-temporal feature vectors can be extracted from video media units through key frame extraction or models such as 3D convolutional networks and video Transformers;

[0077] It should be emphasized that if the video unit is long, average pooling can also be performed on multiple frames to obtain an overall representation.

[0078] In this way, through the above multimodal feature extraction, the present disclosure finally generates corresponding basic feature vectors for each media unit and saves them in memory or temporary storage for subsequent processing.

[0079] Regarding the above S104:

[0080] A simple basic feature vector can only reflect the content of the media unit itself, but cannot reflect "the complementary information that the media unit may need for other media resources". If complex cross-comparisons are directly made between all media units in the online stage, the computational cost is high and the real-time performance is insufficient.

[0081] In a specific implementation, the present disclosure inputs the basic feature vector of each media unit into a pseudo-query module (such as a lightweight Transformer decoder or a self-attention-based sub-network) to generate a pseudo-query vector associated with each media unit. This pseudo-query vector can be understood as a potential expression of "the possible needs or associated points of this media unit for other media".

[0082] In a specific implementation, contrastive learning or reconstruction loss can be adopted in the training stage so that the pseudo-query vector captures the core semantics and potential interaction information of the media unit.

[0083] In a specific implementation, the generation method of the pseudo-query vector can be:

[0084] Input: Basic feature vector;

[0085] Intermediate process: The pseudo-query module performs multi-head self-attention calculation or decoding operation on the basic feature vector according to the trained parameters;

[0086] Output: Pseudo-query vector associated with this media unit.

[0087] Exemplarily, if the media unit is a text describing a cooking process, the pseudo-query vector may include potential needs for related ingredients, cooking temperature, and duration; for an image of a specific scene, the pseudo-query vector can express the association with information such as location, theme, or linked videos.

[0088] As an optional implementation manner, generating the pseudo-query vector associated with each media unit includes:

[0089] Performing an attention mechanism operation on the basic feature vector to generate a potential demand representation;

[0090] Generating a pseudo-query vector based on the potential demand representation and optimizing the pseudo-query vector using reconstruction loss or contrast loss so that the pseudo-query vector can represent the association characteristics between the media unit and other media resources.

[0091] Among them, after multi-modal feature extraction is completed, each media unit has a basic feature vector, denoted as F base . In order to further generate a pseudo-query vector that can represent the association characteristics between the media unit and other media resources, the present disclosure introduces two steps: attention mechanism operation and reconstruction / contrast loss optimization.

[0092] In a specific implementation, F base is input into a lightweight "attention generation sub-module", which can be implemented using a multi-head self-attention structure and at least includes the following key components:

[0093] 1. Linear mapping layer:

[0094] F base is mapped to three vector spaces of query (Q), key (K), and value (V). Specifically, a trainable weight matrix W Q can be set to generate the query vector Q = F base ·W Q , and K and V are generated similarly.

[0095] Since the feature vectors of media units often have a high dimension, the linear mapping layer can ensure that information is not lost while reducing parameters, and can be implemented by means of dimensionality reduction or maintaining the same dimension.

[0096] 2. Attention calculation layer:

[0097] The attention calculation layer performs dot-product attention on Q, K, and V:

[0098]

[0099] Among them, Attention(Q, K, V) is used to allocate learnable weights between input features to highlight the information most relevant to the current task; the softmax function is a normalization function that can map the input vector to the interval (0, 1) and make their sum equal to 1, so as to obtain the attention weights in the form of a probability distribution, and d k is the dimension of the key vector.

[0100] It should be emphasized that for the case of a single vector input (i.e., F base of a single media unit), this implementation can regard F base as batch processing or introduce a virtual sequence length, or introduce fixed position encoding to ensure that there is no conflict in the calculation of the attention mechanism. For example, if F base corresponds to a video key frame sequence, multi-vector attention operations can also be performed at the frame level.

[0101] 3. Multi-head merging layer:

[0102] If multi-head attention is adopted, after the outputs generated by each head are concatenated in the channel dimension, the output vector A is obtained through linear mapping. This output vector A can be regarded as the "latent demand representation", which represents an abstract expression of the media unit and external associated elements after the attention mechanism.

[0103] Through the above attention mechanism operations, the obtained potential demand representation A can reflect the potential cooperation or complementarity requirements of the media unit for other media resources, and initially characterize the feature of "which external information the media unit may be associated with".

[0104] In addition, it is found that in the actual system implementation, a simple self-attention operation cannot ensure that A accurately reflects the cross-media association features. Therefore, the present disclosure further optimizes A through Reconstruction Loss or Contrastive Loss to obtain the final pseudo-query vector (Q pseudo ).

[0105] Regarding the reconstruction loss:

[0106] In a specific implementation, if the system has additional annotations or context information about the media unit (such as key tags, abstracts, corresponding captions, etc.), A can be input into a "small decoder" or classifier to form constraints by predicting or reconstructing the additional information.

[0107] For example, when the media unit is a certain piece of text and the text has manually annotated topic tags, the system can let A predict the topic tags; if the prediction is successful, it means that A captures the core semantics of the text more accurately.

[0108] The loss function can be designed as cross-entropy or mean squared error, etc., and the parameters of the attention generation sub-module and the merging layer are tuned through backpropagation to make A have a higher semantic fit.

[0109] Regarding the contrastive loss:

[0110] In a specific implementation, if the system has positive and negative sample pairs (such as "two video clips of the same scene" and "the text description of the same timestamp pair"), a contrastive learning method can be adopted to make the vector distance between A and the matching samples closer and the vector distance between A and the non-matching samples farther.

[0111] Exemplarily, assume that a media unit A and a media unit B are complementary resources, which can be labeled as a positive sample pair to make the Q pseudo generated by both closer; for a randomly irrelevant media unit C, it is a negative sample pair to make the Q pseudo farther. In this way, A will become stronger and stronger in the ability to distinguish association from non-association.

[0112] Common formulas for contrastive loss such as InfoNCE or TripletLoss, etc., can enable A to learn the ability to distinguish cross-media retrieval or matching requirements after multiple iterations.

[0113] Furthermore, in a specific implementation, A output by the above self-attention mechanism is continuously iteratively trained in backpropagation, and finally the pseudo query vector Q can be stably obtained. pseudo 。

[0114] If the reconstruction loss is used, the system can regard the output A as Q after the training is completed. pseudo ; If the contrastive loss is used, a linear transformation layer (MLP) or a normalization operation can also be added after A to obtain the final Q. pseudo 。

[0115] It should be noted that at the implementable level, the attention module, decoder / contrast network can be trained end-to-end or in stages during the training phase to ensure that each parameter update can reduce the prediction error or increase the discrimination, so that the pseudo query vector not only retains the feature information of the original media unit, but also has the "demand perception" ability for other media.

[0116] It should be emphasized that in the implementation of the present invention, the use of the reconstruction loss or the contrastive loss is not merely a conventional algorithm selection for model training, but is directly deeply coupled with the multimedia fusion application scenario, playing a key technical role in cross-modal data processing and online real-time display.

[0117] For example, relying solely on the "basic feature vector" to index or match media content has the problem of being difficult to fully express the possible complementary needs of a media unit for other modal information. Therefore, large-scale cross-modal interactions or complex comparisons often need to be carried out in the online stage, resulting in a significant increase in latency. In the present invention, the concept of "pseudo query vector" is proposed, so that each media unit can learn the "potential demand for external media content" in the offline or small-batch training stage. In order to make this pseudo query vector more accurately and stably represent cross-media associations or complementary needs, the present invention selects the reconstruction loss or the contrastive loss to cooperate with the attention mechanism to overcome the limitation that pure perceptrons or simple classification losses cannot fully capture the "demand-supply" semantic interaction between multiple modalities.

[0118] Exemplarily, in actual multi-modal services, subtitles, keyword tags, meta-event descriptions, etc. are often configured for video clips or audio segments, or there are cross-annotations of the same content in different modalities (for example, a video clip and its corresponding text description).

[0119] The present invention utilizes the reconstruction loss to make the pseudo-query vector "reconstructible" or "predictable" for these multimodal additional annotations. Once the model learns to reconstruct this auxiliary information through the pseudo-query vector during the training phase, it indicates that the pseudo-query vector has indeed captured the intrinsic semantic connection between the media unit and the external complementary resources.

[0120] In addition, at the application level, the use of the reconstruction loss significantly improves the accuracy of injecting "complementary demand awareness" for each media unit in the offline phase, reducing the redundant inter-modal comparison overhead during online queries, thereby contributing to resource scheduling and fusion presentation in high-concurrency scenarios such as live streaming and online video-on-demand.

[0121] Furthermore, in order to accelerate online retrieval and ensure fusion accuracy, the present invention can distinguish cross-modal positive and negative sample pairs (such as audio and the text explaining it, or video frames and the accompanying image illustrations, etc.) during the training phase.

[0122] As mentioned above, when it is determined that "media units A and B have a complementary relationship in a certain application scenario", it is labeled as a positive sample pair to make the corresponding pseudo-query vectors closer; conversely, if there is no association between them, it is marked as a negative sample pair to widen the vector distance.

[0123] In this way, the contrast loss is no longer just a general machine learning method, but is combined with the specific requirements of multimedia fusion applications, achieving pre-learning of the "cross-modal association degree" during offline training. During online operation, the system only needs to perform vector retrieval on the pseudo-query vector of the target media unit to quickly locate other media resources with the highest "cross-media complementarity". Compared with similar solutions that do not use the contrast loss, the present invention can achieve significant improvements in both cross-modal matching degree and online query speed.

[0124] In the present invention, the "pseudo-query vector" is optimized through the reconstruction loss or the contrast loss, enabling the media unit to learn "how to perform complementary matching with other media resources" during the offline phase or in small-batch training updates. At this time, only simple vector similarity calculations are required during the online phase to determine the best fusion object and presentation method.

[0125] This technical contribution directly serves the online multimedia fusion requirements. For example, in the live e-commerce scenario, the current video clip of the product shown by the anchor can be instantaneously and pop-up fused with the most relevant product poster or the accompanying instruction manual; in the education live streaming scenario, corresponding courseware, exercise questions, or supplementary audio can be recommended in a timely manner, thereby significantly improving the user's acquisition efficiency and interaction experience.

[0126] Different from simply using reconstruction loss or contrast learning to optimize neural networks, the present invention emphasizes integrating this training strategy in the overall process of multi-modal feature extraction, implicit interaction, and vector database retrieval to reduce the online computing burden and improve the efficiency of multi-modal content complementarity. In other words, the present invention utilizes reconstruction / contrast loss not only at the level of simple deep learning algorithms but also combines the system architecture requirements of offline-online fusion of multimedia data, forming specific technical effects of cross-media complementarity: reducing latency, saving bandwidth, improving retrieval accuracy, and user satisfaction, etc.

[0127] Furthermore, this embodiment can train the pseudo-query generation subsystem in an offline environment (such as a GPU / TPU cluster), process existing media data on a large scale, and store the generated pseudo-query vectors in the vector database together;

[0128] When a new media unit is collected by the system and the basic feature extraction is completed, the attention module and the contrast / reconstruction network are immediately called to generate pseudo-query vectors, which are stored in the database for real-time retrieval;

[0129] In addition, if the system discovers that some pseudo-query vectors do not match the actual requirements according to user feedback, the parameters of the attention mechanism can be trained and updated in small batches regularly (or in real time) to enable the system to adapt to the environment and new media.

[0130] In this way, the pseudo-query vectors can more accurately represent the association features between each media unit and other media resources, serving as the basis for subsequent cross-media complementary display. Compared with the traditional single feature vector method, this embodiment injects the attribute of "perceiving the needs of external media" into each media unit at the offline stage, reducing the large-scale computational amount in the online retrieval process and enhancing the robustness and response speed of the system in the multi-modal fusion scenario.

[0131] Regarding the above S105:

[0132] In the offline stage, by further fusing the "basic feature vector" and the "pseudo-query vector", the final representation vector of each media unit (i.e., the "fused media representation vector") can contain both its own features and retain the representation of the needs or associations with other media. In this way, during subsequent online retrieval, there is no need to perform real-time interaction on all media, significantly reducing the computational amount.

[0133] In a specific implementation, the present disclosure inputs the "basic feature vector" and the "pseudo-query vector" of each media unit into the implicit interaction module. Exemplarily, a multi-layer Transformer or a multi-head self-attention network can be used to implement implicit interaction.

[0134] Among them, the operations of this implicit interaction include:

[0135] By means of self-attention or cross-attention mechanism, fuse the "own characteristics of the media unit" with the "pseudo-query information expected or required by the media unit" to generate a fused media representation vector with a stronger "cross-media complementary awareness".

[0136] The purpose of implicit interaction is to inject the perception of the needs for other media forms or content into each media unit during the offline stage, so that large-scale cross-modal complex calculations are no longer required during online retrieval.

[0137] Finally, the fused media representation vector output by the implicit interaction module can better represent the semantic features of the media unit and its potential complementary relationship with other modalities, providing a basis for subsequent retrieval and synthesis.

[0138] Furthermore, the present disclosure stores the fused media representation vector together with the media metadata information of the media unit (such as media type, timestamp, frame index, text paragraph ID, live stream slice ID, etc.) in a vector database.

[0139] Exemplarily, the vector database can adopt structures such as Faiss, Milvus or HNSW, etc., to support large-scale vector similarity search.

[0140] In addition, after the storage is completed, an approximate nearest neighbor (ANN) or hash index can be constructed according to the fused media representation vector to quickly perform subsequent retrieval and matching operations. At this time, each media unit has a unique identifier and index entry in the vector database for online stage lookup.

[0141] As an alternative implementation, the output of the fused media representation vector includes:[[]]

[0142] Concatenate the base feature vector and the pseudo-query vector to form an input sequence;

[0143] Process the input sequence through a multi-head self-attention network to generate a fused media representation vector;

[0144] Output the fused media representation vector for indexing and retrieval.

[0145] Among them, in order to obtain the fused media representation vector and make it convenient for subsequent indexing and retrieval, in addition to the aforementioned generation and optimization of the pseudo-query vector, the following operations can be further performed:

[0146] After the pseudo-query vector Q pseudo and the base feature vector F base are generated, the system concatenates the two to form an input sequence.

[0147] Exemplarily, the following operation steps can be adopted:

[0148] First is the serialization process. When F base and Q pseudo are both single vectors, they can be directly concatenated according to the vector dimension; if both contain temporal information (such as video key frames or audio frame sequences), then the frame vectors can first be merged in chronological or spatial order, and then the pseudo-query vector can be appended to the head or tail of the sequence to form the input sequence (S in ).

[0149] Doing so can ensure that multi-modal information and "requirement / correlation" features are presented in the same input tensor, and subsequent processing layers do not need to perform cumbersome cross-tensor operations.

[0150] Secondly, it is the multi-head self-attention network processing. The system inputs this input sequence (S in ) into the multi-head self-attention network (TransformerEncoder or a similar structure). Each attention head calculates the attention weights between the vectors in the sequence to capture temporal or semantic correlations.

[0151] If the basic feature vectors themselves have position encoding, such as video frame indices or text paragraph IDs, this information can be retained when concatenating with the pseudo-query vector.

[0152] After the multi-head self-attention network finishes execution, a set of context-enhanced vectors (H ctx ) will be produced. In this embodiment, the vector that can best represent the overall information of the sequence (such as the [CLS] position vector or the average pooling result of all vectors) can be selected as the fused media representation vector (F fused ).

[0153] Finally, output this fused media representation vector (F fused ) to the storage or retrieval module for subsequent indexing and retrieval use.

[0154] In specific implementation, for the convenience of subsequent indexing, F fused can be normalized (such as L2 regularization) or dimension-reduced (such as PCA) to generate the representation vector (F final ) finally written into the database. F final Compared with the original basic feature vector F base and the pseudo-query vector Q pseudo , it can better reflect the fusion characteristics of multi-modal information and cross-media requirements, and is suitable for similarity measurement or distance measurement in subsequent retrieval.

[0155] In this way, through the above-mentioned concatenation, multi-head self-attention operation, and vector output processing, the present invention can generate a fused media representation vector, providing a more complementary feature representation for the subsequent indexing and retrieval processes of the vector database.

[0156] See Figure 2 which is the flowchart of the method for building an index in a vector database provided by an embodiment of the present disclosure. As an alternative implementation, storing the media representation vector and its corresponding media metadata in the vector database and building an index for the media representation vector in the vector database includes steps S201 to S203, where:

[0157] S201: Establish a mapping relationship between the media representation vector and the media metadata, where the media metadata includes media type, timestamp, and frame index;

[0158] S202: Store the media representation vector and its media metadata in the vector database;

[0159] S203: Build an index for the media representation vector based on the approximate nearest neighbor search algorithm or the hash index algorithm.

[0160] Regarding the above S201:

[0161] In a specific implementation, when generating the fused media representation vector F fused or F final , synchronously record the corresponding media metadata, including:

[0162] Media type (such as text, video, audio, image, or live stream slice);

[0163] Timestamp (in the video or audio scenario, identifying the start and end times of the segment in the original file; in the live scenario, identifying the actual playback or recording moment);

[0164] Frame index (if it is a video key frame, the frame number or the continuous frame range can be recorded).

[0165] By establishing a "mapping structure" (such as key - value or JSON format) containing the above - mentioned media metadata for each fused media representation vector, refined matching can be achieved through vector similarity plus media metadata condition filtering during future retrieval.

[0166] Regarding the above 202:

[0167] In a specific implementation, a database or engine supporting large - scale vector search can be selected, such as Faiss, Milvus, HNSW, etc. Write F fused and the mapping structure into the database at one time, and assign a unique identifier (such as media_id) to each entry.

[0168] When a user query or a system fusion request occurs, retrieve the stored vector table in a similarity - calculation manner and output several most similar entries.

[0169] Metadata fields (media type, timestamp, frame index) can be used to further perform secondary filtering or sorting in the retrieval results.

[0170] Regarding the above S203:

[0171] In a specific implementation, to improve the retrieval efficiency, an index construction process can be performed on all media representation vectors that have been written into the database:

[0172] ANN (Approximate Nearest Neighbor) index: In a high-dimensional space, partition all vectors or construct a hierarchical graph structure (such as HNSW) to achieve approximate search with O(logN) or sublinear time during querying.

[0173] Hash index: Using algorithms such as LSH (Locality-Sensitive Hashing), perform bucketing on vectors in a high-dimensional space, and during querying, only perform exact comparison within the same or similar hash buckets to improve the query speed.

[0174] This index process is usually carried out offline in batches, and can also be incrementally updated as new media data continuously arrives.

[0175] Exemplarily, if new media units are continuously generated from live data, the system can periodically (or in real-time) insert the newly generated fused media representation vectors into the database and update the ANN or hash index structure to ensure the timeliness of retrieval.

[0176] In this way, through the above operations of establishing, storing, and constructing the mapping relationship, the present invention can achieve fused vector retrieval in a large-scale multimodal media library, and perform precise matching in combination with media metadata, making cross-media fusion or alignment more convenient. Compared with the traditional method of only storing original documents or file names, the method of the present invention significantly improves the cross-modal retrieval speed and accuracy, and can be applied to various application scenarios such as online education, video retrieval, intelligent monitoring, and social media content recommendation.

[0177] Regarding the above S106:

[0178] After receiving a fusion request, it is necessary to promptly provide cross-media recommendations or synthesis results for users or downstream applications. By retrieving the similarity of "fused media representation vectors" and combining media metadata (such as timestamp, frame index), other media that can complement or enhance the content of the target media unit can be quickly located.

[0179] In a specific implementation, when the system receives a fusion request for a certain target media unit, it first retrieves in the vector database other media units whose media representation vectors after fusion with the target media unit have a similarity higher than a preset threshold. Among the retrieval results, it filters out media resources that contain corresponding timestamp information and have a complementary relationship with the target media unit. According to the timestamp or frame information of the obtained media resources, it performs synchronous synthesis with the target media unit.

[0180] For example, if the target media unit is the Nth minute segment of a video, and the matched resource is an audio or subtitle text at the corresponding moment, the audio or subtitle text can be embedded into the video playback stream;

[0181] For another example, if the target media unit is a live stream segment, image / text information with the same time period or similar semantic points is found in the matched resources for real-time overlay display.

[0182] In a specific implementation, after synthesis, the present disclosure performs cross-modal combination or synchronous rendering of the target media unit and the matched media resources.

[0183] For example, the video and text are displayed in split screens on the same playback page; the audio is superimposed on the video stream to form a new multimedia stream; in a live broadcast scenario, corresponding images or text descriptions pop up at critical moments, etc.

[0184] The finally output cross-media presentation result can be played, browsed, or interacted with in the client or front-end interface.

[0185] In this way, through the pseudo-query vector generation and implicit interaction module in the offline stage, the system fuses the potential complementary demands of different media units before storage, reducing the large-scale computational amount in the online stage.

[0186] As an optional implementation manner, the cross-media display includes:

[0187] Receiving a fusion request for the target media unit;

[0188] Retrieving candidate media units in the vector database whose media representation vectors are more similar to the target media unit than a preset threshold;

[0189] Performing splicing and fusion of the target media unit and the candidate media unit according to a set fusion strategy;

[0190] Outputting the fusion result to the front-end display.

[0191] In a specific implementation, in the online operation stage, the system waits for or listens for a fusion request from the client, upper-layer service module, or third-party application. This request includes the following information:

[0192] Target media unit identifier: such as video_segment_id, audio_clip_id, text_paragraph_id, or other unique IDs;

[0193] Fusion preferences: for example, the user specifies that subtitles need to be spliced, images need to be inserted, or multi-camera video segments need to be switched;

[0194] Terminal / platform information: such as information about mobile devices, PCs, or AR / VR devices, etc. Subsequently, different fusion strategies can be processed based on this information.

[0195] In this embodiment, a unified API interface is set on the server side, such as POST / media / fusion_request. When called from the outside, the system parses the request parameters and writes them into a queue or in-memory cache to trigger the next retrieval and fusion process.

[0196] In a specific implementation, after parsing out the target media unit identifier, the following steps are performed in the vector database:

[0197] Find the corresponding "fused media representation vector" (e.g., F fused ) through the target media unit ID. If the database supports fast ID-vector mapping, the vector can be directly obtained; if the system stores the media representation vectors in separate databases or shards, the node storing the target entry needs to be located first.

[0198] Based on the representation vector of the target media unit, perform approximate nearest neighbor (ANN) search or hash bucket search to obtain a set of candidate media units {Candidate 1 , Candidate 2 ,...} with a similarity higher than a preset threshold (e.g., 0.8).

[0199] In addition, coarse ranking (ANN) can be performed first and then fine ranking (exact dot product or cosine similarity) to ensure retrieval efficiency and accuracy.

[0200] For the retrieved candidate media units, the system can sort them in descending order of similarity and perform secondary screening based on metadata information (e.g., timestamp, frame index, media type). For example, in the video subtitle scenario, only text or audio units with timestamps the same as or close to the current video period can be retained.

[0201] In a specific implementation, this application maintains a configurable fusion strategy list in this embodiment, including:

[0202] Splicing methods: such as sequential splicing (merging by timestamp), image / video overlay, or split-screen, etc.;

[0203] Priority or weight: For example, in the video + text scenario, subtitles are preferentially displayed; in the audio + video scenario, the lengths of the audio and video are preferentially aligned.

[0204] Device characteristics: If the client device is a mobile device, segmented preloading can be selected; if it is a PC or a high-performance AR / VR device, more complex 3D overlay or multi-window rendering can be allowed.

[0205] For splicing processing:

[0206] If the candidate media unit contains timestamp information, the system can splice the target media unit with the candidate media unit in the same or a similar time period. For example, in the scenario of live video replay, the video and subtitles at the same moment are merged.

[0207] For multiple retrieved images or videos, they are overlaid or processed in a picture-in-picture manner according to the frame index to generate a new composite media stream.

[0208] If the candidate media unit is text, the text content can be displayed in real time in the form of scrolling subtitles, bubble tips, or bullet screens below or in the sidebar of the video playback screen.

[0209] For audio + video fusion, an audio mixing engine (such as FFmpeg or a self-developed mixing module) can be used to adjust the volume and sampling rate to make them consistent with the target video segment.

[0210] If there are externally set special effects (such as AR filters, specific watermarks, etc.), they can be added according to the preset rules during the synthesis stage.

[0211] This embodiment can use a variety of media processing tools (such as FFmpeg, GStreamer, etc.) or a self-developed multimedia mixing engine to complete the splicing operation in a streaming or file manner and generate the final composite output stream or composite file.

[0212] Furthermore, the spliced multimedia data is encoded or format-encapsulated again. For example, the synthesized video stream is encoded and decoded using H.264 and packaged into MP4 or MPEG-TS format; the text information with overlaid subtitles is rendered to generate a resource link convenient for web or APP clients to call.

[0213] In the case of split-screen rendering, the video stream and text content can also be used as independent windows, and the front end performs layout and playback through HTML5 / JS or a mobile SDK.

[0214] In a specific implementation, the synthesized multimedia file or stream address is returned to the requesting end (such as a client browser, APP). The client performs adaptive playback according to the network conditions (such as bandwidth, latency) and device performance it is in, or performs automatic playback according to system presets.

[0215] In a live broadcast scenario, the fusion request can be executed in a real-time processing pipeline, and the synthesized live stream is pushed to a CDN or media server using protocols such as RTMP / HLS / WebRTC, and the client can achieve synchronous viewing with the host / audience through the playback address.

[0216] In an on-demand scenario, the spliced file can be stored in a media server or object storage, and the access URL is returned for the user to click and play in the front-end browser or APP interface.

[0217] In addition, this implementation can also record the user's click-through rate, dwell time, or satisfaction evaluation of the fusion content at the front end, and transmit this interaction information back to the server for dynamic optimization in subsequent fusion strategies. If it is detected later that the user has a high preference for a certain type of fusion method, the system can appropriately increase the priority of this fusion strategy in the next splicing.

[0218] In this way, the cross-media display process of the present invention for target media units can not only efficiently retrieve candidate media units, but also implement various splicing methods through the set fusion strategy, and finally present the processing results in a visual and diverse manner on the user terminal, so as to meet the cross-media fusion requirements in scenarios such as online education, live e-commerce, video conferencing, and entertainment content aggregation. Its overall process from request to output can significantly improve the complementarity and visualization effect of the content without loss of multimedia quality, and meet the user's expectation of synchronous browsing of multi-modal information with a fast response.

[0219] As an alternative implementation, the cross-media display further includes:

[0220] Slice the live stream data according to time periods, and generate the basic feature vectors and pseudo-query vectors for each time period;

[0221] Incrementally write the fused media representation vectors of the live stream data into the vector database, and use the latest media representation vectors for matching during online retrieval.

[0222] This application further slices the live stream data according to time periods, performs incremental vector writing, and online retrieval matching to adapt to large-scale and real-time data streams in the live broadcast scenario.

[0223] In a specific implementation, after receiving the live stream data, the live content is continuously captured through a pre-deployed real-time acquisition module (such as FFmpeg push stream reception or WebRTC reception), and the live stream is segmented into multiple discrete "live slices" according to a set time period (for example, every 5 seconds or every 10 seconds). After the slicing is completed, the system immediately extracts the basic feature vectors (F base ) on the server side or the GPU cluster side.

[0224] Exemplarily, if the live stream is video content, the system performs key frame extraction or sparse sampling on the 5-second or 10-second video segment, and uses a video feature extraction model (such as 3D CNN or Video Transformer) to perform convolution / self-attention analysis on the key frames or sampled frames to generate basic feature vectors corresponding to the video slice.

[0225] If the live stream contains audio, the audio signal of this period can be further sampled rate unified and Mel spectrum transformed, and an audio coding network (such as CNN+RNN or Audio Transformer) is used to extract audio features and merge them into the feature vector representation of the same live slice.

[0226] After obtaining the basic feature vector of the live slice, the pseudo-query module is called, and based on the multi-head self-attention or decoding operation of the feature vector, a pseudo-query vector (Q pseudo ) of the live slice is generated. Since live data often has high timeliness and no fixed file division, in this implementation, the video or audio segment after time slicing is regarded as a "media unit", so that the subsequent pseudo-query vector generation and cross-media implicit interaction module can both follow the same logic and model structure for offline media.

[0227] Subsequently, the basic feature vector and the pseudo-query vector are input into the implicit interaction module for fusion, and the output fused media representation vector (F fused ) represents the position of the live slice in the cross-media semantic space. Since the live stream has the characteristic of "continuously generating new segments in real time", this implementation adopts an incremental writing mechanism to continuously insert the newly generated fused media representation vector into the vector database.

[0228] In a specific implementation, a long connection with the vector database (such as Milvus, Faiss or HNSW) can be maintained in advance; after each slice is processed, the fused media representation vector (F fused ), corresponding metadata (live room ID, start and end seconds of the time period, frame information) are written into the database to generate new data entries.

[0229] If the database supports online index updates, the index construction module can be triggered after writing to perform ANN index update operations on incremental vector samples; if the database uses micro-batch indexing, the index can be updated in batches within a short time interval to balance real-time performance and throughput.

[0230] When external requests require cross-media display or related retrieval of live streams, the newly inserted live slice representation vector can be directly queried in the vector database. For example, in an educational live broadcast scenario, the fused media representation vector corresponding to a certain period of the teacher's current live broadcast content can be used to retrieve the courseware image or presentation page with the highest similarity.

[0231] For example, based on the current time (such as T 0 ) Find the most recently inserted live slice vector (falling in the time period [T 0 -5s,T 0 ] range), or filter by timestamp field to retrieve only the latest part to ensure that the results are synchronized with the live broadcast progress.

[0232] If other media resources with a similarity higher than a threshold are retrieved, the live slice can be spliced ​​or superimposed with the matched video, audio or text resources according to the fusion strategy, ultimately achieving a multimodal interactive experience of live broadcast and supplementary content on the front end.

[0233] In this way, through the above-mentioned segmented slicing, incremental writing and online retrieval process of live streaming data, this implementation method can achieve real-time synchronization of cross-media fusion in the live broadcast scene. Every time a new slice is collected, segmented, feature extracted and the fused media representation vector is generated, it is immediately written into the vector database and the index is updated, so that subsequent retrieval can obtain the latest live broadcast status or semantic information, thereby greatly improving the user experience in terms of latency and ensuring that the full media fusion solution adapts to the ever-changing content requirements in the live broadcast scene.

[0234] Based on the same inventive concept, the embodiments of the present disclosure also provide an all-media fusion system based on complementary fusion corresponding to the all-media fusion method based on complementary fusion. Since the principle of problem solving by the system in the embodiments of the present disclosure is similar to the above-mentioned all-media fusion method based on complementary fusion in the embodiments of the present disclosure, the implementation of the system can refer to the implementation of the method, and the repeated parts will not be repeated.

[0235] Reference Figure 3 , which is a schematic diagram of an all-media fusion system based on complementary fusion provided by an embodiment of the present disclosure, the system includes: a collection module 10, a first processing module 20, a feature extraction module 30, a second processing module 40, a third processing module 50, and a display module 60;

[0236] The acquisition module 10 is used to acquire different types of media data from multiple media sources and perform preliminary formatting processing on the media data; wherein, the media data includes: text, images, audio, video, and live stream data;

[0237] The first processing module 20 is used to divide the formatted media data into several processable media units according to a preset segmentation strategy; wherein, a media unit refers to a relatively independent media segment within a time or space range, and each media unit includes at least one of the following: a segment of text, one or more frames or segments of video, one or a group of images, and a segment of audio;

[0238] The feature extraction module 30 is used to perform feature extraction on different types of media data respectively to generate basic feature vectors corresponding to each media unit;

[0239] The second processing module 40 is used to use the pseudo-query module to generate pseudo-query vectors associated with each media unit based on the basic feature vectors;

[0240] The third processing module 50 is used to input the basic feature vectors and the pseudo-query vectors into the implicit interaction module, and output the fused media representation vectors; store the media representation vectors and their corresponding media metadata in the vector database, and establish an index for the media representation vectors in the vector database;

[0241] The display module 60 is used to, when receiving a fusion request for a target media unit, retrieve the media resources matching the target media unit from the vector database, and synchronously synthesize the target media unit and the matching media resources according to the timestamp information of the matching media resources for cross-media display of the target media unit.

[0242] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic. It should be understood that determining B based on A does not mean determining B only based on A, but also B can be determined based on A and / or other information.

[0243] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0244] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.

Claims

1. The full media fusion method based on complementary fusion is characterized by: include: Collect different types of media data from multiple media sources and perform preliminary formatting on the media data; wherein the media data includes: text, images, audio, video and live streaming data; The formatted media data is divided into media units according to a preset segmentation strategy; wherein a media unit refers to an independent media segment within a time or space range, and each media unit includes at least one of the following: a text, a frame or a video, an image or a group of images, and an audio segment; Perform feature extraction for different types of media data and generate basic feature vectors corresponding to each media unit; Using a pseudo query module, generating a pseudo query vector associated with each media unit based on the basic feature vector; Input the basic feature vector and the pseudo query vector into an implicit interaction module, and output a fused media representation vector; store the media representation vector and its corresponding media metadata into a vector database, and create an index for the media representation vector in the vector database; In response to receiving a fusion request for a target media unit, the media resources matching the target media unit are retrieved from the vector database, and the target media unit and the matching media resources are synchronously synthesized based on the timestamp information of the matching media resources to perform cross-media display of the target media unit.

2. The method for full media fusion based on complementary fusion according to claim 1, characterized in that: The generating of a pseudo query vector associated with each media unit comprises: Performing an attention mechanism operation on the basic feature vector to generate a potential demand representation; A pseudo query vector is generated based on the potential demand representation, and the pseudo query vector is optimized using a reconstruction loss or a contrastive loss.

3. The method for full media fusion based on complementary fusion according to claim 2, characterized in that: The output fused media representation vector includes: Concatenate the basic feature vector and the pseudo query vector to form an input sequence; Processing the input sequence through a multi-head self-attention network to generate a fused media representation vector; The fused media representation vector is output for indexing and retrieval.

4. The method for full media fusion based on complementary fusion according to claim 3 is characterized in that: The storing the media representation vector and the corresponding media metadata in a vector database and indexing the media representation vector in the vector database comprises: Establishing a mapping relationship between the media representation vector and the media metadata, wherein the media metadata includes a media type, a timestamp, and a frame index; Storing the media representation vector and its media metadata in a vector database; The media representation vectors are indexed based on an approximate nearest neighbor search algorithm or a hash index algorithm.

5. The method for full media fusion based on complementary fusion according to claim 4, characterized in that: The cross-media presentation includes: receiving a fusion request for the target media unit; Retrieving a candidate media unit in the vector database whose media representation vector similarity with the target media unit is higher than a preset threshold; splicing and fusing the target media unit and the candidate media unit according to a set fusion strategy; Output the fusion results to the front end for display.

6. The method for full media fusion based on complementary fusion according to claim 5, characterized in that: The cross-media presentation also includes: Slicing the live streaming data according to time periods, and generating the basic feature vector and pseudo query vector for each time period; The fused media representation vector of the live streaming data is incrementally written into the vector database, and the latest media representation vector is used for matching during online retrieval.

7. The method for full media fusion based on complementary fusion according to claim 6, characterized in that: The feature extraction comprises: For text, image, audio and video, a convolutional neural network, a visual Transformer, a pre-trained language model or a speech recognition model is used to extract the basic feature vector from the original media data.

8. The method for full media fusion based on complementary fusion according to claim 7, characterized in that: The cross-media presentation further includes: embedding the matching media resource into a rendering window corresponding to the target media unit.

9. The full media fusion system based on complementary fusion is characterized by: include: An acquisition module, a first processing module, a feature extraction module, a second processing module, a third processing module, and a display module; The acquisition module is used to acquire different types of media data from multiple media sources and perform preliminary formatting processing on the media data; wherein the media data includes: text, image, audio, video and live streaming data; The first processing module is used to divide the formatted media data into media units according to a preset segmentation strategy; wherein a media unit refers to an independent media segment within a time or space range, and each media unit includes at least one of the following: a text, a frame or a video, an image or a group of images, and an audio segment; The feature extraction module is used to extract features from different types of media data and generate basic feature vectors corresponding to each media unit; The second processing module is used to generate a pseudo query vector associated with each media unit based on the basic feature vector using a pseudo query module; The third processing module is used to input the basic feature vector and the pseudo query vector into the implicit interaction module, and output a fused media representation vector; store the media representation vector and its corresponding media metadata into a vector database, and create an index for the media representation vector in the vector database; The display module is used to, in response to receiving a fusion request for a target media unit, obtain media resources matching the target media unit by retrieving the vector database, and synchronously synthesize the target media unit with the matching media resources based on timestamp information of the matching media resources to perform cross-media display of the target media unit.

Citation Information

Patent Citations

  • Multi-modal feature fusion method and device, computing equipment and medium

    CN113033647A

  • Media information cross-modal retrieval method and system based on semantic alignment and medium

    CN118916529A

  • System and method for automatically recreating personal media through fusion of multimodal features

    US20170201562A1

  • Intelligent cataloging method for all-media news based on multi-modal information fusion understanding

    US20220270369A1

  • Adversarial cross-media retrieving method based on restricted text space

    WO2019148898A1

Cited By

  • Building information intelligent recommendation method and system

    CN120687687A

  • Intelligent media asset management and content production method and system, terminal and storage medium

    CN120950707A