A multi-channel new media data fusion method and system
By constructing a modality missing compensation network and a multivariate complex relationship graph, the problems of modality missing and index distortion in multi-channel new media data fusion were solved, achieving deep fusion and decoupling of cross-modal features, and improving the quality and recall of retrieval results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG VOCATIONAL COLLEGE
- Filing Date
- 2026-04-28
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies for multi-channel new media data fusion suffer from problems such as modality loss, index distortion, fragmented cross-platform spatiotemporal evolution, rigid intent matching, and low cross-modal recall. They cannot effectively overcome modality loss and redundant noise, resulting in fragmented search results.
A modality missing compensation network is used for feature compensation and unified latent space projection alignment to construct a multivariate complex relationship graph. Combined with natural language query and dynamic fusion weights, cross-modal feature deep fusion and decoupling are achieved through graph frequency domain decoupling.
It achieves improved cross-modal recall, accurately aggregates the evolution of events, generates high-quality fusion search result clusters, overcomes the impact of modality loss and redundant noise, and improves the contextual coherence and recall accuracy of search results.
Smart Images

Figure CN122113010A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for multi-channel new media data fusion. Background Technology
[0002] In the current multi-channel new media evolution environment, data exhibits high cross-platform fluidity, multimodal heterogeneity, and unavoidable modality loss. Whether tracing the evolution of specific content or conducting in-depth cross-channel feature retrieval, it is necessary to deal with massive, fragmented, and complexly interconnected underlying data. Overcoming the modality loss bottleneck in real-world scenarios, fully characterizing the spatiotemporal evolutionary dependency topology between multi-channel data, and achieving deep fusion and decoupling of user multidimensional intent and cross-modal graph features on a self-attention-based algorithmic architecture have become core technological barriers that urgently need to be overcome in the field of intelligent new media data fusion and processing.
[0003] Currently, Chinese invention patent application number 202511468698.4 discloses an intelligent processing method and system for new media data. This method calculates usability by extracting thematic relevance and suitability from historical materials, weighting features, and establishing a multi-dimensional classification index based on flat tags. User needs are then transformed into fixed numerical threshold conditions for mechanical filtering, and the filtered basic materials are integrated to generate new media content. However, this technology suffers from the following shortcomings: It lacks a low-level compensation mechanism for modal missing features, leading to feature dimension collapse. Relying solely on surface feature extraction and mechanical filtering, it cannot perform distribution-level inference and projection alignment of missing modalities in the semantic latent space, thus disrupting the independent modal branch structure and rendering subsequent deep cross-modal feature interactions ineffective. Furthermore, the data storage and indexing are extremely flattened, failing to depict the spatiotemporal topological evolution of cross-channel events. The isolated tag-based indexing system severs the graph topological relationships between real data, making it impossible to use mathematical models to characterize complex cross-platform forwarding actions and time dependencies, resulting in a complete loss of contextual coherence in the search results. The system suffers from several drawbacks. Firstly, the rigid intent matching mechanism fails to resolve intent conflicts and underlying attention modulation. It heavily relies on rigid numerical thresholds and lacks the ability to adapt weights to mutually exclusive intents. Furthermore, it cannot directly inject user intent as a bias term into the cross-attention calculation of the underlying large model, resulting in extremely superficial feature fusion. Secondly, it lacks a high-order frequency domain decoupling and posterior verification mechanism for graph signals, making it prone to model loops. It cannot utilize graph signal processing techniques to decouple and separate the core commonalities of events from the differences derived from channels at high and low frequencies. Thirdly, it relies on black-box generation and lacks the ability to introduce delayed-complete real-source data for reverse supervision and verification. Summary of the Invention
[0004] The technical problem solved by this invention is that the underlying fusion of existing technologies mostly adopts rigid static splicing, which cannot overcome the index distortion caused by modality loss and redundant noise, and completely severs the spatiotemporal evolution of events across platforms. In addition, when faced with complex natural language queries, existing technologies are unable to bridge the semantic gap between text and multimodal data, and cannot dynamically allocate feature fusion weights according to query intent, resulting in low cross-modal recall rate and highly fragmented search results.
[0005] To address the aforementioned technical problems, this invention provides the following technical solution: a method for multi-channel new media data fusion, comprising the following steps:
[0006] Step S1: Obtain multi-channel new media data to be fused, use the modality missing compensation network to perform feature compensation and unified latent space projection alignment on the multi-channel new media data, while retaining the independent feature branch encoding of each modality, project and align the multi-modal features to the shared multi-modal semantic latent space, and output the standardized feature representation and the compensation confidence bound to the standardized feature representation.
[0007] Step S2: Using the compensated confidence as a decay penalty term, feature suppression and filtering are performed on the standardized feature representation, core feature nodes are instantiated, cross-platform forwarding actions and time series information between multi-channel new media data are extracted, cross-channel relationship edges connecting each core feature node are constructed, and a multi-dimensional complex relationship graph is generated.
[0008] Step S3: The natural language query command is converted into a retrieval intent vector and a comprehensive modality propensity. The comprehensive modality propensity and the compensated confidence are multiplied in the multivariate complex relationship graph to output dynamic fusion weights. The target graph node feature response values are generated through the dynamic fusion weights.
[0009] Step S4: Decouple the target graph node feature response value from the retrieval intent vector in the graph frequency domain, and calculate the matching degree by combining dynamic fusion weights. Recall the target feature node, traverse the multivariate complex relationship graph with the target feature node as the anchor point, supplement and recall cross-channel context nodes with dependencies, aggregate the target feature node and cross-channel context nodes into events, and output the fusion retrieval result cluster.
[0010] Preferably, step S1 includes the following sub-steps:
[0011] Step S101: Obtain multi-channel new media data to be integrated, input the multi-channel new media data into a preset asymmetric dual-branch multimodal feature extraction network for independent encoding and extraction, and obtain the text modal features and non-text modal features of each channel respectively;
[0012] Text modal features and non-text modal features are concatenated and combined to generate initial multimodal features corresponding to multi-channel new media data;
[0013] Step S102: Detect the modal integrity of the initial multimodal features, extract the known modal features retained in the initial multimodal features, and input the known modal features into the pre-trained modal loss compensation network;
[0014] In the pre-trained modality missing compensation network, a generator with generative adversarial mechanism is used to perform deep semantic encoding on known modal features, and combined with manifold distribution alignment algorithm, the distribution difference of cross-channel multimodal data in the latent space is calculated using maximum mean difference.
[0015] With minimizing the distribution difference as the optimization objective, the generator is constrained to smoothly map the feature distribution of known modal features to the missing modal space, generating a first compensation feature vector corresponding to the missing modality;
[0016] A discriminator using a generative adversarial mechanism is used to perform adversarial discrimination between the first compensated feature vector and the real complete feature distribution, verifying the fidelity of the first compensated feature vector and outputting a fidelity score.
[0017] The first compensation feature vector and the known modal features are input into a pre-trained cross-modal semantic alignment model to calculate the semantic reconstruction error between the first compensation feature vector and the known modal features.
[0018] The realism score and semantic reconstruction error are weighted and fused to output the compensation confidence score that evaluates the generation quality of the first compensation feature vector. The first compensation feature vector that meets the preset quality threshold is selected as the second compensation feature vector.
[0019] Preferably, step S1 further includes the following sub-steps:
[0020] Step S103: Aggregate all complete initial multimodal features corresponding to multi-channel new media data that have not experienced modality loss;
[0021] Extract the second compensation feature vector and the compensation confidence bound to the second compensation feature vector;
[0022] The preset cross-space mapping matrix is invoked. The cross-space mapping matrix is generated in advance based on the cross-modal contrastive learning algorithm and optimized by minimizing the semantic distance between features of different modalities.
[0023] The complete initial multimodal features and the second compensation feature vector are uniformly projected and aligned to a preset shared multimodal semantic latent space using the cross-space mapping matrix. In the shared multimodal semantic latent space, the feature representations of each modality maintain independent feature branch coding structures, and the corresponding standardized feature representations are output.
[0024] The compensated confidence level is used as a quality label, associated and bound with the corresponding standardized feature representation, and then output.
[0025] Preferably, step S2 includes the following sub-steps:
[0026] Step S201 involves inputting the standardized feature representation into a preset multimodal interleaved sparse self-attention network, combining the compensated confidence bound to the standardized feature representation, calculating the feature response weights of each modality in both spatial and channel dimensions, outputting the intermodal correlation through a cross-modal attention matrix, and triggering an adaptive sparse feature selection strategy, specifically including:
[0027] By using the compensated confidence as a decay penalty term, the feature response weight of low confidence features is adaptively reduced, and irrelevant noise information with feature response weights less than the preset response activation value is identified. The irrelevant noise information is suppressed by redundant features, and non-activated redundant features in the single-modal state are filtered out.
[0028] Step S202: Based on the sparse feature space after the redundancy feature suppression, extract feature pairs whose feature response weights are all greater than a preset response threshold and whose intermodal correlation is greater than a preset correlation threshold as complementary features, and retain the complementary features as highly correlated key features.
[0029] The highly correlated key features are spatially aggregated to generate a fusion feature cluster containing independent modal branch encodings;
[0030] Tracing the multi-channel new media data corresponding to the standardized feature representation, obtaining the source data entity identifier corresponding to the standardized feature representation, and jointly binding the fused feature cluster, the pre-bound compensation confidence degree corresponding to the fused feature cluster, and the source data entity identifier, instantiating them as core feature nodes in the graph data structure.
[0031] Preferably, step S2 further includes the following sub-steps:
[0032] Step S203: Based on the source data entity identifier, parse the meta data packets of the source data corresponding to each core feature node, and extract the cross-platform forwarding actions and time series information of the multi-channel new media data across different network platforms;
[0033] Based on the time series information, the derived time difference of events of different core feature nodes is calculated, and the context evolution dependency direction between core feature nodes is determined by combining the cross-platform forwarding action.
[0034] Directed edges connecting each core feature node are constructed based on the evolutionary dependency direction, and the direction of the directed edges is determined by the evolutionary dependency direction.
[0035] The graph edge weights of directed edges are calculated using an exponential decay function based on the derived time difference;
[0036] The topological structure is constructed by all the core feature nodes and weighted directed edges, generating a multivariate complex relationship graph.
[0037] Preferably, step S3 includes the following sub-steps:
[0038] Step S301: Receive the natural language query instruction sent by the client and input the natural language query instruction into the pre-deployed large language model;
[0039] The large language model is used to identify entity information and implicit search intent in natural language query instructions, perform latent logical reasoning and fine-grained decoding on natural language query instructions, and transform single-dimensional natural language query instructions into multi-dimensional search intent vectors, which include multimodal expected features.
[0040] Step S302: Quantify the distribution ratio of each multimodal expected feature in the retrieval intent vector, and transform the retrieval intent vector into a semantic prompt with feature guidance attributes;
[0041] Extract the core intent tags from the semantic prompts and query the pre-built intent modality history mapping table to obtain the prior correlation between the core intent tags and different modality features in the historical retrieval statistics. The intent modality history mapping table stores the prior correlation between the core intent tags and different modality features in the historical retrieval statistics.
[0042] By combining the distribution ratio with prior correlation, the initial modality tendency corresponding to the semantic prompt is analyzed and determined, and the attention conflict detection mechanism is triggered to determine whether there is a core intent tag pointing to a mutually exclusive modality in the semantic prompt.
[0043] Preferably, step S3 further includes the following sub-steps:
[0044] Step S303: When an attention conflict is detected, the conflict resolution module is activated, and the core intent tag is used as the query input to call a preset modality priority rule base for matching, specifically including:
[0045] If a rule priority is matched, the tendency of the high-priority core intent tag is maintained, and the tendency of the low-priority core intent tag is attenuated according to a preset ratio, and the first tendency correction value is output.
[0046] If no rule priority is matched, a preset lightweight conflict matching network is used to adaptively allocate weights to conflicting intent tags and output a second tendency correction value. The lightweight conflict matching network is configured to adaptively allocate weights to conflicting core intent tags.
[0047] The first or second tendency correction value is used as an update parameter to correct the initial modal tendency and output the comprehensive modal tendency.
[0048] The comprehensive modal tendency is extracted as an intent-driven weight, and the compensation confidence bound to the core feature nodes of the multivariate complex relationship graph is extracted as a data quality penalty item.
[0049] The intent-driven weights and data quality penalty terms are multiplied to generate dynamic fusion weights for allocation to different modal branches in the core feature nodes.
[0050] The semantic prompts with feature guidance attributes and the core feature nodes in the multivariate complex relationship graph are input into the preset feature matching network. The semantic prompts are used as query conditions to trigger cross-modal cross-attention calculation between the core feature nodes and the independent modal branch codes contained in the core feature nodes.
[0051] In the cross-modal attention calculation, the dynamic fusion weight is injected as an attention bias term, and the attention scores of different modal branches are adjusted using the dynamic fusion weight to output the feature response values of the target map node.
[0052] Preferably, step S4 includes the following sub-steps:
[0053] Step S401: Combining the preset cross-modal frequency decomposition module, using the preset frequency domain transformation algorithm, the target map node feature response value and the retrieval intent vector are decoupled in the map frequency domain to separate the low-frequency map signal features and the high-frequency map signal features.
[0054] Calculate the low-frequency similarity between low-frequency map signal features and the corresponding low-frequency components of the retrieval intent vector, and calculate the high-frequency similarity between high-frequency map signal features and the corresponding high-frequency components of the retrieval intent vector;
[0055] The low-frequency similarity and high-frequency similarity are weighted and summed to obtain the matching degree of each core feature node;
[0056] Core feature nodes with a matching degree greater than a preset matching threshold are selected as target feature nodes that meet the matching conditions for recall.
[0057] Step S402: Using the recalled target feature node as the initial anchor point, the long-distance perception module is triggered in the multi-dimensional complex relationship graph in combination with the core intent label. The graph is traversed outward along the cross-channel relationship edge, and the semantic decay distance and time evolution step between the adjacent nodes and the target feature node on the propagation path are calculated.
[0058] The adjacent nodes whose semantic decay distance meets the preset distance threshold range and whose time evolution step size meets the preset time threshold range are extracted as associated nodes;
[0059] The associated nodes are treated as cross-channel context nodes that have a strong dependency relationship with the target feature nodes, and supplementary recall is performed.
[0060] Preferably, step S4 further includes the following sub-steps:
[0061] Step S403: parse the target feature node and cross-channel context node, and extract the corresponding source data timestamp and cross-platform derived level;
[0062] Based on the chronological order of the source data timestamps and the topological connectivity logic of the cross-platform derivative levels, the event evolution path is reconstructed, which includes the cause node, the fermentation node, the derivative node, and the reversal node.
[0063] Based on the event evolution, the target feature nodes and cross-channel context nodes are aggregated and sorted in the spatiotemporal dimension to generate a cluster of fused retrieval results;
[0064] Step S404: Extract the real homogeneous data that are delayed in completion as the event evolves from the multi-channel fusion retrieval result cluster as a posterior supervision signal;
[0065] The accuracy of the second compensation feature vector generated in step S102 is reverse-verified using the posterior supervision signal.
[0066] The verification deviation between the second compensation feature vector and the complete cross-channel data is calculated, and the pre-trained modality loss compensation network is updated based on the verification deviation feedback.
[0067] A multi-channel new media data fusion system includes an alignment module, a mapping module, a fusion module, and a retrieval module;
[0068] The alignment module is used to acquire multi-channel new media data to be fused, and to perform feature compensation and unified latent space projection alignment on the multi-channel new media data using a modality missing compensation network. While retaining the independent feature branch encoding of each modality, the multi-modal features are projected and aligned to a shared multi-modal semantic latent space, and the standardized feature representation and the compensation confidence bound to the standardized feature representation are output.
[0069] The graph building module is used to use the compensated confidence as a decay penalty term to suppress and filter the standardized feature representation, instantiate the core feature nodes containing the compensated confidence, extract cross-platform forwarding actions and time series information between multi-channel new media data, construct cross-channel relationship edges connecting each core feature node, and generate a multi-dimensional complex relationship graph.
[0070] The fusion module is used to convert natural language query commands into retrieval intent vectors and comprehensive modality propensity. In the multivariate complex relationship graph, the comprehensive modality propensity and compensated confidence are multiplied to output dynamic fusion weights. The target graph node feature response values are generated through the dynamic fusion weights.
[0071] The retrieval module is used to decouple the target graph node feature response value from the retrieval intent vector in the graph frequency domain, and calculate the matching degree by combining dynamic fusion weights to recall the target feature node. Using the target feature node as the anchor point, it traverses the multivariate complex relationship graph, supplements and recalls cross-channel context nodes with dependencies, aggregates the target feature node and cross-channel context nodes into events, and outputs a fusion retrieval result cluster.
[0072] The beneficial effects of this invention are as follows: It performs feature compensation in a shared multimodal semantic latent space using a manifold distribution alignment algorithm, while preserving the independent feature branch encoding of each modality to overcome feature dimension collapse. It extracts cross-platform forwarding actions and time-series information to construct a multi-dimensional complex relationship graph, fully characterizing the spatiotemporal evolution dependency topology. Facing complex query conflicts, it utilizes a lightweight conflict matching network for adaptive weight allocation, and directly injects the generated dynamic fusion weights as attention bias terms into the cross-modal cross-attention calculation layer for score adjustment. Considering the precise separation of core commonalities and channel-derived differences in events, it uses a graph Laplacian matrix for graph frequency domain decoupling, separating low-frequency graph signal features and high-frequency graph signal features to calculate similarity separately. This enables the system to achieve deep fusion and decoupling of user multidimensional intent and cross-modal graph features on a self-attention-based algorithm architecture. It not only effectively filters redundant features using compensated confidence as a decay penalty term, but also improves cross-modal recall, accurately aggregating fragmented data into an event evolution context and fusion retrieval result cluster containing cause, fermentation, derivation, and reversal nodes. Attached Figure Description
[0073] Figure 1A flowchart illustrating the steps of a multi-channel new media data fusion method according to an embodiment of the present invention;
[0074] Figure 2 This is a basic flowchart of a multi-channel new media data fusion system provided in one embodiment of the present invention. Detailed Implementation
[0075] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0076] Example 1, referring to Figure 1 This paper provides a method for multi-channel new media data fusion, which includes the following steps:
[0077] Step S1: Obtain multi-channel new media data to be fused, and use a modality missing compensation network to perform feature compensation and unified latent space projection alignment on the multi-channel new media data. While retaining the independent feature branch encoding of each modality, the multi-modal features are projected and aligned to the shared multi-modal semantic latent space, and the standardized feature representation and the compensation confidence bound to the standardized feature representation are output.
[0078] Step S2 involves using the compensated confidence as a decay penalty term to suppress and filter the standardized feature representation, instantiating core feature nodes containing the compensated confidence, extracting cross-platform forwarding actions and time series information between multi-channel new media data, constructing cross-channel relationship edges connecting each core feature node, and generating a multivariate complex relationship graph.
[0079] Step S3: The natural language query command is converted into a retrieval intent vector and a comprehensive modality tendency. The comprehensive modality tendency and the compensated confidence are multiplied in the multivariate complex relationship graph to output dynamic fusion weights. The target graph node feature response values are generated through the dynamic fusion weights.
[0080] Step S4: Decouple the target graph node feature response value from the retrieval intent vector in the graph frequency domain, and calculate the matching degree by combining dynamic fusion weights. Recall the target feature node, traverse the multivariate complex relationship graph with the target feature node as the anchor point, supplement and recall cross-channel context nodes with dependencies, aggregate the target feature node and cross-channel context nodes into events, and output the fusion retrieval result cluster.
[0081] In a specific embodiment, step S1 includes the following sub-steps:
[0082] Step S101: Obtain multi-channel new media data to be integrated, input the multi-channel new media data into a preset asymmetric dual-branch multimodal feature extraction network for independent encoding and extraction, and obtain the text modal features and non-text modal features of each channel respectively.
[0083] Text modal features and non-text modal features are concatenated and combined to generate initial multimodal features corresponding to multi-channel new media data.
[0084] Step S102: Detect the modal integrity of the initial multimodal features, extract the known modal features retained in the initial multimodal features, and input the known modal features into the pre-trained modal loss compensation network.
[0085] In the pre-trained modality missing compensation network, a generator with generative adversarial mechanism is used to perform deep semantic encoding on known modal features, and combined with the manifold distribution alignment algorithm, the distribution difference of cross-channel multimodal data in the latent space is calculated using the maximum mean difference.
[0086] With minimizing distribution differences as the optimization objective, the constraint generator smoothly maps the feature distribution of known modal features to the missing modal space, generating the first compensation feature vector corresponding to the missing modality.
[0087] A discriminator using a generative adversarial mechanism is used to perform adversarial discrimination between the first compensated feature vector and the real complete feature distribution, verifying the fidelity of the first compensated feature vector and outputting a fidelity score.
[0088] The first compensated feature vector and the known modal features are input into the pre-trained cross-modal semantic alignment model, and the semantic reconstruction error between the first compensated feature vector and the known modal features is calculated.
[0089] The fidelity score and semantic reconstruction error are weighted and fused to output the compensation confidence score that evaluates the quality of the first compensation feature vector generation. The first compensation feature vector that meets the preset quality threshold is selected as the second compensation feature vector.
[0090] It should be noted that the generator preferably adopts a residual symmetric autoencoder architecture, in which the feature encoding layer consists of several layers of one-dimensional convolutions, used to perform deep semantic extraction on known modal features, and the decoding layer (i.e. the generation layer) maps the features to the missing modal space through transposed convolutions or fully connected layers.
[0091] To prevent gradient vanishing, skip connections are added between the encoding and decoding layers. The discriminator employs a multilayer perceptron classification architecture. The input layer dimension is consistent with the multimodal feature dimension, containing 2 to 3 hidden layers. Each layer incorporates a Dropout mechanism to prevent overfitting. The output layer uses the Sigmoid activation function to map the input features to... The interval is used to determine whether the input features come from true complete sampling or compensation from the generator.
[0092] The training objective function of the modality missing compensation network consists of adversarial loss and manifold alignment loss, with the total loss function being... The mathematical expression is:
[0093] ;
[0094] in, The total loss function of the modality missing compensation network is... To counter the loss function, The maximum mean difference weighting coefficient is used to adjust the ratio of adversarial loss to distribution alignment loss. This is the manifold distribution alignment loss function.
[0095] Adversarial loss function The mathematical expression is:
[0096] ;
[0097] in, This represents the mathematical expectation operation. For data distribution with true and complete features, From The true complete feature vector obtained by sampling in the middle, Given the data distribution with known modal characteristics, Indicates from The known modal feature vector obtained by sampling in the middle, The probability value output by the discriminator. This is the first compensated feature vector output by the generator.
[0098] Manifold Distribution Alignment Loss Function The maximum mean difference is used to calculate the distance between the distribution of known mode transformations and the distribution of the target missing mode in manifold space. The mathematical expression is:
[0099] ;
[0100] in, The number of samples with known modal feature vectors in a training batch. Index for summing samples ( ), For the generator to the first Compensation features generated from known modal feature vectors This represents the number of samples in a training batch that contain true complete feature vectors. Index for summing samples ( ), For the first A true and complete feature vector, The Gaussian kernel mapping function is configured to nonlinearly project the input features onto the reproducing kernel Hilbert space. For the 2-norm square operation in the reproducing kernel Hilbert space.
[0101] Maximum mean difference weighting coefficient This is used to adjust the ratio of adversarial loss to distribution alignment loss, and its preferred value range is... During the pre-training phase, ablation experiments are conducted on an independent validation set using a grid search algorithm. If the generated feature distribution has a high fidelity score but the final map retrieval accuracy is low, the ablation rate is increased within a preset range. Strengthen manifold alignment constraints.
[0102] If the final map retrieval accuracy is high but the generated features exhibit pattern collapse, then reduce the accuracy within a preset range. This is to enhance the diversity of samples generated by adversarial networks.
[0103] Compensation confidence The weighted fusion mathematical expression is as follows:
[0104] ;
[0105] in, To compensate for the confidence level, The realism score output by the discriminator, and , The semantic reconstruction error output by the cross-modal semantic alignment model is calculated using mean squared error. For the natural constant An exponential function with base 0. To balance the weights, the preferred value range is... , The optimal range for the smoothing penalty factor is [missing value]. .
[0106] The balancing weights are dynamically adjusted based on the severity of modality loss in the current channel. If the proportion of missing modalities in a certain channel exceeds the preset missing threshold (e.g., 70%), the system will automatically reduce the threshold. The value of is adjusted to increase the semantic reconstruction error. The penalty weight in the weighting ensures that the generated completion features are semantically rigorous, preventing semantic drift caused by simply pursuing realism.
[0107] The preset quality threshold is obtained through empirical statistical analysis and testing of the quality scores of a large number of historically generated samples during the model training validation phase. With the compensation confidence level normalized to the [0, 1] interval, the preferred value for the preset quality threshold is typically set to 0.8 or 0.85. Here, the preset quality threshold acts as a quality gate for the flow of underlying features, specifically designed to intercept and discard low-quality or semantically distorted fake data generated by the network. This ensures that only high-confidence first-compensation feature vectors can be upgraded to second-compensation feature vectors and enter the true latent space, thereby preventing inferior features from contaminating subsequent graph construction and retrieval matching.
[0108] Step S103: Aggregate all complete initial multimodal features corresponding to multi-channel new media data that have not experienced modality loss.
[0109] Extract the second compensation feature vector and the compensation confidence bound to the second compensation feature vector.
[0110] The pre-defined cross-space mapping matrix is invoked. The cross-space mapping matrix is generated in advance based on the cross-modal contrastive learning algorithm and optimized by minimizing the semantic distance between features of different modalities.
[0111] The complete initial multimodal features and the second compensation feature vector are uniformly projected and aligned to a preset shared multimodal semantic latent space using a cross-space mapping matrix. Within the shared multimodal semantic latent space, the feature representations of each modality maintain independent feature branch encoding structures, and the corresponding standardized feature representations are output.
[0112] The compensated confidence level is used as a quality label, which is then associated and bound to the corresponding standardized feature representation and output.
[0113] This invention addresses the problem of modal loss in real-world, multi-channel new media data, and the feature dimension collapse caused by the forced dimensionality reduction and splicing of traditional fusion methods. It proposes using a generative adversarial mechanism combined with manifold space mapping to reconstruct missing data. Known modal features are input into a pre-trained modality loss compensation network. The maximum mean difference is used to calculate the distribution difference, constraining the generator to smoothly map to the missing modality space to generate the first compensation feature vector. To prevent low-quality generated features from polluting the retrieval latent space, a weighted fusion of the discriminator's fidelity score and semantic reconstruction error is performed, outputting a compensation confidence score to rigorously select the second compensation feature vector. Finally, considering the need to cross different... The semantic gap between modalities is preserved while retaining the boundaries of underlying features. Further, we can call a cross-space mapping matrix to uniformly project and align the complete initial multimodal features and the second compensation feature vector into a preset shared multimodal semantic latent space. This ensures that the standardized feature representation strictly preserves the independent feature branch encoding structure of each modality within the shared multimodal semantic latent space. The compensation confidence is used as a quality label and associated with the standardized feature representation. This not only accurately fills the modal gaps from the algorithmic level, but also avoids the flattening and collapse of multimodal features. It provides a standardized feature representation with underlying data quality penalty basis and absolutely clear modal boundaries for multi-channel new media data fusion.
[0114] In a specific embodiment, step S2 includes the following sub-steps:
[0115] Step S201 involves inputting the standardized feature representation into a pre-defined multimodal interleaved sparse self-attention network. Combining the compensated confidence bound to the standardized feature representation, feature response weights for each modality are calculated in both spatial and channel dimensions. The intermodal correlation is then output through a cross-modal attention matrix, triggering an adaptive sparse feature selection strategy. Specifically, this includes:
[0116] By using the compensated confidence level as the attenuation penalty term, the feature response weight of low confidence features is adaptively reduced, and irrelevant noise information with feature response weights less than the preset response activation value is identified. The irrelevant noise information is then suppressed by redundant features, and non-activated redundant features in the single-modal state are filtered out.
[0117] It should be noted that the multimodal interleaved sparse self-attention network is composed of spatial attention branches and channel attention branches interleaved and concatenated. The spatial attention branch is configured with 8 parallel attention heads, and the feature dimensions of query, key, and value are all set to 512 dimensions. The channel attention branch uses a one-dimensional convolutional layer to extract the global correlation between channels. Interleaving means that the normalized feature representation first enters the spatial attention branch for local semantic enhancement, and the output spatial feature matrix then enters the channel attention branch to perform cross-modal dimension alignment, thereby achieving semantic decoupling and information flow between the spatial and channel dimensions.
[0118] Feature response weights incorporate compensated confidence as a quality penalty term into the computation operator. The mathematical expression for feature response weights is:
[0119] ;
[0120] in, For feature response weights, and These represent the query matrix and key matrix after spatial encoding, respectively. Scaling factor This represents the Sigmoid activation function. The compensation confidence level is the value of the feature node.
[0121] The above formula can be used to map the generated confidence score to a scaling bias of the attention score, thereby suppressing low-quality features.
[0122] Step S202: Based on the sparse feature space after redundant feature suppression, feature pairs with feature response weights greater than a preset response threshold and intermodal correlation greater than a preset correlation threshold are extracted as complementary features, and the complementary features are retained as highly correlated key features.
[0123] Redundant feature filtering is achieved through a hard thresholding gating mechanism, the specific logic of which is as follows:
[0124] For each modal branch's feature vector, if the corresponding feature response weight is less than a preset response threshold, all components of that feature vector are forcibly set to zero. Subsequently, feature pairs whose feature response weights and intermodal correlations both meet the threshold are extracted as complementary features, and max pooling is performed for spatial aggregation. The aggregated feature clusters retain the feature boundaries of independent modal branches and are mapped and bound to the source data entity identifiers, ultimately instantiating into graph data core feature nodes containing multidimensional semantic attributes.
[0125] The preset response threshold is determined by the distribution range of attention scores for non-missing modal features on the statistical validation set. The intermodal correlation is calculated based on cosine similarity, and the preset correlation threshold is determined based on the precision and recall curves of historical labeled data. The correlation value corresponding to a precision of 90% is taken as the threshold, with an optimal range of 0.75 to 0.8.
[0126] It should be noted that the preset response threshold and preset correlation threshold are jointly calibrated by statistically analyzing the feature activation distribution range output by the self-attention network during the training phase and the average attention score of positive sample feature pairs across modalities. Assuming that both the feature response weights and inter-modal correlation are normalized to the [0, 1] interval, the preferred value for the preset response threshold is typically set to 0.6 to 0.7, and the preferred value for the preset correlation threshold is typically set to 0.75 to 0.8. The preset response threshold and preset correlation threshold together act as a dual filtering barrier for extracting complementary features. The preset response threshold directly eliminates invalid single-point features with excessively low activation, while the preset correlation threshold strictly intercepts weakly correlated feature pairs with little semantic overlap, thereby accurately extracting complementary features that are both self-representing and deeply supportive of each other, preventing noise or isolated features from mixing into core feature nodes.
[0127] Highly correlated key features are spatially aggregated to generate fused feature clusters containing independent modal branch encodings, which serve as the feature payloads of the core feature nodes.
[0128] Tracing the multi-channel new media data corresponding to the standardized feature representation, obtaining the source data entity identifier corresponding to the standardized feature representation, and jointly binding the fused feature cluster, the pre-bound compensation confidence corresponding to the fused feature cluster, and the source data entity identifier, instantiating them as core feature nodes in the graph data structure.
[0129] Step S203: Based on the source data entity identifier, parse the metadata data of the source data corresponding to each core feature node, and extract the cross-platform forwarding actions and time series information of multi-channel new media data across different network platforms.
[0130] The derived time difference of events of different core feature nodes is calculated based on time series information, and the context evolution dependency direction between core feature nodes is determined by combining cross-platform forwarding actions.
[0131] Directed edges are constructed to connect each core feature node based on the direction of evolutionary dependency, and the direction of the directed edges is determined by the direction of evolutionary dependency.
[0132] The graph edge weights of directed edges are calculated using an exponential decay function based on the derivation time difference, which makes the graph edge weights of node pairs with smaller derivation time differences larger. The exponential decay function makes the graph edge weights negatively correlated with the derivation time difference.
[0133] The topological structure is constructed by all core feature nodes and weighted directed edges, generating a multivariate complex relationship graph.
[0134] It should be noted that the multi-channel new media data fusion system uses a pre-set acquisition engine and regular expressions to extract relevant fields from source data metadata to identify cross-platform forwarding actions. The pre-set acquisition engine is the system's behavior data monitoring and capture module, responsible for real-time monitoring and recording users' original interactive behaviors such as clicks, dwell times, and forwarding from multiple channels (e.g., social media, websites, and apps). The specific rules for identifying cross-platform forwarding actions are as follows:
[0135] Extract reference links, forwarding sources, or text hash features from the current channel data and match them with existing data in the historical sample library.
[0136] The mathematical expression for the derived time difference is:
[0137] ;
[0138] in, To generate time difference, This is the timestamp for the publication of derived node data on the current platform. This is the timestamp of the source node data being published on the original platform.
[0139] Before calculating the weights, the derived time differences are normalized using the maximum truncation method. Differences exceeding a preset time span (e.g., 168 hours) are uniformly mapped to 1, while the remaining parts are mapped linearly to... Interval.
[0140] Calculate the graph edge weights of directed edges using an exponential decay function based on derived time difference. The mathematical expression for the graph edge weights is:
[0141] ;
[0142] in, Let be the graph edge weights of directed edges, and This is a preset platform collaboration correction factor, with a value range of [value range missing]. This is used to characterize the natural loss of propagation across different social media platforms. This is the time decay sensitivity coefficient, with a value range of [value range missing]. , This is the normalized derived time difference.
[0143] When constructing the topology, a dual threshold for node topology connections was set: a semantic consistency threshold, requiring the cosine similarity between the standardized feature representations of adjacent nodes to be greater than 0.6; and a spatiotemporal constraint threshold, requiring the derived time difference to conform to a preset time threshold range of 0 to 72 hours. Only node pairs that simultaneously meet both thresholds will execute the action of constructing cross-channel relationship edges.
[0144] This invention addresses the problems of feature contamination caused by directly mixing low-quality compensation data during multi-channel new media data fusion, and the flattened feature representation that completely severs the spatiotemporal evolution dependency topology between cross-platform data. It proposes a deep coupling between underlying data quality assessment and graph data structure topology construction. Standardized feature representations are input into a pre-defined multimodal interleaved sparse self-attention network. Compensated confidence is used as a decay penalty term to adaptively reduce the feature response weights of low-confidence features and suppress redundant features, filtering out inactive redundant features in single-modal states. Furthermore, considering the need to preserve independent modal boundaries during feature aggregation, complementary features are extracted based on the sparse feature space. These complementary features are then used as highly correlated key features for spatial aggregation, generating a fused feature cluster containing independent modal branch codes. The fused feature cluster, compensated confidence, and source data entities are then integrated into the fused feature cluster. The identifier is jointly bound and instantiated as the core feature node in the graph data structure. Finally, considering the spatiotemporal dependence and physical attributes of real data cross-platform propagation, the idea is to extract cross-platform forwarding actions and time series information. Combine the cross-platform forwarding actions to determine the context evolution dependency direction between core feature nodes to construct the direction of directed edges. The graph edge weights of directed edges are calculated using an exponential decay function based on the derived time difference. This not only uses the compensated confidence as a decay penalty term to accurately suppress irrelevant noise information and greatly reduce the computational complexity of multimodal fusion, but also strictly preserves the independent modal branch encoding. Under the condition of strictly preserving the independent modal branch encoding, the exponential decay function based on the derived time difference is used to strictly map and connect the originally isolated multi-channel new media data with clear directions. Finally, a topological structure is jointly constructed to generate a multi-dimensional complex relationship graph that can accurately quantify the context evolution dependency direction of cross-platform events.
[0145] In a specific embodiment, step S3 includes the following sub-steps:
[0146] Step S301: Receive the natural language query command sent by the client and input the natural language query command into the pre-deployed large language model.
[0147] By using large language models to identify entity information and implicit search intent in natural language query instructions, we perform latent logical reasoning and fine-grained decoding on natural language query instructions, transforming single-dimensional natural language query instructions into multi-dimensional search intent vectors, which include multimodal expected features.
[0148] Step S302: Quantify the distribution ratio of each multimodal expected feature in the retrieval intent vector, and transform the retrieval intent vector into a semantic prompt with feature guidance attributes.
[0149] Extract the core intent tags from the semantic prompts and query the pre-built intent modality history mapping table to obtain the prior correlation between the core intent tags and different modality features in the historical retrieval statistics. The intent modality history mapping table stores the prior correlation between the core intent tags and different modality features in the historical retrieval statistics.
[0150] By combining the distribution proportion with prior correlation, the initial modality tendency corresponding to the semantic prompt is analyzed and determined, and the attention conflict detection mechanism is triggered to determine whether there is a core intent label pointing to a mutually exclusive modality in the semantic prompt.
[0151] It should be noted that, in order to obtain the prior relevance between the core intent tags and different modal features, this embodiment pre-constructs an intent modality history mapping table to provide historical statistical modality tendency references for the semantic prompts of the current query during real-time retrieval. The intent modality history mapping table is stored in a key-value pair structure, where the key is the core intent tag, i.e., a semantic unit with a clear direction extracted from the user query (e.g., press conference, financial report, and evaluation). The value is a vector with a dimension equal to the number of modalities, representing the historical relevance statistics between the corresponding intent tag and each modal feature. For example, if the system supports four modalities—text, image, video, and audio—then the value is a four-dimensional vector, such as [0.2, 0.7, 0.8, 0.1] representing the relevance of the intent tag to the text, image, video, and audio modalities, respectively. The intent modality history mapping table is pre-constructed through offline statistical methods, specifically including:
[0152] Historical user query commands are collected from system logs. Each query command records the type of search result the user ultimately clicked or interacted with. Each historical query command is input into a large language model, and the same fine-grained decoding method as in step S301 is used to extract the core intent tags contained in the query command. For each core intent tag, the modality distribution corresponding to the search results that led to user interaction in historical searches is statistically analyzed. For example, for the tag "launch event," if the user ultimately clicked on video results (80%), image results (15%), text results (5%), and audio results (0%) in all historical queries containing the tag "launch event," a relevance vector [0.05, 0.15, 0.80, 0.00] is generated. The statistical results are normalized to ensure that the sum of all vector components is 1. Then, the core intent tags are associated with and stored with the corresponding relevance vectors, thus constructing the intent modality history mapping table. During real-time retrieval, after the system extracts the core intent tag from the current semantic prompt, it queries the historical mapping table of intent modalities using the core intent tag as the key. If the query matches (i.e., the core intent tag exists in the mapping table), the corresponding relevance vector is directly read as prior relevance. If the query does not match (i.e., the core intent tag does not exist in the mapping table), the prior relevance is set to a default value, such as equal weighting across modalities or migrating from the closest existing tag based on semantic similarity. The obtained prior relevance is combined with the distribution ratio of the current query to jointly determine the initial modality bias.
[0153] Step S303: When an attention conflict is detected, the conflict resolution module is activated, and the core intent label is used as the query input to call the preset modality priority rule base for matching, specifically including:
[0154] If a rule priority is matched, the bias of the high-priority core intent tag is maintained, and the bias of the low-priority core intent tag is attenuated according to a preset ratio, and the first bias correction value is output.
[0155] If no rule priority is matched, the default lightweight conflict matching network is used to adaptively allocate weights to conflicting intent tags and output a second tendency correction value. The lightweight conflict matching network is configured to adaptively allocate weights to conflicting core intent tags.
[0156] The initial modal propensity is corrected by using either the first or second propensity correction value as an update parameter, and the overall modal propensity is output.
[0157] The comprehensive modal tendency is extracted as the intention-driven weight, and the compensated confidence bound to the core feature nodes of the multivariate complex relationship graph is extracted as the data quality penalty item.
[0158] The intent-driven weights and data quality penalty terms are multiplied to generate dynamic fusion weights for assignment to different modal branches in the core feature nodes.
[0159] The semantic prompts with feature guidance attributes and the core feature nodes in the multivariate complex relationship graph are input into the preset feature matching network. The semantic prompts are used as query conditions to trigger cross-modal cross-attention calculation between the core feature nodes and the independent modal branch codes contained in the core feature nodes.
[0160] In cross-modal attention computation, dynamic fusion weights are injected as attention bias terms. The attention scores of different modal branches are adjusted using dynamic fusion weights, and the target map node feature response values that integrate intention tendency and data quality are output.
[0161] It should be noted that this embodiment introduces a pre-trained lightweight conflict allocation network to automatically output the optimal weight allocation scheme when the mutually exclusive modal conflict rules are not matched in the preset modal priority rule base. Specifically, it includes:
[0162] Network Structure and Training: The lightweight conflict matching network adopts a three-layer multilayer perceptron architecture. The input layer receives a concatenated vector of conflict intent labels transformed by a pre-trained label embedding layer. The hidden layers use ReLU activation function and Dropout mechanism, and the output layer uses Softmax activation function. The output satisfies... The normalized weight allocation is determined. During the training phase, mean squared error is used as the loss function, the concatenated vector of labels from historical conflict queries is used as the input sample, and the optimal weight allocation is used as the supervision label for parameter optimization.
[0163] The training data for the lightweight conflict matching network comes from user retrieval behavior sequences recorded in real time in the system's historical operation logs by a preset acquisition engine. The preset acquisition engine is the system's behavior data monitoring and capture module, responsible for monitoring and recording users' original interaction behaviors such as clicks, dwell times, and forwarding from multiple channels (such as social media, web pages, and apps) in real time. It is the physical source for building the training sample library and posterior supervision signals.
[0164] The sample format for lightweight conflict matching networks is as follows:
[0165] The input sample is from The feature vector is composed of conflict intent labels. Each component is the embedding representation of the corresponding label in a 768-dimensional semantic space.
[0166] The supervision labeling rules for lightweight conflict matching networks are as follows:
[0167] The supervisory labels (i.e., optimal weighting) are generated using a posterior backpropagation method based on user feedback behavior. For historical conflict retrieval tasks, the click-through rate and average dwell time of users on different modalities are statistically analyzed, and a weighted scoring model is used to calculate the contribution score of each modality to satisfying the intent. The supervision weight label corresponding to the sample is generated through normalization. .
[0168] Mean squared error is used as the loss function during the training phase. loss function The mathematical expression is:
[0169] ;
[0170] in, The total number of samples in a training batch. This represents the number of conflicting intent tags in a single query. The first output of the network In the nth sample The predicted weights of each label, This represents the corresponding supervisory label weight.
[0171] To ensure that the sum of the multiple weight components of the output equals 1, the output layer implements constraints by performing a Softmax normalization function. This applies to the raw score of the output of the last layer of the multilayer perceptron. It is converted into the final weight. The calculation process is as follows:
[0172] ;
[0173] By using exponential mapping and global summation normalization, the following conditions are met: Physical constraints.
[0174] When the application logic in step S303 detects a core intent label pointing to a mutually exclusive modality, assuming that label A and label B conflict and the rule is not matched, the system performs the following operations:
[0175] The embedding vectors of label A and label B are concatenated and input into the trained lightweight conflict matching network, which outputs normalized weights. and Normalized weights and As the second tendency correction value, the initial modality tendency vector The components corresponding to label A and label B are multiplied by respectively and Then, it is re-normalized to output the comprehensive modal tendency vector. .
[0176] In step S303, the system injects the dynamically fused weights as attention bias terms into the underlying layer of cross-modal cross-attention calculation, specifically including:
[0177] The standard cross-modal attention calculation formula without introducing a bias term is expressed as follows:
[0178] ;
[0179] in, The query matrix generated for semantic hints, and These are the key matrix and value matrix generated by the independent modal branch encoding in the core feature nodes, respectively. This is the scaling factor.
[0180] In this embodiment, the dynamically fused weights are used as the attention bias matrix. The injection calculation formula is expanded to:
[0181] ;
[0182] Attention bias matrix The generation and modulation of the dynamic fusion weight vector corresponding to each core feature node are obtained. Dynamically fuse weight vectors The attention bias matrix is expanded to have the same dimensions as the attention score matrix through a broadcast mechanism. .
[0183] By extending the formula, the dynamic fusion weights are applied directly to the underlying attention scores before the Softmax normalization operation. Modal branches with higher fusion weights are assigned larger bias terms, thus amplifying their corresponding attention scores, while the attention scores of modal branches with lower fusion weights are suppressed accordingly. After weighted cross-modal attention calculation, the feature matching network directly outputs the feature response values of the target map nodes, which deeply fuse intent tendency (comprehensive modality tendency) and data quality (compensated confidence).
[0184] This invention addresses the challenges of single-dimensional natural language query commands failing to bridge cross-modal semantic gaps, the susceptibility to retrieval conflicts when dealing with core intent tags pointing to mutually exclusive modalities, and the limitations of traditional feature matching which relies on superficial numerical weighting to penetrate the underlying self-attention architecture. The invention utilizes a large language model to perform latent logical reasoning and fine-grained decoding of natural language query commands, transforming them into retrieval intent vectors containing multimodal expected features. It then uses the prior relevance of core intent tags in an intent modality history mapping table to determine the initial modality bias. Furthermore, considering the potential for mutually exclusive intents in complex queries, an attention conflict detection mechanism is implemented. This mechanism calls a pre-defined modality priority rule base or utilizes a lightweight conflict matching network to adaptively allocate weights to conflicting intent tags, outputting a first or second bias correction value to adjust the initial modality bias and thus providing a comprehensive result. To further integrate implicit search intent with underlying data quality, this paper proposes using comprehensive modal tendency as the intent-driven weight and the compensated confidence bound to the core feature nodes of the multivariate complex relationship graph as the data quality penalty term. This is then multiplied to generate dynamic fusion weights. Breaking away from conventional post-processing weighting, the dynamic fusion weights are directly injected into the cross-modal attention calculation as attention bias terms. This approach not only adaptively solves the conflict matching problem of mutually exclusive modal intents through a lightweight conflict matching network, but also achieves strict synchronous adjustment of query intent guidance and underlying data quality penalties at the attention underlying architecture of the feature matching network. By adjusting the attention scores of different modal branches using dynamic fusion weights, the paper outputs the target graph node feature response value that integrates the user's multi-dimensional search intent and the true quality of the core feature nodes.
[0185] In a specific embodiment, step S4 includes the following sub-steps:
[0186] Step S401: Combining the preset cross-modal frequency decomposition module, the preset frequency domain transformation algorithm is used to decouple the target map node feature response value and the retrieval intent vector in the map frequency domain, and separate the low-frequency map signal features and the high-frequency map signal features.
[0187] Calculate the low-frequency similarity between the low-frequency image signal features and the corresponding low-frequency components of the retrieval intent vector, and calculate the high-frequency similarity between the high-frequency image signal features and the corresponding high-frequency components of the retrieval intent vector.
[0188] The matching degree of each core feature node is obtained by weighted summation of low-frequency similarity and high-frequency similarity.
[0189] Core feature nodes with a matching degree greater than a preset matching threshold are selected as target feature nodes that meet the matching conditions for recall.
[0190] It should be noted that the preset matching threshold was determined through joint testing and calibration by comprehensively evaluating the precision and recall of historical retrieval tasks on an offline validation set, seeking the optimal balance point on the precision and recall curves. With the matching degree of each core feature node normalized to the [0, 1] interval, the preferred value for the preset matching threshold is set to 0.75 to 0.8. The preset matching threshold acts as the precision threshold for the entire graph retrieval recall stage, directly eliminating background noise that is weakly or irrelevant to the user's intent from a massive number of nodes. This ensures that only core feature nodes with extremely high matching degrees can be recalled as target feature nodes, providing a reliable initial anchor point for subsequent graph traversal along cross-channel relationship edges.
[0191] Step S402: Using the recalled target feature node as the initial anchor point, the long-distance perception module is triggered in the multivariate complex relationship graph in combination with the core intent label. The graph is traversed outward along the cross-channel relationship edge, and the semantic decay distance and time evolution step between the adjacent nodes and the target feature node on the propagation path are calculated.
[0192] Neighboring nodes whose semantic decay distance and time evolution step size both meet the preset distance threshold range are extracted as associated nodes.
[0193] It should be noted that the preset distance threshold range and preset time threshold range were obtained through statistical analysis of the cross-platform propagation and evolution paths of massive historical public opinion events, and through empirical statistics and boundary calibration of the semantic decay distance and time evolution step between real related nodes. With the semantic decay distance normalized and the time evolution step calculated in hours, the preferred values for the preset distance threshold range are set to [0, 0.3], and the preferred values for the preset time threshold range are set to [0, 72 hours]. The preset distance threshold range and preset time threshold range together serve as the spatiotemporal constraint boundaries for long-distance perception of the graph, used to intercept irrelevant nodes that, although physically connected, have undergone severe semantic distortion or have excessively long time spans during graph traversal. This ensures that only adjacent nodes that are highly homogeneous in content and within the same public opinion propagation cycle can be extracted as related nodes, preventing the overgeneralization of cross-channel context nodes or the recall of meaningless noise data.
[0194] Associated nodes are treated as cross-channel context nodes that have a strong dependency on the target feature node, and supplementary recall is performed.
[0195] Step S403: parse the target feature nodes and cross-channel context nodes, and extract the corresponding source data timestamps and cross-platform derived layers.
[0196] Based on the chronological order of source data timestamps and the topological connectivity logic of cross-platform derivative levels, the event evolution path is reconstructed, which includes the cause node, fermentation node, derivative node, and reversal node.
[0197] Based on the event evolution, the target feature nodes and cross-channel context nodes are aggregated and sorted in the spatiotemporal dimensions to generate a cluster of fused search results.
[0198] Step S404: Extract the real homogeneous data that are delayed in being supplemented as the event evolves from the multi-channel fusion retrieval result cluster as a posterior supervision signal.
[0199] The accuracy of the second compensation feature vector generated in step S102 is verified by using the posterior supervision signal.
[0200] The second compensation feature vector is calculated to verify the deviation between the second compensation feature vector and the complete data across channels. The pre-trained modality loss compensation network is updated based on the verification deviation feedback.
[0201] It should be noted that the posterior supervision signal collection and model update in step S404 are an asynchronous offline process. In actual engineering deployment, steps S1 to S403 constitute a real-time response chain for a single online retrieval. After the online retrieval is completed, multi-channel data sources are continuously and asynchronously monitored. When it is detected that the real source data corresponding to the previously compensated modality is delayed in release due to event evolution, the data is automatically collected and stored in the historical sample library. When the accumulated posterior supervision signals in the sample library reach a preset batch threshold, the system calculates the verification deviation offline in the background and performs batch gradient updates on the modality missing compensation network accordingly, thereby achieving iterative optimization of the model across a single retrieval cycle.
[0202] Decoupling the target map node feature response values and the retrieval intent vector in the map frequency domain, and separating the low-frequency map signal features from the high-frequency map signal features, specifically includes:
[0203] Extracting the graph Laplacian matrix from a multivariate complex relationship graph Graph Laplace matrix The mathematical expression is:
[0204] ;
[0205] in, For degree matrix, It is an adjacency matrix.
[0206] Graph Laplace matrix Perform eigenvalue decomposition to obtain ,in, These are orthogonal eigenvector matrices, i.e., the basis of the graph Fourier transform. This is the corresponding eigenvalue diagonal matrix. In graph signal processing theory, eigenvalues... The magnitude of the eigenvalue corresponds to the frequency of the graph signal. Smaller eigenvalues correspond to signals with gradual changes in the graph (low frequency), while larger eigenvalues correspond to signals with drastic changes in the graph (high frequency). In this embodiment, the preset frequency domain transformation algorithm specifically includes graph Fourier transform and frequency domain filtering operations, for the target graph node feature response value... The graph Fourier transform is defined as follows: To separate the low-frequency and high-frequency signal characteristics, a frequency truncation threshold was preset. In this embodiment, Based on the energy distribution spectrum of the spectral features, the feature value where the cumulative energy percentage reaches 80% can be taken as an adaptive threshold, and a low-pass graph filter can be constructed. With high-frequency graph filters .
[0207] The mathematical expression for the response function of a low-pass filter is:
[0208] ;
[0209] The mathematical expression for the response function of a high-pass filter is:
[0210] ;
[0211] Using low-pass graph filters With high-frequency graph filters By modulating the features in the frequency domain and mapping them back to the spatial domain using the inverse graphical Fourier transform (IGFT), the low-frequency signal features can be obtained. It characterizes the core commonalities of events where adjacent nodes exhibit convergent features, and simultaneously obtains high-frequency graph signal features. The channel-derived differences that characterize the abrupt changes in features of adjacent nodes are represented, and the high- and low-frequency decoupling of the retrieval intent vector is also separated by synchronous projection using an isomorphic Laplacian basis.
[0212] Specifically, before entering the graph frequency domain decoupling, the dimensionality of the retrieval intent vector is first aligned and expanded to match the number of nodes in the multivariate complex relationship graph through linear mapping. Using the existing graph Laplacian eigenvector matrix of the multivariate complex relationship graph as a unified projection basis, the aligned retrieval intent vector is projected onto the same graph frequency domain space and truncated at the same frequency threshold. The low-frequency and high-frequency components corresponding to the search intent are separated.
[0213] This invention addresses the problems of existing retrieval and matching methods, such as the inability to separate commonalities of events from channel differences, the easy loss of cross-platform contextual coherence, and the lack of real feedback in modal compensation networks, which easily leads to model validation dead loops. It proposes to introduce graph signal processing technology, combining a pre-defined cross-modal frequency decomposition module with a pre-defined frequency domain transformation algorithm to decouple the target graph node feature response values and the retrieval intent vector in the graph frequency domain. This separates low-frequency and high-frequency graph signal features, calculating their similarity to recall target feature nodes. Furthermore, considering that isolated nodes cannot reconstruct the full picture of an event, the invention uses the target feature node as an initial anchor point to trigger a long-distance perception module to traverse the graph along cross-channel relationship edges, calculating semantic decay distance and time evolution step size to supplement the recall of cross-channel context nodes with strong dependencies. This reconstructs the event evolution trajectory, including causal nodes, fermentation nodes, derivative nodes, and reversal nodes. Finally, to overcome the bottleneck of the generative model lacking objective verification, the invention extracts real homogeneous data from the multi-channel fusion retrieval result clusters that are delayed and supplemented as the event evolves, as a posterior supervision signal to perform reverse verification on the second compensation feature vector. Not only does it achieve extremely fine-grained and accurate matching through graph-frequency domain decoupling, but it also fully aggregates and generates a cluster of fused retrieval results in the spatiotemporal dimension. At the same time, it uses posterior supervision signals to calculate verification deviations to feed back and update the pre-trained modality missing compensation network, thus constructing a perfect closed loop from data retrieval and recall to the self-optimization of the underlying network on the underlying algorithm architecture.
[0214] Example 2, refer to Figure 2 This paper presents a multi-channel new media data fusion system, including an alignment module, a mapping module, a fusion module, and a retrieval module.
[0215] The alignment module is used to acquire multi-channel new media data to be fused. It uses a modality missing compensation network to perform feature compensation and unified latent space projection alignment on the multi-channel new media data. While preserving the independent feature branch encoding of each modality, it projects and aligns the multi-modal features to a shared multi-modal semantic latent space, and outputs standardized feature representations and compensation confidence scores bound to the standardized feature representations.
[0216] The graph construction module uses the compensated confidence as a decay penalty term to suppress and filter the standardized feature representation, instantiate the core feature nodes containing the compensated confidence, extract cross-platform forwarding actions and time series information between multi-channel new media data, construct cross-channel relationship edges connecting each core feature node, and generate a multi-dimensional complex relationship graph.
[0217] The fusion module is used to convert natural language query commands into retrieval intent vectors and comprehensive modality propensity. In a multivariate complex relation graph, the comprehensive modality propensity and compensated confidence are multiplied to output dynamic fusion weights. The target graph node feature response values are generated through the dynamic fusion weights.
[0218] The retrieval module is used to decouple the target graph node feature response value from the retrieval intent vector in the graph frequency domain, and calculate the matching degree by combining dynamic fusion weights to recall target feature nodes. Using the target feature nodes as anchors, it traverses the multivariate complex relationship graph, supplements the recall of cross-channel context nodes with dependencies, aggregates the target feature nodes and cross-channel context nodes into events, and outputs a fusion retrieval result cluster.
[0219] This invention addresses the problem in existing new media processing system architectures where multimodal feature compensation, graph structure construction, and retrieval matching are isolated, leading to highly fragmented cross-channel data fusion and a tendency to accumulate underlying feature errors. It proposes an end-to-end joint system architecture, setting up alignment, graph construction, fusion, and retrieval modules for deep physical collaboration. Considering the modal gaps and inconsistent quality of the underlying data, the alignment module performs feature compensation and unified latent space projection alignment to output compensation confidence. The graph construction module directly uses the compensation confidence as a decay penalty term to instantiate core feature nodes, and then combines cross-platform forwarding actions and time-series information to generate a multivariate complex relationship graph, aiming to cross natural language processing boundaries. To bridge the semantic gap between spoken and complex graphs, a fusion module was developed to generate dynamic fusion weights to output the feature response values of target graph nodes. Finally, a retrieval module was used to decouple the feature response values of target graph nodes from the retrieval intent vector in the graph frequency domain. Using the recalled target feature nodes as anchors, the system traversed and supplemented the recalled cross-channel context nodes in the multi-dimensional complex relationship graph for event aggregation. This not only broke down the system silos in multi-channel new media data processing, enabling the feature representations that retain the independent feature branches of each modality to achieve low-level integration with the relationship graph, but also realized intent-driven dynamic weight fusion at the system retrieval layer. Ultimately, this provided users with a cluster of fused retrieval results that efficiently and accurately restored the cross-platform dissemination context.
[0220] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the protection scope of the present invention.
Claims
1. A method for fusion of multi-channel new media data, characterized in that, Includes the following steps: Step S1: Obtain multi-channel new media data to be fused, use the modality missing compensation network to perform feature compensation and unified latent space projection alignment on the multi-channel new media data, while retaining the independent feature branch encoding of each modality, project and align the multi-modal features to the shared multi-modal semantic latent space, and output the standardized feature representation and the compensation confidence bound to the standardized feature representation. Step S2: Using the compensated confidence as a decay penalty term, feature suppression and filtering are performed on the standardized feature representation, core feature nodes are instantiated, cross-platform forwarding actions and time series information between multi-channel new media data are extracted, cross-channel relationship edges connecting each core feature node are constructed, and a multi-dimensional complex relationship graph is generated. Step S3: The natural language query command is converted into a retrieval intent vector and a comprehensive modality propensity. The comprehensive modality propensity and the compensated confidence are multiplied in the multivariate complex relationship graph to output dynamic fusion weights. The target graph node feature response values are generated through the dynamic fusion weights. Step S4: Decouple the target graph node feature response value from the retrieval intent vector in the graph frequency domain, and calculate the matching degree by combining dynamic fusion weights. Recall the target feature node, traverse the multivariate complex relationship graph with the target feature node as the anchor point, supplement and recall cross-channel context nodes with dependencies, aggregate the target feature node and cross-channel context nodes into events, and output the fusion retrieval result cluster.
2. The multi-channel new media data fusion method as described in claim 1, characterized in that, Step S1 includes the following sub-steps: Step S101: Obtain multi-channel new media data to be integrated, input the multi-channel new media data into a preset asymmetric dual-branch multimodal feature extraction network for independent encoding and extraction, and obtain the text modal features and non-text modal features of each channel respectively; Text modal features and non-text modal features are concatenated and combined to generate initial multimodal features corresponding to multi-channel new media data; Step S102: Detect the modal integrity of the initial multimodal features, extract the known modal features retained in the initial multimodal features, and input the known modal features into the pre-trained modal loss compensation network; In the pre-trained modality missing compensation network, a generator with generative adversarial mechanism is used to perform deep semantic encoding on known modal features, and combined with manifold distribution alignment algorithm, the distribution difference of cross-channel multimodal data in the latent space is calculated using maximum mean difference. With minimizing the distribution difference as the optimization objective, the generator is constrained to smoothly map the feature distribution of known modal features to the missing modal space, generating a first compensation feature vector corresponding to the missing modality; A discriminator using a generative adversarial mechanism is used to perform adversarial discrimination between the first compensated feature vector and the real complete feature distribution, verifying the fidelity of the first compensated feature vector and outputting a fidelity score. The first compensation feature vector and the known modal features are input into a pre-trained cross-modal semantic alignment model to calculate the semantic reconstruction error between the first compensation feature vector and the known modal features. The realism score and semantic reconstruction error are weighted and fused to output the compensation confidence score that evaluates the generation quality of the first compensation feature vector. The first compensation feature vector that meets the preset quality threshold is selected as the second compensation feature vector.
3. The multi-channel new media data fusion method as described in claim 2, characterized in that, Step S1 further includes the following sub-steps: Step S103: Aggregate all complete initial multimodal features corresponding to multi-channel new media data that have not experienced modality loss; Extract the second compensation feature vector and the compensation confidence bound to the second compensation feature vector; The preset cross-space mapping matrix is invoked. The cross-space mapping matrix is generated in advance based on the cross-modal contrastive learning algorithm and optimized by minimizing the semantic distance between features of different modalities. The complete initial multimodal features and the second compensation feature vector are uniformly projected and aligned to a preset shared multimodal semantic latent space using the cross-space mapping matrix. In the shared multimodal semantic latent space, the feature representations of each modality maintain independent feature branch coding structures, and the corresponding standardized feature representations are output. The compensated confidence level is used as a quality label, associated and bound with the corresponding standardized feature representation, and then output.
4. The multi-channel new media data fusion method as described in claim 3, characterized in that, Step S2 includes the following sub-steps: Step S201 involves inputting the standardized feature representation into a preset multimodal interleaved sparse self-attention network, combining the compensated confidence bound to the standardized feature representation, calculating the feature response weights of each modality in both spatial and channel dimensions, outputting the intermodal correlation through a cross-modal attention matrix, and triggering an adaptive sparse feature selection strategy, specifically including: By using the compensated confidence as a decay penalty term, the feature response weight of low confidence features is adaptively reduced, and irrelevant noise information with feature response weights less than the preset response activation value is identified. The irrelevant noise information is suppressed by redundant features, and non-activated redundant features in the single-modal state are filtered out. Step S202: Based on the sparse feature space after the redundancy feature suppression, extract feature pairs whose feature response weights are all greater than a preset response threshold and whose intermodal correlation is greater than a preset correlation threshold as complementary features, and retain the complementary features as highly correlated key features. The highly correlated key features are spatially aggregated to generate a fusion feature cluster containing independent modal branch encodings; Tracing the multi-channel new media data corresponding to the standardized feature representation, obtaining the source data entity identifier corresponding to the standardized feature representation, and jointly binding the fused feature cluster, the pre-bound compensation confidence degree corresponding to the fused feature cluster, and the source data entity identifier, instantiating them as core feature nodes in the graph data structure.
5. The multi-channel new media data fusion method as described in claim 4, characterized in that, Step S2 further includes the following sub-steps: Step S203: Based on the source data entity identifier, parse the meta data packets of the source data corresponding to each core feature node, and extract the cross-platform forwarding actions and time series information of the multi-channel new media data across different network platforms; Based on the time series information, the derived time difference of events of different core feature nodes is calculated, and the context evolution dependency direction between core feature nodes is determined by combining the cross-platform forwarding action. Directed edges connecting each core feature node are constructed based on the evolutionary dependency direction, and the direction of the directed edges is determined by the evolutionary dependency direction. The graph edge weights of directed edges are calculated using an exponential decay function based on the derived time difference; The topological structure is constructed by all the core feature nodes and weighted directed edges, generating a multivariate complex relationship graph.
6. The multi-channel new media data fusion method as described in claim 5, characterized in that, Step S3 includes the following sub-steps: Step S301: Receive the natural language query instruction sent by the client and input the natural language query instruction into the pre-deployed large language model; The large language model is used to identify entity information and implicit search intent in natural language query instructions, perform latent logical reasoning and fine-grained decoding on natural language query instructions, and transform single-dimensional natural language query instructions into multi-dimensional search intent vectors, which include multimodal expected features. Step S302: Quantify the distribution ratio of each multimodal expected feature in the retrieval intent vector, and transform the retrieval intent vector into a semantic prompt with feature guidance attributes; Extract the core intent tags from the semantic prompts and query the pre-built intent modality history mapping table to obtain the prior correlation between the core intent tags and different modality features in the historical retrieval statistics. The intent modality history mapping table stores the prior correlation between the core intent tags and different modality features in the historical retrieval statistics. By combining the distribution ratio with prior correlation, the initial modality tendency corresponding to the semantic prompt is analyzed and determined, and the attention conflict detection mechanism is triggered to determine whether there is a core intent tag pointing to a mutually exclusive modality in the semantic prompt.
7. The multi-channel new media data fusion method as described in claim 6, characterized in that, Step S3 further includes the following sub-steps: Step S303: When an attention conflict is detected, the conflict resolution module is activated, and the core intent tag is used as the query input to call a preset modality priority rule base for matching, specifically including: If a rule priority is matched, the tendency of the high-priority core intent tag is maintained, and the tendency of the low-priority core intent tag is attenuated according to a preset ratio, and the first tendency correction value is output. If no rule priority is matched, a preset lightweight conflict matching network is used to adaptively allocate weights to conflicting intent tags and output a second tendency correction value. The lightweight conflict matching network is configured to adaptively allocate weights to conflicting core intent tags. The first or second tendency correction value is used as an update parameter to correct the initial modal tendency and output the comprehensive modal tendency. The comprehensive modal tendency is extracted as an intent-driven weight, and the compensation confidence bound to the core feature nodes of the multivariate complex relationship graph is extracted as a data quality penalty item. The intent-driven weights and data quality penalty terms are multiplied to generate dynamic fusion weights for allocation to different modal branches in the core feature nodes. The semantic prompts with feature guidance attributes and the core feature nodes in the multivariate complex relationship graph are input into the preset feature matching network. The semantic prompts are used as query conditions to trigger cross-modal cross-attention calculation between the core feature nodes and the independent modal branch codes contained in the core feature nodes. In the cross-modal attention calculation, the dynamic fusion weight is injected as an attention bias term, and the attention scores of different modal branches are adjusted using the dynamic fusion weight to output the feature response values of the target map node.
8. The multi-channel new media data fusion method as described in claim 7, characterized in that, Step S4 includes the following sub-steps: Step S401: Combining the preset cross-modal frequency decomposition module, using the preset frequency domain transformation algorithm, the target map node feature response value and the retrieval intent vector are decoupled in the map frequency domain to separate the low-frequency map signal features and the high-frequency map signal features. Calculate the low-frequency similarity between low-frequency map signal features and the corresponding low-frequency components of the retrieval intent vector, and calculate the high-frequency similarity between high-frequency map signal features and the corresponding high-frequency components of the retrieval intent vector; The low-frequency similarity and high-frequency similarity are weighted and summed to obtain the matching degree of each core feature node; Core feature nodes with a matching degree greater than a preset matching threshold are selected as target feature nodes that meet the matching conditions for recall. Step S402: Using the recalled target feature node as the initial anchor point, the long-distance perception module is triggered in the multi-dimensional complex relationship graph in combination with the core intent label. The graph is traversed outward along the cross-channel relationship edge, and the semantic decay distance and time evolution step between the adjacent nodes and the target feature node on the propagation path are calculated. The adjacent nodes whose semantic decay distance meets the preset distance threshold range and whose time evolution step size meets the preset time threshold range are extracted as associated nodes; The associated nodes are treated as cross-channel context nodes that have a strong dependency relationship with the target feature nodes, and supplementary recall is performed.
9. The multi-channel new media data fusion method as described in claim 8, characterized in that, Step S4 further includes the following sub-steps: Step S403: parse the target feature node and cross-channel context node, and extract the corresponding source data timestamp and cross-platform derived level; Based on the chronological order of the source data timestamps and the topological connectivity logic of the cross-platform derivative levels, the event evolution path is reconstructed, which includes the cause node, the fermentation node, the derivative node, and the reversal node. Based on the event evolution, the target feature nodes and cross-channel context nodes are aggregated and sorted in the spatiotemporal dimension to generate a cluster of fused retrieval results; Step S404: Extract the real homogeneous data that are delayed in completion as the event evolves from the multi-channel fusion retrieval result cluster as a posterior supervision signal; The accuracy of the second compensation feature vector generated in step S102 is reverse-verified using the posterior supervision signal. The verification deviation between the second compensation feature vector and the complete cross-channel data is calculated, and the pre-trained modality loss compensation network is updated based on the verification deviation feedback.
10. A multi-channel new media data fusion system, applied in a multi-channel new media data fusion method as described in any one of claims 1-9, characterized in that, It includes alignment, mapping, fusion, and retrieval modules; The alignment module is used to acquire multi-channel new media data to be fused, and to perform feature compensation and unified latent space projection alignment on the multi-channel new media data using a modality missing compensation network. While retaining the independent feature branch encoding of each modality, the multi-modal features are projected and aligned to a shared multi-modal semantic latent space, and the standardized feature representation and the compensation confidence bound to the standardized feature representation are output. The graph building module is used to use the compensated confidence as a decay penalty term to suppress and filter the standardized feature representation, instantiate the core feature nodes containing the compensated confidence, extract cross-platform forwarding actions and time series information between multi-channel new media data, construct cross-channel relationship edges connecting each core feature node, and generate a multi-dimensional complex relationship graph. The fusion module is used to convert natural language query commands into retrieval intent vectors and comprehensive modality propensity. In the multivariate complex relationship graph, the comprehensive modality propensity and compensated confidence are multiplied to output dynamic fusion weights. The target graph node feature response values are generated through the dynamic fusion weights. The retrieval module is used to decouple the target graph node feature response value from the retrieval intent vector in the graph frequency domain, and calculate the matching degree by combining dynamic fusion weights to recall the target feature node. Using the target feature node as the anchor point, it traverses the multivariate complex relationship graph, supplements and recalls cross-channel context nodes with dependencies, aggregates the target feature node and cross-channel context nodes into events, and outputs a fusion retrieval result cluster.