Network content propagation method and system based on fusion cognitive understanding and intelligent treatment

By adopting the network content dissemination method based on integrated cognitive understanding and intelligent governance in the information dissemination era, the problems of multimodal information integration and hotspot/abnormal content recognition are solved, and efficient and accurate content dissemination and distribution are achieved.

CN119939229AActive Publication Date: 2025-05-06NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP +2

Patent Information

Application Number
CN202510446313.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-06
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

In the current era of all media, the heterogeneity and complex interactivity of information dissemination make it difficult for traditional technologies to effectively integrate multimodal information, identify hot spots and abnormal content, and the propagation path and effect are difficult to predict and control.

Method used

A network content propagation method based on fusion cognitive understanding and intelligent governance is proposed. By decoupling data features of different modalities, a sub-graph network is constructed to convolutionize cross-modal features, identify hot spots and abnormal contents, simulate the propagation process to predict the content propagation link, and adjust the weight of key nodes to determine the optimal delivery strategy.

Benefits of technology

It effectively improves the dissemination efficiency and quality of all-media content, realizes the unified expression of multimodal information, accurately identifys hot spots and abnormal content, and improves the accuracy of propagation path prediction and the efficiency of content distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939229A_ABST
    Figure CN119939229A_ABST
Patent Text Reader

Abstract

The invention discloses a network content propagation method and system based on fusion of cognitive understanding and intelligent governance, and the method comprises the steps: constructing a sub-graph network of each mode, and generating cross-modal feature graph representation; performing topic modeling on texts in the text knowledge base, extracting potential topic words, and constructing an event graph; nodes related to the hot spots are retrieved from the event atlas, a hot spot mode is formed, content features similar to the hot spot mode are matched in cross-modal feature graph representation, and potential hot spot content is recognized; constructing a cross-modal content graph, learning an abnormal mode by using the labeled abnormal data, training an abnormal detection model according to the abnormal mode, and identifying abnormal content in the cross-modal content graph; predicting a content propagation link of the hot content and the abnormal content; by adjusting the weight of the key node in the content propagation link, the content delivery time, the content delivery position and the target audience are determined. According to the method and the device, the accurate personalized recommendation and the maximization of the propagation effect are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of deep learning technology, and in particular to a network content dissemination method and system based on the integration of cognitive understanding and intelligent governance. Background Art

[0002] In the current omnimedia era, the speed and breadth of information dissemination have increased unprecedentedly, which has put forward higher requirements for accurate perception, intelligent dissemination and effective governance of content. First, the heterogeneity and complex interactivity of omnimedia content are the main technical challenges currently faced. Data in different modes such as audio, video, image and text have significant differences in structure and semantics. Traditional single-modality processing methods cannot effectively integrate these multimodal information, resulting in incomplete and inaccurate information. In addition, the dynamic spatiotemporal coupling of information dissemination networks also increases the uncertainty of content dissemination, making it difficult to predict and control the dissemination path and effect. Secondly, existing content dissemination technologies are insufficient in identifying hot spots and abnormal content. Although traditional keyword matching and frequency statistics methods are simple, they have limited effects when processing complex texts and deep semantic information. For example, on social media, the formation of hot topics is often the result of the combined action of multiple factors, while abnormal content such as false information and sensitive images are hidden and diverse, and difficult to accurately identify through traditional methods. In addition, the prediction of the dissemination effect and distribution guidance of network content also face many challenges. Existing communication models mostly use static analysis, which cannot capture the dynamic changes in the communication process in real time, resulting in insufficient accuracy and reliability of prediction results. At the same time, the optimization of content distribution strategies also lacks effective technical support, making it difficult to achieve accurate personalized recommendations and maximize communication effects. Summary of the invention

[0003] In order to solve the above problems, this application proposes a network content dissemination method and system based on the integration of cognitive understanding and intelligent governance, which can effectively improve the dissemination efficiency and quality of all media content.

[0004] The present application discloses a network content dissemination method based on integrated cognitive understanding and intelligent governance, which includes: Step 1: Decouple data features of different modalities and output the decoupled data features; different modalities include text, video, audio, and image; Step 2: Based on the decoupled data features of audio, video, image, and text, a subgraph network for each modality is constructed, and convolution operations on cross-modal features are performed through multiple subgraph networks to generate cross-modal feature graph representations; Step 3: Perform topic modeling on the text in the text knowledge base, extract potential topic words, and extract event triplets from the potential topic words to construct an event graph; retrieve nodes related to hot spots from the event graph to form a hot spot pattern, match content features similar to the hot spot pattern in the cross-modal feature graph representation, and identify potential hot content; Step 4: Combine the cross-modal feature graph representation to build a cross-modal content map, use the labeled abnormal data to learn abnormal patterns, train an anomaly detection model based on the abnormal patterns, and identify abnormal content in the cross-modal content map through the trained anomaly detection model; abnormal content includes false information, sensitive images, and dense videos; Step 5: Simulate the process of information propagation among users and predict the content propagation links of hot content and abnormal content; Step 6: Adjust the weights of key nodes in the content dissemination chain to determine the best time, location, and target audience for content delivery.

[0005] Furthermore, the step 1 comprises: Step 11: Extract data features of text, video, audio, and image and input them into the encoder, extract features of the corresponding modality, and align video features using dynamic time warping; Step 12: Model the extracted cross-modal data features; Step 13: Extract the cross-modal decoupled data features from the modeled cross-modal features through an attention mechanism-based method.

[0006] Furthermore, the step 11 comprises: Extract text data features, video features, audio data features and image data features from text, video, audio and image respectively; The video features are calculated by the following formula And the distance between different time steps of audio data feature A, we get the cumulative cost matrix:

[0007] Based on this distance, the cumulative cost matrix cost is obtained by the following formula:

[0008] in, represents the distance between the i-th time step of the video and the j-th time step of the audio, represents the features of the video at the i-th time step, express The i-th element in represents the features of the audio at the jth time step, represents the jth element in A, represents the Euclidean distance, represents the cumulative cost matrix, From the starting point arrive The minimum cumulative cost of From the end point of the cumulative cost matrix , that is, starting from the last time step of the video and audio, by backtracking to its starting point, accumulating the initial position D(0,0) of the cost matrix, and calculating the optimal matching path P by the following formula:

[0009]

[0010] Among them, P represents an optimal path for aligning the video and audio time steps, indicates that the last time step of video and audio is aligned, represents the total number of time steps of video features, represents the total number of time steps of audio features, is a point on the path, Represents the video time step and the audio time step The alignment relationship, Indicates choosing the path with the smallest cumulative cost. represents the path number, K represents the maximum path number, Represents the previous video time step and the previous audio time step The cumulative cost to the current position, taking into account the total time steps of video and audio and As an adjustment factor; Through the optimal matching path P, the aligned video features are obtained.

[0011] Further, the step 12 comprises: The data features extracted from all modes are integrated by linear weighted summation to generate a weighted global data feature combination. , the formula is:

[0012] in, represents the global data features after post-weighted combination, represents the feature vector after decoding of the i-th modality, i represents the modality type, and m represents the number of modalities of the cross-modal data features. Representation characteristics The weight vector of , and satisfy the constraints ; By splicing , retaining global and fine-grained features and forming the final cross-modal data feature vector representation:

[0013] Where X represents the vector of the final cross-modal data features, [·] is the concatenation operation; The step 13 comprises: The data feature vector X is decoupled by a method based on the attention mechanism to obtain the decoupled data features.

[0014] Furthermore, the step 2 comprises: Step 21: Use the cross entropy loss function to optimize the node graph representation within each subgraph network; Step 22: KL divergence is used as the loss constraint condition between different subgraph networks to quantify the similarity of node distribution between modalities, establish the similarity matrix of nodes between modalities, and calculate the total loss function between modalities; Step 23: Fuse each subgraph network according to the similarity matrix, and combine the total loss function and use graph convolution operations to integrate each subgraph network into a global cross-modal feature map representation.

[0015] Furthermore, the step 21 comprises: Initialize the subgraph network constructed for each modality i , using loss function and graph convolution operation to optimize ; Point set Represents the nodes and edge sets in mode i Represents the connection relationship between the internal nodes of modality i, which is initially defined as node similarity:

[0016] in, is the cosine similarity, represents the eigenvector of the i-th mode of the j-th node after cross-modal decoupling, represents the eigenvector of the i-th mode of the k-th node after cross-modal decoupling, represents the cosine similarity between node j and node k in the i-th mode; Then the node graph representation inside each subgraph network is optimized using the cross entropy loss function through the following formula:

[0017]

[0018] Use gradient descent to optimize the node graph representation in each subgraph network:

[0019]

[0020] in, represents the normalized probability distribution of the jth node in the i-th modal subgraph network, represents the total number of nodes in the i-th modal subgraph network, Representation Node The characteristic vector of , k represents the node number in the i-th modal subgraph network, represents the exponential map, It means calculating the sum of the eigenvalues ​​of all nodes in the i-th modal subgraph network after exponential mapping. represents the total distribution loss within the subgraph network, m represents the total number of modality types, and η is the learning rate.

[0021] Furthermore, in step 22, the KL divergence of nodes between different modes is obtained by the following formula:

[0022] in, , Indicates the mode number, Representing modality Midpoint j and mode The KL divergence between nodes k, Representing modality The normalized probability distribution of node j in is, Representing modality Normalized probability distribution of node k; Will Mapped to the corresponding matrix elements , build a similarity matrix :

[0023] in, Representing modality Midpoint j and mode The KL divergence between nodes k, Representing modality Midpoint n and mode KL divergence between nodes n; The similarity loss between modalities is calculated by the following formula :

[0024] Combining similarity loss between modalities And the total distribution loss within the subgraph network , and get the total loss function :

[0025] in, Represents the similarity matrix of nodes between modalities, storing the KL divergence between node pairs, represents the KL divergence of the node pair (j, k) between modalities, represents the KL divergence loss between modalities, Represents the balance coefficient.

[0026] Further, in step 23: In the process of inter-modal correlation fusion, based on the similarity matrix Determine the similarity of nodes between modalities and establish additional cross-modal connections based on the original subgraph network:

[0027] in, Representing modality Node j and mode in The adjacency relationship between nodes k in Set thresholds; For each subgraph Update, perform standard graph convolution operation through the following formula:

[0028] in, is the normalized adjacency moment, which is composed of the original connection within the modality and the cross-modal connection constrained by KL divergence; is the feature matrix of the Lth layer, are the trainable parameters of graph convolution, is the activation function; During the optimization process, the loss function is combined with the gradient descent algorithm to integrate each subgraph network into a global cross-modal feature graph representation:

[0029] in, represents the final cross-modal feature map, Represents the subgraph network corresponding to modality i after optimization.

[0030] Furthermore, the step 3 comprises: Step 31: Combine the mainstream value vertical field text knowledge base, use the LDA model to perform topic modeling on the text in the text knowledge base, and extract potential topic words; Step 32: Based on the extracted keywords, the RoBERTa-CRF model is used to extract event triplets, an event graph is constructed, and nodes related to hot spots are retrieved from the event graph through a graph retrieval engine; Step 33: Construct a subgraph network from related nodes, and convert the subgraph network into a vector representation to form a hotspot pattern; Step 34: By calculating the Jaccard similarity, the content features similar to the hotspot mode are matched in the cross-modal feature graph representation obtained in step 2 to identify potential hotspot content in the cross-modal feature graph.

[0031] Furthermore, in step 31: The following formula is used to implement topic modeling of texts in the text knowledge base using the LDA model:

[0032] in, For terms On topic Distribution under, documentation The topic distribution Subject to the Dirichlet prior, the term In Theme The distribution under Construct a topic-word matrix.

[0033] Further, the step 32 comprises: Through the RoBERTa layer to the input text Encode and get the vector representation of each word ; Learning Sequence Probabilities of Word Labels Based on Vector Representations ,in is a label sequence, according to the maximum conditional probability To predict the optimal label sequence for each word:

[0034] in, is the normalization factor, is the weight of the transferred feature, It is The word label, is the label of the i-th word, represents the index of the tag, is an exponential function, is the length of the word sequence, is the number of label types; The event triples are converted into event graphs, and the association relationship of events is represented by a graph structure; the subject and object in the event triples are used as nodes in the graph, and the predicate is used as the edge connecting the nodes:

[0035] in, Represents the event graph, is a collection of nodes, is the edge set; Each entity pair in the event triple By predicate connected.

[0036] Further, the step 33 comprises: For each node , Is a node set, generating a fixed-length node sequence , as the context of the event graph, where q is the walk length; Node sequence Embedding learning is performed to maximize the co-occurrence probability between a node and its context by the following formula:

[0037] in, Is a node The context node set of Is a node Generate context node probability; Each node The embedding vector That is its representation in d-dimensional vector space; for the entire subgraph , the overall vector representation of the subgraph is generated by aggregating the embedding vectors of all its nodes. The subgraph vector It is expressed as:

[0038] in, is an aggregate function, Represents the subgraph vector, that is, the hotspot pattern finally formed.

[0039] Further, the step 34 includes: By calculating the Jaccard similarity between the sub-graph vector and the cross-modal graph representation obtained in step 2, content features similar to the hot spot pattern are found in the obtained cross-modal graph representation, so as to identify potential hot content; by comparing the Jaccard similarity, potential hot content in the cross-modal feature graph is identified.

[0040] Furthermore, the combining of the cross-modal feature graph representation to construct a cross-modal content graph includes: Combined with the global feature map representation obtained in step 2 , mapping all modal graph representations to a high-dimensional semantic space; the dimension of the semantic space is d, and the graph representation of each modality is Define a mapping function :

[0041] Mapping Function The graph representation of each modality is mapped to a high-dimensional semantic space, and the generated semantic vector is recorded as ,in:

[0042] Get the set of modal feature vectors after semantic alignment , that is, the cross-modal content graph.

[0043] Furthermore, the learning of abnormal patterns using the labeled abnormal data and training of an abnormality detection model according to the abnormal patterns include: The features of different modalities are semantically aligned using graph embedding technology. The nodes of the cross-modal graph are , each node corresponds to the characteristics of a mode; By constructing a semantic similarity matrix :

[0044] in, Represents the similarity between modality i and modality j, using cosine similarity Calculate the similarity of each pair of modes, and They are the feature vectors of modality i and modality j after being mapped to the high-dimensional semantic space; The annotated anomaly dataset is And use it to train anomaly detection model: According to the semantic similarity matrix , using graph neural networks to detect abnormal nodes in cross-modal graphs:

[0045] in, Nodes in the graph The characteristic representation of node The abnormal probability of node Is it an abnormal node? represents the sigmoid activation function, represents the weight of the feature in the model, represents the parameters learned during the training process, Representation Node The features and their neighbor nodes The weighted sum of the features of When an anomaly detection model is trained, the anomaly detection loss function can be defined by the following formula: :

[0046] Get the loss function After that, by minimizing To optimize the anomaly detection model.

[0047] Furthermore, identifying abnormal content in the cross-modal content graph by using the trained anomaly detection model includes: Calculate the abnormal probability of all nodes in the modal content graph, filter out abnormal nodes according to preset thresholds, and process or alarm the corresponding modal content; identify abnormal nodes through the graph neural network model.

[0048] Furthermore, the step 5 comprises: Step 51: Construct a SKIR model to simulate the process of information dissemination among users; Step 52: Use a dual-factor coupled long-tail information cascade method to predict the propagation links of hot content and abnormal content.

[0049] Furthermore, the step 51 comprises: The SKIR model is constructed, and the informed state is introduced to expand the propagation mechanism of the model: SKIR model definition: unknown information state, informed state, propagation state, unaware state:

[0050]

[0051]

[0052]

[0053] in, represents the number of individuals whose information is unknown at time t, represents the number of individuals who are in the informed state at time t, It represents the number of users in the propagation state at time t. represents the number of individuals in the propagation state at time t, represents how the number of users with unknown information status at time t changes over time, represents how the number of informed users at time t changes over time, represents the change in the number of users in the propagation state at time t, Indicates the dynamic change of the number of users in the unaware state at time t, is the decay rate corresponding to different states, is the transmission intensity, is the propagation diffusion coefficient associated with the informed state, controlling the rate at which information propagates from the informed state to the propagated state, is the rate of transition from the informed state to the unaware state, indicating that informed users stop spreading information because they lose interest in it. is the hesitant informed rate, which indicates the proportion of informed users who stop spreading information due to doubts about the information; The stability condition is used to determine whether the SKIR model has reached a state of equilibrium and whether information propagation will stop. The stability condition is defined by the following formula:

[0054] in, Indicates the stability of the equilibrium point of the SKIR model. When Δ<0, the model will reach equilibrium and stop information propagation. It is the resistance coefficient of information dissemination.

[0055] Further, the step 52 includes: The sample selection of the model is optimized by long-tail distribution sampling and class-balanced sampling. The calculation formula of the sampling probability is:

[0056]

[0057] And by weighted combination of long-tail distribution sampling and class-balanced sampling, we get the mixed sampling probability:

[0058] in, is the long-tail distribution sampling probability of node j, is the eigenvalue of node j, is the total number of the i-th node, R is the sampling parameter, C represents the number of categories, is the probability of class-balanced sampling, where the sampling probability of each class is equal and is and is the mixed sampling probability, e is the balance factor, which is used to adjust the weights of different sampling strategies; The calculation formula of global propagation intensity is:

[0059]

[0060] in, represents the propagation intensity at time t, is the global propagation parameter, is the time at which the current event was propagated, is the time of the last propagation event, δ is the global decay factor, It is the sampling probability calculated by the sampling strategy, which reflects the influence of the propagation intensity of different nodes on the sampling probability; is the propagation intensity prediction within a local time window, and It is a function calculated from the time difference or other input features, which is used to indicate the influence of a node in the propagation process. and are the weighting coefficients of the global and local propagation models, Affects the calculation of local propagation intensity; Combining the stability conditions of the SKIR model , and the final fusion of global and local propagation strength is obtained:

[0061] in, is the final content dissemination link, that is, the final dissemination strength, is the propagation strength from different modules, is the weight coefficient of each module, It is a stability adjustment factor used to control the impact of stability conditions on the strength of the propagation link.

[0062] Further, the step 6 comprises: Step 61: Introduce dynamic adjustment factors based on the SKIR model framework and node weight , construct the control equation, and control the state transition in the information propagation process through the adjustment factors and node weights. The state transition includes the probability changes of susceptible state, informed state and propagation state; Step 62: Monitor and update parameters of the propagation link in step 5, and use the Kalman filter algorithm to update model parameters, where the updated model parameters include state prediction and observation correction; Step 63: perform key node weight adjustment, and calculate the node adjustment weight based on the node betweenness centrality and real-time influence index; Step 64: Optimize the propagation direction and rate to maximize the coverage of the target content and establish an optimal control problem; Step 65: Generate a dynamic distribution strategy, input the optimized parameters, divide users into multiple groups through the spectral clustering algorithm, define group sensitivity, and calculate the optimal delivery time window.

[0063] Furthermore, in step 61, based on the SKIR model framework, a dynamic adjustment factor is introduced. and node weight , construct the control equations, including:

[0064] in, For Node The neighbor set of It is a dynamic diffusion adjustment factor, ranging from [0,1], which is used to control the intensity of information penetration across nodes. For Node The propagation weight ranges from [0,2], and is adjusted by It can suppress or amplify the propagation influence of nodes. represents the rate at which the unknown information state changes to the propagation state, is the propagation diffusion coefficient associated with the informed state, controlling the rate at which information propagates from the informed state to the propagated state, It is the transition rate from the informed state to the unaware state, indicating that informed users stop spreading information because they lose interest in it.

[0065] Further, the step 62 includes: State prediction based on historical parameters To predict the current parameters :

[0066] in, is the state transfer matrix, is the control input matrix, For external intervention signals, such as manual control instructions, is the process noise, is the final propagation link outputted in step 5; Observation correction through real-time observation data To modify the parameters:

[0067] in, is the observation matrix, is the Kalman gain matrix, Represents the updated state prediction parameters.

[0068] Further, the step 63 includes: Based on node betweenness centrality and real-time impact indicators , calculate the node adjustment weight:

[0069] in, is the learning rate, which is used to control the weight update amplitude. Finally, the adjusted node weight set is output .

[0070] Further, the step 64 includes: Optimize the direction and rate of dissemination to maximize the coverage of target content As the goal, establish the optimal control problem:

[0071] in, For the delivery cycle, is the weight regularization coefficient. The Pontryagin maximum principle is used to solve the Hamiltonian function , derive the optimal control law:

[0072] in, is a covariate variable, represents the optimal output weight.

[0073] Further, the step 65 includes: By entering the optimized parameters , using spectral clustering algorithm to divide users into Groups , defining group sensitivity , calculate the optimal delivery time window:

[0074] in, represents the best delivery time window, that is, the best content delivery time for a group. is the attenuation coefficient, It represents the average number of users in the propagation state at time s, the triple , indicating that In time By weight Deliver content.

[0075] The present application also discloses a network content dissemination system based on integrated cognitive understanding and intelligent governance, which implements the above-mentioned network content dissemination method based on integrated cognitive understanding and intelligent governance, and includes: A decoupling module is used to decouple data features of different modalities and output the decoupled data features; different modalities include text, video, audio and image; The generation module is used to construct a subgraph network for each modality based on the decoupled data features of audio, video, image, and text, and to perform convolution operations on cross-modal features through multiple subgraph networks to generate cross-modal feature graph representations; The hot content identification module is used to perform topic modeling on the text in the text knowledge base, extract potential topic words, and extract event triplets from the potential topic words to construct an event graph; retrieve nodes related to hot spots from the event graph to form a hot spot pattern, match content features similar to the hot spot pattern in the cross-modal feature graph representation, and identify potential hot content; The abnormal content identification module is used to combine the cross-modal feature graph representation, build a cross-modal content map, use the labeled abnormal data to learn the abnormal pattern, train the abnormal detection model according to the abnormal pattern, and identify the abnormal content in the cross-modal content map through the trained anomaly detection model; abnormal content includes false information, sensitive images and dense videos; The prediction module is used to simulate the process of information dissemination among users and predict the content dissemination links of hot content and abnormal content; The determination module is used to adjust the weights of key nodes in the content dissemination chain and determine the best time, location and target audience for content delivery.

[0076] Due to the adoption of the above technical solution, the present application has the following advantages: 1. This application realizes the unified expression of multimodal information, breaks through the difficulty of mapping traditional image data to high-dimensional representation in a unified space, and effectively improves the ability to integrate heterogeneous information under complex networks.

[0077] 2. Accurate perception technology for abnormal content improves the accuracy of content dissemination by accurately identifying potential hot content and abnormal information, and successfully increases the accuracy of dissemination content detection to 91.2% and 88.0%.

[0078] 3. The content distribution technology based on propagation prediction, combined with the SKIR information propagation method, realizes accurate content propagation path prediction and propagation strategy adjustment, and the prediction accuracy is improved by 13.5% compared with the existing technology.

[0079] 4. Adaptive dynamic communication control and optimization technology, which regulates the allocation of communication budget by optimizing node weights, implements precise content delivery strategies, significantly improves the efficiency and coverage of information dissemination, and ensures that content can be disseminated at the most appropriate time, location and target audience. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0081] Figure 1 A flowchart of a network content dissemination method based on the integration of cognitive understanding and intelligent governance according to an embodiment of the present application. DETAILED DESCRIPTION

[0082] The present application is further described in conjunction with the accompanying drawings and embodiments, and the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field should fall within the scope of protection of the embodiments of the present application.

[0083] This application realizes the coordinated processing and deep integration of different modal data such as audio, video, image and text, effectively solves the inconsistency of information structure between modalities, and uses the accurate perception technology of potential hot spots and abnormal content to quickly identify hot topics and abnormal information in the network, providing strong support for public opinion monitoring and guidance. Through the intelligent prediction and distribution guidance technology of network content dissemination, it accurately predicts the information dissemination path and effect, and optimizes the content distribution strategy to achieve accurate personalized recommendation and maximize the dissemination effect, thereby improving the intelligence level of all-media content dissemination.

[0084] See also Figure 1 , the embodiment of the present application provides a network content dissemination method based on integrated cognitive understanding and intelligent governance, which includes S1 to S6: S1: Extract data features of text, video, audio and image, that is, process text data to extract its semantic information and contextual features; for video data, extract key frame image features and temporal dynamic information; extract key information such as spectral features, pitch and duration from audio data; extract visual features such as color, texture, edge from image data, and then pass it into the method based on the attention mechanism to decouple the data features of different modalities (i.e. text, video, audio and image), and output the decoupled data features; S2: Decouple the audio, video, image, and text data features output by S1 to construct a subgraph network for each modality, and perform convolution operations on cross-modal features through multiple independent subgraph networks to generate a cross-modal feature graph representation.

[0085] S3: Combined with the mainstream value vertical field text knowledge base, this knowledge base collects and organizes hot information, keywords, event triples and other data in different fields. The text in the text knowledge base is subject modeled through the LDA model to extract potential keywords. The RoBERTa-CRF model is used to extract event triples from potential keywords to construct an event graph. Nodes related to hot spots are retrieved from the event graph through the graph retrieval engine. The subgraph composed of related nodes is converted into a vector representation through graph embedding technology to form a hot spot pattern (that is, a structured feature representation extracted from the event graph and associated with the current hot event). By calculating the Jaccard similarity, content features similar to the hot spot pattern are matched in the cross-modal feature graph representation obtained in S2 to identify potential hot content; S4: Combine the cross-modal feature graph representation obtained in S2 to construct a cross-modal content map, use the labeled abnormal data to learn abnormal patterns, optimize the anomaly detection model based on the abnormal patterns, and identify abnormal content in the cross-modal content map through the trained anomaly detection model; abnormal content includes false information, sensitive images and dense videos.

[0086] S5: Based on the classic propagation model SIR and the mean field theory, the informed state (K) is introduced and the SKIR model is proposed. The SKIR model simulates the information propagation process between users and adopts the long-tail information cascade modeling and prediction method with dual-factor coupling to achieve high-precision prediction of the propagation links of hot content and abnormal content identified by S3 and S4.

[0087] S6: Through the dynamic propagation control equation, the content propagation link obtained by S5 is monitored in real time and adjusted dynamically. The direction and speed of propagation are controlled by adjusting the weights of key nodes in the content propagation link, and the content propagation strategy is optimized, that is, the best time, location and target audience for content delivery are determined. The distribution and guidance of hot content and abnormal content identified by S3 and S4 are achieved.

[0088] Optionally, S1 includes S11 to S13: S11: Extract data features of text, video, audio, and images and input them into the encoder to extract features of the corresponding modality and align video features using dynamic time warping (DTW); S12: Model the extracted cross-modal data features to obtain a unified representation; S13: Extract cross-modal decoupled data features from the modeled cross-modal features through an attention-based method.

[0089] Optionally, S11 includes: Text data feature extraction usually uses a pre-trained language model, such as BERT or GPT, to convert the input text data into a high-dimensional feature vector so that it can effectively capture the semantic information of the text; Image data feature extraction uses contrastive learning and coarse-grained and fine-grained fusion Transformer architecture image encoder to extract its feature vector; Audio data feature extraction uses Mel spectrum or MFCC (Mel frequency cepstral coefficient) to obtain audio feature vectors through their corresponding encoders; Video feature extraction uses a Transformer architecture image encoder with contrastive learning and coarse-grained fusion to encode each frame of the video in the multimodal data, and then clusters them to form video feature vectors of key frames. Since videos usually contain audio features, they need to be feature extracted separately before dynamic time warping (DTW) alignment. DTW calculates video features using the following formula The distance between different time steps of audio data feature A is used to obtain the cumulative cost matrix:

[0090] Based on this distance, the cumulative cost matrix cost is obtained by the following formula:

[0091] in, represents the distance between the i-th time step of the video and the j-th time step of the audio, represents the features of the video at the i-th time step, express The i-th element in represents the features of the audio at the jth time step, represents the jth element in A, represents the Euclidean distance, represents the cumulative cost matrix, From the starting point arrive The minimum cumulative cost.

[0092] From the end point of the cumulative cost matrix , that is, starting from the last time step of the video and audio, by backtracking to its starting point, the initial position D(0,0) of the cumulative cost matrix, that is, the first time step of the video and audio, the optimal matching path P is calculated by the following formula:

[0093]

[0094] Among them, P represents an optimal path for aligning the video and audio time steps, indicates that the last time step of video and audio is aligned, represents the total number of time steps of video features, represents the total number of time steps of audio features, is a point on the path, Represents the video time step and the audio time step The alignment relationship, Indicates choosing the path with the smallest cumulative cost. represents the path number, K represents the maximum path number, Represents the previous video time step and the previous audio time step The cumulative cost to the current position, taking into account the total time steps of video and audio and as an adjustment factor.

[0095] By obtaining the optimal matching path P, the aligned video features can be obtained.

[0096] Optionally, S12 includes: After extracting the data feature information of each single mode, the data features extracted from all modes are integrated through linear weighted summation to generate a weighted global data feature combination. , the formula is:

[0097] in, represents the global data features after post-weighted combination, represents the feature vector after decoding of the i-th modality, i represents the modality type, and m represents the number of modalities of the cross-modal data features. Representation characteristics The weight vector of , and satisfy the constraints ; By splicing , retaining global and fine-grained features and forming the final cross-modal data feature vector representation:

[0098] Where X represents the vector of the final cross-modal data features, [·] is the concatenation operation; S13 includes: The data feature vector X after input feature modeling is decoupled by the method based on the attention mechanism. The implementation process of the method based on the attention mechanism is as follows: First, X is linearly transformed to obtain the total query vector Q, key vector K and value vector V, and then the attention mechanism is used to calculate the attention weight matrix:

[0099] in, represents the attention weight calculation result, represents the normalization function, represents the dimension of the key vector K, T represents the transpose calculation, represents the dot product computation of the transposed query and key; Through the multi-head attention mechanism, the weighted output of each attention head is calculated independently:

[0100] in, represents the attention output of the i-th head; The outputs of multiple heads are concatenated and linearly transformed to obtain the final multi-head output:

[0101] in, Represents the final multi-head output, The weight matrix representing the linear transformation; The final output is obtained through multi-head output and linear transformation:

[0102] in, Represents the output of the attention mechanism method, that is, the decoupled data features, which only contain the independent representation of each modality, but not the global data features , to avoid information redundancy, it can be expressed as:

[0103] in, It represents the feature vector of the i-th modality after decoupling by the attention mechanism, i represents the modality type, and m represents the number of modalities of cross-modal data features.

[0104] Optionally, S2 includes S21 to S23: S21: Use the cross entropy loss function to optimize the node graph representation within each subgraph network; S22: KL divergence is used as the loss constraint between different subgraph networks to quantify the similarity of node distribution between modalities and establish a similarity matrix of nodes between modalities , and calculate the total loss function between modalities; S23: Based on the similarity matrix By fusing each subgraph network and combining the total loss function with graph convolution operations, each subgraph network is integrated into a global cross-modal feature graph representation to fully capture the semantic associations and interactions between multimodal contents.

[0105] Optionally, S21 includes: First, the data features after cross-modal decoupling of input S1 , m represents the total number of modalities, and initializes the subgraph network constructed for each modality i , and then optimize by using the loss function and graph convolution operation .

[0106] Among them, the point set represents the nodes in mode i, Indicates the first Nodes, edge sets Represents the connection relationship between the internal nodes of modality i, which is initially defined as node similarity:

[0107] is the cosine similarity, represents the eigenvector of the i-th mode of the j-th node after cross-modal decoupling, represents the eigenvector of the i-th mode of the k-th node after cross-modal decoupling, Represents the cosine similarity between nodes j and k in the i-th mode.

[0108] Then the node graph representation inside each subgraph network is optimized using the cross entropy loss function through the following formula:

[0109]

[0110] Use gradient descent to optimize the node graph representation in each subgraph network:

[0111]

[0112] in, represents the normalized probability distribution of the jth node in the ith modal subgraph network, ni represents the total number of nodes in the ith modal subgraph network, Representation Node The characteristic vector of , k represents the node number in the i-th modal subgraph network, represents the exponential map, It means calculating the sum of the eigenvalues ​​of all nodes in the i-th modal subgraph network after exponential mapping. Represents the total distribution loss within the subgraph network, m represents the total number of modal types, and η is the learning rate. By continuously iteratively adjusting the node feature vector, the goal of minimizing the cross entropy loss is finally achieved, thereby optimizing the node graph representation in each subgraph network.

[0113] Optionally, in S22, in step 22, the KL divergence of nodes between different modes is obtained by the following formula:

[0114] in, , Indicates the mode number, Representing modality Midpoint j and mode The KL divergence between nodes k, Representing modality The normalized probability distribution of node j in is, Representing modality Normalized probability distribution of node k; To unify the representation, Mapped to the corresponding matrix elements (Coming soon Composition similarity matrix ), build a similarity matrix :

[0115] in, Representing modality Midpoint j and mode The KL divergence between nodes k, Representing modality Midpoint n and mode KL divergence between nodes n; The similarity loss between modalities is calculated by the following formula :

[0116] Combining similarity loss between modalities And the total distribution loss within the subgraph network , and get the total loss function :

[0117] in, Represents the similarity matrix of nodes between modalities, storing the KL divergence between node pairs, represents the KL divergence of the node pair (j, k) between modalities, represents the KL divergence loss between modalities, Represents the balance coefficient.

[0118] Optionally, in S23, In the process of inter-modal correlation fusion, firstly based on the similarity matrix Determine the similarity of nodes between modalities and establish additional cross-modal connections based on the original subgraph network. Specifically:

[0119] in, Representing modality Node j and mode in The adjacency relationship between the nodes k in , that is, whether there is a cross-modal edge, Set the threshold if KL divergence Less than , then the two nodes are considered similar enough to establish a connection in the graph, otherwise they are not connected.

[0120] Then, the graph convolution operation is introduced so that the intra-modal and inter-modal features can be optimized synchronously during the information propagation process. Update, using standard graph convolution operations:

[0121] in, is the normalized adjacency moment, which is composed of the original connection within the modality and the cross-modal connection constrained by KL divergence. The cross-modal edge is only valid when the KL divergence is less than the threshold. When adding, is the feature matrix of the Lth layer, are the trainable parameters of graph convolution, is an activation function. In this way, each subgraph can not only maintain its own structural information, but also realize mutual propagation of features through cross-modal edges, thereby enhancing the feature consistency between modalities.

[0122] At the same time, during the optimization process, it is also necessary to combine the calculated loss function and use the gradient descent algorithm again to integrate each subgraph network into a global cross-modal feature graph representation:

[0123] in represents the final cross-modal feature map, Represents the subgraph network corresponding to modality i after optimization.

[0124] Optionally, S3 includes S31 to S34: S31: Combined with the mainstream value vertical field text knowledge base, the LDA model is used to perform topic modeling on the text in the text knowledge base and extract potential topic words; S32: Based on the extracted keywords, the RoBERTa-CRF model is used to extract event triplets, build an event graph, and retrieve nodes related to hot spots from the event graph through the graph retrieval engine; S33: Construct a subgraph network from related nodes, and convert the subgraph network into a vector representation to form a hotspot pattern; S34: By calculating the Jaccard similarity, the content features similar to the hotspot mode are matched in the cross-modal feature graph representation obtained in S2, and the potential hotspot content in the cross-modal feature graph is identified.

[0125] Optionally, S31 includes: Combined with the hot information in the mainstream value vertical domain knowledge base, the LDA model is used to perform topic modeling on the data in the text knowledge base and extract the key words; the LDA topic modeling process is formalized as follows:

[0126] in, For terms In Theme Distribution under, documentation The topic distribution Subject to the Dirichlet prior, the term In Theme The distribution under Construct a topic-word matrix; For each document, Select a topic distribution , for each word in the document , first select a topic from the topic distribution of the document , and then select words from the word distribution corresponding to the topic ;in, Is the theme The word distribution of is a dimensional vector, is the number of topics; is a dimensional vector, is the size of the vocabulary; In mathematical expression, given Topics and documents ,document The generation probability of the word for:

[0127] in, Indicated in the topic Next word The probability of Representation Document Medium Theme The probability of is a topic-word distribution, is a document-topic distribution, Representation Document The number of words in .

[0128] Optionally, S32 includes: Through the RoBERTa layer to the input text Encode and get the vector representation of each word ; Learning the sequence probability of word labels based on vector representation through CRF layer ,in is a label sequence, and the CRF layer maximizes the conditional probability To predict the optimal label sequence for each word:

[0129] in, is the normalization factor, is the weight of the transferred feature, It is The word label, is the label of the i-th word, Indicates the index of the tag; is an exponential function, is the length of the word sequence, is the number of label types.

[0130] The event triples are converted into event graphs, and the association relationship of events is represented by a graph structure; the subject and object in the event triples are used as nodes in the graph, and the predicate is used as the edge connecting the nodes:

[0131] in, Represents the event graph, is a collection of nodes, is the edge set; Each entity pair in the event triple By predicate Connected, expressions include:

[0132] in, is the set of all possible predicate relations.

[0133] Optionally, S33 includes: For each node , Is a node set, generating a fixed-length node sequence , as the context of the event graph, where q is the walk length; The node sequence is analyzed by Skip-Gram model Perform embedding learning to maximize the co-occurrence probability between a node and its context:

[0134] in, Is a node The context node set of Is a node Generate context node probability.

[0135] The transition probability of the random walk method is:

[0136] in, Is a node arrive The transfer weight, is a normalization constant; the Node2Vec model introduces two hyperparameters and Control the depth-first and breadth-first tendencies of the walk so that the transfer weight Defined as:

[0137] in, Representation Node and Shortest path distance in a graph; hyperparameters Control the probability of returning a node, a hyperparameter Control the probability of exploring new nodes to achieve a balanced modeling of the local and global structures of the graph.

[0138] node Generate context node Probability , expressed using the Softmax function:

[0139] in, Representation Node and The dot product of the embedding vector is used to measure the similarity; if negative sampling is used for optimization, the objective function is approximated as:

[0140] in, is a Sigmoid function, and ; is the negative sampling distribution, is the number of negative samples; represents the mathematical expectation operation, is a negative sampling distribution The elements sampled from ; After training, each node The embedding vector That is its representation in d-dimensional vector space; for the entire subgraph , can generate the overall vector representation of the subgraph by aggregating the embedding vectors of all its nodes. The aggregation methods include averaging, weighted summation, and graph pooling-based methods; finally, the subgraph vector It is expressed as:

[0141] in, is an aggregate function, Represents the subgraph vector, that is, the hotspot pattern finally formed.

[0142] Optionally, S34 includes: By calculating the Jaccard similarity between the subgraph vector and the cross-modal graph representation obtained by S2, we search for content features similar to the hotspot mode in the obtained cross-modal graph representation, thereby identifying potential hotspot content. The calculation formula of Jaccard similarity is:

[0143] in, represents the cross-modal graph representation obtained in S2, B represents the hotspot mode feature set to be compared, , from the subgraph vector Extract features from yes and The number of elements in the intersection, and yes and The number of elements in the union; the value range of Jaccard similarity is between 0 and 1, and the larger the value, the higher the similarity between the two sets; Finally, by comparing If the similarity exceeds the threshold , the content will be identified as potential hot content.

[0144] Optionally, S4 includes S41 to S43: S41: constructing a cross-modal content map according to the cross-modal feature map obtained in S2; S42: Use graph embedding technology for semantic alignment and use labeled anomaly datasets to learn potential anomaly patterns and optimize anomaly detection models; S43: Use the trained anomaly detection model to perform abnormal content detection.

[0145] Optionally, S41 includes: In S2, the global feature map representation has been obtained Then, the graph representations of all modalities are mapped to a high-dimensional semantic space. The dimension of the semantic space is d, and the graph representation of each modality is Define a mapping function

[0146]

[0147] This function maps the graph representation of each modality into a high-dimensional semantic space, and the generated semantic vector is recorded as ,in:

[0148] All modal graph representations will be aligned in the high-dimensional semantic space through this process, and a set of modal feature vectors after semantic alignment will be obtained. , that is, the cross-modal content graph.

[0149] Optionally, S42 includes: The features of different modalities are further semantically aligned using graph embedding technology. The nodes of the cross-modal graph are , these nodes represent different entities or events, and each node corresponds to a specific modal feature. The goal of graph embedding is to capture the semantic relationship between modalities by calculating the similarity between modal features. Specifically, by constructing the semantic similarity matrix S:

[0150] in, Represents the similarity between modality i and modality j, and uses cosine similarity to calculate the similarity of each pair of modalities. They are the feature vectors of modality i and modality j after being mapped to the high-dimensional semantic space.

[0151] Using the labeled anomaly data, we can learn the potential anomaly patterns. , which contains the feature representation of abnormal patterns. Use these labeled data to train an anomaly detection model. First, the model uses a graph neural network (GNN) to detect abnormal nodes in the cross-modal graph based on the semantic similarity matrix S:

[0152] Among them Nodes in the graph The characteristic representation of node The abnormal probability of node Is it an abnormal node? Through the calculation of the graph neural network, the output value will be between [0,1], indicating the probability of abnormality. represents the sigmoid activation function, represents the weight of the feature in the model, Represents parameters that are usually learned during training and are used to adjust the model output. Representation Node The features and their neighbor nodes Weighted sum of features.

[0153] Finally, the anomaly detection loss function is defined by :

[0154] Get the loss function After that, you can minimize To optimize the anomaly detection model.

[0155] Optionally, S43 includes: Through the trained anomaly detection model, potential abnormal nodes in the cross-modal content graph can be marked as abnormal content. These abnormal nodes may correspond to abnormal content such as false information, sensitive images or dense videos. The detection results of abnormal nodes can be achieved through the following process: 1) Calculate the abnormal probability of all nodes in the modal content graph.

[0156] 2) According to the preset threshold Filter abnormal nodes. , then the node Determined as an abnormal node.

[0157] 3) Mark abnormal nodes and further process or alarm the corresponding modal content.

[0158] Finally, the abnormal nodes identified by the graph neural network model are abnormal contents such as false information, sensitive images, and dense videos.

[0159] Optionally, S5 includes S51 and S52: S51: Construct SKIR model to simulate the process of information dissemination among users; S52: A dual-factor coupled long-tail information cascade method is used to accurately predict the propagation links of hot content and abnormal content.

[0160] Optionally, S51 includes: Construct the SKIR model. The SKIR model is based on the classic SIR (susceptible-informed-propagating) model and introduces the informed state (K) to expand the propagation mechanism of the model. The specific construction process is as follows: The model defines four states: unknown information state (S), informed state (K), propagation state (I), and unaware state (R):

[0161]

[0162]

[0163]

[0164] in represents the number of individuals whose information is unknown at time t, represents the number of individuals who are in the informed state at time t, It represents the number of users in the propagation state at time t. represents the number of individuals in the propagation state at time t, represents how the number of users with unknown information status at time t changes over time, represents how the number of informed users at time t changes over time, represents the change in the number of users in the propagation state at time t, Indicates the dynamic change of the number of users in the unaware state at time t, is the decay rate corresponding to different states, is the transmission intensity, is the propagation diffusion coefficient associated with the informed state, controlling the rate at which information propagates from the informed state to the propagated state, is the rate of transition from the informed state to the unaware state, indicating that informed users stop spreading information because they lose interest in it. is the hesitant informed rate, which indicates the proportion of informed users who stop spreading information due to doubts about the information.

[0165] Then, by setting the stability condition, we can determine whether the SKIR model reaches equilibrium and whether information propagation stops. This stability condition is defined by the following formula:

[0166] In the formula It represents the stability of the equilibrium point of the SKIR model. When Δ<0, the model will reach equilibrium and stop information propagation. is the resistance coefficient of information propagation. By adjusting these propagation parameters, the scope and speed of information propagation can be controlled, thereby accurately intervening in the direction of information diffusion in the system.

[0167] Through these precise formulas and parameter adjustments, the SKIR model can capture the fine-grained dynamic characteristics of information propagation between users, improve the interpretability of the propagation path, and provide strong support for propagation prediction. The propagation parameters in the model determine the scope and speed of information propagation. Further adjustment of these parameters can accurately predict and intervene in the diffusion direction of unknown states during the information propagation process.

[0168] Optionally, S52 includes: Combining the long-tail information cascade modeling and prediction method of dual-factor coupling to predict the propagation link, the following is the specific implementation process of the long-tail information cascade method of dual-factor coupling: First, the sample selection of the model is optimized by long-tail distribution sampling and class-balanced sampling. The calculation formula of the sampling probability is:

[0169]

[0170] And by weighted combination of long-tail distribution sampling and class-balanced sampling, we get the mixed sampling probability:

[0171] in, is the long-tail distribution sampling probability of node j, is the eigenvalue of node j, is the total number of the i-th node, R is the sampling parameter, C represents the number of categories, is the probability of class-balanced sampling, where the sampling probability of each class is equal and is and is the mixed sampling probability, and e is the balance factor, which is used to adjust the weights of different sampling strategies.

[0172] Secondly, the propagation intensity is closely related to the sampling probability. The model captures the information diffusion relationship in the propagation process through the global dependency module and the local dependency module. The calculation formula for the global propagation intensity is:

[0173]

[0174] in, represents the propagation intensity at time t, is the global propagation parameter, is the time at which the current event was propagated, is the time of the last propagation event, δ is the global decay factor, It is the sampling probability calculated by the sampling strategy, which reflects the influence of the propagation intensity of different nodes on the sampling probability; is the propagation intensity prediction within a local time window, and It is usually a function calculated from the time difference or other input features to indicate the influence of a node in the propagation process. and are the weighting coefficients of the global and local propagation models, This further affects the calculation of local propagation intensity and ensures that the sampling probability is reflected in different modules.

[0175] Finally, combined with the stability conditions of the above SKIR model , and the final fusion of global and local propagation strength is obtained:

[0176] in, is the final content dissemination link, that is, the final dissemination strength, is the propagation strength from different modules, is the weight coefficient of each module, It is a stability adjustment factor used to control the impact of stability conditions on the strength of the propagation link.

[0177] Optionally, S6 includes S61 to S65: S61: Based on the SKIR model framework, dynamic adjustment factors are introduced and node weight , the control equation is constructed to control the state transition in the information propagation process through adjustment factors and node weights, including the probability changes of susceptible state, informed state and propagation state.

[0178] S62: Monitor and update the parameters of the propagation link of S5, and use the Kalman filter algorithm to update the model parameters, including state prediction and observation correction. State prediction predicts current parameters based on historical parameters.

[0179] S63: Perform key node weight adjustment and calculate node adjustment weight based on node betweenness centrality and real-time influence index.

[0180] S64: Optimize the propagation direction and rate to maximize the target content coverage and establish the optimal control problem.

[0181] S65: Generate a dynamic distribution strategy, input the optimized parameters, divide users into multiple groups through the spectral clustering algorithm, define group sensitivity, and calculate the optimal delivery time window.

[0182] Optionally, S61 includes: This equation is based on the SKIR model framework and introduces a dynamic adjustment factor and node weight , construct the dynamic propagation control equation, the specific implementation process is as follows:

[0183] in, For Node The neighbor set of It is a dynamic diffusion adjustment factor, ranging from [0,1], which is used to control the intensity of information penetration across nodes. For Node The propagation weight ranges from [0,2], and is adjusted by It can suppress or amplify the propagation influence of nodes. represents the rate at which the unknown information state changes to the propagation state, is the propagation diffusion coefficient associated with the informed state, controlling the rate at which information propagates from the informed state to the propagated state, It is the transition rate from the informed state to the unaware state, indicating that informed users stop spreading information because they lose interest in it.

[0184] Optionally, S62 includes: The propagation link of S5 is monitored and updated in real time, and the Kalman filter algorithm is used to update the model parameters. State prediction is based on historical parameters To predict the current parameters :

[0185] in, is the state transfer matrix, is the control input matrix, For external intervention signals, such as manual control instructions, is the process noise, It is the final propagation link output by S5.

[0186] Observation correction through real-time observation data To modify the parameters:

[0187] in, is the observation matrix, is the Kalman gain matrix, Represents the updated state prediction parameters.

[0188] Optionally, S63 includes: Based on node betweenness centrality and real-time impact indicators , calculate the node adjustment weight:

[0189] in, is the learning rate, which is used to control the weight update amplitude. Finally, the adjusted node weight set is output .

[0190] Optionally, S64 includes: Optimize the direction and rate of dissemination to maximize the coverage of target content As the goal, establish the optimal control problem:

[0191] in, For the delivery cycle, is the weight regularization coefficient. The Pontryagin maximum principle is used to solve the Hamiltonian function , derive the optimal control law:

[0192] in, is a covariate variable, represents the optimal output weight.

[0193] Optionally, S65 includes: Generate a dynamic distribution strategy that determines the best time, location, and target audience for content delivery. By entering the optimized parameters , using spectral clustering algorithm to divide users into Groups , defining group sensitivity , calculate the optimal delivery time window:

[0194] in, represents the best delivery time window, that is, the best content delivery time for a group. is the attenuation coefficient, It represents the average number of users in the propagation state at time s, the triple , indicating that In time By weight Deliver content.

[0195] The embodiment of the present application further provides a network content dissemination system based on integrated cognitive understanding and intelligent governance, which implements the network content dissemination method based on integrated cognitive understanding and intelligent governance described in the above embodiment, and includes: A decoupling module is used to decouple data features of different modalities and output the decoupled data features; different modalities include text, video, audio and image; The generation module is used to construct a subgraph network for each modality based on the decoupled data features of audio, video, image, and text, and to perform convolution operations on cross-modal features through multiple subgraph networks to generate cross-modal feature graph representations; The hot content identification module is used to perform topic modeling on the text in the text knowledge base, extract potential subject words, and extract event triplets from the potential subject words to construct an event graph; retrieve nodes related to hot spots from the event graph, convert the subgraph composed of related nodes into a vector representation through graph embedding technology to form a hot spot pattern, match content features similar to the hot spot pattern in the cross-modal feature graph representation, and identify potential hot content; The abnormal content identification module is used to combine the cross-modal feature graph representation, build a cross-modal content map, use the labeled abnormal data to learn the abnormal pattern, train the abnormal detection model according to the abnormal pattern, and identify the abnormal content in the cross-modal content map through the trained anomaly detection model; abnormal content includes false information, sensitive images and dense videos; The prediction module is used to simulate the process of information dissemination among users and predict the content dissemination links of hot content and abnormal content; The determination module is used to obtain the direction and speed of dissemination by adjusting the weights of key nodes in the content dissemination link, and to determine the best time, location and target audience for content delivery.

[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present application rather than to limit it. Although the present application has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present application can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present application should be included in the scope of protection of the claims of the present application.

Claims

1. A network content dissemination method based on integrated cognitive understanding and intelligent governance, characterized in that: include: Step 1: Decouple data features of different modalities and output the decoupled data features; different modalities include text, video, audio, and image; Step 2: Based on the decoupled data features of audio, video, image, and text, a subgraph network for each modality is constructed, and convolution operations on cross-modal features are performed through multiple subgraph networks to generate cross-modal feature graph representations; Step 3: Perform topic modeling on the text in the text knowledge base, extract potential topic words, and extract event triplets from the potential topic words to construct an event graph; retrieve nodes related to hot spots from the event graph to form hot spot patterns, match content features related to hot spot patterns in the cross-modal feature graph representation, and identify potential hot content; Step 4: Combine the cross-modal feature graph representation to build a cross-modal content map, use the labeled abnormal data to learn abnormal patterns, train an anomaly detection model based on the abnormal patterns, and identify abnormal content in the cross-modal content map through the trained anomaly detection model; abnormal content includes false information, sensitive images, and dense videos; Step 5: Simulate the process of information propagation among users and predict the content propagation links of hot content and abnormal content; Step 6: Adjust the weights of key nodes in the content dissemination chain to determine the time, location, and target audience for content delivery.

2. The network content dissemination method based on integrated cognitive understanding and intelligent governance according to claim 1 is characterized in that: The step 1 comprises: Step 11: Extract data features of text, video, audio, and image and input them into the encoder, extract features of the corresponding modality, and align video features using dynamic time warping; Step 12: Model the extracted cross-modal data features; Step 13: Extract the cross-modal decoupled data features from the modeled cross-modal features through an attention mechanism-based method.

3. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 2 is characterized in that: The step 11 comprises: Extract text data features, video features, audio data features and image data features from text, video, audio and image respectively; The video features are calculated by the following formula And the distance between different time steps of audio data feature A, we get the cumulative cost matrix: Based on this distance, the cumulative cost matrix cost is obtained by the following formula: in, Represents the distance between the i-th time step of the video and the j-th time step of the audio; yes The i-th element in represents the feature of the i-th time step of the video; is the jth element in A, representing the features of the jth time step of the audio; represents the Euclidean distance; represents the cumulative cost matrix, From the starting point arrive The minimum cumulative cost of From the end point of the cumulative cost matrix , that is, starting from the last time step of the video and audio, by backtracking to its starting point, accumulating the initial position D(0,0) of the cost matrix, and calculating the optimal matching path P by the following formula: Among them, P represents an optimal path for aligning the video and audio time steps, indicates that the last time step of video and audio is aligned, represents the total number of time steps of video features, represents the total number of time steps of audio features, is a point on the path, Represents the video time step and the audio time step The alignment relationship, Indicates choosing the path with the smallest cumulative cost. represents the path number, K represents the maximum path number, Represents the previous video time step and the previous audio time step The cumulative cost to the current position, taking into account the total time steps of video and audio and As an adjustment factor; Through the optimal matching path P, the aligned video features are obtained.

4. The network content dissemination method based on integrated cognitive understanding and intelligent governance according to claim 2 is characterized in that: The step 12 comprises: The data features extracted from all modes are integrated by linear weighted summation to generate a weighted global data feature combination. , the formula is: in, represents the global data features after post-weighted combination, represents the feature vector after decoding of the i-th modality, i represents the modality type, and m represents the number of modalities of the cross-modal data features. Representation characteristics The weight vector of , and satisfy the constraints ; By splicing , retaining global and fine-grained features and forming the final cross-modal data feature vector representation: Where X represents the vector of the final cross-modal data features, [·] is the concatenation operation; The step 13 comprises: The data feature vector X is decoupled by a method based on the attention mechanism to obtain the decoupled data features.

5. The network content dissemination method based on integrated cognitive understanding and intelligent governance according to claim 1 is characterized in that: The step 2 comprises: Step 21: Use the cross entropy loss function to optimize the node graph representation within each subgraph network; Step 22: KL divergence is used as the loss constraint condition between different subgraph networks to quantify the similarity of node distribution between modalities, establish the similarity matrix of nodes between modalities, and calculate the total loss function between modalities; Step 23: Fuse each subgraph network according to the similarity matrix, and combine the total loss function and use graph convolution operations to integrate each subgraph network into a global cross-modal feature map representation.

6. The network content dissemination method based on integrated cognitive understanding and intelligent governance according to claim 5 is characterized in that: The step 21 comprises: Initialize the subgraph network constructed for each modality i , using loss function and graph convolution operation to optimize ; Point set represents the nodes in mode i, Indicates the first Nodes, edge sets Represents the connection relationship between the internal nodes of modality i, which is initially defined as node similarity: in, is the cosine similarity, represents the eigenvector of the i-th mode of the j-th node after cross-modal decoupling, represents the eigenvector of the i-th mode of the k-th node after cross-modal decoupling, represents the cosine similarity between node j and node k in the i-th mode; Then the node graph representation inside each subgraph network is optimized using the cross entropy loss function through the following formula: Use gradient descent to optimize the node graph representation in each subgraph network: in, represents the normalized probability distribution of the jth node in the i-th modal subgraph network, represents the total number of nodes in the i-th modal subgraph network, Representation Node The characteristic vector of , k represents the node number in the i-th modal subgraph network, represents the exponential map, It means calculating the sum of the eigenvalues ​​of all nodes in the i-th modal subgraph network after exponential mapping. represents the total distribution loss within the subgraph network, m represents the total number of modality types, and η is the learning rate.

7. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 5 is characterized in that: In step 22, the KL divergence of nodes between different modes is obtained by the following formula: in, , Indicates the mode number, Representing modality Midpoint j and mode The KL divergence between nodes k, Representing modality The normalized probability distribution of node j in is, Representing modality Normalized probability distribution of node k; Will Mapped to the corresponding matrix elements , build a similarity matrix : in, Representing modality Midpoint j and mode The KL divergence between nodes k, Representing modality Midpoint n and mode KL divergence between nodes n; The similarity loss between modalities is calculated by the following formula : Combining similarity loss between modalities And the total distribution loss within the subgraph network , and get the total loss function : in, Represents the similarity matrix of nodes between modalities, storing the KL divergence between node pairs, represents the KL divergence of the node pair (j, k) between modalities, represents the KL divergence loss between modalities, Represents the balance coefficient.

8. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 5 is characterized in that: In step 23: In the process of inter-modal correlation fusion, based on the similarity matrix Determine the similarity of nodes between modalities and establish additional cross-modal connections based on the original subgraph network: in, Representing modality Node j and mode in The adjacency relationship between nodes k in Set thresholds; For each subgraph Update, perform standard graph convolution operation through the following formula: in, is the normalized adjacency moment, which is composed of the original connection within the modality and the cross-modal connection constrained by KL divergence; is the feature matrix of the Lth layer, are the trainable parameters of graph convolution, is the activation function; During the optimization process, the loss function is combined with the gradient descent algorithm to integrate each subgraph network into a global cross-modal feature graph representation: in, represents the final cross-modal feature map, Represents the subgraph network corresponding to modality i after optimization.

9. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 1 is characterized in that: The step 3 comprises: Step 31: Combine the mainstream value vertical field text knowledge base, use the LDA model to perform topic modeling on the text in the text knowledge base, and extract potential topic words; Step 32: Based on the extracted keywords, the RoBERTa-CRF model is used to extract event triplets, an event graph is constructed, and nodes related to hot spots are retrieved from the event graph through a graph retrieval engine; Step 33: Construct a subgraph network from related nodes, and convert the subgraph network into a vector representation to form a hotspot pattern; Step 34: By calculating the Jaccard similarity, the content features similar to the hotspot mode are matched in the cross-modal feature graph representation obtained in step 2 to identify potential hotspot content in the cross-modal feature graph.

10. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 9 is characterized in that: In step 31: The following formula is used to implement topic modeling of texts in the text knowledge base using the LDA model: in, For terms On topic Distribution under, documentation The topic distribution Subject to the Dirichlet prior, the term On topic The distribution under Construct a topic-word matrix.

11. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 9 is characterized in that: The step 32 comprises: Through the RoBERTa layer to the input text Encode and get the vector representation of each word ; Learning Sequence Probabilities of Word Labels Based on Vector Representations ,in is a label sequence, according to the maximum conditional probability To predict the optimal label sequence for each word: in, is the normalization factor, is the weight of the transferred feature, It is The word label, is the label of the i-th word, represents the index of the tag, is an exponential function, is the length of the word sequence, is the number of label types; The event triples are converted into event graphs, and the association relationship of events is represented by a graph structure; the subject and object in the event triples are used as nodes in the graph, and the predicate is used as the edge connecting the nodes: in, Represents the event graph, is a collection of nodes, is the edge set; Each entity pair in the event triple By predicate connected.

12. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 9 is characterized in that: The step 33 comprises: For each node , Is a node set, generating a fixed-length node sequence , as the context of the event graph, where q is the walk length; Node sequence Embedding learning is performed to maximize the co-occurrence probability between a node and its context by the following formula: in, Is a node The context node set of Is a node Generate context node probability; Each node The embedding vector That is its representation in d-dimensional vector space; for the entire subgraph , the overall vector representation of the subgraph is generated by aggregating the embedding vectors of all its nodes. The subgraph vector It is expressed as: in, is an aggregate function, Represents the subgraph vector, that is, the hotspot pattern finally formed.

13. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 9 is characterized in that: The step 34 comprises: By calculating the Jaccard similarity between the sub-graph vector and the cross-modal graph representation obtained in step 2, content features similar to the hot spot pattern are found in the obtained cross-modal graph representation, so as to identify potential hot content; by comparing the Jaccard similarity, potential hot content in the cross-modal feature graph is identified.

14. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 1 is characterized in that: The method of combining the cross-modal feature graph representation to construct a cross-modal content graph includes: Combined with the global feature map representation obtained in step 2 , mapping all modal graph representations to a high-dimensional semantic space; the dimension of the semantic space is d, and the graph representation of each modality is Define a mapping function : Mapping Function The graph representation of each modality is mapped to a high-dimensional semantic space, and the generated semantic vector is recorded as ,in: Get the set of modal feature vectors after semantic alignment , that is, the cross-modal content graph.

15. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 14 is characterized in that: The method of learning anomaly patterns using labeled anomaly data and training anomaly detection models according to the anomaly patterns includes: The features of different modalities are semantically aligned using graph embedding technology. The nodes of the cross-modal graph are , each node corresponds to the characteristics of a mode; By constructing a semantic similarity matrix : in, Represents the similarity between modality i and modality j, using cosine similarity Calculate the similarity of each pair of modes, and They are the feature vectors of modality i and modality j after being mapped to the high-dimensional semantic space; The annotated anomaly dataset is And use it to train anomaly detection model: According to the semantic similarity matrix , using graph neural networks to detect abnormal nodes in cross-modal graphs: in, Nodes in the graph The characteristic representation of node The abnormal probability of node Is it an abnormal node? represents the sigmoid activation function, represents the weight of the feature in the model, represents the parameters learned during the training process, Representation Node The features and their neighbor nodes The weighted sum of the features of When the anomaly detection model is trained, the loss function of the anomaly detection model is defined by the following formula: : Get the loss function After that, by minimizing To optimize the anomaly detection model.

16. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 1 is characterized in that: The identifying abnormal content in the cross-modal content graph by using the trained anomaly detection model includes: Calculate the abnormal probability of all nodes in the modal content graph, filter out abnormal nodes according to preset thresholds, and process or alarm the corresponding modal content; identify abnormal nodes through the graph neural network model.

17. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 1 is characterized in that: The step 5 comprises: Step 51: Construct a SKIR model to simulate the process of information dissemination among users; Step 52: Use a dual-factor coupled long-tail information cascade method to predict the propagation links of hot content and abnormal content.

18. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 17 is characterized in that: The step 51 comprises: The SKIR model is constructed, and the informed state is introduced to expand the propagation mechanism of the model: SKIR model definition: unknown information state, informed state, propagation state, unaware state: in, represents the number of individuals whose information is unknown at time t, represents the number of individuals who are in the informed state at time t, It represents the number of users in the propagation state at time t. represents the number of individuals in the propagation state at time t, represents how the number of users with unknown information status at time t changes over time, represents how the number of informed users at time t changes over time, represents the change in the number of users in the propagation state at time t, Indicates the dynamic change of the number of users in the unaware state at time t, is the decay rate corresponding to different states, is the transmission intensity, is the propagation diffusion coefficient associated with the informed state, controlling the rate at which information propagates from the informed state to the propagated state, is the rate of transition from the informed state to the unaware state, indicating that informed users stop spreading information because they lose interest in it. is the hesitant informed rate; The stability condition is used to determine whether the SKIR model has reached a state of equilibrium and whether information propagation will stop. The stability condition is defined by the following formula: in, Indicates the stability of the equilibrium point of the SKIR model. When Δ<0, the model will reach equilibrium and stop information propagation. It is the resistance coefficient of information dissemination.

19. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 18 is characterized in that: The step 52 comprises: The sample selection of the model is optimized by long-tail distribution sampling and class-balanced sampling. The calculation formula of the sampling probability is: And by weighted combination of long-tail distribution sampling and class-balanced sampling, we get the mixed sampling probability: in, is the long-tail distribution sampling probability of node j, is the eigenvalue of node j, is the total number of the i-th node, R is the sampling parameter, C represents the number of categories, is the probability of class-balanced sampling, where the sampling probability of each class is equal and is and is the mixed sampling probability, e is the balance factor, which is used to adjust the weights of different sampling strategies; The calculation formula of global propagation intensity is: in, represents the propagation intensity at time t, is the global propagation parameter, is the time at which the current event was propagated, is the time of the last propagation event, δ is the global decay factor, It is the sampling probability calculated by the sampling strategy, which reflects the influence of the propagation intensity of different nodes on the sampling probability; is the propagation intensity prediction within a local time window, and It is a function calculated from the time difference or other input features, which is used to indicate the influence of a node in the propagation process. and are the weighting coefficients of the global and local propagation models, Affects the calculation of local propagation intensity; Combining the stability conditions of the SKIR model , and the final fusion of global and local propagation strength is obtained: in, is the final content dissemination link, that is, the final dissemination intensity, is the propagation strength from different modules, is the weight coefficient of each module, It is a stability adjustment factor used to control the impact of stability conditions on the strength of the propagation link.

20. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 1 is characterized in that: The step 6 comprises: Step 61: Introduce dynamic adjustment factors based on the SKIR model framework and node weight , construct the control equation, and control the state transition in the information propagation process through the adjustment factors and node weights. The state transition includes the probability changes of susceptible state, informed state and propagation state; Step 62: Monitor and update the parameters of the propagation link in step 5, and use the Kalman filter algorithm to update the model parameters, where the updated model parameters include state prediction and observation correction; Step 63: perform key node weight adjustment, and calculate the node adjustment weight based on the node betweenness centrality and real-time influence index; Step 64: Optimize the propagation direction and rate to maximize the target content coverage and establish an optimal control problem; Step 65: Generate a dynamic distribution strategy, input the optimized parameters, divide users into multiple groups through the spectral clustering algorithm, define group sensitivity, and calculate the delivery time window.

21. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 20 is characterized in that: In step 61, based on the SKIR model framework, a dynamic adjustment factor is introduced and node weight , construct the control equations, including: in, For Node The neighbor set of is a dynamic diffusion adjustment factor, ranging from [0,1], which is used to control the information penetration intensity across nodes. For Node The propagation weight ranges from [0,2], and is adjusted by It can suppress or amplify the propagation influence of nodes. represents the rate at which the unknown information state changes to the propagation state, is the propagation diffusion coefficient associated with the informed state, controlling the rate at which information propagates from the informed state to the propagated state, It is the transition rate from the informed state to the unaware state, indicating that informed users stop spreading information because they lose interest in it.

22. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 21 is characterized in that: The step 62 comprises: State prediction based on historical parameters To predict the current parameters : in, is the state transfer matrix, is the control input matrix, For external intervention signals, is the process noise, is the final propagation link outputted in step 5; Observation correction through real-time observation data To modify the parameters: in, is the observation matrix, is the Kalman gain matrix, Represents the updated state prediction parameters.

23. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 20 is characterized in that: The step 63 comprises: Based on node betweenness centrality and real-time impact indicators , calculate the node adjustment weight: in, is the learning rate, which is used to control the weight update amplitude, and finally outputs the adjusted node weight set .

24. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 20 is characterized in that: The step 64 comprises: Optimize the direction and rate of dissemination to maximize the coverage of target content As the goal, establish the optimal control problem: in, For the delivery cycle, is the weight regularization coefficient, and the Pontryagin maximum principle is used to solve the Hamiltonian function , derive the optimal control law: in, is a covariate variable, represents the optimal output weight.

25. The network content dissemination method based on integrated cognitive understanding and intelligent management according to claim 21 is characterized in that: The step 65 comprises: By entering the optimized parameters , using spectral clustering algorithm to divide users into Groups , defining group sensitivity , calculate the delivery time window: in, It represents the best delivery time window, that is, the best time for content delivery to a group. is the attenuation coefficient, It represents the average number of users in the propagation state at time s, the triple , indicating that In time By weight Deliver content.

26. A network content dissemination system based on integrated cognitive understanding and intelligent management, which implements the network content dissemination method based on integrated cognitive understanding and intelligent management as described in any one of claims 1 to 25, characterized in that: include: A decoupling module is used to decouple data features of different modes and output the decoupled data features; Different modalities include text, video, audio, and images; The generation module is used to construct a subgraph network for each modality based on the decoupled data features of audio, video, image, and text, and to perform convolution operations on cross-modal features through multiple subgraph networks to generate cross-modal feature graph representations; The hot content identification module is used to perform topic modeling on the text in the text knowledge base, extract potential topic words, and extract event triplets from the potential topic words to construct an event graph; retrieve nodes related to hot spots from the event graph to form hot spot patterns, match content features related to hot spot patterns in the cross-modal feature graph representation, and identify potential hot content; The abnormal content identification module is used to combine the cross-modal feature graph representation, construct a cross-modal content map, use the annotated abnormal data to learn abnormal patterns, and train an anomaly detection model based on the abnormal patterns to identify abnormal content in the cross-modal content map; abnormal content includes false information, sensitive images, and dense videos; The prediction module is used to simulate the process of information dissemination among users and predict the content dissemination links of hot content and abnormal content; The determination module is used to adjust the weights of key nodes in the content dissemination chain and determine the time, location and target audience of content delivery.

Citation Information

Patent Citations

  • Long-tail cascade popularity prediction model, training method and prediction method

    CN113887806A

  • SKIR information spreading method based on online social network

    CN114628038A

  • Method and system for constructing double-layer coupled network infectious disease transmission model based on game strategy

    CN119207827A

  • Potential hot content identification method based on mainstream value vertical domain knowledge base

    CN119557460A

  • Cross-modal data decoupling method, characterization method and characterization system

    CN119557844A

Cited By

  • Wireless bridge fixed link channel feature decoupling and abnormal state detection method and system based on double-domain comparative learning

    CN120639227A

  • Multi-mode AI content security risk traceability and identification detection system

    CN120832695A

  • A Multimodal AI Content Security Risk Tracing and Identification System

    CN120832695B

  • Network news false information spreading screening method

    CN121092809A

  • Multi-modal video public opinion monitoring method and system based on hotspot identification and tracking

    CN122116240A