A multi-modal video public opinion monitoring method and system based on hotspot identification tracking

By constructing a multimodal feature fusion and evolution network, the problems of passive response and single criterion in video public opinion monitoring are solved, enabling proactive early warning and multi-dimensional analysis of public opinion events, and improving the intelligence level and early warning accuracy of public opinion monitoring.

CN122116240APending Publication Date: 2026-05-29BEIJING SHIXIN INTERNET TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING SHIXIN INTERNET TECH CO LTD
Filing Date
2026-03-18
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies in video public opinion monitoring suffer from problems such as passive response, single risk criteria, and flat analysis. They lack the ability to proactively predict future development trends and cannot accurately locate the source and evolution of risks.

Method used

A multimodal feature acquisition and processing module is constructed. Through intramodal attention enhancement, cross-modal interaction association, and multimodal deep fusion, a multimodal fusion representation is generated to identify public opinion hotspots and construct an evolution network. This enables the dynamic evolution process of public opinion events to be abstracted into a structured, multi-dimensional evolution trajectory, allowing for current situation monitoring and future development risk prediction.

Benefits of technology

It has improved the intelligence level and proactive early warning capabilities of public opinion monitoring, achieved a technological leap from passive response to proactive early warning, improved the accuracy and timeliness of identifying public opinion hotspots, and can accurately identify abnormal fluctuations and locate the source of risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116240A_ABST
    Figure CN122116240A_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal video public opinion monitoring method and system based on hotspot identification tracking, belong to public opinion monitoring technical field, is by extracting basic modal feature and extended modal feature, by modal internal attention enhancement, cross-modal interaction correlation and multi-modal depth fusion, generate the multi-modal fusion representation of representation video content comprehensive feature, provide accurate, robust feature basis for hotspot identification, based on fusion representation identification public opinion hotspot event and extract feature package, identify the evolution content in subsequent video stream and construct evolution network, the dynamic evolution process of this network constitutes the evolution trajectory of event, based on the tracking of evolution trajectory carries out current situation monitoring analysis and future development risk prediction to public opinion hotspot event, to solve the limitation of traditional technology to rely on single isolated coefficient to carry out state assessment, and the technical defects that cannot track in-depth content semantics, significantly improve the timeliness and accuracy of multi-modal video public opinion monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of public opinion monitoring technology, and more specifically, to a multimodal video public opinion monitoring method and system based on hotspot identification and tracking. Background Technology

[0002] With the explosive growth of internet video content, video content has become the main carrier of online public opinion. Efficient and accurate monitoring and analysis of video public opinion is of great significance for maintaining cyberspace security and social stability.

[0003] For example, Chinese patent application CN120951245A proposes a multimodal video public opinion intelligent monitoring method and system. This method integrates visual, voice, and text features to construct a three-dimensional indicator system of "identification-trend-propagation," and generates a three-dimensional propagation map for visualization, enabling quantitative assessment and tracking of public opinion events. However, this solution mainly focuses on monitoring the status and describing the historical trajectory of already occurred public opinion events, which is a passive monitoring approach. When facing rapidly changing online public opinion, it still has the following technical limitations:

[0004] First, the monitoring mode is passive: it calculates the public opinion judgment coefficient based on the current state and responds after the fact by comparing the public opinion judgment coefficient with the preset threshold, lacking the ability to proactively predict future development trends.

[0005] Second, the risk assessment criteria are too simplistic: compressing multimodal features into a public opinion judgment coefficient makes it difficult to reflect the internal structure and evolution of public opinion, and it cannot pinpoint the source of risk. When the coefficient exceeds the threshold and triggers an early warning, the system can only indicate "risk exists" without revealing the specific source of the risk—is it increased emotional polarization? A sudden change in network structure? Or a conflict of positions at key nodes? This results in decision-makers lacking targeted intervention criteria.

[0006] Third, the evolutionary analysis is flattened: the three-dimensional propagation map is only used as a visualization tool and fails to abstract the process of public opinion evolution into a structured data model that can be quantified and analyzed, resulting in the map information being unquantifiable and unable to reveal the evolutionary laws.

[0007] The aforementioned technical deficiencies mean that the system still needs improvement in terms of early warning accuracy, decision support capabilities, and intelligence level. To address this, we propose a multimodal video public opinion monitoring method and system based on hotspot identification and tracking. This method involves constructing a quantifiable evolutionary network model to abstract the dynamic evolution of public opinion events into a structured, multi-dimensional evolutionary trajectory. Based on this, it achieves an organic unity between current situation monitoring and future development risk prediction, thereby improving the intelligence level and proactive early warning capabilities of public opinion monitoring. Summary of the Invention

[0008] The purpose of this invention is to address practical technical deficiencies. It provides a multimodal video public opinion monitoring method and system based on hotspot identification and tracking, achieving a technological leap from passive response to proactive early warning, from single-point judgment to multidimensional analysis, and from planar display to structured modeling.

[0009] The objective of this invention can be achieved through the following technical solution: a multimodal video public opinion monitoring system based on hotspot identification and tracking, including a multimodal feature acquisition and processing module, used to acquire multimodal feature data in parallel from video data streams, the multimodal feature data including basic modal features and extended modal features, and to perform time axis alignment and vectorized embedding on feature data of different modalities to generate a set of multimodal feature vectors of a unified dimension;

[0010] The cross-fusion analysis module is used to perform intramodal attention enhancement, cross-modal interaction correlation, and multimodal deep fusion on the multimodal feature vector set to generate a multimodal fusion representation that represents the comprehensive features of video content;

[0011] The hot topic event identification module is used to identify public opinion hot topics from video data streams based on multimodal fusion representation and through preset hot topic detection algorithms, and to construct the feature package of the public opinion hot topic event;

[0012] The evolution trajectory tracking module is used to identify the evolution content related to the public opinion hotspot event in the subsequent video data stream based on feature packets, construct an evolution network based on the evolution content, and the dynamic evolution process of the evolution network over time constitutes the evolution trajectory of the event.

[0013] The evolution trajectory early warning module is used to monitor and analyze the current situation and predict future development risks of public opinion hot events based on the evolution trajectory, and generate corresponding early warning signals based on the analysis results.

[0014] Furthermore, the basic modal features include visual features, speech features, and text semantic features, while the extended modal features include metadata features, user-generated content sentiment features, dissemination behavior features, and spatiotemporal environment features.

[0015] Furthermore, the cross-fusion analysis module includes:

[0016] The intramodal attention unit is used to perform self-attention calculation on each type of modality feature in the multimodal feature vector set, extract the deep correlation features and temporal dependencies within the modality, and output the intramodal enhanced feature sequence;

[0017] The cross-modal interaction unit is used to perform interactive calculations on intra-modal enhanced feature sequences of different modalities, quantify the correlation between features of different modalities, and output cross-modal correlation information;

[0018] The multimodal fusion unit is used to deeply fuse intramodal enhanced feature sequences with cross-modal correlation information to generate the multimodal fusion representation.

[0019] Furthermore, the process of identifying trending public opinion events from video data streams includes:

[0020] The multimodal fusion representation is input into a preset hotspot classifier, which outputs the confidence score of the video content belonging to a public opinion hotspot event. When the confidence score exceeds the preset hotspot judgment threshold, the video content is marked as a public opinion hotspot event. The public opinion hotspot events are aggregated and analyzed. Multiple video contents with multimodal fusion representation similarity exceeding the preset clustering threshold and overlapping time windows are aggregated into the same public opinion hotspot event. The start time, key video segments, core dissemination nodes and initial sentiment tendency of the public opinion hotspot event are extracted to form the feature package of the public opinion hotspot event.

[0021] Furthermore, the process of identifying evolving content in subsequent video data streams includes:

[0022] The feature package of the hot public opinion event is obtained, and the multimodal fusion representation of multiple key video segments contained in the feature package is aggregated to generate the event center representation vector.

[0023] Using the event center representation vector as the retrieval benchmark, the system continuously scans subsequent video data streams, calculates the similarity between the multimodal fusion representation of new video content and the event center representation vector, and classifies new video content into different association levels based on the similarity level, assigns corresponding relevance weights, and uses video content that meets the preset association conditions as evolutionary content.

[0024] Furthermore, the process of constructing evolutionary networks based on evolutionary content and tracking evolutionary trajectories includes:

[0025] Each piece of evolved content is treated as a network node, and each network node is assigned attributes including at least the publication time, publication platform, sentiment value, and the aforementioned relevance weight.

[0026] Edges are established between nodes based on the propagation relationship or content similarity between evolved content, forming an evolutionary network. The dynamic evolution process of the evolutionary network over time includes the addition of new nodes, the establishment of new edges, and changes in node attributes, which together constitute the evolutionary trajectory of public opinion hotspot events.

[0027] Furthermore, the process of monitoring and analyzing the current situation of public opinion hotspots includes:

[0028] Obtain the multidimensional temporal characteristics of the evolutionary network at the current moment and the multidimensional temporal characteristics in its historical evolution record, including propagation heat characteristics, network structure characteristics, sentiment characteristics and spatial distribution characteristics;

[0029] Historical baseline values ​​of various time series features are extracted from the historical multidimensional time series feature sequence. The current multidimensional time series features are compared with the corresponding historical baseline values. The current risk values ​​of the spread heat risk index, network structure risk index, sentiment risk index and spatial diffusion risk index are calculated respectively. Based on the preset current risk judgment rules, the current risk values ​​of each feature are compared with the threshold. If the current risk value is greater than the current risk threshold, the corresponding warning signal is generated.

[0030] Furthermore, the process of predicting and analyzing the future development risks of public opinion hot topics includes: inputting the multidimensional time series feature sequence into a pre-trained time series prediction model, and outputting the predicted multidimensional time series feature sequence within a future preset time window;

[0031] Each predicted feature in the predicted multidimensional time series is compared and analyzed with the corresponding historical benchmark or preset risk threshold. The predicted risk values ​​of the spread heat risk index, network structure risk index, sentiment risk index and spatial diffusion risk index are calculated. Based on the preset future risk judgment rules, the predicted risk values ​​are compared with the threshold. If the predicted risk value is greater than the predicted risk threshold, the corresponding early warning signal is generated.

[0032] This invention also proposes a multimodal video public opinion monitoring method based on hotspot identification and tracking, comprising the following steps:

[0033] S1. Collect multimodal feature data from the video data stream, perform time-axis alignment and vectorized embedding, and generate a set of multimodal feature vectors of uniform dimension;

[0034] S2. Perform intra-modal attention enhancement, cross-modal interaction association, and multimodal deep fusion on the multimodal feature vector set to generate a multimodal fusion representation that represents the comprehensive features of the video content;

[0035] S3. Based on multimodal fusion representation, identify public opinion hotspots from video data streams and extract the feature packages of these public opinion hotspots.

[0036] S4. Based on feature packets, identify the evolutionary content related to public opinion hot topics in subsequent video data streams, construct an evolutionary network based on the evolutionary content, and track the evolutionary trajectory of the event based on the evolutionary network;

[0037] S5. Based on the evolution trajectory, conduct current situation monitoring and analysis and future development risk prediction analysis of public opinion hot events, and generate corresponding early warning signals based on the analysis results.

[0038] Compared with the prior art, the advantages of this invention are:

[0039] 1. This invention constructs a comprehensive, multi-layered multimodal feature system by extracting basic modal features (visual, speech, and text semantics) and extended modal features (metadata, user-generated content sentiment, dissemination behavior, and spatiotemporal environment). Based on this, it extracts deep correlation features and temporal dependencies within various modalities through intramodal attention mechanisms. It quantifies the semantic correlation, dissemination dynamics correlation, and subject correlation between different modalities through cross-modal interactive computation. Finally, it generates a multimodal fusion representation that represents the comprehensive features of video content through deep multimodal fusion. This multimodal fusion representation serves as the input to the hot topic event identification module, significantly improving the accuracy, robustness, and early detection capability of public opinion hot topic identification.

[0040] 2. This invention also abstracts the dynamic evolution of public opinion hotspots into a structured, multi-dimensional evolutionary trajectory by constructing a quantifiable evolutionary network. The evolutionary network records key information in the event propagation process in the form of nodes and edges, including release time, release platform, sentiment tendency, and propagation relationship, providing a data foundation for in-depth quantitative analysis of public opinion trends. It completes the technical leap from "visual display" to "structured modeling" and effectively overcomes the limitations of traditional technologies that can only perform state monitoring and planar display.

[0041] 3. This invention also achieves the organic unity of current situation monitoring and future development risk prediction based on the evolution trajectory. By transforming the public opinion evolution process into quantifiable and analyzable multi-dimensional time-series characteristics and comparing them with historical benchmark values, it accurately identifies abnormal fluctuations and locates the source of risks, avoiding the risk of misjudgment due to single threshold judgment. Through time-series prediction models, it predicts future evolution trends and combines prediction features with risk judgment rules to predict trends such as heat peaks, emotional polarization, and structural mutations in advance, winning a golden window for public opinion handling. It significantly improves the timeliness and accuracy of multimodal video public opinion monitoring, realizes in-depth quantification of the public opinion evolution process, and effectively avoids the limitations of traditional technologies that rely solely on isolated coefficients for state assessment. Attached Figure Description

[0042] Figure 1 This is a block diagram illustrating the system module principle of the present invention;

[0043] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0045] Example 1: This invention discloses a multimodal video public opinion monitoring system based on hotspot identification and tracking. Please refer to [link / reference]. Figures 1-2 It includes a multimodal feature acquisition and processing module, a cross-fusion analysis module, a hotspot event identification module, an evolution trajectory tracking module, and an evolution trajectory early warning module.

[0046] The multimodal feature acquisition module is used to acquire multimodal feature data in parallel from the video data stream. The multimodal feature data includes basic modal features and extended modal features.

[0047] Here, the basic modal features include visual features, speech features, and text semantic features. Specifically, keyframes are extracted from the video at intervals using keyframe extraction technology, and scene, object, facial expression, and action features are extracted as visual features using a pre-trained visual model. Acoustic features are extracted from the video audio stream, and speech is converted into text using an automatic speech recognition model to extract speech features. Text content is extracted from video subtitles, OCR-recognized on-screen text, video titles, and descriptions, and semantic vectors are generated using a pre-trained language model as text semantic features.

[0048] Extended modal features include metadata features, user-generated content sentiment features, dissemination behavior features, and spatiotemporal environment features. Among them, structured data such as video title, tags, upload time, channel identifier, video duration, and resolution are obtained through API interfaces or web crawlers as metadata features. Bullet comments, user comments, likes, and dislikes are collected in real time and sentiment analysis and popularity calculation are performed to obtain user-generated content sentiment features. Completion rate, number of reposts and their growth rate, like rate, comment rate, cross-platform dissemination volume, and 24-hour information increment dissemination behavior indicators are collected as dissemination behavior features. User geographical location distribution, video release time, and time decay factor on the dissemination path are analyzed as spatiotemporal environment features.

[0049] This module collects features from multiple dimensions and aligns and embeds them into vectorized feature data from different modalities along the time axis, generating a unified set of multimodal feature vectors. This provides a comprehensive and rich data foundation for subsequent analysis, ensuring that the system can fully perceive public opinion information from multiple perspectives such as content, interaction, dissemination, and spatiotemporal dimensions.

[0050] The cross-fusion analysis module, connected to the multimodal feature acquisition and processing module, is used to perform intramodal attention enhancement, cross-modal interaction correlation, and multimodal deep fusion on the multimodal feature vector set to generate a multimodal fusion representation that represents the comprehensive features of video content.

[0051] Specifically, the cross-integration analysis module includes:

[0052] The intramodal attention unit is used to perform self-attention calculation on each type of modality feature in the multimodal feature vector set, extract the deep correlation features and temporal dependencies within the modality, and output the intramodal enhanced feature sequence;

[0053] The cross-modal interaction unit is used to perform interactive calculations on intra-modal enhanced feature sequences of different modalities, quantify the correlation between features of different modalities, and output cross-modal correlation information;

[0054] The multimodal fusion unit is used to deeply fuse intramodal enhanced feature sequences with cross-modal correlation information to generate a multimodal fusion representation;

[0055] The specific process for obtaining the intra-modal enhanced feature sequence includes:

[0056] For the feature data of each modality, the attention weight between each time point within the feature vector sequence is calculated through a multi-head self-attention mechanism. Based on the attention weight, the feature data of all time points in the sequence are weighted and summed to obtain the enhanced feature vector of the current time point that incorporates the sequence context information. Based on the time series, the enhanced feature sequence within the modality is output.

[0057] The intramodal augmented feature sequence carries the deep correlation features and temporal dependencies within the modality extracted through self-attention computation. The deep correlation features refer to the non-linear, high-order semantic or physical correlations that exist between different time points within the same modality, while the temporal dependencies refer to the patterns, trends, and stages of feature evolution over time.

[0058] The intramodal attention unit fully integrates temporal contextual information into each modal feature through a self-attention mechanism, enhancing the representational ability within the modality, such as capturing the trajectory of emotional evolution in visual images and the temporal fluctuation pattern of bullet screen emotions.

[0059] The specific process of obtaining the cross-modal correlation vector sequence includes:

[0060] For any two intramodal augmented feature sequences of different modalities, a cross-modal interaction matrix is ​​constructed. Each element in the cross-modal interaction matrix represents the correlation between the feature vector of the first modality at the first time point and the feature vector of the second modality at the second time point. The correlation is obtained by calculating the similarity between the two feature vectors. The similarity calculation method includes cosine similarity, vector dot product or cross-modal attention function with learnable parameterization.

[0061] The cross-modal interaction matrix is ​​pooled along the time axis to obtain a cross-modal association vector that reflects the overall degree of association between the two modes. A corresponding cross-modal association vector sequence is generated for each pair of modes participating in the interaction.

[0062] The correlation calculation includes one or more of the following: the semantic correlation between visual features and the emotional features of user-generated content, the correlation between dissemination behavior features and spatiotemporal environment features, and the subject correlation between metadata features and interaction relationship features;

[0063] Cross-modal interaction units reveal implicit correlations between modalities by calculating the similarity between features of different modalities, such as the synchronicity between visual content and audience reactions, and the coupling relationship between dissemination behavior and spatiotemporal environment. They can discover public opinion signals that cannot be captured by a single modality.

[0064] The specific generation process of the multimodal fusion representation, which characterizes the comprehensive features of video content, includes:

[0065] Receive the intramodal augmentation feature sequence of all modalities and the cross-modal association vector sequence of all modal pairs. For each time point, concatenate the augmentation feature vector of all modalities corresponding to that time point with the cross-modal association vector of that time point (including at least direct vector concatenation, weighted concatenation, extended concatenation, and hierarchical concatenation) to form a joint feature vector that integrates intramodal information and intermodal association.

[0066] Temporal pooling is applied to the joint feature vector at all time points to obtain the global representation vector of the entire video data stream as a multimodal fusion representation;

[0067] Temporal pooling operations can be one of average pooling, max pooling, or attention-based weighted pooling. Alternatively, a hierarchical fusion strategy can be adopted: first, the features of some modalities are fused to obtain a primary joint representation, and then the primary joint representation is gradually fused with the features of the remaining modalities to finally output a fixed-dimensional multimodal fusion representation.

[0068] Intramodal attention units organically integrate intramodal enhancement features with intermodal correlation information through splicing, pooling, or hierarchical fusion strategies, forming a holographic encoding of video content. This provides an accurate, robust, and early-detection-sensitive feature foundation for subsequent hotspot identification.

[0069] The hot topic event identification module, connected to the cross-fusion analysis module, is used to identify public opinion hot topics from video data streams based on multimodal fusion representation and a preset hot topic detection algorithm, and to construct the feature package of the public opinion hot topic event.

[0070] The process of identifying trending public opinion events from video data streams includes:

[0071] The multimodal fusion representation is input into a preset hotspot classifier, which is trained based on historical public opinion data, and outputs the confidence level that the video content belongs to a public opinion hotspot event.

[0072] When the confidence level exceeds the preset hotspot determination threshold, the video content is marked as a hot topic event in public opinion.

[0073] Aggregate analysis of public opinion hotspots, and aggregate multiple video contents with multimodal fusion representation similarity exceeding the preset clustering threshold and overlapping time windows into the same public opinion hotspot;

[0074] Extract the start time, key video clips, core dissemination nodes, and initial sentiment of the public opinion hotspot event to form a feature package of the public opinion hotspot event.

[0075] The evolution trajectory tracking module, connected to the hot event identification module, is used to identify the evolution content related to public opinion hot events in subsequent video data streams based on feature packets, construct an evolution network based on the evolution content, and the dynamic evolution process of the evolution network over time constitutes the evolution trajectory of the event, and tracks the evolution trajectory of the event in time, space and relation dimensions in real time.

[0076] The process of identifying evolving content in subsequent video data streams based on feature packets includes:

[0077] The feature package of the hot public opinion event is obtained. This feature package contains multiple key video segments identified in the initial stage of the event and their corresponding multimodal fusion representations. The multimodal fusion representations of the multiple key video segments in the feature package are aggregated to generate the event center representation vector. The event center representation vector comprehensively reflects the overall semantic features of the hot event in the initial stage.

[0078] Using the event center representation vector as the retrieval benchmark, subsequent video data streams are continuously scanned. The similarity between the multimodal fusion representation of new video content and the event center representation vector is calculated. Based on the similarity level, the new video content is divided into different association levels and assigned corresponding relevance weights. Here, the relevance weights are determined according to the similarity value between the evolved content and the event center representation vector, following a preset mapping rule. Video content that meets the preset association conditions is considered evolved content.

[0079] Specifically, when the similarity exceeds a preset high threshold, the new video content is marked as core evolution content. For example, a similarity of 0.8 or higher is considered core related content, and a weight of 1.0 is assigned, giving it a high weight. This confirms it as newly added related video content. Based on the multimodal fusion representation of the newly added related video content, the event center representation vector is updated for continuous identification of content related to the event in subsequent video streams. This process, by dynamically updating the event center representation vector, can adapt to semantic drift during the event evolution process and continuously recall related content.

[0080] When the similarity is between the preset high threshold and the preset low threshold, the new video content is marked as edge evolution content. The similarity between 0.5 and 0.8 is edge related content. The weight is assigned to 0.3-0.8 by linear interpolation, and it is given a medium to low weight. Its propagation information is recorded but not used to update the event center representation vector.

[0081] When the similarity is lower than a preset low threshold, the new video content is marked as background noise and not included in the evolutionary network construction;

[0082] All new video content that is tagged as core-related content and peripheral-related content is collectively referred to as evolutionary content, and is output for subsequent evolutionary network construction.

[0083] The process of constructing evolutionary networks based on evolutionary content and tracking evolutionary trajectories includes:

[0084] Each evolved content is treated as a network node, and each network node is assigned attributes including at least the release time, release platform, sentiment value, and relevance weight. The process of obtaining the sentiment value is as follows: the multimodal fusion representation of each evolved content is input into a pre-trained sentiment classification model, which is trained on a multimodal video dataset labeled with sentiment polarity and outputs a sentiment value in the range of [-1,1], where -1 represents extreme negative, 0 represents neutral, and 1 represents extreme positive.

[0085] Edges between nodes are established based on the propagation relationship or content similarity between evolved content, forming an evolutionary network. The dynamic evolution process of the evolutionary network over time includes the addition of new nodes, the establishment of new edges, and the changes in node attributes, which together constitute the evolutionary trajectory of public opinion hotspot events displayed in a three-dimensional map.

[0086] This three-dimensional map uses time as the horizontal axis, space and relationships as the vertical axis, and depth as the evolutionary trajectory. The evolutionary trajectory includes changes in the intensity of propagation in the time dimension, cross-platform diffusion paths in the spatial dimension, and the evolution of the positions of key nodes in the relationship dimension. This dimensional information is directly reflected in the dynamic evolution of the evolutionary network, visually displaying the entire process of public opinion hotspots from outbreak to decline. The evolutionary trajectory is a structured, multi-dimensional, and interpretable abstract representation of the event evolution process.

[0087] Example 2: Evolutionary trajectory early warning module, connected to the evolutionary trajectory tracking module, is used to monitor and analyze the current situation and predict future development risks of public opinion hot events based on the evolutionary trajectory, and generate corresponding early warning signals based on the analysis results;

[0088] The process of monitoring and analyzing the current situation of public opinion hot topics includes:

[0089] The evolutionary trajectory tracking module generates multidimensional temporal features of the evolutionary network at the current moment and its historical evolution records. The multidimensional temporal features include propagation heat features, network structure features, sentiment features and spatial distribution features.

[0090] Extract historical benchmark values ​​of various time series features from the historical multidimensional time series feature sequence, including at least the historical mean, historical trend line or historical fluctuation range;

[0091] The multidimensional time-series characteristics at the current moment are compared and analyzed with the corresponding historical benchmark values ​​to calculate the current risk values ​​of the spread heat risk index, network structure risk index, sentiment risk index and spatial diffusion risk index respectively.

[0092] Based on the preset current risk judgment rules, the current risk values ​​of each item are compared with the threshold. If the current risk value is greater than the current risk threshold, a warning signal corresponding to the current risk value is generated.

[0093] Among them, the characteristics of the spread popularity include the curves of the number of core related content, the number of peripheral related content, and the total number of nodes changing over time; the characteristics of the network structure include the indicators of the density, average degree, number of connected components, and size of the core subgraph of the evolved network changing over time; the characteristics of sentiment tendency include the weighted distribution of the sentiment tendency values ​​of all nodes and its evolution; and the characteristics of spatial distribution include the distribution of nodes on the platform or in the region and its diffusion map. These characteristics not only contain quantitative information, but also contain topological relationships and evolutionary laws, which can reveal the internal mechanism of event propagation.

[0094] This process uses multidimensional deviation analysis to accurately identify abnormal fluctuations and pinpoint the source of risk, avoiding the risk of misjudgment from a single threshold.

[0095] The process of predicting and analyzing the future development risks of trending public opinion events includes:

[0096] Input the multidimensional time series feature sequence into the pre-trained time series prediction model, and output the predicted multidimensional time series feature sequence within the future preset time window;

[0097] Each predicted feature in the predicted multidimensional time series feature sequence is compared and analyzed with the corresponding historical benchmark or preset risk threshold to calculate the predicted risk values ​​of the spread heat risk index, network structure risk index, sentiment risk index and spatial diffusion risk index.

[0098] Based on the preset future risk judgment rules, the predicted risk values ​​are compared with thresholds. If the predicted risk value is greater than the predicted risk threshold, an early warning signal corresponding to the predicted risk value is generated.

[0099] This process uses a time-series prediction model to anticipate trends such as peak popularity, emotional polarization, and structural abrupt changes in advance, thus winning a golden window of opportunity for public opinion management and achieving a technological leap from passive response to proactive early warning.

[0100] The risk indicators upon which the current and future risk assessment rules are based include:

[0101] The risk index of dissemination popularity is calculated based on the degree of deviation between the current or predicted characteristic values ​​of the number of core related content, the number of peripheral related content, and the total number of nodes and their historical benchmark values.

[0102] Network structure risk indicators are calculated based on the degree of deviation or rate of change of the current or predicted characteristic values ​​of the density, average degree, number of connected components, and core subgraph size of the evolved network from their historical baseline values.

[0103] The sentiment tendency risk index is calculated by comparing the current or predicted characteristic value of the weighted distribution of sentiment tendency values ​​of all nodes and their degree of dispersion with the historical benchmark value or polarization threshold.

[0104] The spatial diffusion risk index is calculated by comparing the current or predicted characteristic values ​​of the distribution breadth and diffusion rate of nodes on the platform or region with historical benchmark values ​​or diffusion thresholds.

[0105] The current risk assessment rules and future risk assessment rules may include: the method of setting thresholds (static thresholds, dynamic thresholds, historical benchmark deviation thresholds), the combination method of risks in each dimension (single trigger or comprehensive weighting), and the threshold range corresponding to different risk levels.

[0106] It should also be added that the various threshold comparisons mentioned in the article, such as thresholds, preset values, and preset ranges, are set for result comparison and analysis to determine good or bad. The magnitude of these thresholds is determined by a combination of large-scale model analysis of sample data and human experience, and can also be appropriately adjusted based on seasonal or common-sense influencing factors.

[0107] As can be seen from Examples 1 and 2, this invention also proposes a multimodal video public opinion monitoring method based on hotspot identification and tracking. Please refer to [link / reference]. Figure 2 It includes the following steps:

[0108] S1. Multimodal feature acquisition and processing: Acquire multimodal feature data from video data streams, perform time-axis alignment and vectorized embedding, and generate a set of multimodal feature vectors of uniform dimension;

[0109] S2. Cross-fusion analysis: Intramodal attention enhancement, cross-modal interaction correlation and multimodal deep fusion are performed on the multimodal feature vector set to generate a multimodal fusion representation that represents the comprehensive features of video content;

[0110] S3. Hotspot event identification and feature package extraction: Based on multimodal fusion representation, identify public opinion hotspot events from video data streams, perform aggregate analysis on public opinion hotspot events, and extract feature packages of the same public opinion hotspot event;

[0111] S4. Evolutionary Trajectory Tracking: Based on feature packets, identify the evolutionary content related to public opinion hot topics in subsequent video data streams, construct an evolutionary network based on the evolutionary content, and track the evolutionary trajectory of the event through the dynamic evolution of the evolutionary network;

[0112] S5. Evolutionary Trajectory Early Warning: Based on the evolutionary trajectory, the current situation of public opinion hotspots is monitored and analyzed, and future development risks are predicted and analyzed. Corresponding early warning signals are generated based on the analysis results.

[0113] In summary, this system includes: a multimodal feature acquisition and processing module, a cross-fusion analysis module, a hotspot event identification module, an evolution trajectory tracking module, and an evolution trajectory early warning module. By extracting basic and extended modal features, and through intramodal attention enhancement, cross-modal interaction correlation, and multimodal deep fusion, a multimodal fusion representation representing the comprehensive features of video content is generated, providing an accurate and robust feature foundation for hotspot identification. Based on the fusion representation, public opinion hotspot events are identified and feature packages are extracted. Then, the evolving content in subsequent video streams is identified and an evolution network is constructed. The dynamic evolution process of this network constitutes the evolution trajectory of the event. Based on the evolution trajectory, the current situation monitoring and analysis and future development risk prediction analysis of public opinion hotspot events are performed, and early warning signals are generated.

[0114] Furthermore, by constructing a structured evolutionary trajectory, the evolution of public opinion is transformed into multi-dimensional temporal characteristics that can be quantified and analyzed. On this basis, the organic unity of accurate monitoring of the current situation and proactive prediction of future risks is achieved, which solves the technical defects of traditional technologies that can only make single-point threshold judgments, lack prediction capabilities and closed-loop feedback. This significantly improves the intelligence level of public opinion monitoring, the accuracy of early warning and the initiative of risk prevention and control, and provides efficient and intelligent technical means for the governance of public opinion in cyberspace.

[0115] The above description is merely a preferred embodiment of the present invention; however, the scope of protection of the present invention is not limited thereto; any equivalent substitutions or modifications made by those skilled in the art within the technical scope disclosed in the present invention, based on the technical solution and its improved concept, should be covered within the scope of protection of the present invention.

Claims

1. A multimodal video public opinion monitoring system based on hotspot identification and tracking, characterized in that: include: The multimodal feature acquisition and processing module is used to acquire multimodal feature data in parallel from the video data stream. The multimodal feature data includes basic modal features and extended modal features. The module performs time axis alignment and vectorized embedding on the feature data of different modalities to generate a set of multimodal feature vectors of a unified dimension. The cross-fusion analysis module is used to perform intramodal attention enhancement, cross-modal interaction correlation, and multimodal deep fusion on the multimodal feature vector set to generate a multimodal fusion representation that represents the comprehensive features of video content; The hot topic event identification module is used to identify public opinion hot topics from video data streams based on multimodal fusion representation and through preset hot topic detection algorithms, and to construct the feature package of the public opinion hot topic event; The evolution trajectory tracking module is used to identify the evolution content related to the public opinion hotspot event in the subsequent video data stream based on feature packets, construct an evolution network based on the evolution content, and the dynamic evolution process of the evolution network over time constitutes the evolution trajectory of the event. The evolution trajectory early warning module is used to monitor and analyze the current situation and predict future development risks of public opinion hot events based on the evolution trajectory, and generate corresponding early warning signals based on the analysis results.

2. The multimodal video public opinion monitoring system based on hotspot identification and tracking according to claim 1, characterized in that: The basic modal features include visual features, speech features, and text semantic features, while the extended modal features include metadata features, user-generated content emotional features, dissemination behavior features, and spatiotemporal environment features.

3. The multimodal video public opinion monitoring system based on hotspot identification and tracking according to claim 2, characterized in that: The cross-fusion analysis module includes: The intramodal attention unit is used to perform self-attention calculation on each type of modality feature in the multimodal feature vector set, extract the deep correlation features and temporal dependencies within the modality, and output the intramodal enhanced feature sequence; The cross-modal interaction unit is used to perform interactive calculations on intra-modal enhanced feature sequences of different modalities, quantify the correlation between features of different modalities, and output cross-modal correlation information; The multimodal fusion unit is used to deeply fuse intramodal enhanced feature sequences with cross-modal correlation information to generate the multimodal fusion representation.

4. The multimodal video public opinion monitoring system based on hotspot identification and tracking according to claim 3, characterized in that: The process of identifying trending public opinion events from video data streams includes: The multimodal fusion representation is input into a preset hotspot classifier, which outputs the confidence score of whether the video content belongs to a public opinion hotspot event. When the confidence score exceeds the preset hotspot judgment threshold, the video content is marked as a public opinion hotspot event. The public opinion hotspot events are aggregated and analyzed. Multiple video contents with multimodal fusion representation similarity exceeding the preset clustering threshold and overlapping time windows are aggregated into the same public opinion hotspot event. The start time, key video segments, core dissemination nodes and initial sentiment tendency of the public opinion hotspot event are extracted to form the feature package of the public opinion hotspot event.

5. A multimodal video public opinion monitoring system based on hotspot identification and tracking according to claim 4, characterized in that: The process of identifying evolving content in subsequent video data streams includes: The feature package of the hot public opinion event is obtained, and the multimodal fusion representation of multiple key video segments contained in the feature package is aggregated to generate the event center representation vector. Using the event center representation vector as the retrieval benchmark, the system continuously scans subsequent video data streams, calculates the similarity between the multimodal fusion representation of new video content and the event center representation vector, and classifies new video content into different association levels based on the similarity level, assigns corresponding relevance weights, and uses video content that meets the preset association conditions as evolutionary content.

6. The multimodal video public opinion monitoring system based on hotspot identification and tracking according to claim 5, characterized in that: The process of constructing evolutionary networks based on evolutionary content and tracking evolutionary trajectories includes: Each piece of evolved content is treated as a network node, and each network node is assigned attributes including at least the publication time, publication platform, sentiment value, and the aforementioned relevance weight. Edges are established between nodes based on the propagation relationship or content similarity between evolved content, forming an evolutionary network. The dynamic evolution process of the evolutionary network over time includes the addition of new nodes, the establishment of new edges, and changes in node attributes, which together constitute the evolutionary trajectory of public opinion hotspot events.

7. A multimodal video public opinion monitoring system based on hotspot identification and tracking according to claim 6, characterized in that: The process of monitoring and analyzing the current situation of public opinion hot topics includes: Obtain the multidimensional temporal characteristics of the evolutionary network at the current moment and the multidimensional temporal characteristics in its historical evolution record, including propagation heat characteristics, network structure characteristics, sentiment characteristics and spatial distribution characteristics; Historical baseline values ​​of various time series features are extracted from the historical multidimensional time series feature sequence. The current multidimensional time series features are compared with the corresponding historical baseline values. The current risk values ​​of the spread heat risk index, network structure risk index, sentiment risk index and spatial diffusion risk index are calculated respectively. Based on the preset current risk judgment rules, the current risk values ​​of each feature are compared with the threshold. If the current risk value is greater than the current risk threshold, the corresponding warning signal is generated.

8. The multimodal video public opinion monitoring method and system based on hotspot identification and tracking according to claim 7, characterized in that: The process of predicting and analyzing the future development risks of public opinion hotspots includes: inputting the multidimensional time series feature sequence into a pre-trained time series prediction model, and outputting the predicted multidimensional time series feature sequence within a future preset time window; Each predicted feature in the predicted multidimensional time series is compared and analyzed with the corresponding historical benchmark or preset risk threshold. The predicted risk values ​​of the spread heat risk index, network structure risk index, sentiment risk index and spatial diffusion risk index are calculated. Based on the preset future risk judgment rules, the predicted risk values ​​are compared with the threshold. If the predicted risk value is greater than the predicted risk threshold, the corresponding early warning signal is generated.

9. A multimodal video public opinion monitoring method based on hotspot identification and tracking, employing a multimodal video public opinion monitoring system based on hotspot identification and tracking as described in any one of claims 1-8, characterized in that, Includes the following steps: S1. Collect multimodal feature data from the video data stream, perform time-axis alignment and vectorized embedding, and generate a set of multimodal feature vectors of uniform dimension; S2. Perform intra-modal attention enhancement, cross-modal interaction association, and multimodal deep fusion on the multimodal feature vector set to generate a multimodal fusion representation that represents the comprehensive features of the video content; S3. Based on multimodal fusion representation, identify public opinion hotspots from video data streams and extract the feature packages of these public opinion hotspots. S4. Based on feature packets, identify the evolutionary content related to public opinion hot topics in subsequent video data streams, construct an evolutionary network based on the evolutionary content, and track the evolutionary trajectory of the event based on the evolutionary network; S5. Based on the evolution trajectory, conduct current situation monitoring and analysis and future development risk prediction analysis of public opinion hot events, and generate corresponding early warning signals based on the analysis results.