A big data-based digital new media information management method and system
By filtering nominal features and dynamically adjusting information situation categories, the problems of limited real-time text data analysis capabilities and time-consuming video processing in existing systems have been solved, enabling efficient information situation discovery and processing, and ensuring timely discovery of hot information and decision support.
Patent Information
- Application Number
- CN202510455244.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-04-11
AI Technical Summary
Existing digital new media information management systems lack part-of-speech filtering and multi-dimensional evaluation when selecting keywords, which leads to the blockage of real-time text data analysis capabilities, long video processing time, delayed discovery of trending information, and fixed resource allocation strategies that cause delays in the processing of high-value data.
By selecting noun-related features to construct documents, calculating information status value scores based on page views, forwards, and comments, and establishing a dynamic threshold-based three-level classification model for popularity, the system performs clustering on text data and dynamically adjusts information status categories, prioritizing high-popularity information status data to achieve intelligent resource allocation.
It improves the efficiency of constructing information situation categories, reduces the efficiency of discovering similar hot information situations, prioritizes the efficiency of hot spot discovery, ensures the timely discovery and processing of hot information situations, and provides timely guarantee for information situation early warning and response decision-making.
Smart Images

Figure CN120448864B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of data analysis and management, in particular to an information management method and system for digital new media based on big data. BACKGROUND
[0002] In the era of digital new media, information management faces three core challenges: massive data, real-time requirements and multi-source heterogeneity. To cope with these challenges, a systematic management method needs to be built to achieve efficient information processing through technology integration and process optimization.
[0003] At the same time, with the increasing complexity of information, the value of information situation monitoring is increasingly prominent, and it has become a key support for organizational strategic decision-making. There is a deep synergy between the information management system and the information situation monitoring system: the former provides the basic capabilities of data collection, cleaning and storage for the latter, and the latter feeds back the intelligent upgrade of the former through semantic analysis, trend prediction and other functions. The two form an organic whole of dynamic feedback, and together improve the risk response and decision-making efficiency of the organization.
[0004] The existing digital new media information management system generally adopts a fixed cycle batch processing mode, which has the following defects: the existing system lacks part-of-speech filtering and multi-dimensional system evaluation when selecting keywords, adopts a text and video data synchronous processing mechanism, and due to the long time-consuming of video processing, the real-time analysis capability of text data is blocked, delaying the discovery of hot information situation, and adopts a fixed resource allocation strategy, resulting in delay in processing high-value data. Therefore, we propose an information management method and system for digital new media based on big data to solve this problem. SUMMARY
[0005] To solve the above technical problems, an information management method and system for digital new media based on big data are provided, which solves the problem of the existing system lacking part-of-speech filtering and multi-dimensional system evaluation when selecting keywords, adopting a text and video data synchronous processing mechanism, and due to the long time-consuming of video processing, the real-time analysis capability of text data is blocked, delaying the discovery of hot information situation, and adopting a fixed resource allocation strategy, resulting in delay in processing high-value data.
[0006] To achieve the above purpose, the following technical solutions are implemented: the present application provides an information management method for digital new media based on big data, which comprises:
[0007] S1. Determine the target data source according to the analysis requirements, obtain the digital media data in the specified time range, including text data and video data, and synchronously obtain the number of views, the number of forwards and the number of comments of each data;
[0008] S2. Select noun features from text data to construct documents, and achieve text data clustering through keyword extraction and similarity comparison to form initial information situation categories;
[0009] S3. Based on the number of views, reposts, and comments, calculate the information situation value score of each text data, and sum them to generate the total score of the information situation category;
[0010] S4. Establish a three-level classification model for information trend popularity based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously.
[0011] S5. Classify and process video data, and dynamically adjust the information status category.
[0012] Furthermore, a fire monitoring and early warning system based on multi-source image data is proposed to implement the information management method based on big data and digital new media described above, including:
[0013] The digital media data acquisition module is used to determine the target data source according to the analysis requirements, acquire text and video data of digital media within a specified time period, and simultaneously acquire the number of views, reposts and comments.
[0014] The information situation category construction module is used to filter noun features from text data to construct documents, and to cluster text data through keyword extraction and similarity comparison to form initial information situation categories;
[0015] The Information Situation Value Assessment module is used to calculate the data information situation value score based on the number of views, reposts, and comments, and to accumulate and generate a total score for the information situation category.
[0016] The hotspot identification and information situation analysis module is used to identify the hotspot information situation category through a preset threshold mechanism and extract its multi-dimensional feature information.
[0017] The video information situation fusion module calculates the situation value score of video data information. If it can be included in the existing category, the score is accumulated and the hotspot status is updated. If it cannot be included, a new category is created and the total score is established.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0019] 1. This invention standardizes text data by part-of-speech tagging, retaining only noun words, thus avoiding ambiguity and reducing the number of candidate words, improving the efficiency of keyword selection. It comprehensively considers the semantic representativeness of nouns, cross-category distribution differences, and basic weights to select keywords with strong representativeness and high distinguishability, thereby improving the efficiency of information situation category construction and reducing similar categories.
[0020] 2. This invention forms initial information situation categories by clustering text data, uses the sum of information situation value scores of text data as the total score of the information situation category, and compares it with the information situation score threshold to achieve hotspot prediction. It also classifies video data, prioritizes matching potential hotspots, and enables categories close to the information situation score threshold to be identified as hotspot categories more quickly, thereby comprehensively improving the efficiency of hotspot information situation discovery.
[0021] 3. By monitoring the hotspot status of each information situation category in real time, dynamically controlling the start and stop of data parsing, and intelligently allocating system resources, the main computing resources are concentrated on a few hotspot categories, prioritizing the processing of high-intensity information situation data, improving parsing efficiency, and implementing hierarchical resource allocation based on the total score of the information situation of each category, with more resources for higher scores, ensuring that hotspot information situations are processed first, accelerating analysis and providing timely support for information situation early warning and response decisions. Attached Figure Description
[0022] Figure 1 This is a flowchart illustrating an information management method for digital new media based on big data, as proposed in this invention.
[0023] Figure 2 This is a flowchart of the method for identifying hotspot information situation categories in this invention;
[0024] Figure 3 This is a flowchart of the method for dynamically adjusting information situation categories in this invention;
[0025] Figure 4 This is a structural block diagram of an information management system for digital new media based on big data proposed in this invention. Detailed Implementation
[0026] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0027] Reference Figures 1-4 As shown, this invention provides an information management method for digital new media based on big data, the method comprising:
[0028] S1. Determine the target data source based on the analysis requirements, acquire digital media data within a specified time range, including text data and video data, and simultaneously acquire the number of views, reposts, and comments for each piece of data;
[0029] S2. Select noun features from text data to construct documents, and achieve text data clustering through keyword extraction and similarity comparison to form initial information situation categories;
[0030] Specifically, the process of constructing documents by filtering noun-like features from text data, and clustering text data through keyword extraction and similarity comparison to form initial information status categories includes:
[0031] Text data is processed using part-of-speech tagging, retaining all lexical elements with noun part-of-speech tags, which are then recorded as documents. Word frequency statistics are performed on these documents, and inverse document frequency is combined to calculate noun frequencies. Basic weights ;
[0032] It should be noted that, compared to verbs and adjectives, nouns have the characteristics of semantic stability and clear reference, which can effectively avoid ambiguity. Processing only nouns can significantly reduce the number of candidate words and improve the efficiency of keyword screening.
[0033] Nouns are generated through pre-trained word vectors And convert all document nouns into semantic vectors, calculate the cosine similarity between the vector of noun t and the vector of each document noun, using the formula Analysis of nouns semantic representativeness score ,in and Each is a noun The vector of and the vectors of other nouns in the document;
[0034] Unsupervised topic modeling is performed on all documents to generate multiple temporary topic clusters, and nouns are calculated. Document frequency in each temporary cluster and take the maximum value Statistical terms Number of occurrences in all documents From the formula Analysis of nouns Cross-category distributional dissimilarity;
[0035] It should be noted that the semantic representativeness score When the value approaches 1, it indicates that the noun is highly related to multiple nouns in the document, and is likely a core topic term, with cross-category distribution dissimilarity. When the value approaches 1, it indicates that the term is highly concentrated in a single topic cluster and is suitable as a classification tag, with a basic weight. In addition, high-frequency words can directly reflect the importance of nouns in a document.
[0036] Nouns are calculated using a semantic statistical joint weight formula. Keyword scoring;
[0037] The formula for the joint weight of semantic statistics is:
[0038] ;
[0039] In the formula, Scoring keywords These are the weighting coefficients;
[0040] For example, the above Specifically, the values can be 0.6, 0.3, or 0.1. The semantic representativeness score directly determines and dominates the fit between the noun and the information situation category. Cross-category difference effectively isolates similar information situations as a key auxiliary factor. Excessively high scores... This can lead to interference from high-frequency common words in long documents.
[0041] The top five nouns by keyword score are retained to construct a document keyword set. All documents are clustered based on the similarity of the document keyword sets to form initial information situation categories. The keywords of all documents in each information situation category are counted, and the top three most frequent keywords are taken as category labels.
[0042] S3. Based on the number of views, reposts, and comments, calculate the information situation value score of each text data, and sum them to generate the total score of the information situation category;
[0043] Specifically, the calculation of the information situation value score for each text data based on page views, reposts, and comments includes: a percentile-based information situation value scoring process to obtain the information situation value score for the text data, the process of which includes S31-S34:
[0044] S31. Perform validity checks on page views, reposts, and comments, remove zero values and abnormal negative values, and fill missing values with default values;
[0045] S32. Based on historical data from the data source platform, the value formula... Calculate the percentile values for page views, shares, and comments. For the total amount of data, These correspond to page views, reposts, and comments, respectively. The ranking is based on the number of views, shares, and comments.
[0046] S33. If the values of page views, reposts, and comments are... Exceeding the maximum value of historical data Then the value formula is set as follows:
[0047] ;
[0048] S34. Set the weighting metrics for page views, shares, and comments. From the formula Calculate the information situation value score of text data.
[0049] For example, the above can Specifically, the values are 0.5, 0.3, and 0.2. By converting percentiles, the number of views, reposts, and comments are unified to the same dimension, making it easier to calculate the weighted average directly. This ensures that high-scoring content has a significant impact on the final score, preventing it from being diluted by ordinary content, and allowing top-performing content to effectively differentiate itself and truly reflect the current information trends and hot topics.
[0050] S4. Establish a three-level classification model for information trend popularity based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously.
[0051] See Figure 2 As shown, the establishment of a three-level classification model for information trend popularity based on dynamic thresholds, which triggers a recalculation of the total score and synchronously updates the category status when text or video data is added or deleted, specifically includes:
[0052] Set information situation score threshold And similarity judgment threshold, the total score The information situation category is marked as the hot topic category, and the total score is placed in... The information situation categories within the interval are marked as potential hotspot categories, and the total score is calculated accordingly. The information situation category is marked as an unpopular category;
[0053] It should be noted that the information situation threshold setting needs to consider both data volume and data quality. First, the total amount of digital media data N within the time period is statistically analyzed, and the data percentage P of hot topic categories is determined (15%~30%). Then, a value between 40 and 60 is taken as the data quality Y, based on the information situation value scoring process. Finally, the threshold is determined through... The information situation score threshold was calculated.
[0054] When text or video data is added to or removed from an information situation category, the total information situation score for that category is recalculated, the latest total score is compared with the information situation score threshold in real time, and the hotspot status of the information situation category is updated.
[0055] It should be noted that this invention uses text and video data as the core analysis objects to construct an information situation monitoring system. Based on the main proportion of text and video data in the total amount of information situation data, this solution uses them as basic analysis units. The architecture design can be extended in the embodiments to be compatible with the collaborative parsing of heterogeneous data sources such as images and voice.
[0056] This invention, through real-time monitoring of the hotspot status of each information situation category and dynamic control of the start and stop mechanism of data parsing, achieves intelligent allocation of system resources. The system automatically concentrates the main computing resources on a few hotspot categories, prioritizing the processing of high-intensity information situation data, significantly improving parsing efficiency. Based on the total score of the information situation category, a hierarchical resource allocation is implemented, with categories with higher scores receiving more computing resources, ensuring that hotspot information situations receive priority processing rights. This mechanism, through dynamic resource allocation, not only accelerates the analysis and processing speed of hotspot information situations but also provides timely assurance for subsequent information situation early warning and response decisions.
[0057] S5. Classify and process video data, and dynamically adjust the information status category.
[0058] See Figure 3 As shown, the process of classifying and dynamically adjusting the information status categories for video data includes: dynamically extracting key frames from video data, performing image analysis on key frames to generate a set of keywords for each frame, aggregating the keyword sets of each frame, and generating a set of keywords for the video data.
[0059] Specifically, the dynamic extraction of keyframes from video data includes: performing motion activity analysis on the video data and dynamically adjusting the sampling interval based on motion intensity: using a larger sampling interval in low motion intervals and a smaller sampling interval in high motion intervals to generate a preliminary keyframe set; identifying scene switching points through shot boundary detection and including the first and last frames of each new scene into the preliminary keyframe set; based on the preliminary keyframe set, reducing the keyframe resolution to generate thumbnails, calculating the average pixel difference between adjacent thumbnails, and marking the keyframe corresponding to the thumbnail as a frame to be examined when the difference exceeds a preset threshold; constructing a three-channel feature evaluation model for brightness, edge, and motion, and evaluating the comprehensive difference between the frames to be examined; triggering a keyframe supplementation mechanism when the comprehensive difference exceeds an adaptive threshold; and combining the preliminary keyframe set and the supplemented keyframes to perform temporal consistency verification on the final keyframe set and outputting a single-frame keyword set for the video data.
[0060] It should be noted that the preset threshold is set to 20~30, which can be adjusted proportionally based on the thumbnail resolution; the baseline value of the adaptive threshold is set to 30, and the motion intensity is calculated by the inter-frame difference method. During high motion, the threshold is increased to 1.3 times the baseline value, and during low motion, it is reduced to 0.8 times the baseline value.
[0061] Based on the priority order of potential hot topic category, hot topic category, and unpopular category, the semantic similarity between the keyword set of video data and the category tags is calculated in turn. If the similarity is greater than or equal to the similarity judgment threshold, the video data is included in the information situation category.
[0062] The information situation value score of video data is calculated based on the percentile information situation value scoring process.
[0063] If the video data can be included in an existing information situation category, its information situation value score is added to the total score of that category, and the hotspot status is updated. If it cannot be included, a new information situation category is created and its information situation value score is established as the total score.
[0064] In this embodiment of the invention, the system first matches video data with potential hotspot categories, rather than traversing all information situation categories. When a match is successful, the system accumulates the information situation value score of the video into the total score of the corresponding category. This optimization strategy enables categories close to the information situation score threshold to be identified as hotspot categories more quickly. Through the dynamic evaluation system of information situation value scores, the system updates the accumulated scores of each category in real time. When the total score of an information situation category is close to the information situation score threshold, the system marks it as a potential hotspot category. This mechanism ensures that hotspot information situations can be discovered in a timely manner, while avoiding the performance bottleneck of full data processing.
[0065] See Figure 4 As shown, this solution proposes an information management system for digital new media based on big data, used to implement the aforementioned information management method for digital new media based on big data, including:
[0066] The digital media data acquisition module is used to determine the target data source according to the analysis requirements, acquire text and video data of digital media within a specified time period, and simultaneously acquire the number of views, reposts and comments.
[0067] The information situation category construction module is used to filter noun features from text data to construct documents, and to cluster text data through keyword extraction and similarity comparison to form initial information situation categories;
[0068] The Information Situation Value Assessment module is used to calculate the data information situation value score based on the number of views, reposts, and comments, and to accumulate and generate a total score for the information situation category.
[0069] The hotspot identification and information situation analysis module is used to establish a three-level classification model of information situation popularity based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously.
[0070] The video information situation fusion module calculates the situation value score of video data information. If it can be included in the existing category, the score is accumulated and the hotspot status is updated. If it cannot be included, a new category is created and the total score is established.
[0071] The information situation category construction module specifically includes:
[0072] The noun feature extraction unit is used to perform part-of-speech tagging analysis on text data, filter and retain all noun vocabulary elements, count the word frequency of nouns in the document, and calculate the basic weight value of each noun.
[0073] The semantic representativeness calculation unit is used to load the pre-trained word vector model, convert nouns into semantic vector representations, calculate the cosine similarity between the noun vectors and other nouns in the document, and obtain the semantic representativeness score by combining the semantic statistics joint weight formula.
[0074] The cross-topic distribution difference calculation unit is used to perform topic modeling on the document set, generate a preset number of temporary topic clusters, count the distribution frequency of nouns in each topic cluster, calculate the cross-topic distribution difference of nouns, and output the cross-category distribution difference of nouns.
[0075] The keyword scoring calculation unit is used to set the weight coefficients of semantic and distribution features, integrate semantic representativeness scores and distribution differences, and calculate the final keyword score of the noun.
[0076] The information situation category generation unit is used to construct a keyword similarity matrix between documents, execute a hierarchical clustering algorithm, merge documents with similarity reaching a threshold, generate labels by statistically analyzing high-frequency keywords within each category, and output a set of labeled information situation categories.
[0077] The information situation value assessment module specifically includes:
[0078] The data cleaning and verification unit is used to verify the validity of page views, reposts, and comments, filter out outliers less than or equal to zero, and fill in missing data with default values.
[0079] The Information Situation Value Score Calculation Unit is used to calculate the percentile values of page views, reposts, and comments based on historical data distribution, and to perform a weighted sum of the percentile values of the three dimensions according to preset weight coefficients to generate a standardized Information Situation Value Score.
[0080] The video information situational fusion module specifically includes:
[0081] The video keyword extraction unit is used to extract key frames from the video, analyze the content of key frames using an image recognition model, generate single-frame keywords, and aggregate the keywords and text information of each frame to form a keyword set of the video data.
[0082] A multi-level category matching unit is used to establish a three-level matching priority queue and calculate the semantic similarity between the set of video keywords and the tags of each category.
[0083] The Information Situation Value Calculation Unit is used to calculate the Information Situation Value Score of video data based on the percentile-based Information Situation Value Scoring Process.
[0084] The dynamic fusion decision unit is used to determine whether video data can be included in the existing information situation category. If it can, the information situation value score of the video data is added to the total score of the category, and the hotspot status is updated. If it cannot be included, a new information situation category is created and the information situation value score of the video data is established as the total score.
[0085] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. An information management method for digital new media based on big data, characterized in that, The method includes: S1. Determine the target data source based on the analysis requirements, acquire digital media data within a specified time range, including text data and video data, and simultaneously acquire the number of views, reposts, and comments for each piece of data; S2. Select noun features from text data to construct documents, and achieve text data clustering through keyword extraction and similarity comparison to form initial information situation categories; S3. Based on the number of views, reposts, and comments, calculate the information situation value score of each text data, and sum them to generate the total score of the information situation category; S4. Establish a three-level classification model for information trend popularity based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously. S5. Classify and process video data, and dynamically adjust information status categories; the step of selecting noun features from text data to construct documents, and achieving text data clustering through keyword extraction and similarity comparison to form initial information status categories specifically includes: Text data is processed by part-of-speech tagging, all lexical elements with noun part-of-speech tags are retained and recorded as documents, word frequency statistics are performed on the documents, and the basic weight Base(t) of noun t is calculated by combining the inverse document frequency; The noun t and all document nouns are transformed into semantic vectors by pre-training word vectors. The cosine similarity between the vector of noun t and the vector of each document noun is calculated one by one. The semantic representativeness score R(t) of noun t is analyzed by formula R(t)=avg(cos_sim(v_t,v_n)), where v_t and v_n are the vector of noun t and the vector of other nouns in the document, respectively. Unsupervised topic modeling is performed on all documents to generate multiple temporary topic clusters. The document frequency of noun t in each temporary cluster is calculated and the maximum value max_DF(t) is taken. The number of occurrences of noun t in all documents, DF(t), is counted. The cross-category distribution difference of noun t is analyzed by formula S(t)=1-max_DF(t) / DF(t). The keyword score of noun t is calculated by semantic statistical joint weight formula. The formula for the joint weight of semantic statistics is: F(t)=α·R(t)+β·S(t)+γ·Base(t) In the formula, F(t) is the keyword score, and α, β, γ are the weight coefficients; The top five nouns in the keyword scores are retained to construct a document keyword set. All documents are clustered based on the similarity of the document keyword sets to form an initial information situation category. The keywords of all documents in each information situation category are counted, and the top three most frequent keywords are taken as category labels. The establishment of a three-level classification model for information trend popularity based on dynamic thresholds, which triggers a recalculation of the total score and synchronously updates the category status when text or video data is added or deleted, specifically includes: Set an information situation score threshold T and a similarity judgment threshold. Information situation categories with a total score ≥ T are marked as hot categories, information situation categories with a total score in the range [T / 2, T) are marked as potential hot categories, and information situation categories with a total score ≤ T are marked as unpopular categories. When text or video data is added to or removed from an information situation category, the total information situation score for that category is recalculated, the latest total score is compared with the information situation score threshold in real time, and the hotspot status of the information situation category is updated. The process of classifying and dynamically adjusting the information status categories for video data specifically includes: Dynamically extract keyframes from video data, perform image analysis on the keyframes to generate a single-frame keyword set, aggregate the keyword sets of each frame, and generate a keyword set for the video data. Based on the priority order of potential hot topic category, hot topic category, and unpopular category, the semantic similarity between the keyword set of video data and the category tags is calculated in turn. If the similarity is greater than or equal to the similarity judgment threshold, the video data is included in the information situation category. The information situation value score of video data is calculated based on the percentile information situation value scoring process. If the video data can be included in an existing information situation category, its information situation value score is added to the total score of that category, and the hotspot status is updated. If it cannot be included, a new information situation category is created and its information situation value score is established as the total score.
2. The method according to claim 1, characterized in that, The calculation of the information situation value score for each text data based on page views, reposts, and comments specifically includes: The information situation value scoring process based on percentiles is used to obtain the information situation value score of text data. The process includes S31-S34: S31. Perform validity checks on page views, reposts, and comments, remove zero values and abnormal negative values, and fill missing values with default values; S32. Based on historical data from the data source platform, the value formula P... i =100×(rank(i)-1) / (N-1) calculates the percentile values of views, reposts and comments, where N is the total number of data, i=1,2,3 correspond to views, reposts and comments respectively, and rank(i) is the ranking position corresponding to views, reposts and comments. S33. If the values of page views, reposts, and comments are Φ i If the value exceeds the maximum value max_H in historical data, then the value formula is set as follows: P i =100+(Φ i / max_H)×10 S34. Set the weighting metrics for page views, shares, and comments ∑w i =1, from the formula Calculate the information situation value score of text data.
3. The method according to claim 2, characterized in that, The keyframes for dynamically extracting video data specifically include: Motion activity analysis is performed on video data, and the sampling interval is dynamically adjusted based on motion intensity: a larger sampling interval is used in low motion intervals and a smaller sampling interval is used in high motion intervals to generate a preliminary set of keyframes. Scene transition points are identified by lens boundary detection, and the first and last frames of each new scene are included in the initial keyframe set. Based on the initial set of keyframes, the resolution of the keyframes is reduced to 1 / 8 to generate thumbnails. The average pixel difference between adjacent thumbnails is calculated. When the difference is greater than a preset threshold, the keyframe corresponding to the thumbnail is marked as a frame to be examined. A three-channel feature evaluation model for brightness, edge, and motion is constructed, and the overall difference between the frames under examination is evaluated. When the overall difference exceeds the adaptive threshold, a keyframe supplementation mechanism is triggered. By combining the initial keyframe set and the supplementary keyframes, the final keyframe set is subjected to temporal consistency verification, and the single-frame keyword set of the video data is output.
4. An information management system for digital new media based on big data, characterized in that, The information management method for digital new media based on big data as described in any one of claims 1-3 includes: The digital media data acquisition module is used to determine the target data source according to the analysis requirements, acquire text and video data of digital media within a specified time period, and simultaneously acquire the number of views, reposts and comments. The information situation category construction module is used to filter noun features from text data to construct documents, and to cluster text data through keyword extraction and similarity comparison to form initial information situation categories; The Information Situation Value Assessment module is used to calculate the data information situation value score based on the number of views, reposts, and comments, and to accumulate and generate a total score for the information situation category. The hotspot identification and information situation analysis module is used to establish a three-level classification model of information situation popularity based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously. The video information situation fusion module calculates the situation value score of video data information. If it can be included in the existing category, the score is accumulated and the hotspot status is updated. If it cannot be included, a new category is created and the total score is established.
5. The information management system for digital new media based on big data according to claim 4, characterized in that, The information situation category construction module specifically includes: The noun feature extraction unit is used to perform part-of-speech tagging analysis on text data, filter and retain all noun vocabulary elements, count the word frequency of nouns in the document, and calculate the basic weight value of each noun. The semantic representativeness calculation unit is used to load the pre-trained word vector model, convert nouns into semantic vector representations, calculate the cosine similarity between the noun vectors and other nouns in the document, and obtain the semantic representativeness score by combining the semantic statistics joint weight formula. The cross-topic distribution difference calculation unit is used to perform topic modeling on the document set, generate a preset number of temporary topic clusters, count the distribution frequency of nouns in each topic cluster, calculate the cross-topic distribution difference of nouns, and output the cross-category distribution difference of nouns. The keyword scoring calculation unit is used to set the weight coefficients of semantic and distribution features, integrate semantic representativeness scores and distribution differences, and calculate the final keyword score of the noun. The information situation category generation unit is used to construct a keyword similarity matrix between documents, execute a hierarchical clustering algorithm, merge documents with similarity reaching a threshold, generate labels by statistically analyzing high-frequency keywords within each category, and output a set of labeled information situation categories.
6. The information management system for digital new media based on big data according to claim 4, characterized in that, The information situation value assessment module specifically includes: The data cleaning and verification unit is used to verify the validity of page views, reposts, and comments, filter out outliers less than or equal to zero, and fill in missing data with default values. The Information Situation Value Score Calculation Unit is used to calculate the percentile values of page views, reposts, and comments based on historical data distribution, and to perform a weighted sum of the percentile values of the three dimensions according to preset weight coefficients to generate a standardized Information Situation Value Score.
7. The information management system for digital new media based on big data according to claim 4, characterized in that, The video information situational fusion module specifically includes: The video keyword extraction unit is used to extract key frames from the video, analyze the content of key frames using an image recognition model, generate single-frame keywords, and aggregate the keywords and text information of each frame to form a keyword set of the video data. A multi-level category matching unit is used to establish a three-level matching priority queue and calculate the semantic similarity between the set of video keywords and the tags of each category. The Information Situation Value Calculation Unit is used to calculate the Information Situation Value Score of video data based on the percentile-based Information Situation Value Scoring Process. The dynamic fusion decision unit is used to determine whether video data can be included in the existing information situation category. If it can, the information situation value score of the video data is added to the total score of the category, and the hotspot status is updated. If it cannot be included, a new information situation category is created and the information situation value score of the video data is established as the total score.
Citation Information
Patent Citations
A food safety public opinion monitoring method and a food safety public opinion monitoring system
CN109241429A
Short video data processing system and processing method based on image processing technology monitoring
CN117235343A