Information management method and system for digital new media based on big data
By screening noun characteristics and dynamic resource allocation, the problem of blocked real-time analysis capabilities of text data in existing systems and delayed processing of high-value data in existing systems is solved, efficient information situation discovery and processing is achieved, and real-time and decision-making support capabilities of the information management system are improved.
Patent Information
- Application Number
- CN202510455244.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The existing digital new media information management system lacks part-of-speech filtering and multi-dimensional system evaluation when selecting keywords, resulting in the blockage of real-time analysis capabilities of text data, long video processing, delaying the discovery timeliness of hot spot information situations, and fixed resource allocation strategies lead to delays in processing high-value data.
By filtering noun characteristics, the document is constructed, keyword extraction and similarity comparison are used to achieve text data clustering, and the initial information trend category is formed, and the information trend value score is calculated based on the number of views, forwarding and comments, a three-level classification model of information trend popularity with dynamic thresholds is established, and resource allocation is dynamically adjusted to prioritize high-hot information.
It improves the efficiency of information situation category construction, reduces the number of candidate words, improves the efficiency of discovering hot information situations, ensures the timely processing of high-value data, and provides information situation warning and timely guarantee for responding to decisions.
Smart Images

Figure CN120448864A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis and management, and in particular to an information management method and system based on big data digital new media. Background Art
[0002] In the digital new media era, information management faces three core challenges: massive data, real-time requirements, and multi-source heterogeneity. To meet these challenges, it is necessary to build a systematic management method to achieve efficient information processing through technology integration and process optimization.
[0003] At the same time, with the increase in information complexity, the value of information situation monitoring has become increasingly prominent and has become a key support for organizational strategic decision-making. There is a deep synergistic relationship between the information management system and the information situation monitoring system: the former provides the latter with the basic capabilities of data collection, cleaning and storage, and the latter feeds back the former's intelligent upgrade through functions such as semantic analysis and trend prediction. The two form an organic whole with dynamic feedback, jointly improving the organization's risk response and decision-making efficiency.
[0004] Existing digital new media information management systems generally adopt a fixed-cycle batch processing model, which has the following defects: the existing system lacks part-of-speech filtering and multi-dimensional system evaluation when selecting keywords, and adopts a synchronous processing mechanism for text and video data. Since video processing takes a long time, the real-time analysis capability of text data is blocked, delaying the discovery of hot information trends. In addition, the use of a fixed resource allocation strategy leads to delays in the processing of high-value data. Therefore, we propose an information management method and system for digital new media based on big data to solve this problem. Summary of the Invention
[0005] In order to solve the above technical problems, a method and system for information management of digital new media based on big data are provided. This technical solution solves the problem that the existing system proposed in the above background technology lacks part-of-speech filtering and multi-dimensional system evaluation when selecting keywords, and adopts a synchronous processing mechanism for text and video data. Since video processing takes a long time, the real-time analysis capability of text data is blocked, delaying the discovery time of hot information trends, and adopts a fixed resource allocation strategy, which leads to delays in high-value data processing.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: The present invention provides an information management method for digital new media based on big data, the method comprising: S1. Determine the target data source based on analysis requirements and obtain digital media data within a specified time range, including text and video data, and simultaneously obtain the number of views, forwarding, and comments for each piece of data; S2. Filter noun features from text data to construct documents, cluster text data through keyword extraction and similarity comparison, and form initial information situation categories; S3. Calculate the information situation value score for each text data item based on the number of views, forwarding, and comments, and add them up to generate a total score for the information situation category; S4. Establish a three-level classification model for information situation heat based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated simultaneously. S5. Classify and process the video data, and dynamically adjust the information situation category.
[0007] Furthermore, a fire monitoring and early warning system based on multi-source image data is proposed, which is used to implement any of the above-mentioned digital new media information management methods based on big data, including: The digital media data acquisition module is used to determine the target data source based on analysis requirements, obtain text and video data of digital media within a specified time period, and simultaneously obtain the number of views, forwarding volumes, and comments; The information situation category construction module is used to filter noun features from text data to construct documents, cluster text data through keyword extraction and similarity comparison, and form initial information situation categories; The information situation value assessment module is used to calculate the data information situation value score based on the number of views, forwarding volumes and comments, and accumulate the total score of the information situation category; Hotspot identification and information situation analysis module, which is used to identify hotspot information situation categories through a preset threshold mechanism and extract their multi-dimensional feature information; The video information situation fusion module calculates the video data information situation value score. If it can be included in the existing category, the score is accumulated and the hot spot status is updated. If not, a new category is created and the total score is established.
[0008] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention performs part-of-speech standardization on text data, retaining only noun words, avoiding ambiguity while reducing the number of candidate words, improving the efficiency of keyword screening, and comprehensively considering the semantic representativeness, cross-category distribution differences and basic weights of nouns to screen out keywords with strong representativeness and high discrimination, thereby improving the efficiency of information situation category construction and reducing similar categories.
[0009] 2. The present invention forms initial information situation categories by clustering text data, uses the sum of the information situation value scores of the text data as the total score of the information situation category, and compares it with the information situation score threshold to achieve hotspot prediction, classify and process the video data, and give priority to matching potential hotspots, so that categories close to the information situation score threshold can be identified as hotspot categories more quickly, thereby comprehensively improving the efficiency of discovering hotspot information situations.
[0010] 3. By real-time monitoring of the hotspot status of each information situation category, dynamically controlling the start and stop of data analysis, and intelligently allocating system resources, the main computing resources are concentrated in a few hotspot categories, giving priority to processing high-heat information situation data, improving analysis efficiency, and implementing hierarchical resource allocation based on the total score of the category information situation. The higher the score, the more resources are allocated, ensuring that hotspot information situations are processed first, and accelerating analysis to provide timeliness guarantee for information situation warning and response decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 This is a flow chart of an information management method for digital new media based on big data proposed by the present invention; Figure 2 This is a flow chart of the method for identifying hotspot information situation categories in the present invention; Figure 3 This is a flow chart of the method for dynamically adjusting information situation categories in the present invention; Figure 4 This is a structural block diagram of an information management system based on big data digital new media proposed by the present invention. DETAILED DESCRIPTION
[0012] The following description is intended to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are merely examples, and those skilled in the art may conceive of other obvious variations.
[0013] Reference Figure 1-4 As shown, the present invention provides an information management method for digital new media based on big data, the method comprising: S1. Determine the target data source based on analysis requirements and obtain digital media data within a specified time range, including text and video data, and simultaneously obtain the number of views, forwarding, and comments for each piece of data; S2. Filter noun features from text data to construct documents, cluster text data through keyword extraction and similarity comparison, and form initial information situation categories; Specifically, the process of screening noun features from text data to construct documents, clustering text data through keyword extraction and similarity comparison, and forming initial information situation categories specifically includes: Process text data through part-of-speech tagging, retain all lexical elements with noun part-of-speech tags, record them as documents, perform word frequency statistics on documents, and calculate noun frequency by combining inverse document frequency. The basic weight ; It should be noted that compared with verbs and adjectives, nouns have the characteristics of stable semantics and clear reference, which can effectively avoid ambiguity. Processing only nouns can significantly reduce the number of candidate words and improve the efficiency of keyword screening.
[0014] Nouns are transformed into And all document nouns are converted into semantic vectors, and the cosine similarity between the vector of noun t and each document noun vector is calculated one by one, according to the formula Analyze nouns The semantic representativeness score of ,in and Nouns The vector of and the vectors of other nouns in the document; Perform unsupervised topic modeling on all documents, generate multiple temporary topic clusters, and calculate nouns The document frequency in each temporary cluster and the maximum value , statistical terms The number of occurrences in all documents , according to the formula Analyzing nouns The difference in cross-category distribution; It should be noted that the semantic representativeness score When it approaches 1, it means that the noun is highly correlated with multiple nouns in the document and is likely to be a core keyword. When it approaches 1, it indicates that the noun is highly concentrated in a single topic cluster and is suitable as a classification label word. The basic weight As a supplement, high-frequency words can directly reflect the importance of nouns in the document.
[0015] Calculate nouns by semantic statistical joint weight formula Keyword ratings; The semantic statistical joint weight formula is: ; Where, Score keywords. is the weight coefficient; For example, the above It can be specifically 0.6, 0.3, and 0.1. The semantic representativeness score directly determines the degree of fit between the noun and the information situation category and plays a dominant role. The cross-category difference can effectively isolate similar information situations as a key auxiliary. Too high This will cause interference from high-frequency common words in long documents.
[0016] The top five nouns with the highest keyword scores are retained to construct a document keyword set. All documents are clustered according to the similarity of the document keyword set to form initial information situation categories. The keywords of all documents in each information situation category are counted, and the top three high-frequency keywords are taken as category labels.
[0017] S3. Calculate the information situation value score for each text data item based on the number of views, forwarding, and comments, and add them up to generate a total score for the information situation category; Specifically, based on the number of views, forwardings, and comments, calculating the information situation value score of each text data specifically includes: an information situation value scoring process based on percentiles, obtaining the information situation value score of the text data, the process includes S31-S34: S31. Perform validity checks on page views, forwarding numbers, and comment counts, remove zero and abnormal negative values, and fill in missing values with default values. S32. Based on the historical data of the data source platform, the value formula Calculate the percentile values of views, reposts, and comments, where is the total amount of data, Corresponding to the number of views, forwarding and comments respectively, The rankings corresponding to the number of views, reposts, and comments; S33. If the values of views, forwarding and comments Exceeds the maximum value of historical data , then the value formula is set to: ; S34. Set weights for page views, forwarding, and comment counts , according to the formula Calculate the information situation value score of text data.
[0018] For example, the above Specifically, they are 0.5, 0.3, and 0.2. Through percentile conversion, the number of views, reposts, and comments are unified into the same dimension, which is convenient for direct weighted calculation. It also ensures that high-scoring content has a significant impact on the final score to avoid being diluted by ordinary content, so that the top content can effectively widen the gap and truly reflect the hot spots of the information situation.
[0019] S4. Establish a three-level classification model for information situation heat based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated simultaneously. See Figure 2 As shown, the three-level classification model of information situation heat based on dynamic thresholds is established. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously. Specifically, it includes: Setting information situation score thresholds and similarity judgment threshold, the total score The information situation category is marked as a hot category, and the total score is The information situation categories within the interval are marked as potential hotspot categories, and the total score The information situation category is marked as an unpopular category; It should be noted that the information situation threshold setting needs to consider both data volume and data quality. First, the total amount of digital media data N in the time period is counted, and the proportion of data in the hot category P (15%~30%) is determined. Combined with the information situation value scoring process, a value between 40 and 60 is taken as the data quality Y. Finally, the threshold is passed. Calculate the information situation score threshold.
[0020] When text data or video data is added to or removed from an information situation category, the total information situation score of the category is recalculated, the latest total score is compared with the information situation score threshold in real time, and the hot status of the information situation category is updated; It should be noted in particular that the present invention uses text and video data as the core analysis objects to construct an information situation monitoring system. Based on the main proportion of text and video data in the total amount of information situation data, this solution uses them as basic analysis units. The architecture design can be expanded in the embodiments to be compatible with the collaborative analysis of heterogeneous data sources such as pictures and voice.
[0021] The embodiment of the present invention realizes intelligent allocation of system resources by real-time monitoring of the hotspot status of each information situation category and dynamically controlling the start and stop mechanism of data analysis. The system automatically concentrates the main computing resources on a few hotspot categories, gives priority to processing high-heat information situation data, and significantly improves the analysis efficiency. Based on the total score of the category information situation, hierarchical resource allocation is implemented. The higher the score, the more computing resources the category obtains, ensuring that the hotspot information situation obtains priority processing rights. This mechanism not only accelerates the analysis and processing speed of the hotspot information situation through dynamic resource allocation, but also provides timeliness guarantee for subsequent information situation warning and response decision-making.
[0022] S5. Classify and process the video data, and dynamically adjust the information situation category.
[0023] See Figure 3 As shown, the video data is classified and processed, and the information situation category is dynamically adjusted, specifically including: dynamically extracting key frames of the video data, performing image analysis on the key frames to generate a single-frame keyword set, aggregating the keyword sets of each frame, and generating a keyword set of the video data; Specifically, the dynamic extraction of key frames of video data includes: performing motion activity analysis on video data, dynamically adjusting the sampling interval based on motion intensity: using a larger sampling interval in low motion intervals and a smaller sampling interval in high motion intervals, and generating a preliminary key frame set; identifying scene switching points through shot boundary detection, and incorporating the first frame and the last frame of each new scene into the preliminary key frame set; based on the preliminary key frame set, reducing the key frame resolution to generate thumbnails, calculating the average pixel difference between adjacent thumbnails, and when the difference is greater than a preset threshold, marking the key frame corresponding to the thumbnail as a frame to be examined; constructing a three-channel feature evaluation model of brightness, edge and motion, and evaluating the comprehensive difference between the frames to be examined, and triggering the key frame supplement mechanism when the comprehensive difference exceeds an adaptive threshold; combining the preliminary key frame set and the supplementary key frame, performing a temporal consistency check on the final key frame set, and outputting a single-frame keyword set of the video data.
[0024] It should be noted that the preset threshold is set at 20~30 and can be adjusted proportionally based on the thumbnail resolution; the baseline value of the adaptive threshold is set to 30, and the motion intensity is calculated by the inter-frame difference method. It is increased to 1.3 times the baseline value during high motion and reduced to 0.8 times the baseline value during low motion.
[0025] According to the priority order of potential hot categories, hot categories, and unpopular categories, the semantic similarity between the keyword set of the video data and the labels of each category is calculated in turn. If the similarity is greater than or equal to the similarity judgment threshold, the video data is included in the information situation category; Calculate the information situation value score of video data based on the percentile-based information situation value scoring process; If the video data can be included in the existing information situation category, its information situation value score will be added to the total score of the category and the hotspot status will be updated. If it cannot be included, a new information situation category will be created and its information situation value score will be established as the total score.
[0026] In an embodiment of the present invention, the system will first match video data with potential hotspot categories instead of traversing all information situation categories. When the match is successful, the system will accumulate the information situation value score of the video into the total score of the corresponding category. This optimization strategy enables categories close to the information situation score threshold to be identified as hotspot categories more quickly; through the dynamic evaluation system of the information situation value score, the system will update the cumulative score of each category in real time. When the total score of an information situation category is close to the information situation score threshold, the system will mark it as a potential hotspot category. This mechanism ensures that hot information situations can be discovered in a timely manner while avoiding the performance bottleneck of full data processing.
[0027] See Figure 4As shown, this solution proposes an information management system for digital new media based on big data, which is used to implement the above-mentioned information management method for digital new media based on big data, including: The digital media data acquisition module is used to determine the target data source based on analysis requirements, obtain text and video data of digital media within a specified time period, and simultaneously obtain the number of views, forwarding volumes, and comments; The information situation category construction module is used to filter noun features from text data to construct documents, cluster text data through keyword extraction and similarity comparison, and form initial information situation categories; The information situation value assessment module is used to calculate the data information situation value score based on the number of views, forwarding volumes and comments, and accumulate the total score of the information situation category; The hotspot identification and information situation analysis module is used to establish a three-level classification model for information situation heat based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously. The video information situation fusion module calculates the video data information situation value score. If it can be included in the existing category, the score is accumulated and the hot spot status is updated. If not, a new category is created and the total score is established.
[0028] The information situation category building module specifically includes: Noun feature extraction unit, used to perform part-of-speech tagging analysis on text data, filter and retain all noun vocabulary elements, count the frequency of nouns in the document and calculate the basic weight value of each noun; The semantic representativeness calculation unit is used to load the pre-trained word vector model, convert nouns into semantic vector representations, calculate the cosine similarity between the noun vector and other nouns in the document, and combine the semantic statistical joint weight formula to obtain the semantic representativeness score; A cross-topic distribution difference calculation unit is used to perform topic modeling on the document set, generate a preset number of temporary topic clusters, count the distribution frequency of nouns in each topic cluster, calculate the cross-topic distribution difference of nouns, and output the cross-category distribution difference of nouns; The keyword score calculation unit is used to set the weight coefficients of semantic and distribution features, integrate the semantic representativeness score and distribution difference, and calculate the final keyword score of the noun; The information situation category generation unit is used to construct a keyword similarity matrix between documents, execute a hierarchical clustering algorithm, merge documents whose similarity reaches a threshold, generate labels based on high-frequency keywords within a category, and output a set of labeled information situation categories.
[0029] The information situation value assessment module specifically includes: The data cleaning and verification unit is used to verify the validity of page views, forwarding volume, and comment counts, filter out abnormal values less than or equal to zero, and fill in missing data with default values; The information situation value score calculation unit is used to calculate the percentile values of views, forwarding volumes and comment numbers based on historical data distribution, and to perform weighted summation of the percentile values of the three dimensions according to preset weight coefficients to generate a standardized information situation value score.
[0030] The video information situation fusion module specifically includes: The video keyword extraction unit is used to extract video key frames, analyze the key frame content using an image recognition model, generate single-frame keywords, and aggregate the keywords and text information of each frame to form a keyword set for the video data; A multi-level category matching unit is used to establish a three-level matching priority queue and calculate the semantic similarity between the video keyword set and each category label; An information situation value calculation unit, configured to calculate the information situation value score of the video data based on a percentile information situation value scoring process; The dynamic fusion decision unit is used to determine whether the video data can be included in the existing information situation category. If it can, the information situation value score of the video data will be added to the total score of the category and the hotspot status will be updated. If it cannot be included, a new information situation category will be created and the information situation value score of the video data will be established as the total score.
[0031] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions merely illustrate the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for information management of digital new media based on big data, characterized in that: The method comprises: S1. Determine the target data source based on analysis requirements and obtain digital media data within a specified time range, including text and video data, and simultaneously obtain the number of views, forwarding, and comments for each piece of data; S2. Filter noun features from text data to construct documents, cluster text data through keyword extraction and similarity comparison, and form initial information situation categories; S3. Calculate the information situation value score for each text data item based on the number of views, forwarding, and comments, and add them up to generate a total score for the information situation category; S4. Establish a three-level classification model for information situation heat based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated simultaneously. S5. Classify and process the video data, and dynamically adjust the information situation category.
2. The method according to claim 1, characterized in that The process of screening noun features from text data to construct documents, clustering text data through keyword extraction and similarity comparison, and forming initial information situation categories specifically includes: Process text data through part-of-speech tagging, retain all lexical elements with noun part-of-speech tags, record them as documents, perform word frequency statistics on documents, and calculate noun frequency by combining inverse document frequency. The basic weight ; Nouns are transformed into And all document nouns are converted into semantic vectors, and the cosine similarity between the vector of noun t and each document noun vector is calculated one by one. Analyzing nouns The semantic representativeness score of ,in and Nouns The vector of and the vectors of other nouns in the document; Perform unsupervised topic modeling on all documents, generate multiple temporary topic clusters, and calculate nouns The document frequency in each temporary cluster and the maximum value , statistical terms The number of occurrences in all documents , according to the formula Analyzing nouns The difference in cross-category distribution; Calculate nouns by semantic statistical joint weight formula Keyword ratings; The semantic statistical joint weight formula is: ; Where, Score keywords. is the weight coefficient; The top five nouns with the highest keyword scores are retained to construct a document keyword set. All documents are clustered according to the similarity of the document keyword set to form initial information situation categories. The keywords of all documents in each information situation category are counted, and the top three high-frequency keywords are taken as category labels.
3. The method according to claim 1, characterized in that The calculation of the information situation value score of each text data based on the number of views, forwarding and comments specifically includes: The information situation value scoring process based on percentiles is used to obtain the information situation value score of text data. The process includes S31-S34: S31. Perform validity checks on page views, forwarding numbers, and comment counts, remove zero and abnormal negative values, and fill in missing values with default values. S32. Based on the historical data of the data source platform, the value formula Calculate the percentile values of views, reposts, and comments, where is the total amount of data, Corresponding to the number of views, forwarding and comments respectively, The rankings corresponding to the number of views, reposts, and comments; S33. If the values of views, forwarding and comments Exceeds the maximum value of historical data , then the value formula is set to: ; S34. Set weights for page views, forwarding, and comment counts , according to the formula Calculate the information situation value score of text data.
4. The method according to claim 3, characterized in that The three-level classification model of information situation heat based on dynamic thresholds is established. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously. Specifically, it includes: Setting information situation score thresholds and similarity judgment threshold, the total score The information situation category is marked as a hot category, and the total score is The information situation categories within the interval are marked as potential hotspot categories, and the total score The information situation category is marked as an unpopular category; When text data or video data is added to or removed from an information situation category, the total information situation score of the category is recalculated, the latest total score is compared with the information situation score threshold in real time, and the hotspot status of the information situation category is updated.
5. The method according to claim 4, characterized in that The categorization of video data and dynamic adjustment of information situation categories specifically include: Dynamically extract key frames from video data, perform image analysis on the key frames to generate a single-frame keyword set, aggregate the keyword sets of each frame, and generate a keyword set for the video data; According to the priority order of potential hot categories, hot categories, and unpopular categories, the semantic similarity between the keyword set of the video data and the labels of each category is calculated in turn. If the similarity is greater than or equal to the similarity judgment threshold, the video data is included in the information situation category; Calculate the information situation value score of video data based on the percentile-based information situation value scoring process; If the video data can be included in the existing information situation category, its information situation value score will be added to the total score of the category and the hotspot status will be updated. If it cannot be included, a new information situation category will be created and its information situation value score will be established as the total score.
6. The method according to claim 5, characterized in that The key frames of the dynamic extraction of video data specifically include: Perform motion activity analysis on the video data and dynamically adjust the sampling interval based on the motion intensity: use a larger sampling interval in low-motion intervals and a smaller sampling interval in high-motion intervals to generate a preliminary set of key frames; Identify scene switching points through shot boundary detection and include the first and last frames of each new scene into the preliminary keyframe set; Based on the preliminary set of keyframes, reduce the keyframe resolution to Generate thumbnails, calculate the average pixel difference between adjacent thumbnails, and when the difference is greater than a preset threshold, mark the key frame corresponding to the thumbnail as a frame to be examined; Construct a three-channel feature evaluation model for brightness, edge, and motion, and evaluate the comprehensive difference between the frames to be examined. When the comprehensive difference exceeds the adaptive threshold, the key frame supplement mechanism is triggered; Combined with the preliminary key frame set and the supplementary key frame, the final key frame set is checked for temporal consistency, and a single-frame keyword set of the video data is output.
7. An information management system for digital new media based on big data, characterized in that: A method for implementing information management of digital new media based on big data as claimed in any one of claims 1 to 6, comprising: The digital media data acquisition module is used to determine the target data source based on analysis requirements, obtain text and video data of digital media within a specified time period, and simultaneously obtain the number of views, forwarding volumes, and comments; The information situation category construction module is used to filter noun features from text data to construct documents, cluster text data through keyword extraction and similarity comparison, and form initial information situation categories; The information situation value assessment module is used to calculate the data information situation value score based on the number of views, forwarding volumes and comments, and accumulate the total score of the information situation category; The hotspot identification and information situation analysis module is used to establish a three-level classification model for information situation heat based on dynamic thresholds. When text data or video data is added or deleted, the total score is recalculated and the category status is updated synchronously. The video information situation fusion module calculates the video data information situation value score. If it can be included in the existing category, the score is accumulated and the hotspot status is updated. If not, a new category is created and the total score is established.
8. The information management system for digital new media based on big data according to claim 7, characterized in that: The information situation category building module specifically includes: Noun feature extraction unit, used to perform part-of-speech tagging analysis on text data, filter and retain all noun vocabulary elements, count the frequency of nouns in the document and calculate the basic weight value of each noun; The semantic representativeness calculation unit is used to load the pre-trained word vector model, convert nouns into semantic vector representations, calculate the cosine similarity between the noun vector and other nouns in the document, and combine the semantic statistical joint weight formula to derive the semantic representativeness score; A cross-topic distribution difference calculation unit is used to perform topic modeling on the document set, generate a preset number of temporary topic clusters, count the distribution frequency of nouns in each topic cluster, calculate the cross-topic distribution difference of nouns, and output the cross-category distribution difference of nouns; The keyword score calculation unit is used to set the weight coefficients of semantic and distribution features, integrate the semantic representativeness score and distribution difference, and calculate the final keyword score of the noun; The information situation category generation unit is used to construct a keyword similarity matrix between documents, execute a hierarchical clustering algorithm, merge documents whose similarity reaches a threshold, generate labels based on high-frequency keywords within a category, and output a set of labeled information situation categories.
9. The information management system for digital new media based on big data according to claim 7, characterized in that: The information situation value assessment module specifically includes: The data cleaning and verification unit is used to verify the validity of page views, forwarding volume, and comment counts, filter out abnormal values less than or equal to zero, and fill in missing data with default values; The information situation value score calculation unit is used to calculate the percentile values of the number of views, forwarding volumes and comments based on the historical data distribution, and to perform weighted summation of the percentile values of the three dimensions according to the preset weight coefficient to generate a standardized information situation value score.
10. The information management system for digital new media based on big data according to claim 7, characterized in that: The video information situation fusion module specifically includes: The video keyword extraction unit is used to extract video key frames, analyze the key frame content using an image recognition model, generate single-frame keywords, and aggregate the keywords and text information of each frame to form a keyword set for the video data; A multi-level category matching unit is used to establish a three-level matching priority queue and calculate the semantic similarity between the video keyword set and each category label; An information situation value calculation unit, configured to calculate the information situation value score of the video data based on a percentile information situation value scoring process; The dynamic fusion decision unit is used to determine whether the video data can be included in the existing information situation category. If it can, the information situation value score of the video data will be added to the total score of the category and the hotspot status will be updated. If it cannot be included, a new information situation category will be created and the information situation value score of the video data will be established as the total score.
Citation Information
Patent Citations
A food safety public opinion monitoring method and a food safety public opinion monitoring system
CN109241429A
Video classification method and training method and device of video classification model
CN116310556A
Short video data processing system and processing method based on image processing technology monitoring
CN117235343A
Learning data creation device, classification model learning device, and category assignment device
JP2020052779A