A news hotspot prediction and scheduling method and system based on a big data platform

By using multiple means to acquire and preprocess news data on the big data platform, multi-dimensional feature extraction and fusion, and using adaptive deep learning models for training, the existing news hotspot prediction methods are solved, and higher prediction accuracy and user experience are achieved.

CN119721410BActive Publication Date: 2025-05-23金华市新闻传媒中心
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510249872.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-23
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing news hotspot prediction methods have low prediction accuracy and poor scheduling strategies, resulting in poor user experience.

Method used

The news hotspot prediction and scheduling method under the big data platform is adopted to obtain news data through multiple means, perform preprocessing and perform multi-dimensional feature extraction and feature fusion. The adaptive deep learning model is used to combine cross entropy loss function and reinforcement learning reward function for training, and intelligently adjust the news scheduling strategy.

Benefits of technology

It improves the accuracy and reliability of news hotspot predictions, optimizes news scheduling strategies, and improves the timeliness of user experience and news push.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721410B_ABST
    Figure CN119721410B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and system for predicting and scheduling news hot spots under a big data platform. It belongs to the technical fields of big data processing, artificial intelligence, machine learning and information prediction and scheduling. The method includes: obtaining news data from relevant data sources by multiple means; the news data includes historical news data and real-time news data; and preprocessing the obtained news data; extracting features of multi-dimensional data with different features after preprocessing by different means; and fusing features of different dimensions through feature fusion strategies to generate fusion feature vectors; and dynamically adjusting network structures and parameters according to historical data and real-time data of news hot spots through adaptive deep learning models. Acquiring and preprocessing news data by multiple means can effectively integrate information from different sources and improve data quality and processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention proposes a news hotspot prediction and scheduling method and system under a big data platform, belonging to the technical field of big data processing, artificial intelligence, machine learning and information prediction and scheduling. Background Art

[0002] Existing news hotspot prediction methods mostly rely on traditional statistical analysis and keyword matching technology, which has limited prediction accuracy and is difficult to capture the dynamic changes of news hotspots. At the same time, news hotspot scheduling often lacks intelligence and automation, resulting in low scheduling efficiency and poor user experience. Therefore, there is an urgent need for a news hotspot prediction and scheduling method and system that can combine advanced technologies such as big data processing, feature fusion, and adaptive deep learning to improve prediction accuracy, optimize scheduling strategies, and enhance user experience. Summary of the invention

[0003] The present invention provides a news hotspot prediction and scheduling method and system under a big data platform, which is used to solve the problems of low prediction accuracy and unintelligent scheduling strategy of existing news hotspot prediction and scheduling methods:

[0004] The present invention proposes a news hotspot prediction and scheduling method under a big data platform, the method comprising:

[0005] S1. Acquire news data from relevant data sources by multiple means; the news data includes historical news data and real-time news data; and pre-process the acquired news data;

[0006] S2, extracting features of preprocessed multi-dimensional data with different features by different means; and fusing features of different dimensions by feature fusion strategy to generate a fused feature vector;

[0007] S3, dynamically adjust the network structure and parameters according to the historical data and real-time data of news hotspots through adaptive deep learning models;

[0008] S4, inputting the fused feature vector and the corresponding news hotspot label into the adaptive deep learning model, and using the cross entropy loss function and the reinforcement learning reward function to jointly guide the model training;

[0009] S5. Use the trained adaptive deep learning model to predict real-time news hotspot data, and intelligently adjust the news scheduling strategy based on the prediction results and dynamic changes in news hotspots.

[0010] The present invention proposes a news hotspot prediction and scheduling system under a big data platform, the system comprising:

[0011] Data acquisition module: acquires news data from relevant data sources through multiple means; the news data includes historical news data and real-time news data; and pre-processes the acquired news data;

[0012] Feature extraction module: extracts features of preprocessed multi-dimensional data with different features through different means; and fuses features of different dimensions through feature fusion strategy to generate fused feature vectors;

[0013] Parameter adjustment module: dynamically adjusts network structure and parameters based on historical and real-time data of news hotspots through adaptive deep learning models;

[0014] Model training module: Input the fused feature vector and the corresponding news hotspot label into the adaptive deep learning model, and use the cross entropy loss function and reinforcement learning reward function to jointly guide the model training;

[0015] Strategy adjustment module: Use the trained adaptive deep learning model to predict real-time news hotspot data, and intelligently adjust the news scheduling strategy based on the prediction results and dynamic changes in news hotspots.

[0016] The invention has the following beneficial effects: by acquiring and preprocessing news data through multiple means, information from different sources can be effectively integrated to improve data quality and processing efficiency; by using multiple technologies to extract multi-dimensional features of news data, and by using feature fusion strategies, a more representative fusion feature vector is generated, which helps to improve the accuracy of prediction; the adaptive deep learning model can dynamically adjust the network structure and parameters according to historical and real-time data, making the model more flexible and adaptable; by jointly guiding model training through the cross entropy loss function and the reinforcement learning reward function, the accuracy and reliability of news hot spot prediction can be effectively improved; according to the prediction results and the dynamic changes of news hot spots, the news scheduling strategy is intelligently adjusted, which helps to improve the timeliness and user satisfaction of news push; by combining user interest portraits and news hot spot prediction results, personalized news recommendations are made to improve user experience and reduce information overload; by continuously tracking user feedback and A / B test results, the push strategy is iteratively optimized to ensure that the system remains competitive in a changing market environment; by analyzing user social network data and evaluating the influence of users in information dissemination, it helps to identify key nodes and community structures and optimize news dissemination paths. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a step diagram of the method of the present invention;

[0018] Figure 2 This is a system module diagram of the present invention. DETAILED DESCRIPTION

[0019] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0020] One embodiment of the present invention, as Figure 1 As shown, a news hotspot prediction and scheduling method under a big data platform includes:

[0021] S1. Acquire news data from relevant data sources through multiple means; the multiple means include PAI interface and web crawler; the relevant data sources include news websites, social media platforms and user behavior logs; the news data include historical news data and real-time news data; the news data include multi-dimensional data such as text, pictures, videos and user comments; and pre-process the acquired news data;

[0022] S2. Extract features of pre-processed multi-dimensional data with different features through different means; including extracting text features through natural language processing technology (such as BERT, GPT, etc.), extracting image and video features using deep learning models (such as ResNet, VGG, etc.), and extracting user interest features in combination with user behavior log analysis; and through feature fusion strategies (such as attention mechanism, graph neural network, etc.), fuse features of different dimensions to generate fused feature vectors;

[0023] S3, dynamically adjust the network structure and parameters according to the historical data and real-time data of news hotspots through adaptive deep learning models;

[0024] S4, inputting the fused feature vector and the corresponding news hotspot label into the adaptive deep learning model, and using the cross entropy loss function and the reinforcement learning reward function to jointly guide the model training;

[0025] S5. Use the trained adaptive deep learning model to predict real-time news hotspot data, and intelligently adjust the news scheduling strategy based on the prediction results and the dynamic changes of news hotspots. The scheduling strategy includes push order, push frequency and push channel; at the same time, introduce a user feedback mechanism to evaluate the scheduling effect based on indicators such as user click rate, number of comments, and number of shares.

[0026] The working principle of the above technical solution is as follows: using PAI interface, web crawlers and other means to collect news data from data sources such as news websites, social media platforms and user behavior logs; the collected data includes historical news data and real-time news data, covering multi-dimensional information such as text, pictures, videos and user comments; pre-processing the acquired news data, such as data cleaning, denoising, formatting, etc., to ensure the quality and consistency of the data; using natural language processing technology (such as BERT, GPT, etc.) to extract text features, these technologies can deeply understand the semantics of text; using deep learning models (such as ResNet, VGG, etc.) to extract image and video features, these models are good at capturing key information in images and videos; combining user behavior log analysis to extract user interest features and understand user preferences and needs; through feature fusion strategies (such as attention mechanism, graph neural network, etc.), features of different dimensions are fused to generate a fused feature vector. This vector integrates information from multiple aspects such as news content and user interests; an adaptive deep learning model is constructed, which includes an input layer, a feature extraction layer, an adaptive learning layer, and an output layer; in the feature extraction layer, deep learning models such as convolutional neural networks (CNN) and recurrent neural networks (RNN) are used to further extract the deep features of the fused feature vector; the adaptive learning layer introduces reinforcement learning algorithms (such as Q-learning, DeepQ-Network, etc.) so that the model can self-optimize according to the prediction results and scheduling effects, and dynamically adjust the network structure and parameters; the fused feature vector and the corresponding news hotspot label are input into the adaptive deep learning model; the cross entropy loss function and the reinforcement learning reward function are used together to Guide model training to ensure that the model can optimize the scheduling strategy while accurately predicting news hotspots; during the training process, adopt dynamic learning rate adjustment strategy, gradient clipping technology and early stopping method and other techniques to prevent model overfitting and improve training efficiency; through continuous iterative training until the model converges, obtain the optimal model parameters; use the trained adaptive deep learning model to predict real-time news hotspot data; according to the prediction results and the dynamic changes of news hotspots, intelligently adjust the news scheduling strategy, including push order, push frequency and push channel, etc.; introduce user feedback mechanism to evaluate the scheduling effect according to indicators such as user click-through rate, number of comments, and number of shares; further optimize the scheduling strategy according to the evaluation results to achieve accurate push and efficient scheduling of news hotspots.

[0027] The effects of the above technical solutions are as follows: through PAI interface, web crawlers and other means, news data is obtained from multiple data sources such as news websites, social media platforms and user behavior logs, ensuring the diversity and comprehensiveness of the data; pre-processing the obtained news data improves the quality and availability of the data, laying a solid foundation for subsequent feature extraction and model training; using natural language processing technology, deep learning models and other technical means to extract features of multi-dimensional data such as text, pictures, videos and user comments, fully capturing the rich information of news content; through feature fusion strategies, features of different dimensions are fused to generate fused feature vectors, improving the accuracy and completeness of feature expression, and providing strong support for model prediction; the adaptive deep learning model can dynamically adjust the network structure and parameters according to the historical data and real-time data of news hotspots, improving The model is adaptable and flexible; a reinforcement learning algorithm is introduced to enable the model to self-optimize according to the prediction results and scheduling effects, and continuously improve the prediction accuracy and scheduling efficiency; in the model training process, dynamic learning rate adjustment strategy, gradient clipping technology and early stopping method are used to effectively prevent model overfitting and improve training efficiency and model performance; through continuous iterative training until the model converges, the optimal model parameters are obtained to ensure the stability and accuracy of the model; the trained adaptive deep learning model is used to predict real-time news hotspot data, and the news scheduling strategy is intelligently adjusted according to the prediction results and the dynamic changes of news hotspots, which improves the accuracy and timeliness of news push; a user feedback mechanism is introduced to evaluate the scheduling effect according to indicators such as user click-through rate, number of comments, and number of shares, and the scheduling strategy is further optimized to achieve accurate matching of user needs and personalized push.

[0028] In one embodiment of the present invention, the S1 includes:

[0029] S11. Use the platform's news API to regularly pull historical news data from authoritative news sources, including news titles, text, release time, sources, etc., and set request strategies based on the API access frequency limit;

[0030] S12. For different social media platforms and news websites, crawl data in real time through web crawlers, and dynamically adjust crawler crawling tasks and priorities based on the crawler scheduling system according to the frequency and demand of data updates;

[0031] S13. Collect the user's browsing history, click behavior, comment content, sharing behavior and other log data from the application, website and other corresponding channels with the user's authorization; and perform desensitization processing on the collected log data;

[0032] S14. Use hash algorithms to deduplicate news data, and filter out useless data such as junk information and advertising information based on preset keyword blacklists and regular expressions; complete missing key fields such as news titles, texts, and release times through other sources or historical data; and mark data that cannot be completed;

[0033] S15. Use natural language processing tools to perform text processing such as word segmentation, stop word removal, stem extraction, and part-of-speech tagging on the text data; and perform spelling check on the text data to correct possible spelling errors;

[0034] S16. Perform image processing such as resizing, format conversion, and quality compression on the image and video data, unify the input standard, and perform image enhancement processing on the image data, such as rotation, scaling, and cropping.

[0035] The working principle of the above technical solution is: through the platform's news API interface, according to the set request strategy (considering the API access frequency limit, for example, no more than ten times a minute can be accessed), historical news data is regularly pulled from authoritative news sources. These data usually include key information such as news title, text, release time and source; for specific social media platforms and news websites, web crawlers are deployed to crawl data in real time. The crawler scheduling system dynamically adjusts the crawling tasks and priorities of the crawler according to the frequency and needs of data updates to ensure the timeliness and accuracy of the data. It is assumed here that all platforms and websites provide crawler permissions; through user authorization, log data such as user browsing history, click behavior, comment content and sharing behavior are collected from applications, websites and other channels. In order to protect user privacy, the collected log data is desensitized to remove or blur personal sensitive information; the news data is deduplicated using a hash algorithm to ensure the uniqueness of the data. At the same time, useless data such as spam and advertising information are filtered out according to the preset keyword blacklist and regular expressions. For missing key fields (such as news titles, text, and release time), try to complete them through other sources or historical data; mark the data that cannot be completed for identification during subsequent analysis; use natural language processing tools to segment text data, remove stop words, extract stems, and tag parts of speech to improve the accuracy of text analysis. At the same time, perform spelling checks to correct possible spelling errors and ensure the accuracy of text data; resize, convert formats, and compress quality for images and video data to unify input standards. Perform image enhancement processing (such as rotation, scaling, cropping, etc.) on image data to improve the accuracy and robustness of image recognition. These preprocessing steps help with subsequent feature extraction and model training.

[0036] The effects of the above technical solutions are as follows: through the platform's news API interface, historical news data from authoritative news sources can be regularly pulled, covering multiple dimensions such as news title, text, release time, source, etc., ensuring the authority and comprehensiveness of the data; for different social media platforms and news websites, real-time data capture is carried out using web crawlers, further broadening the data sources and making the collected news data more abundant and diverse; the combined use of API interfaces and web crawlers can realize real-time collection and updating of news data, ensuring the timeliness and accuracy of the data; setting request strategies according to the access frequency limit of the API, and dynamically adjusting the capture tasks and priorities based on the crawler scheduling system, further improving the efficiency and real-time performance of data collection; desensitizing the collected log data to effectively protect user privacy and data security; using hash algorithms to deduplicate news data to avoid data duplication and redundancy; Regular expressions filter junk information and advertising information, improving the purity and availability of data; missing key fields such as news titles, texts, and release times are completed to ensure the integrity and accuracy of the data; natural language processing tools are used to perform word segmentation, stop word removal, stem extraction, and part-of-speech tagging on text data, improving the readability and analysis efficiency of text data; image and video data are resized, format converted, and quality compressed to unify input standards for subsequent analysis and processing; possible spelling errors are corrected through spelling checking to further improve the accuracy and readability of text data; image data is enhanced by image enhancement processing such as rotation, scaling, and cropping to improve the quality and availability of image data; data that cannot be completed is marked to facilitate identification and processing during subsequent analysis; data that has been standardized and accurately processed provides a reliable foundation for subsequent news hotspot predictions, trend analysis, and user behavior analysis.

[0037] In one embodiment of the present invention, the S2 includes:

[0038] S21. Use the pre-trained BERT or GPT model to understand the semantics of news texts and extract advanced features such as keywords, topics, and sentiment. Fine-tune the model based on the characteristics of the news texts.

[0039] S22. Based on word embedding technology, the semantic similarity between news texts is calculated to identify related news clusters; and clustering algorithms such as K-means and DBSCAN are used to perform cluster analysis on news texts to further explore the relevance between news; the semantic similarity is calculated using the following formula:

[0040]

[0041] in, and Respectively represent text and The word embedding vector of Represents the dimension of the word embedding vector; and They are respectively and The vector is The value of the dimension; the numerator represents the dot product of two vectors; the denominator represents the product of the Euclidean norms of the two vectors;

[0042] S23. Use deep learning models such as ResNet or VGG to perform object recognition and scene classification on news pictures, extract image features, and perform dimensionality reduction on image features;

[0043] S24, extracting key information of the video content through video frame extraction and key frame recognition technology, and extracting and analyzing features of the key frames in combination with image feature extraction technology;

[0044] S25. Conduct statistical analysis on user historical behavior data to build user interest profiles, including preferred topics, active time periods, reading habits, etc.; and use recommendation algorithms such as collaborative filtering and matrix decomposition to explore user potential interests;

[0045] S26. Analyze the interaction of users in social networks and evaluate the influence of users in information dissemination; use social network analysis algorithms to identify key nodes and community structures;

[0046] S27. Based on the attention mechanism, dynamically adjust the contribution of different features to the prediction results, and capture the complex relationship between different features through the multi-head attention mechanism;

[0047] S28. Build a news-user relationship graph, use graph neural network to capture the correlation between news and the mutual influence between users, and combine with graph embedding technology to represent news and users as low-dimensional vectors.

[0048] The working principle of the above technical solution is as follows: Use pre-trained BERT or GPT models to perform deep semantic understanding of news texts. These models are trained with a large amount of text data and can accurately capture the semantic information of the text. By fine-tuning the model parameters to make it more suitable for the characteristics of news texts, high-level features such as keywords, topics, and emotional tendencies can be extracted. These features help understand the core content and emotional color of the news; based on word embedding technology, the news text is converted into points in a high-dimensional vector space, and the semantic similarity between news texts is evaluated by calculating metrics such as cosine similarity or Euclidean distance between vectors. Then, clustering algorithms such as K-means and DBSCAN are used to perform cluster analysis on these vectors to form related news clusters. This helps to explore the correlation between news and discover news hotspots and trends; deep learning models such as ResNet or VGG are applied to perform object recognition and scene classification on news pictures. These models can automatically learn features in pictures and extract useful image features. In order to reduce computational complexity and improve processing efficiency, the extracted image features are processed by dimensionality reduction, such as using PCA (principal component analysis) or t-SNE and other dimensionality reduction algorithms; key frames are extracted from the video content through video frame extraction and key frame recognition technology. Key frames are frames in the video that can represent its main content. Then, combined with image feature extraction technology, key frames are extracted and analyzed, such as color histogram, edge detection, texture analysis, etc. This helps to understand the main features and emotional color of the video content; statistical analysis of user historical behavior data, including browsing history, click behavior, comment content, etc., to build user interest portraits. The portrait includes the user's preferred topics, active time periods, reading habits, etc. Then, collaborative filtering, matrix decomposition and other recommendation algorithms are used to explore the user's potential interests, that is, news content that the user may be interested in but has not yet been exposed to; the interaction of users on social networks is analyzed, including likes, comments, forwarding and other behaviors, to evaluate the user's influence in information dissemination. Social network analysis algorithms, such as PageRank and HITS, are used to identify key nodes and community structures. These key nodes and community structures play an important role in news dissemination and provide social influence support for news scheduling. Based on the attention mechanism, the contribution of different features to the prediction results is dynamically adjusted. The attention mechanism can give higher weights to important features and reduce the weights of secondary features, thereby improving the accuracy of predictions. At the same time, through the multi-head attention mechanism, the complex relationships between different features are captured, such as the association between text and images, the match between user interests and news content, etc.; a news-user relationship graph is constructed, with news and users as nodes in the graph, and the association between them (such as a user reading a piece of news) as edges in the graph. Then, graph neural networks are used to capture the correlation between news and the mutual influence between users. Graph neural networks can learn complex structural information in graphs and represent them as low-dimensional vectors.These vectors can be used for subsequent tasks such as news recommendation and user profile update.

[0049] The effects of the above technical solutions are as follows: using the pre-trained BERT or GPT model to perform deep semantic understanding of news texts, it is possible to accurately extract advanced features such as keywords, themes, and sentiment tendencies. These features provide a rich information basis for subsequent news classification, recommendation, and other tasks; fine-tuning the model according to the characteristics of the news text makes the model more suitable for specific tasks in the news field, and improves the accuracy and pertinence of feature extraction; calculating the semantic similarity between news texts based on word embedding technology helps to identify related news clusters, and provides strong support for news aggregation, topic tracking, etc.; clustering analysis of news texts using clustering algorithms such as K-means and DBSCAN can further explore the correlation between news and discover news hotspots and trends; applying deep learning models such as ResNet or VGG to perform object recognition and scene classification on news images, extract image features, and reduce computational complexity through dimensionality reduction processing, thereby improving the efficiency of image processing; extracting key information of video content through video frame extraction and key frame recognition technology, and combining image feature extraction technology to conduct in-depth analysis of key frames, providing an important basis for the recommendation and understanding of video news; and performing clustering analysis on user historical behavior data. Statistical analysis is used to construct user interest portraits, including preferred topics, active time periods, reading habits, etc., which provides a basis for personalized news recommendations; recommendation algorithms such as collaborative filtering and matrix decomposition are used to explore user potential interests, improving the diversity and accuracy of news recommendations; the interaction of users' social networks is analyzed to evaluate the influence of users in information dissemination, and social network analysis algorithms are used to identify key nodes and community structures, providing strong support for news scheduling and dissemination strategies; based on the attention mechanism, the contribution of different features to the prediction results is dynamically adjusted, and the complex relationship between different features is captured through the multi-head attention mechanism, which improves the depth and breadth of the model's understanding of news content; a news-user relationship graph is constructed and a graph neural network is used to capture the correlation between news and the mutual influence between users, providing more detailed and comprehensive information for news recommendation, user portrait update, etc.; combined with graph embedding technology, news and users are represented as low-dimensional vectors, which is convenient for subsequent similarity calculation, cluster analysis and other tasks, improving processing efficiency and accuracy. The above semantic similarity calculation formula uses the product form of dot product and Euclidean norm, which is simple and fast to calculate. This calculation method is not only easy to implement, but also has high computational efficiency; the semantic similarity results calculated by this formula are accurate and reliable. It can reflect the true semantic relationship between texts and provide strong support for tasks such as clustering and classification of news texts.

[0050] In one embodiment of the present invention, the S26 includes:

[0051] Collect and clean user social network data, including user follow-up relationships, forwarding, commenting, liking and other interactive behavior data, and extract key features in user social networks, such as number of followers, interaction frequency, interactive content, etc.; build a user influence evaluation model based on user social network characteristics and interactive behavior data; train and verify the model, and at the same time, fine-tune and optimize the model according to actual application scenarios and needs;

[0052] Use social network analysis algorithms to identify key nodes and community structures in social networks, and further analyze the identified key nodes, including their scope of influence, dissemination efficiency, and duration of influence. At the same time, analyze the characteristics of community structures and the interactive relationships between members to provide targeted strategic support for news scheduling.

[0053] Use visualization techniques (such as network diagrams and heat maps) to intuitively display the influence distribution, key nodes and community structure in the user's social network, and generate detailed reports based on the analysis results, including user influence rankings, specific information on key nodes and community structure, and strategic recommendations for news scheduling;

[0054] Through the dynamic monitoring mechanism of social influence, the interactive behavior and influence changes in users' social networks are tracked in real time, and the influence evaluation model and key node identification algorithm are regularly updated and optimized based on the monitoring results.

[0055] The working principle of the above technical solution is as follows: First, collect the user's social network data, which includes the user's follow-up relationship, forwarding, commenting, liking and other interactive behavior data. Then, clean the data to remove invalid, redundant or abnormal data to ensure the accuracy and reliability of the data, and provide a high-quality data foundation for subsequent feature extraction and model construction; extract the key features of the user's social network from the cleaned data, such as the number of followers, interaction frequency, and interactive content. Based on these features, build a user influence evaluation model, and through feature extraction, transform the complex information of the user's social network into quantifiable indicators, and then evaluate the user's influence through the model; train and verify the constructed user influence evaluation model to ensure the accuracy and stability of the model. At the same time, according to the actual application scenarios and needs, the model is fine-tuned and optimized. Through training and verification, the generalization ability and adaptability of the model are improved, so that it can better evaluate the influence of users; social network analysis algorithms (such as K-means, DBSCAN and other clustering algorithms, or Girvan-Newman algorithm and Louvain algorithm based on community detection, etc.) are used to identify key nodes and community structures in social networks. By identifying key nodes and community structures, the core strength and group characteristics in the user's social network are revealed, providing targeted strategic support for news scheduling; the identified key nodes are further analyzed, including their influence range, dissemination efficiency, influence duration, etc. At the same time, the characteristics of the community structure and the interactive relationship between members are analyzed, and more specific and detailed strategic suggestions are provided for news scheduling through in-depth analysis of key nodes and community structures; visualization technology (such as network diagrams, heat maps, etc.) is used to intuitively display the influence distribution, key nodes and community structure in the user's social network. A detailed report is generated based on the analysis results, including user influence rankings, specific information on key nodes and community structures, and strategic recommendations for news scheduling. Through visual display and report generation, the analysis results are more intuitive and easy to understand, making it easier for decision makers to quickly understand and apply them. Through the dynamic monitoring mechanism of social influence, the interactive behavior and influence changes of users in social networks are tracked in real time. The influence assessment model and key node identification algorithm are regularly updated and optimized based on the monitoring results. Through dynamic monitoring and optimization, the accuracy and timeliness of the assessment model and identification algorithm are ensured so that they can adapt to the ever-changing social network environment.

[0056] The effects of the above technical solutions are as follows: by comprehensively collecting and cleaning users' social network data, including attention relationships, interactive behaviors, etc., the accuracy and completeness of the data are ensured, providing a solid foundation for building a user influence evaluation model; extracting key features from the data and building a user influence evaluation model based on these features can more accurately reflect the actual influence of users in social networks; rigorously training and verifying the model, while fine-tuning and optimizing it according to actual application scenarios and needs, further improving the accuracy and adaptability of the model; using advanced social network analysis algorithms, such as K-means, DBSCAN, Girvan-Newman, Louvain, etc., it is possible to accurately identify key nodes and community structures in social networks, providing strong support for news scheduling; in-depth analysis of the identified key nodes, including the scope of influence, dissemination efficiency, duration of influence, etc., is helpful Understand the importance and role of these nodes in social networks; analyze the characteristics of community structure and the interactive relationship between members, which will help reveal the internal structure and dynamic changes of the community and provide more refined strategic recommendations for news scheduling; use visualization techniques such as network diagrams and heat maps to intuitively display the influence distribution, key nodes and community structure in the user's social network, making the analysis results easier to understand and apply; generate detailed reports based on the analysis results, including user influence rankings, specific information on key nodes and community structure, and strategic recommendations for news scheduling, providing a comprehensive reference for decision makers; through the social influence dynamic monitoring mechanism, track the interactive behavior and influence changes in the user's social network in real time, ensuring continuous attention and updating of user influence; regularly update and optimize the influence evaluation model and key node identification algorithm based on the monitoring results to ensure the timeliness and accuracy of these models and algorithms.

[0057] In one embodiment of the present invention, the S27 includes:

[0058] The attention weight of each feature for the prediction result is calculated by multiplying the feature vector with the attention matrix and normalizing it through the softmax function; the attention weight is calculated by the following formula:

[0059]

[0060] Where X represents the feature vector (dimension is × , is the number of features, is the dimension of the feature vector); and Indicates The query and key weight matrices corresponding to each head, represents the dimension of the key vector (usually );

[0061] According to the calculated attention weight, each feature vector is weighted, where features with larger weights will occupy a larger proportion in the subsequent prediction process. The multi-head attention mechanism is introduced to divide the feature vector into multiple heads (i.e., multiple sub-vectors), and the attention weight is calculated for each head separately; the attention weight of each head is calculated by the following formula:

[0062]

[0063] in, Indicates The value weight matrix corresponding to each head;

[0064] Multiply the attention weight of each head by the corresponding feature sub-vector and perform weighted summation to obtain a feature representation that integrates multi-head attention.

[0065] The feature representation that integrates multi-head attention is integrated with other features; according to the integrated feature representation, the contribution of different features to the prediction results is dynamically adjusted;

[0066] Visualize the attention weights by generating heat maps or attention weight distribution maps; combine the visualization results to explain and analyze the model's predictive behavior.

[0067] The working principle of the above technical solution is as follows: First, the feature vector is multiplied by the predefined attention matrix. This step is to calculate the original attention score of each feature for the prediction result; then, the original attention score is normalized by the softmax function to obtain the attention weight of each feature. The softmax function ensures that the sum of the attention weights of all features is 1, so that each weight represents a relative importance; each feature vector is weighted according to the calculated attention weight. This means that features with larger weights will occupy a larger proportion in the subsequent prediction process, thus having a greater impact on the prediction results; in order to capture the complex relationship between features, a multi-head attention mechanism is introduced. This mechanism divides the feature vector into multiple heads (i.e., multiple sub-vectors) and calculates the attention weight for each head separately. Each head can focus on different aspects of the feature, thereby providing more comprehensive information; the attention weight of each head is multiplied by the corresponding feature sub-vector and weighted summed. This step obtains a feature representation that combines the attention information from different heads; the feature representation that combines the multi-head attention is fused with other features (such as text features, image features, user interest features, etc.). This step aims to combine information from different sources to enhance the model's predictive ability; dynamically adjust the contribution of different features to the prediction results based on the fused feature representation. This means that the model can flexibly adjust the weights of each feature according to the actual situation to optimize the prediction results; visualize the attention weights by generating heat maps or attention weight distribution maps. This step helps to intuitively understand the model's attention to different features during the prediction process; combined with the visualization results, the model's predictive behavior is explained and analyzed. This helps to understand the model's decision-making process, find potential improvements, and further optimize model performance.

[0068] The effect of the above technical solution is: by calculating the attention weight of each feature for the prediction result and weighting the feature vector, the features with larger weights occupy a larger proportion in the prediction process. This refined weighting method helps the model pay more attention to those features that have an important impact on the prediction results, thereby improving the accuracy of the prediction; introducing a multi-head attention mechanism, splitting the feature vector into multiple heads (i.e., multiple sub-vectors), and calculating the attention weight for each head separately. This method can capture the complex relationship between features and pay attention to features from multiple angles, further enhancing the prediction ability of the model; by generating a heat map or an attention weight distribution map, the attention weight is visualized. This allows us to intuitively see the degree of attention of the model to different features during the prediction process, thereby enhancing the interpretability of the model; combined with the visualization results, the prediction behavior of the model is explained and analyzed. This helps us understand the decision-making process of the model, find potential improvement points, and further optimize the model performance. At the same time, this also provides valuable reference information for researchers in related fields; according to the fused feature representation, the contribution of different features to the prediction results is dynamically adjusted. This dynamic adjustment method enables the model to flexibly adjust its prediction strategy according to different input data, thereby improving the flexibility and adaptability of the model; the feature representation that integrates multi-head attention is integrated with other features (such as text features, image features, user interest features, etc.). This multi-feature fusion method can make full use of information from different sources, further enhance the prediction ability of the model, and enable it to cope with more complex prediction tasks; the introduced multi-head attention mechanism provides new ideas and methods for the study of attention mechanism. This helps to promote research and development in related fields and provide strong support for future technological innovation; this technical solution has significant advantages in improving model prediction accuracy, enhancing interpretability, and improving flexibility and adaptability. Therefore, it has broad application prospects in many fields such as natural language processing, image recognition, and recommendation systems. This helps to expand the application scenarios in these fields and promote the further development of related technologies. The above formula multiplies the feature vector with the attention matrix, calculates the attention weight by normalization through the softmax function, and introduces the multi-head attention mechanism to weight the feature vector. The model can more effectively utilize the information in the input data, improve the prediction performance, and enhance the flexibility and robustness of the model.

[0069] In one embodiment of the present invention, S3 includes:

[0070] S31, receiving the fused feature vector as input, including text features, multimedia features, user interest features, etc., combining deep learning models such as CNN and RNN, to perform deep feature extraction on the fused feature vector; through multi-scale convolution kernels, capturing feature information of different scales;

[0071] S32. Use recurrent neural network models such as LSTM and GRU to capture time series dependencies, design state space, action space and reward function, and build a reinforcement learning framework;

[0072] S33. Output the predicted probability or classification result of news hotspots, and perform fine-grained classification of news hotspots through multiple classifiers; dynamically adjust network structure parameters such as convolution kernel size and number of RNN layers according to the type and scale of news data.

[0073] The working principle of the above technical solution is as follows: receiving fused feature vectors, which contain various types of data such as text features, multimedia features, and user interest features; combining deep learning models such as CNN (convolutional neural network) and RNN (recurrent neural network) to perform deep feature extraction on the fused feature vectors; capturing feature information of different scales through multi-scale convolution kernels. Multi-scale convolution kernels can process features of different sizes and shapes, thereby extracting feature information more comprehensively; extracting richer and deeper feature information, providing a basis for subsequent processing and analysis; using recurrent neural network models such as LSTM (long short-term memory network) and GRU (gated recurrent unit) to capture time series dependencies. These models can handle time dependency and long-term dependency in sequence data, so as to more accurately understand the dynamic changes of data; introducing reinforcement learning algorithms such as Q-learning and DeepQ-Network, so that the model can self-optimize according to the prediction results and scheduling effects. By designing the state space, action space and reward function, a reinforcement learning framework is constructed to enable the model to continuously learn and improve; the model can more accurately capture time series dependencies, and self-adjust and optimize according to actual results to improve the accuracy of prediction and scheduling; output the predicted probability or classification results of news hotspots; and use multiple classifiers to perform fine-grained classification of news hotspots. Multiple classifiers can process multiple types of news data and classify them according to their characteristics; dynamically adjust network structure parameters such as convolution kernel size and number of RNN layers according to the type and scale of news data. This dynamic adjustment enables the model to better adapt to news data of different types and scales, improve the accuracy and efficiency of prediction, and obtain the predicted probability or classification results of news hotspots, which can be used in multiple fields such as news recommendation.

[0074] The effect of the above technical solution is: by combining deep learning models such as CNN and RNN, S31 can perform deep feature extraction on the fused feature vector. This multi-level and multi-dimensional feature extraction method helps the model to understand the intrinsic characteristics of news data more comprehensively and deeply, thereby improving the accuracy of prediction; the introduction of multi-scale convolution kernels enables the model to capture feature information of different scales. This helps the model to more accurately identify key information when processing complex and changeable news data, and further improves the accuracy of prediction; using recurrent neural network models such as LSTM and GRU, it is possible to capture the time series dependencies in news data. This is of great significance for understanding the evolution trend of news hotspots and predicting future hotspots; the introduction of reinforcement learning algorithms enables the model to self-optimize according to prediction results and scheduling effects. This self-adjustment ability helps the model maintain a high prediction accuracy in a constantly changing news environment; dynamically adjust network structure parameters such as convolution kernel size and number of RNN layers according to the type and scale of news data. This dynamic adjustment capability enables the model to better adapt to different types of news data and improve the generalization ability of the model. At the same time, dynamic adjustment of network structure parameters also helps the model maintain high operating efficiency when processing large-scale news data and improves the practicality of the model. Fine-grained classification of news hotspots through multiple classifiers can provide more accurate and specific classification results. This helps users understand the types and properties of news hotspots more accurately and improves the user experience. By designing the state space, action space and reward function, a reinforcement learning framework is constructed. This enables the model to clearly display the decision-making process and learning effect during the self-optimization process. The construction of the reinforcement learning framework helps users better understand the model's prediction behavior and decision-making basis, and improves the model's interpretability and transparency. Although the technical solution does not directly mention the visualization of the feature extraction and classification process, combined with the characteristics of the deep learning model and actual application needs, the feature extraction and classification process can be displayed and analyzed through visualization tools. This visual analysis helps users understand the model's workflow and output results more intuitively, further improving the model's interpretability and transparency.

[0075] In one embodiment of the present invention, the S4 includes:

[0076] Preprocess the received feature vectors such as text features, multimedia features, user interest features, etc., and fuse the preprocessed feature vectors to form a unified feature representation;

[0077] Construct a CNN model and design a network structure, which includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer; wherein the convolutional layer is used to extract local features in the feature vector;

[0078] Introduce multi-scale convolution kernels in the convolution layer, and capture feature information of different scales by setting convolution kernels of different sizes; downsample the feature map output by the convolution layer through the pooling layer;

[0079] The feature map output by the pooling layer is input into the fully connected layer for further feature extraction and classification; the fully connected layer maps the feature map to the classification space and outputs the prediction result;

[0080] Construct an RNN model and design a network structure, which includes an input layer, a hidden layer, and an output layer; wherein the hidden layer will be used to capture the time series dependency in the feature vector;

[0081] LSTM (Long Short-Term Memory Network) is introduced into the RNN model, and the feature vector extracted by the CNN model is used as the input of the RNN model. The time series dependency in the feature vector is captured through the hidden layer;

[0082] The feature results extracted by the CNN model and the RNN model are fused to form the final deep-level feature representation; the fused feature representation is input into the classifier for classification or regression prediction, and the prediction probability or classification result of the news hotspot is output;

[0083] At the same time, according to the type and scale of news data, the network structure parameters such as convolution kernel size and number of RNN layers are dynamically adjusted.

[0084] The working principle of the above technical solution is as follows: the received feature vectors such as text features, multimedia features, user interest features, etc. are first preprocessed. The preprocessing step may include standardization, denoising, normalization and other operations to ensure the consistency and effectiveness of the feature vectors in subsequent processing; the preprocessed feature vectors are fused to form a unified feature representation. This step aims to integrate feature vectors from different sources into a whole so that the subsequent model can fully and accurately capture feature information; build a CNN model and design the network structure. The network structure includes input layer, convolution layer, pooling layer and fully connected layer.

[0085] Input layer: receives the fused feature vector as input.

[0086] Convolutional layer: used to extract local features from feature vectors. By introducing multi-scale convolution kernels, convolution kernels of different sizes can be set to capture feature information of different scales. This multi-scale feature extraction method helps the model to more fully understand the internal structure of feature vectors.

[0087] Pooling layer: downsamples the feature map output by the convolutional layer. The pooling layer can retain the key information in the feature map while reducing the dimension and amount of calculation of the feature map, thus improving the robustness and operation efficiency of the model.

[0088] Fully connected layer: maps the feature map output by the pooling layer to the classification space and outputs the prediction result. The fully connected layer converts the feature map into a specific prediction result through further feature extraction and classification operations.

[0089] Build an RNN model and design the network structure, which includes input layer, hidden layer and output layer.

[0090] Input layer: receives the feature vector extracted by the CNN model as input.

[0091] Hidden layer: used to capture time series dependencies in feature vectors. By introducing LSTM (Long Short-Term Memory Network), the gradient vanishing and gradient exploding problems of traditional RNN models when capturing long-distance dependencies can be solved. The introduction of LSTM helps improve the model's ability to process time series data.

[0092] Output layer: Outputs the prediction results or feature representations after further processing.

[0093] The feature results extracted by the CNN model and the RNN model are fused to form the final deep feature representation. This step aims to integrate the feature representations from different models into a whole so that the subsequent classifier can perform classification or regression prediction more accurately; the fused feature representation is input into the classifier for classification or regression prediction. The classifier outputs the predicted probability or classification result of the news hotspot based on the feature representation; according to the type and scale of the news data, the network structure parameters such as the convolution kernel size and the number of RNN layers are dynamically adjusted. This step aims to enable the model to better adapt to different types of news data and improve the generalization ability and prediction accuracy of the model.

[0094] The effect of the above technical solution is as follows: by introducing multi-scale convolution kernels in the convolution layer, the technical solution can capture feature information of different scales. This helps the model to understand the intrinsic characteristics of news data more comprehensively, thereby improving the accuracy of prediction; the RNN model, especially after the introduction of the LSTM network, can effectively capture the time series dependency in the feature vector. This is of great significance for understanding the evolution trend of news hot spots and predicting future hot spots. The LSTM network solves the gradient vanishing and gradient explosion problems of the traditional RNN model when capturing long-distance dependencies, and further improves the prediction ability of the model; the feature results extracted by the CNN model and the RNN model are fused to form the final deep feature representation. This fusion method combines the advantages of the two models and can more comprehensively capture the feature information of news data, thereby improving the accuracy and robustness of prediction; the technical solution can dynamically adjust the network structure parameters such as the convolution kernel size and the number of RNN layers according to the type and scale of news data. This dynamic adjustment capability enables the model to better adapt to different types of news data and improve the generalization ability and adaptability of the model; due to the diversity and complexity of news data, traditional fixed model structures are often difficult to cope with. This technical solution can flexibly cope with various complex scenarios and improve the practicality and application value of the model by dynamically adjusting the network structure parameters and fusing multiple feature vectors; by reasonably designing the network structure of CNN and RNN models, this technical solution can improve the computational efficiency while ensuring the prediction accuracy. For example, by introducing a pooling layer to downsample the feature map output by the convolution layer, the dimension and amount of computation of the feature map can be reduced, thereby improving the operation efficiency of the model; this technical solution can reduce redundant computation and improve the utilization of computing resources by fusing multiple feature vectors and dynamically adjusting the network structure parameters. This helps to reduce the operation cost of the model and improve the computational efficiency; although the visualization of the feature extraction process is not directly mentioned in the technical solution, the feature extraction process can be displayed and analyzed through visualization tools in combination with the characteristics of the deep learning model and the actual application requirements. This helps users better understand the prediction behavior and decision-making basis of the model and improves the interpretability of the model; this technical solution enables the model to self-optimize according to the prediction results and scheduling effects by introducing a reinforcement learning algorithm (although it is not directly mentioned in the description, similar ideas can be applied to model optimization) or other optimization strategies. This self-tuning process is transparent and helps users understand the model's optimization process and performance improvement.

[0095] In one embodiment of the present invention, the S4 includes:

[0096] S41, distinguish the difference between the model prediction probability and the true label, and guide the model to learn in the right direction based on the discrimination result; and use weighted cross entropy loss to perform imbalanced processing on news hotspots of different categories;

[0097] S42. Design a reward function based on the effects of news scheduling, such as increased click-through rate and user satisfaction, and encourage continuous optimization of the model during the prediction and scheduling process based on a multi-stage reward mechanism;

[0098] S43, adopt the learning rate decay strategy to balance the convergence speed and stability of the model, and based on the learning rate warm-up mechanism, gradually increase the learning rate in the early stage of training to prevent the model from falling into the local optimal solution;

[0099] S44. Limit the gradient norm to prevent gradient explosion, and standardize the gradient through the gradient normalization algorithm; monitor the performance of the validation set and terminate the training early when the performance no longer improves.

[0100] The working principle of the above technical solution is as follows: First, the difference between the model prediction probability and the true label is discriminated. This is the basis of model learning. By comparing the prediction results with the actual situation, the loss value can be calculated to guide the model to adjust; for the problem of imbalanced news hotspot categories, a weighted cross entropy loss function is used. By assigning different weights to news hotspots of different categories, the model's attention to different categories can be balanced, thereby improving the model's prediction accuracy on a few categories; according to the news scheduling effect, such as increased click-through rate and improved user satisfaction, a reward function is designed. These indicators can directly reflect the performance of the model in practical applications, and therefore are an important basis for optimizing the model; a multi-stage reward mechanism is introduced to encourage the model to continuously optimize during the prediction and scheduling process. This means that the model should not only show good prediction ability in the initial stage, but also continue to improve in subsequent stages to adapt to the changing news environment and user needs; a learning rate decay strategy is adopted to balance the convergence speed and stability of the model. In the early stage of training, a larger learning rate helps the model converge quickly; as the training progresses, the learning rate is gradually reduced, which helps the model to make subtle adjustments to avoid overfitting; in the early stage of training, the learning rate is gradually increased. This step helps the model explore a wider parameter space in the initial stage and avoid falling into the local optimal solution; limit the gradient norm to prevent gradient explosion. Gradient explosion is one of the common problems in deep learning training, which will cause the model parameters to be updated too much, thus affecting the stability of the model. By limiting the gradient norm, the stability of the model during training can be ensured; the gradient normalization algorithm is used to standardize the gradient. This step helps to speed up the convergence of the model and improve the generalization ability of the model; during the training process, the performance of the validation set is continuously monitored. When the performance of the validation set no longer improves, the training is terminated early. This step helps to avoid overfitting and ensure the performance of the model in practical applications.

[0101] The effects of the above technical solutions are as follows: by discriminating the difference between the model prediction probability and the true label, and guiding the model learning based on this discriminant result, it can ensure that the model continuously adjusts parameters during the training process to get closer to the prediction results of the true label; using the weighted cross entropy loss function to perform imbalanced processing for news hotspots of different categories, it helps the model pay more attention to a few categories, thereby improving the accuracy of the overall prediction; designing a reward function based on the news scheduling effect, and encouraging the model to continuously optimize during the prediction and scheduling process through a multi-stage reward mechanism. This mechanism can stimulate the potential of the model, so that it maintains optimization momentum in multiple stages, thereby continuously improving performance; using a learning rate decay strategy to balance the convergence speed and stability of the model, and using a learning rate warm-up mechanism to gradually increase the learning rate in the early stage of training, it helps the model to converge quickly in the exploration stage, and stably optimize in subsequent stages to avoid falling into a local optimal solution; limiting the gradient norm to prevent gradient explosion and maintain the stability of the training process. At the same time, standardizing the gradient through the gradient normalization algorithm helps the model to update parameters more stably during training; monitoring the performance of the validation set and terminating the training early when the performance no longer improves can avoid overfitting of the model in the later stages of training, thereby ensuring the generalization ability of the model; this technical solution comprehensively considers multiple aspects such as model prediction accuracy, optimization ability, training stability and generalization ability, so that the model can better adapt to complex news hotspot prediction scenarios; measures such as weighted cross entropy loss to deal with category imbalance problems, learning rate adjustment strategy to avoid local optimal solutions, and gradient norm restriction to prevent gradient explosion are jointly enhanced to enhance the robustness of the model, so that it can maintain stable prediction performance when facing different news hotspots and data distributions.

[0102] In one embodiment of the present invention, S5 includes:

[0103] S51. Process and analyze real-time news data through stream processing technologies, such as Apache Flink and Storm, and generate fused feature vectors based on real-time feature extraction and fusion algorithms;

[0104] S52, applying the prediction results to the news push strategy, intelligently adjusting the push order, frequency and channel according to the hot trend, and based on the multi-level push strategy, performing personalized push according to the urgency of the news hotspot and the user's interest;

[0105] S53, continue to track user click-through rate, number of comments, number of shares and other feedback indicators, evaluate the scheduling effect, and compare user feedback and performance indicators under different push strategies through the A / B testing framework;

[0106] S54. Based on user feedback and A / B test results, iteratively optimize the push strategy, combine user interest portraits and news hotspot prediction results to make more accurate personalized news recommendations; and use recommendation diversity algorithms to prevent users from falling into information cocoons.

[0107] The working principle of the above technical solution is as follows: Apache Flink, Storm and other stream processing technologies are used to efficiently process and analyze real-time news data streams. These technologies can capture, process and analyze news data in real time to ensure the timeliness and accuracy of information; in the process of stream processing, real-time feature extraction algorithms are applied to extract key features from news data, such as title, content, release time, source, etc. Subsequently, these features are fused with other related features (such as user behavior data, historical news data, etc.) through feature fusion algorithms to generate fused feature vectors. These feature vectors will serve as the basis for subsequent predictions and recommendations; the prediction results of the news hotspot prediction model (as described in S4) are applied to the news push strategy. According to the hot trend, the order, frequency and channel of news push are intelligently adjusted to ensure that users can first receive the most interesting and valuable news; according to the urgency of news hotspots and user interests, a multi-level push strategy is implemented. Urgent and important news are pushed first; for news with high user interest, personalized push is carried out according to the user's active time period and preferences; the user's feedback indicators such as click-through rate, number of comments, and number of shares on news are continuously tracked to evaluate the effectiveness of the push strategy. These indicators can directly reflect the user's interest in news and the pros and cons of the push strategy; use the A / B test framework to compare user feedback and performance indicators under different push strategies. By comparing user behavior data and performance indicators under different strategies, we can objectively evaluate the effectiveness of the push strategy and find the optimal strategy; based on user feedback and A / B test results, we continuously iterate and optimize the push strategy. By adjusting parameters such as push order, frequency, and channels, and introducing new features and algorithms, we can continuously improve the accuracy and effectiveness of the push strategy; combine user interest portraits and news hotspot prediction results to make more accurate personalized news recommendations. At the same time, apply recommendation diversity algorithms to prevent users from falling into information cocoons and improve the freshness and satisfaction of recommendations. By recommending diverse news content, we can meet users' different needs and interests and improve user experience and loyalty.

[0108] The effects of the above technical solutions are: by adopting stream processing technologies such as Apache Flink and Storm, news data can be captured, processed and analyzed in real time to ensure the timeliness and accuracy of news information. This is crucial for news recommendation systems because the value of news often decays rapidly over time; real-time feature extraction algorithms can quickly extract key information from news data, while feature fusion algorithms can effectively integrate this information with other relevant features (such as user behavior data) to generate fused feature vectors with rich information. This provides a solid foundation for subsequent predictions and recommendations; applying the prediction results to news push strategies can intelligently adjust the order, frequency and channels of news push to ensure that users receive the most interesting and valuable news first. This greatly improves the pertinence and effectiveness of push; based on a multi-level push strategy, personalized push is performed according to the urgency of news hotspots and user interests. This strategy not only takes into account the timeliness and importance of news, but also fully respects the personalized needs of users and improves user satisfaction; continuously tracks feedback indicators such as user click-through rate, number of comments, and number of shares, and conducts real-time evaluation of scheduling effects. This helps to promptly identify problems in the push strategy and make targeted adjustments; through the A / B testing framework, user feedback and performance indicators under different push strategies are compared to find the optimal strategy. At the same time, based on user feedback and test results, the push strategy is iteratively optimized to continuously improve the performance of the recommendation system; combined with user interest portraits and news hotspot prediction results, more accurate personalized news recommendations are made. At the same time, through the recommendation diversity algorithm, users are prevented from falling into information cocoons and the freshness and satisfaction of recommendations are improved. This helps to meet the diverse needs of users and improve user experience; because the push strategy is more intelligent, personalized and diverse, users can continue to receive news content that meets their interests and needs. This will enhance users' trust and reliance on the recommendation system and improve user loyalty.

[0109] In one embodiment of the present invention, the S51 includes:

[0110] Use distributed message queue systems such as Apache Kafka to access streaming data provided by major news sources in real time, including news titles, content summaries, release time, sources, tags, etc., and pre-process the data;

[0111] Based on the TF-IDF model, we extract text features such as keywords and topic vectors from news texts to reflect the core content of the news. We also introduce time features, such as news release time, update time, and the rate of change of hot topics based on time windows, to capture the time sensitivity and trend of news. We use sentiment analysis algorithms to evaluate the sentiment tendency (positive, negative, neutral) of news content as sentiment features.

[0112] Combine social media data, user behavior data, and historical news data to build a multi-dimensional feature system; use deep learning models to fuse multi-source data and extract advanced features;

[0113] Based on the feature fusion algorithm, a fused feature vector is generated by feature concatenation.

[0114] The working principle of the above technical solution is as follows: using distributed message queue systems such as Apache Kafka to access streaming data provided by major news sources in real time; the accessed data includes news titles, content summaries, release times, sources, tags, etc.; performing pre-processing operations such as cleaning, deduplication, and formatting on the accessed raw data to ensure data accuracy and consistency; it may include steps such as missing value filling, outlier processing, and data standardization; based on the TF-IDF model, extracting text features such as keywords and topic vectors from news texts to reflect the core content of the news; introducing time features such as news release time, update time, and the rate of change of hot topics based on time windows to capture the time sensitivity and Trend; Use sentiment analysis algorithms (such as machine learning or deep learning-based methods) to evaluate the sentiment tendency (positive, negative, neutral) of news content as sentiment features; Integrate data such as the heat of relevant discussions and sentiment tendencies on social media platforms (such as Weibo and Twitter) to reflect the public's attention and attitude towards news events; Analyze user browsing history, click preferences and other behavioral data to understand user interest preferences and reading habits; Use historical news data to analyze the evolution trend of news events, reporting angles, etc., to provide reference for the analysis of current news events; Apply deep learning models (such as convolutional neural networks CNN and long short-term memory networks LSTM) to fuse multi-source data and extract advanced features. These models can capture complex patterns and associations in data and improve the accuracy and efficiency of feature extraction; Based on feature fusion algorithms (such as splicing, weighted summation, etc.), fuse text features, sentiment features, social media data features, user behavior data features, and historical news data features; Generate a fusion feature vector through feature fusion algorithms. This vector fully reflects the comprehensive attributes of news and user preferences, providing strong support for subsequent news analysis, recommendation and other applications.

[0115] The effect of the above technical solution is: by using distributed message queue systems such as Apache Kafka, S51 can access streaming data provided by major news sources in real time. This means that once the news data is released, it can be quickly captured and processed, thereby ensuring the timeliness and freshness of news processing; the distributed message queue system not only provides the ability to access real-time data, but also ensures the efficiency and stability of data access. Even in the face of large-scale data streams, the system can maintain smooth operation without data congestion or loss; based on the TF-IDF model, it can accurately extract text features such as keywords and topic vectors in news texts, which can truly reflect the core content of the news; by introducing time features such as news release time, update time, and the rate of change of hot topics based on time windows, it can capture the time sensitivity and trend of news, providing richer information for news analysis and recommendation; using sentiment analysis algorithms, it can evaluate the sentiment tendency of news content, which helps to understand the public's attitude and emotions towards news events; combining social media data, user behavior data and historical news data, a multi-dimensional feature system is constructed. This system can fully reflect the comprehensive attributes of news and user preferences, and provide strong support for news recommendation and personalized services; by applying deep learning models such as convolutional neural network CNN and long short-term memory network LSTM, it can fuse multi-source data and extract advanced features. These advanced features can capture complex patterns and associations in the data and improve the accuracy of news analysis and recommendation; based on the feature fusion algorithm, a fusion feature vector is generated by feature splicing. This vector fully integrates news text features, time features, emotional features and advanced features of multi-source data, providing comprehensive information support for news recommendation and personalized services; by fully reflecting the comprehensive attributes of news and user preferences, it can provide users with more personalized news recommendation services. This not only improves user satisfaction and loyalty, but also brings more traffic and revenue to news platforms.

[0116] In one embodiment of the present invention, the S52 includes:

[0117] Using the fusion feature vector output by the real-time news data processing system and the results of the news hotspot prediction model, we can conduct an in-depth analysis of news hotspot trends;

[0118] Based on the results of hot trend analysis, the news push priority is intelligently set; news hot spots that emerge quickly and have great influence should be given higher push priority;

[0119] Build user interest profiles, integrating multi-dimensional data such as users' historical browsing records, click preferences, comments and sharing behaviors, and characterize users' news preferences;

[0120] Develop personalized news push strategies based on news hotspot prediction results and user interest profiles, including pushing relevant news based on user interest preferences, and adjusting push frequency and channels based on the urgency of news hotspots and user attention;

[0121] Design a multi-level push mechanism to divide news push into different levels such as instant push, priority push, and regular push according to the urgency of news hotspots and different user interests;

[0122] Implement a multi-level push strategy and use intelligent push algorithms to dynamically adjust push order and channels based on factors such as the user's online status, device type, and reading habits;

[0123] The working principle of the above technical solution is as follows: the real-time news data processing system (such as Apache Flink or Storm) is responsible for processing and analyzing the real-time news data stream, extracting key information and generating fused feature vectors. These feature vectors contain multiple dimensional information of news content, such as topics, keywords, emotional tendencies, etc.; the news hotspot prediction model uses these fused feature vectors, combined with historical data and machine learning algorithms, to predict key indicators such as the speed of rise, duration, and potential influence of news hotspots; through the analysis of the prediction results, the system can deeply understand the trend of news hotspots and provide a basis for the formulation of subsequent push strategies; based on the results of the news hotspot trend analysis, the system will intelligently set the push priority of news according to the speed of rise and influence of the news; as the news hotspots develop and change, the system will update the push priority in real time to ensure that important news can be pushed quickly; the system integrates multi-dimensional data such as users' historical browsing records, click preferences, comments and sharing behaviors to build user interest portraits; user interests are not fixed. The system will track changes in user interests and dynamically update user portraits. Combining news hotspot prediction results and user interest portraits, the system can push news related to users' interest preferences. According to the urgency of news hotspots and user attention, the system will adjust the push frequency and channels to ensure that users can obtain important information in a timely manner. According to the urgency of news hotspots and user interests, the system will divide news push into different levels such as instant push, priority push, and regular push. For news of different levels, the system will adopt corresponding push strategies to ensure that important news can reach users quickly. The system uses intelligent push algorithms to dynamically adjust the push order and channels according to factors such as the user's online status, device type, and reading habits. The system will collect user feedback on push content and continuously optimize push strategies based on feedback results.

[0124] The effects of the above technical solutions are as follows: through the fusion feature vector output by the real-time news data processing system, combined with the news hot spot prediction model, it is possible to deeply analyze the key indicators such as the rise speed, duration, potential influence, etc. of news hot spots, and improve the accuracy of news hot spot trend prediction; based on the hot spot trend analysis results, the news push priority is intelligently set to ensure that important news can reach users quickly; build user interest portraits, integrate multi-dimensional user data, characterize users' news preferences, and thus formulate personalized news push strategies to improve user satisfaction and stickiness; design a multi-level push mechanism, and divide news push into different levels such as instant push, priority push, and regular push according to the urgency of news hot spots and different user interests, to avoid users Being overwhelmed by a large amount of irrelevant information; implementing a multi-level push strategy, and dynamically adjusting the push order and channels according to factors such as the user's online status, device type, reading habits, etc., to further improve push efficiency and user experience; this technical solution can push relevant news content that meets the user's interests according to the user's personalized needs, and promote the diversification and precision of news content; at the same time, by adjusting the push frequency and channels, it can more effectively meet the user's news needs in different scenarios; for news media, this technical solution can help them better understand user needs and market trends, so as to formulate more targeted news production and push strategies; by improving the accuracy and efficiency of news push, it can enhance the competitiveness and brand influence of news media.

[0125] In one embodiment of the present invention, the S54 includes:

[0126] Collect and analyze user feedback data on news push, including but not limited to click-through rate, dwell time, comment content, sharing behavior, etc., and obtain user preferences and satisfaction with push content based on the analysis results;

[0127] Use the A / B testing framework to compare user feedback and performance indicators under different push strategies, such as user activity, retention rate, conversion rate, etc., to identify the best performing push strategy and its key elements;

[0128] Iterate and optimize the push strategy based on user feedback and A / B test results; including adjusting the push time, frequency, and channel, optimizing the selection and presentation of news content, introducing new personalized recommendation algorithms, and automatically adjusting the push strategy using machine learning algorithms;

[0129] Deepen the construction of user interest portraits, introduce more dimensional user data, such as social media behavior, search records, purchase history, etc., deeply integrate news hotspot prediction results with user interest portraits, and use deep learning algorithms to accurately match news content with user interests;

[0130] Through the recommendation algorithm based on content diversity, users are prevented from being trapped in information cocoons. The application effect of the recommendation diversity algorithm is continuously tracked. Through user feedback, click-through rate, dwell time and other related indicators, the role of the algorithm in improving recommendation quality and user experience is evaluated;

[0131] Establish a continuous optimization mechanism for the intelligent recommendation system, and use online learning algorithms to continuously learn and adapt to changes in user behavior.

[0132] The working principle of the above technical solution is as follows: through various channels and tools, the system comprehensively collects user feedback data on news push, including but not limited to key indicators such as click-through rate, dwell time, comment content, sharing behavior, etc.; using advanced data analysis technology and algorithms, these feedback data are deeply mined and analyzed to obtain user preferences and satisfaction with the pushed content; based on the analysis results, the system can more accurately understand user needs and expectations, and provide strong support for the subsequent optimization of push strategies; using the A / B testing framework, the system can compare user feedback and performance indicators under different push strategies; by randomly dividing users into experimental and control groups, the system can push different news content to them respectively. Or adopt different push methods; then, the system collects and compares feedback data from the two groups of users, such as user activity, retention rate, conversion rate, etc., to identify the best performing push strategy and its key elements; based on user feedback and A / B test results, the system can continuously iterate and optimize the push strategy; this includes adjusting parameters such as push time, frequency, and channel to better adapt to user needs and habits; at the same time, the system will also optimize the selection and presentation of news content, such as improving the design of elements such as titles, summaries, and pictures to increase the attractiveness and readability of news; in addition, the system will also introduce new personalized recommendation algorithms to further improve the accuracy and effectiveness of push; with the help of machine learning algorithms, the system can automatically adjust The system will improve the push strategy to adapt to the ever-changing user behavior and market environment; the system will deepen the construction of user interest portraits and introduce more dimensional user data; these data may include social media behavior, search records, purchase history, etc., which can more comprehensively reflect the user's interests, needs and preferences; by deeply integrating the news hotspot prediction results with the user interest portrait, the system can more accurately predict the news content that the user may be interested in; using deep learning algorithms, the system can achieve accurate matching of news content with user interests, thereby improving the satisfaction and effectiveness of push; in order to prevent users from falling into information cocoons, the system will adopt a recommendation algorithm based on content diversity; this algorithm can ensure the accuracy of recommendations while Increase the diversity and breadth of recommended content; by continuously tracking the application effect of the recommendation diversity algorithm and evaluating it with relevant indicators such as user feedback, click-through rate, and dwell time, the system can continuously optimize the performance and effect of the algorithm; the system will establish a continuous optimization mechanism for the intelligent recommendation system, and use online learning algorithms to continuously learn and adapt to changes in user behavior; by real-time monitoring and analysis of user behavior data, the system can promptly discover and respond to changes and trends in user interests; at the same time, the system will also regularly update and optimize the recommendation algorithms and models to maintain their accuracy and effectiveness; this continuous optimization mechanism can ensure that the intelligent recommendation system always stays at the forefront of the industry and provides users with better quality and personalized news push services.

[0133] The effects of the above technical solution are as follows: by collecting and analyzing user feedback data on news push, including click-through rate, dwell time, comment content and sharing behavior, the solution can deeply understand the user's preference and satisfaction with the pushed content. Based on these data, the system can more accurately push news that users are interested in, thereby improving the user's reading experience and satisfaction; using the A / B testing framework to compare user feedback and performance indicators under different push strategies, such as user activity, retention rate and conversion rate, the solution can identify the best performing push strategy and its key elements. By iteratively optimizing the push strategy, such as adjusting the push time, frequency and channel, optimizing the selection and presentation of news content, and introducing new personalized recommendation algorithms, the system can further improve user activity and participation; by deepening the construction of user interest portraits and introducing more dimensional user data, such as social media behavior, search records and purchase history, the solution can deeply integrate the news hotspot prediction results with user interest portraits. By using deep learning algorithms to accurately match news content with user interests, the system can provide users with more personalized news recommendation services and enhance user stickiness and loyalty; by using a recommendation algorithm based on content diversity, this solution can prevent users from being trapped in information cocoons and ensure that users can receive diverse news content. By continuously tracking the application effect of the recommendation diversity algorithm and evaluating it using relevant indicators such as user feedback, click-through rate, and dwell time, the system can continuously optimize the recommendation algorithm, improve the quality of recommendations and user experience; establish a continuous optimization mechanism for the intelligent recommendation system, and use online learning algorithms to continuously learn and adapt to changes in user behavior. This solution can ensure that the intelligent recommendation system always stays at the forefront of the industry, responds to changes and trends in user interests in a timely manner, and provides users with more high-quality and personalized news push services.

[0134] One embodiment of the present invention, as Figure 2 As shown, a news hotspot prediction and scheduling system under a big data platform includes:

[0135] Data acquisition module: acquires news data from relevant data sources through multiple means; the news data includes historical news data and real-time news data; and pre-processes the acquired news data;

[0136] Feature extraction module: extracts features of preprocessed multi-dimensional data with different features through different means; fuses features of different dimensions to generate a fused feature vector;

[0137] Parameter adjustment module: dynamically adjusts network structure and parameters based on historical and real-time data of news hotspots through adaptive deep learning models;

[0138] Model training module: Input the fused feature vector and the corresponding news hotspot label into the adaptive deep learning model, and use the cross entropy loss function and reinforcement learning reward function to jointly guide the model training;

[0139] Strategy adjustment module: Use the trained adaptive deep learning model to predict real-time news hotspot data, and intelligently adjust the news scheduling strategy based on the prediction results and dynamic changes in news hotspots.

[0140] The working principle of the above technical solution is as follows: using PAI interface, web crawlers and other means to collect news data from data sources such as news websites, social media platforms and user behavior logs; the collected data includes historical news data and real-time news data, covering multi-dimensional information such as text, pictures, videos and user comments; pre-processing the acquired news data, such as data cleaning, denoising, formatting, etc., to ensure the quality and consistency of the data; using natural language processing technology (such as BERT, GPT, etc.) to extract text features, these technologies can deeply understand the semantics of text; using deep learning models (such as ResNet, VGG, etc.) to extract image and video features, these models are good at capturing key information in images and videos; combining user behavior log analysis to extract user interest features and understand user preferences and needs; through feature fusion strategies (such as attention mechanism, graph neural network, etc.), features of different dimensions are fused to generate a fused feature vector. This vector integrates information from multiple aspects such as news content and user interests; an adaptive deep learning model is constructed, which includes an input layer, a feature extraction layer, an adaptive learning layer, and an output layer; in the feature extraction layer, deep learning models such as convolutional neural networks (CNN) and recurrent neural networks (RNN) are used to further extract the deep features of the fused feature vector; the adaptive learning layer introduces reinforcement learning algorithms (such as Q-learning, DeepQ-Network, etc.) so that the model can self-optimize according to the prediction results and scheduling effects, and dynamically adjust the network structure and parameters; the fused feature vector and the corresponding news hotspot label are input into the adaptive deep learning model; the cross entropy loss function and the reinforcement learning reward function are used together to Guide model training to ensure that the model can optimize the scheduling strategy while accurately predicting news hotspots; during the training process, adopt dynamic learning rate adjustment strategy, gradient clipping technology and early stopping method and other techniques to prevent model overfitting and improve training efficiency; through continuous iterative training until the model converges, obtain the optimal model parameters; use the trained adaptive deep learning model to predict real-time news hotspot data; according to the prediction results and the dynamic changes of news hotspots, intelligently adjust the news scheduling strategy, including push order, push frequency and push channel, etc.; introduce user feedback mechanism to evaluate the scheduling effect according to indicators such as user click-through rate, number of comments, and number of shares; further optimize the scheduling strategy according to the evaluation results to achieve accurate push and efficient scheduling of news hotspots.

[0141] The effects of the above technical solutions are as follows: through PAI interface, web crawlers and other means, news data is obtained from multiple data sources such as news websites, social media platforms and user behavior logs, ensuring the diversity and comprehensiveness of the data; pre-processing the obtained news data improves the quality and availability of the data, laying a solid foundation for subsequent feature extraction and model training; using natural language processing technology, deep learning models and other technical means to extract features of multi-dimensional data such as text, pictures, videos and user comments, fully capturing the rich information of news content; through feature fusion strategies, features of different dimensions are fused to generate fused feature vectors, improving the accuracy and completeness of feature expression, and providing strong support for model prediction; the adaptive deep learning model can dynamically adjust the network structure and parameters according to the historical data and real-time data of news hotspots, improving The model is adaptable and flexible; a reinforcement learning algorithm is introduced to enable the model to self-optimize according to the prediction results and scheduling effects, and continuously improve the prediction accuracy and scheduling efficiency; in the model training process, dynamic learning rate adjustment strategy, gradient clipping technology and early stopping method are used to effectively prevent model overfitting and improve training efficiency and model performance; through continuous iterative training until the model converges, the optimal model parameters are obtained to ensure the stability and accuracy of the model; the trained adaptive deep learning model is used to predict real-time news hotspot data, and the news scheduling strategy is intelligently adjusted according to the prediction results and the dynamic changes of news hotspots, which improves the accuracy and timeliness of news push; a user feedback mechanism is introduced to evaluate the scheduling effect according to indicators such as user click-through rate, number of comments, and number of shares, and the scheduling strategy is further optimized to achieve accurate matching of user needs and personalized push.

[0142] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. A news hotspot prediction and scheduling method based on a big data platform, characterized in that: The method comprises: S1. Acquire news data from relevant data sources by multiple means; the news data includes historical news data and real-time news data; and pre-process the acquired news data; S2, extracting features of preprocessed multi-dimensional data with different features by different means; and fusing features of different dimensions by feature fusion strategy to generate a fused feature vector; S3, dynamically adjust the network structure and parameters according to the historical data and real-time data of news hotspots through adaptive deep learning models; S4, inputting the fused feature vector and the corresponding news hotspot label into the adaptive deep learning model, and using the cross entropy loss function and the reinforcement learning reward function to jointly guide the model training; S5. Use the trained adaptive deep learning model to predict real-time news hotspot data, and intelligently adjust the news scheduling strategy based on the prediction results and dynamic changes of news hotspots; The S5 comprises: S51, generating a fused feature vector based on real-time feature extraction and fusion algorithm; S52, applying the prediction results to the news push strategy, and performing personalized push based on the multi-level push strategy; S53, evaluating the scheduling effect; S54, iterative optimization push strategy; The S51 includes: Access streaming data provided by major news sources in real time and pre-process the data; Extract text features based on the TF-IDF model; capture the time sensitivity and trend of news; obtain sentiment features; Build a multi-dimensional feature system; fuse multi-source data and extract advanced features; Based on the feature fusion algorithm, a fused feature vector is generated by feature concatenation.

2. According to the method for predicting and scheduling news hot spots under a big data platform as described in claim 1, it is characterized in that: Said S1 comprises: S11. Regularly pull historical news data from news sources through the platform's news API interface; and set request strategies based on the API access frequency limit; S12. For different social media platforms and news websites, crawl data in real time through web crawlers, and dynamically adjust crawler crawling tasks and priorities based on the crawler scheduling system according to the frequency and demand of data updates; S13. Collect the user's log data from the corresponding channels with the user's authorization; and perform desensitization processing on the collected log data; S14. Use hash algorithms to deduplicate news data and filter out useless data based on preset keyword blacklists and regular expressions; complete missing key fields through other sources or historical data; and mark data that cannot be completed; S15. Use natural language processing tools to process the text data; and perform spelling check on the text data to correct spelling errors; S16. Perform image processing on the image and video data, unify the input standard, and perform image enhancement processing on the image data.

3. According to the method for predicting and scheduling news hot spots under a big data platform as described in claim 1, it is characterized in that: The S2 comprises: S21. Use the pre-trained BERT or GPT model to understand the semantics of news text and extract high-level features; and fine-tune the model according to the characteristics of the news text; S22. Based on word embedding technology, the semantic similarity between news texts is calculated to identify related news clusters; clustering algorithms are used to perform cluster analysis on news texts to further explore the correlation between news; S23. Apply deep learning models such as ResNet or VGG to process news pictures and extract image features; and perform dimensionality reduction on image features; S24, extracting key information of the video content through video frame extraction and key frame recognition technology, and extracting and analyzing features of the key frames in combination with image feature extraction technology; S25. Statistically analyze the user's historical behavior data, build a user interest profile, and use a recommendation algorithm to explore the user's potential interests; S26. Analyze the interaction of users in social networks and evaluate the influence of users in information dissemination; use social network analysis algorithms to identify key nodes and community structures; S27. Based on the attention mechanism, dynamically adjust the contribution of different features to the prediction results, and capture the complex relationship between different features through the multi-head attention mechanism; S28. Construct a news-user relationship graph, use graph neural networks to capture the correlation between news and the mutual influence between users, and combine graph embedding technology to represent news and users as low-dimensional vectors.

4. According to the method for predicting and scheduling news hot spots under a big data platform as described in claim 3, it is characterized in that: The S26 comprises: Collect and clean user social network data, extract key features from user social networks, and build a user influence evaluation model based on user social network features and interactive behavior data; Use social network analysis algorithms to identify key nodes and community structures in social networks, and further analyze the identified key nodes; Use visualization technology to intuitively display the influence distribution, key nodes and community structure in the user's social network, and generate detailed reports based on the analysis results; Through the dynamic monitoring mechanism of social influence, the interactive behavior and influence changes in users' social networks are tracked in real time, and the influence evaluation model and key node identification algorithm are regularly updated and optimized based on the monitoring results.

5. According to the method for predicting and scheduling news hot spots under a big data platform as described in claim 3, it is characterized in that: The S27 comprises: The attention weight of each feature for the prediction result is calculated by multiplying the feature vector with the attention matrix and normalizing it through the softmax function; According to the calculated attention weight, each feature vector is weighted, and a multi-head attention mechanism is introduced to split the feature vector into multiple heads, and the attention weight is calculated for each head separately; Multiply the attention weight of each head by the corresponding feature sub-vector and perform weighted summation to obtain a feature representation that integrates multi-head attention. The feature representation that integrates multi-head attention is integrated with other features; according to the integrated feature representation, the contribution of different features to the prediction results is dynamically adjusted; Visualize the attention weights by generating heat maps or attention weight distribution maps; combine the visualization results to explain and analyze the model's predictive behavior.

6. According to the method for predicting and scheduling news hot spots on a big data platform as described in claim 1, it is characterized in that: The S3 includes: S31, receiving the fused feature vector as input, combining with the deep learning model, performing deep feature extraction on the fused feature vector; capturing feature information of different scales through multi-scale convolution kernels; S32. Use the recurrent neural network model to capture time series dependencies, design state space, action space and reward function, and build a reinforcement learning framework; S33. Output the predicted probability or classification result of news hotspots, and perform fine-grained classification of news hotspots through multiple classifiers; dynamically adjust the network structure parameters according to the type and scale of news data.

7. The news hotspot prediction and scheduling method under the big data platform according to claim 6 is characterized in that: The S4 comprises: Preprocess the received feature vectors and fuse the preprocessed feature vectors to form a unified feature representation; Construct a CNN model and design a network structure, which includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer; wherein the convolutional layer is used to extract local features in the feature vector; Introduce multi-scale convolution kernels in the convolution layer, and capture feature information of different scales by setting convolution kernels of different sizes; downsample the feature map output by the convolution layer through the pooling layer; The feature map output by the pooling layer is input into the fully connected layer for further feature extraction and classification; the fully connected layer maps the feature map to the classification space and outputs the prediction result; Construct an RNN model and design a network structure, which includes an input layer, a hidden layer, and an output layer; wherein the hidden layer will be used to capture the time series dependency in the feature vector; Introduce LSTM into the RNN model; use the feature vector extracted by the CNN model as the input of the RNN model, and capture the time series dependency in the feature vector through the hidden layer; The feature results extracted by the CNN model and the RNN model are fused to form the final deep-level feature representation; the fused feature representation is input into the classifier for classification or regression prediction, and the prediction probability or classification result of the news hotspot is output; At the same time, the network structure parameters are dynamically adjusted according to the type and scale of news data.

8. The news hotspot prediction and scheduling method under the big data platform according to claim 1 is characterized in that: The S4 comprises: S41, distinguish the difference between the model prediction probability and the true label, and guide the model to learn in the right direction based on the discrimination result; and use weighted cross entropy loss to perform imbalanced processing on news hotspots of different categories; S42. Design a reward function based on the news scheduling effect, and based on a multi-stage reward mechanism, encourage the model to continuously optimize during the prediction and scheduling process; S43, adopt the learning rate decay strategy to balance the convergence speed and stability of the model, and gradually increase the learning rate in the early stage of training based on the learning rate warm-up mechanism; S44. Limit the gradient norm to prevent gradient explosion, and standardize the gradient through the gradient normalization algorithm; monitor the performance of the validation set and terminate the training early when the performance no longer improves.

9. The news hotspot prediction and scheduling method under the big data platform according to claim 1 is characterized in that: The S5 comprises: S51, processing and analyzing real-time news data through stream processing technology, and generating fusion feature vectors based on real-time feature extraction and fusion algorithms; S52, applying the prediction results to the news push strategy, intelligently adjusting the push order, frequency and channel according to the hot trend, and based on the multi-level push strategy, performing personalized push according to the urgency of the news hotspot and the user's interest; S53. Continue to track user feedback indicators, evaluate the scheduling effect, and compare user feedback and performance indicators under different push strategies through the A / B testing framework; S54. Based on user feedback and A / B test results, iteratively optimize the push strategy, combine user interest portraits and news hotspot prediction results to make personalized news recommendations; and use the recommendation diversity algorithm to prevent users from falling into information cocoons.

10. A system for implementing the news hotspot prediction and scheduling method under the big data platform as claimed in claim 1, characterized in that: The system comprises: Data acquisition module: acquires news data from relevant data sources through multiple means; the news data includes historical news data and real-time news data; and pre-processes the acquired news data; Feature extraction module: extracts features of preprocessed multi-dimensional data with different features through different means; and fuses features of different dimensions through feature fusion strategy to generate fused feature vectors; Parameter adjustment module: dynamically adjusts network structure and parameters based on historical and real-time data of news hotspots through adaptive deep learning models; Model training module: Input the fused feature vector and the corresponding news hotspot label into the adaptive deep learning model, and use the cross entropy loss function and reinforcement learning reward function to jointly guide the model training; Strategy adjustment module: Use the trained adaptive deep learning model to predict real-time news hotspot data, and intelligently adjust the news scheduling strategy based on the prediction results and dynamic changes in news hotspots.

Citation Information

Patent Citations

  • Method for distinguishing important goals and community groups of social network

    CN103024017A

  • Social network-oriented hot event prediction method

    CN113806534A

  • News event classification method

    CN119128155A