News transmission method based on big data
By constructing dynamic hypergraph models and deep learning technology, the timeliness of static hypergraph technology in the field of news communication is solved, real-time analysis of user behavior and news semantics is realized, and the accuracy and adaptability of news communication feature extraction is improved.
Patent Information
- Application Number
- CN202510773046.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The existing static hypergraph technology cannot update the semantic relationship between user behavior and news topics in the field of news communication, resulting in the lack of timeliness of dynamic interactions and the inability to synchronize the semantic drift rate of news topics.
Build a dynamic hypergraph model, dynamically update nodes and hyper-edges through Apache Kafka, analyze news dissemination feature vectors using hypergraph neural network and deep learning model, and generate optimized dissemination strategies.
It realizes dynamic structured modeling of user-news behavior intensity correlation and news-topic semantic correlation, improves the comprehensiveness and timeliness of relational modeling, and enhances the accuracy and adaptability of news communication feature extraction.
Smart Images

Figure CN120296174A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a news dissemination method based on big data. Background Art
[0002] With the complexity of the Internet information ecosystem, news dissemination has evolved from traditional content distribution to the stage of personalized recommendation based on user behavior; however, existing technologies generally adopt collaborative filtering algorithms and matrix factorization methods to construct user-news association models, and predict users' potential interests through users' historical behavior data; in recent years, hypergraph theory has been introduced into the field of social network analysis, and the high-order association characteristics of hypergraph theory can effectively model the non-linear relationship between multi-modal data. Some studies have attempted to apply static hypergraphs to news topic clustering and user group division; in terms of real-time data processing, distributed message systems have also achieved dynamic collection and transmission of data streams, providing the news dissemination system with the ability of incremental update with millisecond-level delay.
[0003] The deficiencies of current static hypergraph technology in the field of news dissemination are mainly reflected in the lack of timeliness of dynamic interaction: existing technologies mostly adopt a batch processing mechanism with a fixed time window, resulting in a delay in the update of the association of user behavior intensity, and being unable to synchronize the semantic drift rate of news topics; typical methods adopted by existing technologies such as Hypergraph Neural Network mostly rely on offline full-scale data to reconstruct the hypergraph topology, and although the news feature extraction algorithm based on TF-IDF can capture static semantics, it is difficult to eliminate the interference of expired interaction records on the current user feature vector. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a news dissemination method based on big data to solve the deficiencies of current static hypergraph technology in the field of news dissemination.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a news dissemination method based on big data, which includes collecting news data and performing preprocessing to generate a structured news data set. The structured news data set includes hashed user identifiers, user feature vectors, interaction records, news content feature vectors, and topic feature vectors; constructing a dynamic hypergraph to perform structured modeling on the structured news data set, capturing the behavioral intensity association between users and news and the semantic association between news and topics through the connection of nodes and hyperedges, and using Apache Kafka to dynamically update nodes and hyperedges and remove expired interaction records; using a hypergraph neural network model to perform multi-layer aggregation and non-linear transformation on the node features of the updated dynamic hypergraph to generate news dissemination feature vectors; and analyzing the spatio-temporal dynamic characteristics of the news dissemination feature vectors through a deep learning model to generate an optimized dissemination strategy.
[0007] As a preferred solution of the news dissemination method based on big data according to the present invention, wherein: constructing a dynamic hypergraph to perform structured modeling on the structured news data set means reading structured news data from the structured news data set, defining user nodes, news nodes, and topic nodes, and creating user-news hyperedges and news-topic hyperedges to form a dynamic hypergraph.
[0008] As a preferred solution of the news dissemination method based on big data according to the present invention, wherein: capturing the behavioral intensity association between users and news means extracting interaction records from the structured news data, creating user-news hyperedges to connect user nodes and news nodes, setting hyperedge weights according to the interaction type, and if there are multiple interaction records between a user and a news, updating the weight of a single user-news hyperedge by aggregating and accumulating interaction weights through an aggregation operation; The semantic association between news and topics means obtaining news content feature vectors and topic feature vectors from the structured news data, reducing the dimensionality of the news content feature vectors using the fully connected layer of the BERT model, calculating the cosine similarity between the reduced-dimensional news content feature vectors and the topic feature vectors through the built-in operation function of Python, and if the cosine similarity is greater than the similarity threshold, creating a news-topic hyperedge to connect the news node and the topic node, and using the cosine similarity as the weight of the news-topic hyperedge to reflect the semantic association between the news and the topic.
[0009] As a preferred solution of the news dissemination method based on big data according to the present invention, wherein: using Apache Kafka to dynamically update nodes and hyperedges means obtaining new structured news data through Apache Kafka, updating the structured news data set, adding user nodes, news nodes, and topic nodes, updating the weights of user-news hyperedges based on new interaction records, and calculating the cosine similarity based on news content feature vectors and topic feature vectors to update news-topic hyperedges to form an updated dynamic hypergraph.
[0010] As a preferred solution of the news dissemination method based on big data according to the present invention, wherein: the multi-layer aggregation of the node features of the updated dynamic hypergraph by using the hypergraph neural network model is specifically carried out as follows: Load the user feature vector of the user node, the news content feature vector of the news node, and the topic feature vector of the topic node from the dynamic hypergraph; Reduce the dimensions of the user feature vector and the news content feature vector through the fully connected layer of the BERT model; Perform weighted fusion on the dimension-reduced user feature vector and news content feature vector and the non-dimension-reduced topic feature vector to form a unified node feature vector; Based on the dynamic hypergraph, construct a hypergraph neural network model and set multiple convolutional layers; The first layer forms an initial aggregated feature vector by weighted aggregation of the unified node feature vectors of the neighborhood nodes through the user-news hyperedge and the news-topic hyperedge; The second layer forms an intermediate aggregated feature vector by weighted aggregation of the initial aggregated feature vectors of the neighborhood nodes through the user-news hyperedge and the news-topic hyperedge; The third layer forms a final aggregated feature vector by weighted aggregation of the intermediate aggregated feature vectors of the neighborhood nodes through the user-news hyperedge and the news-topic hyperedge.
[0011] As a preferred solution of the news dissemination method based on big data according to the present invention, wherein: the non-linear transformation is specifically carried out as follows: Apply the ReLU activation function to the initial aggregated feature vector generated by the first-layer convolution for non-linear mapping to generate an initial enhanced feature vector; Apply the ReLU activation function to the intermediate aggregated feature vector generated by the second-layer convolution for non-linear mapping to generate an intermediate enhanced feature vector; Apply the ReLU activation function to the final aggregated feature vector generated by the third-layer convolution for non-linear mapping to generate a news dissemination feature vector.
[0012] As a preferred solution of the news dissemination method based on big data according to the present invention, wherein: the analysis of the spatio-temporal dynamic characteristics of the news dissemination feature vector by using the deep learning model is specifically carried out as follows: Obtain the news dissemination feature vector from Apache Kafka; Process the news dissemination feature vector through the attention mechanism layer of the deep learning model to generate attention weights; Adjust the news dissemination feature vector based on the attention weights to form a weighted news dissemination feature vector; The time step processing is performed on the weighted news dissemination feature vector through the LSTM layer of the deep learning model to capture the dissemination dynamic characteristics of the weighted news dissemination feature vector; The gating mechanism of the LSTM layer is used to update the hidden state of the LSTM layer to generate a spatio-temporal dynamic characteristic representation of the news dissemination feature vector.
[0013] As a preferred solution of the news dissemination method based on big data according to the present invention, wherein: the optimized dissemination strategy includes news identification, dissemination priority, dissemination channel and dissemination frequency.
[0014] In a second aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the news dissemination method based on big data as described in the first aspect of the present invention is implemented.
[0015] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by the processor, any step of the news dissemination method based on big data as described in the first aspect of the present invention is implemented.
[0016] The beneficial effects of the present invention are as follows: by constructing a dynamic hypergraph based on a structured news data set, defining user nodes, news nodes and topic nodes, and creating user-news hyperedges and news-topic hyperedges by using interaction records and cosine similarity, and at the same time dynamically updating nodes and hyperedges and removing expired interaction records through Apache Kafka, the dynamic structured modeling of user-news behavior intensity association and news-topic semantic association is realized; capturing multi-dimensional complex relationships in the form of a hypergraph and maintaining real-time performance, which is suitable for analyzing the dynamic changes of user behavior and news semantics, and finally achieving the effect of improving the comprehensiveness and timeliness of relationship modeling, and enhancing the accuracy and adaptability of news dissemination feature extraction. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0018] Figure 1 It is a flowchart of a news dissemination method based on big data.
[0019] Figure 2 It is a flowchart of a big data news dissemination architecture.
[0020] Figure 3Schematic diagram for dynamic hypergraph modeling.
[0021] Figure 4 Flowchart of feature aggregation for hypergraph neural network. Detailed implementation manners
[0022] To make the above objects, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific implementation manners of the present invention will be given in conjunction with the accompanying drawings of the specification.
[0023] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the spirit of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.
[0025] Refer to Figures 1 to 4 , this embodiment provides a news dissemination method based on big data, including the following steps: S1. Collect news data and perform preprocessing to generate a structured news data set.
[0026] Furthermore, the news data includes news text content, user behavior data, and social media interaction data, specifically as follows: Obtain news text content (news title, news body, news keywords, and news release time) from Xinhuanet; Obtain social media interaction data (post content, reposting records, comments, and topic tags) from Weibo; Obtain user behavior data (user identification, news identification, interaction type (click, like, comment, share), timestamp, and browsing duration) from news applications; Develop a crawler using the Scrapy framework, customize collection rules for each type of data in the news data (news text content, user behavior data, and social media interaction data), and ensure the acquisition of real-time news data. The collection rules are: the collection frequency is set to 1 second, and the 1-second frequency balances real-time performance and server load (set with reference to the upper limit of requests per second of the Weibo API); Configure 3 news storage topics (including news_content topic, social_interaction topic, and user_behavior topic) through Apache Kafka (a distributed message queue); Among them, the news_content topic stores news text content; The social_interaction topic stores social media interaction data; The user_behavior topic stores user behavior data; The Scrapy crawler writes the collected news data into the 3 news storage topics in Apache Kafka. The Apache Kafka consumer reads the news data from the 3 news storage topics and stores it in the MongoDB distributed database. Apache Kafka is a distributed message queue used for efficient data transmission, divided into producer and consumer roles. The Scrapy crawler acts as the producer and writes the collected news data into the 3 news storage topics in Apache Kafka. The consumer is a pre-configured application (for example, an Apache Kafka client implemented in Python) responsible for reading the news data from the 3 news storage topics and storing it in the MongoDB distributed database.
[0027] Perform preprocessing tasks on news text content, user behavior data, and social media interaction data respectively, as follows: Use the hashlib hash calculation library in Python to calculate the MD5 hash value of the news text content, compare the hash value of the current news text content with the stored hash value set (stored in the MongoDB distributed database), and remove duplicate news text content records. The repetition rate is set to 5% (based on the reprint ratio of news aggregation platforms such as Xinhua News Agency); Encode news titles, news texts, post contents, and comments to UTF-8 to ensure consistent character formats and compatibility with multilingual characters; Set default values for empty fields in news text content, social media interaction data, and user behavior data: for example, fill in "unknown" if the news keyword is empty, fill in "0" if the comment is empty, and fill in "0.00" if the browsing duration is empty; Use the re library in Python to remove HTML tags from news titles, news texts, post contents, and comments, and filter special characters (such as emojis) to generate pure text data; Using the TF-IDF algorithm of the scikit-learn machine learning library, we extract the top 10 keywords with high TF-IDF values from news titles and news texts (the top 10 are based on information retrieval standards, and the top 10 are sorted in descending order of TF-IDF values) to generate news content feature vectors (300 dimensions), as follows: Get the cleaned news title and news body (from text cleaning) from the MongoDB distributed database; Use scikit-learn's TfidfVectorizer (TF-IDF vectorization tool) to merge the news title and news body into a document, calculate the TF-IDF value, and select the top 10 keywords with high TF-IDF values; The keywords are mapped into a 300-dimensional sparse vector, and the first 10 dimensions are filled with the TF-IDF values of the keywords at fixed positions, and the remaining 290 dimensions are zero to generate the news content feature vector.
[0028] The news title and news body are encoded through the pre-trained BERT (language) model to form a news content feature vector (300 dimensions), as follows: Get news title and news body from MongoDB distributed database; Merge the news title and news body into a complete text by string concatenation to capture semantic information; Use the built-in tokenizer of the BERT model to tokenize the complete text (for example, "AI", "technology", "breakthrough") and generate a token ID sequence and attention mask. The token ID sequence and attention mask are input into the encoder of the BERT model, and a 768-dimensional CLS vector is generated through forward propagation to capture the overall semantics of the news title and news body. Through the fully connected layer, linear transformation is applied to the 768-dimensional CLS vector to generate the news content feature vector.
[0029] The news content feature vectors generated by the TF-IDF algorithm and the BERT model are merged into the final news content feature vector through weighted averaging (weight is 0.7:0.3).
[0030] Collect the user's recent click records (such as 100 times), use the pre-trained text classification model (based on the BERT model, classified into 10 categories of user interest distribution such as technology, sports, and entertainment), and generate a user feature vector (128 dimensions), as follows: Get the news title and news text of the click record, use the BERT model encoder to generate a 768-dimensional CLS vector; The 768-dimensional CLS vector is input into the fully connected layer, and a 10-dimensional probability vector is output through softmax; Use NumPy (a numerical computing library) to average the 10-dimensional probability vectors of 100 clicks to generate a 10-dimensional user interest distribution; Perform a linear transformation on the user interest distribution through a fully connected layer to generate a user feature vector.
[0031] Use the Word2Vec (word embedding) model to tokenize the topic tags of social media interaction data to generate topic feature vectors (64-dimensional), as follows: Obtain the topic tags in the social media interaction data from the MongoDB distributed database; Use the jieba tokenization tool to tokenize the topic tags (e.g., "AI technology" is tokenized into "AI" and "technology"); The input layer defines the tokenization of the topic tags as one-hot vectors; The hidden layer projects the one-hot vectors into a 64-dimensional space through an embedding matrix to generate 64-dimensional hidden layer vectors; The output layer uses context words to predict the context word probability distribution of the 64-dimensional hidden layer vectors, and optimizes the embedding matrix through backpropagation to generate 64-dimensional word embedding vectors; Load the pre-trained Word2Vec (word embedding) model; If the topic tag contains a single word, load the pre-trained Word2Vec model and use the 64-dimensional word embedding vector as the topic feature vector; If the topic tag contains multiple words (such as "AI" and "technology"), load the pre-trained Word2Vec model and take the average of the 64-dimensional word embedding vectors to generate topic feature vectors.
[0032] Apply the differential privacy algorithm to the user behavior data to protect the privacy of the user behavior data, as follows: Obtain the user behavior data from the MongoDB distributed database; Normalize the click records of the user behavior data (min-max normalization) to generate standardized user behavior data; Use the diffprivlib differential privacy library in Python to generate Laplace noise (privacy parameter ε = 0.1, which controls the privacy protection intensity of the click records of the user behavior data. A smaller value provides stronger protection, and the sensitivity is 1.0. ε and the sensitivity are set with reference to the differential privacy standard); Add the Laplace noise to the click records of the standardized user behavior data (e.g., the normalized click count 0.5 becomes 0.5 + noise); Use SHA-256 hashing to anonymize the user identifiers of the user behavior data to generate irreversible hashed user identifiers; Output anonymized user behavior data (including hashed user identifiers and noisy click records), and store it in the MongoDB distributed database.
[0033] Use PySpark (a distributed computing framework) to load news text content, social media interaction data, and user behavior data from the MongoDB distributed database, generate a PySpark DataFrame, and perform preprocessing tasks (SHA-256 hashing, BERT classification, TF-IDF, Word2Vec inference) in parallel through PySpark. Aggregate and generate a structured news dataset (including hashed user identifiers, user feature vectors, interaction records, news content feature vectors, and topic feature vectors), and write it to the MongoDB distributed database. Among them, a PySpark DataFrame is a distributed data structure in PySpark, which includes hashed user identifiers, user feature vectors, interaction records (hashed user identifiers, news identifiers, interaction types: click, like, comment, share, timestamp), news content feature vectors, and topic feature vectors.
[0034] S2. Construct a dynamic hypergraph to perform structured modeling on the structured news dataset. Through the connection of nodes and hyperedges, capture the behavioral intensity association between users and news and the semantic association between news and topics. Use Apache Kafka to dynamically update nodes and hyperedges and remove expired interaction records.
[0035] Furthermore, read the structured news dataset from the MongoDB distributed database; Based on the structured news dataset, construct a dynamic hypergraph; Define the node types of the dynamic hypergraph as user nodes, news nodes, and topic nodes; Associate the user nodes with the hashed user identifiers in the structured news dataset; Associate the news nodes with the news identifiers in the structured news dataset; Associate the topic nodes with the topic labels in the structured news dataset; Match the node attributes through PySpark join. Match the hashed user identifiers with the user feature vectors, the news identifiers with the news content feature vectors, and the topic labels with the topic feature vectors. Use the matched feature vectors (user feature vectors, news content feature vectors, and topic feature vectors) as node attributes; Initialize the graph object in NetworkX (a graph analysis library), write the user nodes, news nodes, and topic nodes into the graph object in NetworkX, and construct the node set of the hypergraph; The interaction records extracted from the PySpark DataFrame are used to create user-news hyperedges in the graph object in NetworkX. Each hyperedge connects a user node (hashed user identifier) to a news node (news identifier). The hyperedge weight is set according to the interaction type (click: 0.5, like: 1.0, comment: 1.5, share: 2.0. The hyperedge weights are predefined and reflect the interaction intensity). If there is already a hyperedge for the same user-news pair or there are multiple new interaction records, the PySpark groupBy operation is used to group by user and news, and the sum operation is used to add up the weight values of multiple interactions (the weight value refers to the numerical value mapped according to the interaction type for each interaction (e.g., like is 1, comment is 2, share is 3)) to update the hyperedge weight, thus generating or updating a single user-news hyperedge. The hyperedge weight represents the total interaction intensity of the user with the news; The 300-dimensional news content feature vector is reduced to a 64-dimensional news content feature vector through the fully connected layer of the BERT model; The cosine similarity between the reduced news content feature vector and the topic feature vector generated by the Word2Vec model is calculated using the general computing methods in Python (standard mathematical operation functions built into Python, such as addition, multiplication, division, and square root) to generate the semantic similarity for each pair of news-topic; If the cosine similarity of the news-topic is greater than 0.7 (a commonly used similarity threshold in the field of information retrieval, indicating high semantic relevance), a news-topic hyperedge is created in the graph object in NetworkX, connecting the news node and the topic node, with the cosine similarity as the news-topic hyperedge weight, representing the semantic association intensity; The user nodes, news nodes, topic nodes, user-news hyperedges, and news-topic hyperedges are stored in the initialized graph object in NetworkX to form a dynamic hypergraph; Using the Apache Kafka consumer, subscribe to the news_content, social_interaction, and user_behavior topics to obtain new news data; Repeat the preprocessing tasks to generate new hashed user identifiers, user feature vectors, news content feature vectors, and topic feature vectors, and update the structured news dataset; If a new hashed user identifier is added to the updated structured news dataset, a user node is added to the graph object in NetworkX, with the new hashed user identifier as the unique identifier, and the generated user feature vector is bound; If a new news identifier is added to the updated structured news dataset, a news node is added to the graph object in NetworkX, with the new news identifier as the unique identifier, and the generated news content feature vector is bound; If new topic tags are added to the updated structured news dataset, topic nodes are added to the graph object in NetworkX, with the new topic tags as the unique identifiers, and the generated topic feature vectors are bound; Based on the new interaction records in the updated structured news dataset, the hyperedge weights are updated to generate or update a single user-news hyperedge; For the new news text content in the updated structured news dataset, use the fully connected layer of the BERT model to reduce the original 300-dimensional news content feature vector to a 64-dimensional news content feature vector, and calculate the cosine similarity between the reduced news content feature vector and the topic feature vector generated by the Word2Vec model. If the cosine similarity is greater than 0.7 (the threshold in the field of information retrieval), a news-topic hyperedge is added to the graph object in NetworkX, connecting the news node and the topic node, with the cosine similarity as the news-topic hyperedge weight; Use PySpark for incremental processing of the updated structured news dataset, and repeat the steps of adding new nodes and new hyperedges to update the graph object in NetworkX; Set a 7-day time window (a common period to balance news timeliness and data volume), and remove interaction records with timestamps exceeding 7 days from the interaction records; Based on the removed interaction records, delete the corresponding user-news hyperedges from the graph object in NetworkX; If a news node or a topic node is no longer connected to other nodes because all user-news hyperedges and news-topic hyperedges are removed, it becomes an isolated node, and the isolated news node or topic node is deleted from the graph object in NetworkX; An isolated node refers to a news node or a topic node that is no longer connected to user nodes, news nodes, and topic nodes in the dynamic hypergraph through user-news hyperedges and news-topic hyperedges; Use the scheduled task of PySpark (such as executing daily) to remove interaction records, corresponding user-news hyperedges, and isolated nodes that exceed 7 days, and regularly update the graph object in NetworkX; Export the graph object in NetworkX to JSON format, including node data and hyperedge data, and store it in the MongoDB distributed database as a graph structure data collection; Use NetworkX (a graph analysis library) to count the number of nodes and hyperedges in the graph object in NetworkX, perform graph analysis using the NetworkX graph object, calculate the average hyperedge weight, and generate metadata on the hypergraph scale and connection strength (the number of nodes to the number of hyperedges); Store the generated metadata into the metadata collection of the MongoDB distributed database for monitoring the dynamic hypergraph scale and connection quality.
[0036] S3. Use the hypergraph neural network model to perform multi-layer aggregation and non-linear transformation on the node features of the updated dynamic hypergraph to generate news propagation feature vectors.
[0037] Furthermore, read the graph structure data collection of the dynamic hypergraph from the MongoDB distributed database; Load the pre-trained hypergraph neural network model, set 3 convolutional layers, and each convolutional layer aggregates node and neighborhood information through hyperedge weighting; The input of the hypergraph neural network model is the node feature vector, hyperedge connection relationship, and hyperedge weight as the convolutional layer aggregation weight; Through the fully connected layer of the BERT model, take the user feature vector as the input, perform matrix multiplication and ReLU activation, and reduce the dimension to a 64-dimensional user feature vector, which is aligned with the topic feature vector dimension of the topic node; Through the fully connected layer of the BERT model, take the news content feature vector as the input, perform matrix multiplication and ReLU activation, and reduce the dimension to a 64-dimensional news content feature vector, which is aligned with the topic feature vector dimension of the topic node; Weightedly fuse the topic feature vector of the non-dimension-reduced topic node, the dimension-reduced user feature vector, and the news content feature vector dimensions to form a unified node feature vector, which is used as the input of the hypergraph neural network model; among them, the topic feature vector of the topic node remains unchanged and is consistent with the dimension-reduced user feature vector and news content feature vector dimensions; The first convolutional layer finds neighborhood nodes based on user-news hyperedges and news-topic hyperedges and weightedly aggregates the unified-dimensional node feature vectors of neighborhood nodes (the neighborhood nodes are also nodes in the dynamic hypergraph, so they also have the above-generated unified node feature vectors), generates an initial aggregated feature vector, applies ReLU activation to the initial aggregated feature vector for non-linear transformation, generates an initial enhanced feature vector to enhance the expression ability of the initial aggregated feature vector, and then applies Dropout to prevent the hypergraph neural network model from overfitting, generating an intermediate feature vector; The second convolutional layer finds neighborhood nodes based on user-news hyperedges and news-topic hyperedges, and weightedly aggregates the 64-dimensional intermediate feature vectors of neighborhood nodes, generates an intermediate aggregated feature vector, applies ReLU activation to the intermediate aggregated feature vector for non-linear transformation, generates an intermediate enhanced feature vector to enhance the expression ability of the second-layer aggregated feature vector, and then applies Dropout to prevent the hypergraph neural network model from overfitting, generating an updated intermediate feature vector; The third convolutional layer finds neighborhood nodes based on user-news hyperedges and news-topic hyperedges, and weighted aggregates the updated 64-dimensional intermediate feature vectors of the neighborhood nodes to generate the third-layer aggregated feature vector. Then, the ReLU activation is applied to the third-layer aggregated feature vector for non-linear transformation to enhance the expressive ability of the third-layer aggregated feature vector. There is no Dropout processing, and the final news propagation feature vector (i.e., the final aggregated feature vector) is generated; Add a time dimension to the news propagation feature vector through the timestamp of the interaction record to generate a news propagation feature vector with time information; Store the news propagation feature vector with time information in the MongoDB distributed database as a news propagation feature vector set, which is updated every hour, consistent with the update frequency of the dynamic hypergraph; Use an Apache Kafka consumer to subscribe to the news_content topic, social_interaction topic, and user_behavior topic to obtain new news data (generate new user nodes, news nodes, topic nodes, user-news hyperedges, and news-topic hyperedges based on news text content, social media interaction data, and user behavior data); If the dynamic hypergraph is updated due to new news data, repeat the steps of updating the news propagation feature vector with time information; If the dynamic hypergraph removes expired interaction records, repeat the steps of deleting the news nodes, topic nodes, and user-news and news-topic hyperedges of the dynamic hypergraph, and remove the corresponding news propagation feature vectors with time information to maintain the data consistency between the news propagation feature vector set and the dynamic hypergraph; Use a PySpark scheduled task (executed daily) to check the news propagation feature vector set (including news propagation feature vectors with time information) in the MongoDB distributed database, and remove the news propagation feature vectors in the news propagation feature vector set whose timestamps exceed 7 days to be consistent with the timeliness of the dynamic hypergraph; Use the general computing method of Python to count the scale of the news propagation feature vector set (the number of news propagation feature vectors of news nodes) and the average norm of the news propagation feature vectors (take the average after calculating the Euclidean norm of the news propagation feature vectors of news nodes), generate the metadata of the news propagation feature vectors, and store them in the metadata set of the MongoDB distributed database; The output is a news propagation feature vector with time information, stored in the MongoDB distributed database, and transmitted to the next step through Apache Kafka for generating an optimized propagation strategy; It should be noted that the hypergraph neural network model is an existing model and a well-known model in the field. Its basic architecture is implemented through multiple layers of convolutions. Each layer of convolution aggregates the node feature vectors based on the hyperedge connection relationship to generate more expressive feature representations, and is widely used in the field of propagation analysis.
[0038] S4. Analyze the spatio-temporal dynamic characteristics of the news propagation feature vectors through the deep learning model to generate an optimized propagation strategy.
[0039] Furthermore, receive the news propagation feature vectors with time information from Apache Kafka; Load the news propagation feature vectors with time information into the pre-trained deep learning model, which includes an attention mechanism layer, an LSTM layer, and a fully connected layer, for analyzing the spatio-temporal dynamic characteristics of the news propagation feature vectors; The attention mechanism layer inputs the news propagation feature vectors into the fully connected layer to generate attention scores through matrix multiplication; then, normalize the attention scores through the Softmax function to generate attention weights (indicating the relative importance of each dimension in the news propagation feature vectors to the propagation strategy, ranging from 0 to 1, and the sum is 1); then, multiply the news propagation feature vectors element-wise with the attention weights to generate the weighted news propagation feature vectors; The LSTM layer sets the hidden dimension to 128 and the time step to 24 (indicating the propagation dynamics within 24 hours, based on the propagation cycle); then, use the weighted news propagation feature vectors as input, perform time step processing using the time stamps (time series within 24 hours), and update the hidden state of the LSTM layer through the gating mechanism of the LSTM (input gate, forget gate, output gate) to generate the time series feature vectors (with a dimension of 128, indicating the spatio-temporal dynamic characteristics of the news propagation within 24 hours); The fully connected layer receives the time series feature vectors and maps them to a 3-dimensional propagation strategy vector through matrix multiplication (indicating the propagation priority, propagation channel, and propagation frequency, with a total of 3 dimensions); then, normalize the 3-dimensional propagation strategy vector through the Softmax function (ranging from 0 to 1, and the sum is 1) to generate the final propagation strategy vector (indicating the normalized propagation strategy); Associate the propagation strategy vector with the news identifier through index matching to generate an optimized propagation strategy, in the format of (news identifier, propagation priority, propagation channel, and propagation frequency), for guiding news propagation practice.
[0040] It should be noted that the attention mechanism layer involves the operation of the fully connected layer because the fully connected layer generates attention scores through matrix mapping to measure the importance of the dimensions of the news propagation feature vector. However, the core operation of the attention mechanism layer is not limited to the fully connected layer, but also includes normalization processing and weighted calculation, which together achieve the dimensionality weighting enhancement of the news propagation feature vector.
[0041] This embodiment also provides a computer device applicable to the situation of the news propagation method based on big data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the news propagation method based on big data as proposed in the above embodiment.
[0042] This computer device can be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of this computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of this computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of this computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0043] This embodiment also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the news propagation method based on big data as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0044] In summary, the present invention constructs a dynamic hypergraph based on a structured news dataset, defines user nodes, news nodes, and topic nodes, creates user-news hyperedges and news-topic hyperedges by using interaction records and cosine similarity, and dynamically updates nodes and hyperedges and removes expired interaction records through Apache Kafka, realizing dynamic structured modeling of user-news behavior intensity association and news-topic semantic association; capturing multi-dimensional complex relationships in the form of a hypergraph and maintaining real-time performance, being applicable to analyzing the dynamic changes of user behavior and news semantics, and finally achieving the effect of improving the comprehensiveness and timeliness of relationship modeling, and enhancing the accuracy and adaptability of news dissemination feature extraction.
[0045] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A news dissemination method based on big data, characterized in that: Including, collecting news data, preprocessing it to generate a structured news dataset, where the structured news dataset includes hashed user identifiers, user feature vectors, interaction records, news content feature vectors, and topic feature vectors; constructing a dynamic hypergraph to perform structured modeling on the structured news dataset, capturing the behavioral intensity association between users and news and the semantic association between news and topics through the connection of nodes and hyperedges, and dynamically updating nodes and hyperedges using Apache Kafka and removing expired interaction records; using a hypergraph neural network model to perform multi-layer aggregation and non-linear transformation on the node features of the updated dynamic hypergraph to generate news propagation feature vectors; analyzing the spatio-temporal dynamic characteristics of the news propagation feature vectors through a deep learning model to generate an optimized propagation strategy.
2. The news dissemination method based on big data according to claim 1, characterized in that: The constructing of a dynamic hypergraph to perform structured modeling on the structured news dataset means reading the structured news data from the structured news dataset, defining user nodes, news nodes, and topic nodes, and creating user-news hyperedges and news-topic hyperedges to form a dynamic hypergraph.
3. The news dissemination method based on big data according to claim 1, characterized in that: The capturing of the behavioral intensity association between users and news means extracting interaction records from the structured news data, creating user-news hyperedges to connect user nodes and news nodes, setting hyperedge weights according to the interaction type, and if there are multiple interaction records between a user and a news, updating the weight of a single user-news hyperedge by aggregating and accumulating the interaction weights. The semantic association between news and topics means obtaining the news content feature vector and the topic feature vector from the structured news data, reducing the dimension of the news content feature vector using the fully connected layer of the BERT model, calculating the cosine similarity between the reduced-dimensional news content feature vector and the topic feature vector using the built-in arithmetic functions in Python, and if the cosine similarity is greater than the similarity threshold, creating a news-topic hyperedge to connect the news node and the topic node, and using the cosine similarity as the weight of the news-topic hyperedge to reflect the semantic association between the news and the topic.
4. The news dissemination method based on big data according to claim 1, characterized in that: The using of Apache Kafka to dynamically update nodes and hyperedges means obtaining new structured news data through Apache Kafka, updating the structured news dataset, adding user nodes, news nodes, and topic nodes, updating the weights of user-news hyperedges based on new interaction records, calculating the cosine similarity based on the news content feature vector and the topic feature vector to update the news-topic hyperedges, and forming an updated dynamic hypergraph.
5. The news dissemination method based on big data according to claim 1, wherein: The using of a hypergraph neural network model to perform multi-layer aggregation on the node features of the updated dynamic hypergraph, the specific steps are as follows: loading the user feature vector of the user node, the news content feature vector of the news node, and the topic feature vector of the topic node from the dynamic hypergraph; reducing the dimensions of the user feature vector and the news content feature vector through the fully connected layer of the BERT model; weightedly fusing the reduced-dimensional user feature vector and news content feature vector with the non-reduced topic feature vector to form a unified node feature vector; based on the dynamic hypergraph, constructing a hypergraph neural network model and setting multi-layer convolutional layers; The first layer forms an initial aggregated feature vector by weighted aggregating the neighborhood node's unified node feature vectors through user-news hyperedges and news-topic hyperedges; The second layer forms an intermediate aggregated feature vector by weighted aggregating the initial aggregated feature vectors of the neighborhood nodes through user-news hyperedges and news-topic hyperedges; The third layer forms a final aggregated feature vector by weighted aggregating the intermediate aggregated feature vectors of the neighborhood nodes through user-news hyperedges and news-topic hyperedges.
6. The news dissemination method based on big data according to claim 1, characterized in that: For the non-linear transformation, the specific steps are as follows. Apply the ReLU activation function to the initial aggregated feature vector generated by the first layer of convolution for non-linear mapping to generate an initial enhanced feature vector; Apply the ReLU activation function to the intermediate aggregated feature vector generated by the second layer of convolution for non-linear mapping to generate an intermediate enhanced feature vector; Apply the ReLU activation function to the final aggregated feature vector generated by the third layer of convolution for non-linear mapping to generate a news propagation feature vector.
7. The news dissemination method based on big data according to claim 1, wherein: For analyzing the spatio-temporal dynamic characteristics of the news propagation feature vector by the deep learning model, the specific steps are as follows. Obtain the news propagation feature vector from Apache Kafka; Process the news propagation feature vector through the attention mechanism layer of the deep learning model to generate attention weights; Adjust the news propagation feature vector based on the attention weights to form a weighted news propagation feature vector; Perform time-step processing on the weighted news propagation feature vector through the LSTM layer of the deep learning model to capture the propagation dynamic characteristics of the weighted news propagation feature vector; Use the gating mechanism of the LSTM layer to update the hidden state of the LSTM layer to generate a spatio-temporal dynamic characteristic representation of the news propagation feature vector.
8. The news dissemination method based on big data according to claim 1, characterized in that: The optimized propagation strategy includes news identification, propagation priority, propagation channels, and propagation frequency.
Citation Information
Patent Citations
News recommendation method fusing spatio-temporal characteristics
CN115618107A
Contrast learning news recommendation method based on hypergraph enhancement
CN117033763A
Social platform recommendation method based on hypergraph reconstruction
CN117370782A
Information propagation prediction method of sequence hypergraph neural network based on common attention fusion
CN118364185A
Online public opinion propagation and offline event association semantic generation method and system
CN118536510A