Method and system for realizing intelligent search and display of Internet important information
By generating information feature vectors through semantic segmentation and noise removal, a dual-path recommendation model and information analysis network are constructed, solving the problems of capturing the semantic connotation of information and tracking the propagation rules in existing technologies, and realizing personalized recommendation and dynamic information propagation analysis.
Patent Information
- Application Number
- CN202511472303.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2025-11-14
AI Technical Summary
Existing internet information search and recommendation systems struggle to accurately capture the semantic meaning and importance of information, lack effective feature fusion mechanisms, fail to fully grasp the complex relationship between user interests and information content, and lack the ability to dynamically track and predict the patterns and trends of information dissemination.
Information feature vectors are generated through semantic segmentation and noise removal. Information entropy and gain ratio are calculated to determine information weight distribution. A dual-path recommendation model is constructed for feature learning. An information analysis network is built using conditional random fields and probabilistic graphical models to generate an information evolution path graph for dynamic tracking.
It enables personalized information recommendations, improves the accuracy and efficiency of information processing, dynamically tracks the laws and evolution trends of information dissemination, and provides users with comprehensive information dissemination links and development trend analysis.
Smart Images

Figure CN120950771A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to Internet technology, and in particular to a method and system for intelligently searching and displaying important information on the Internet. Background Technology
[0002] With the explosive growth of internet information, users often struggle to efficiently obtain and filter important information amidst the massive amount of data. Intelligent internet information search and display technologies aim to achieve intelligent identification, accurate recommendation, and dynamic tracking of important information by analyzing and mining information content and user behavior characteristics. Traditional information search and recommendation systems, relying primarily on keyword matching and simple user behavior statistics, are no longer sufficient to meet users' demands for high-quality, personalized information services.
[0003] Existing technologies are relatively simple in terms of information feature extraction and weight allocation, making it difficult to accurately capture the semantic connotation and importance of information. Most systems rely only on surface text features and simple statistical models, failing to deeply understand the core value and semantic structure of information, resulting in low-quality information annotation and affecting the accuracy of subsequent recommendations.
[0004] Current recommendation systems typically consider only user interests or content features, lacking effective feature fusion mechanisms. This fragmented feature learning approach struggles to fully grasp the complex relationship between user interests and information content, often resulting in homogenized and simplistic recommendation outcomes that fail to meet users' diverse information needs.
[0005] Existing technologies have limited ability to analyze the laws and evolutionary trends of information dissemination, lacking dynamic tracking and prediction capabilities. Most systems focus only on static information recommendations, ignoring the propagation paths and evolutionary trends of information within the network. This fails to provide users with a comprehensive overview and in-depth insights into information development, making it difficult for users to grasp the value and direction of information. Summary of the Invention
[0006] The embodiments of the present invention provide a method and system for intelligently searching and displaying important information on the Internet, which can solve the problems in the prior art.
[0007] A first aspect of this invention provides a method for intelligently searching and displaying important information on the Internet, comprising:
[0008] Acquire Internet information, perform semantic segmentation and noise removal on the Internet information, generate information feature vectors, calculate information entropy value and information gain ratio on the information feature vectors, determine information weight distribution based on the information entropy value and the information gain ratio, and obtain information labeling data based on the information weight distribution;
[0009] A dual-path recommendation model is constructed based on the information-annotated data. The dual-path recommendation model performs feature learning from the user interest dimension and the information content dimension, respectively, and integrates the user behavior features and content semantic features in the information-annotated data to output personalized recommendation results.
[0010] Based on the personalized recommendation results and the information annotation data, an information analysis network is constructed using conditional random fields and probabilistic graphical models, and the information propagation rules in the personalized recommendation results are calculated based on the information analysis network.
[0011] Based on the aforementioned information dissemination patterns, an information evolution path diagram is generated. The personalized recommendation results are then dynamically tracked using this information evolution path diagram, generating a tracking analysis report that includes the dissemination links and evolution trends.
[0012] The information entropy and information gain ratio are calculated for the information feature vector. The information weight distribution is determined based on the information entropy and information gain ratio. Information annotation data is obtained based on the information weight distribution, including:
[0013] A multidimensional feature tensor is constructed from the information feature vector, and the information entropy value of each dimension is calculated based on the multidimensional feature tensor. Dynamic weights are calculated based on the information entropy values of each dimension, and the dynamic weight distribution is obtained through iterative optimization with information classification accuracy as the evaluation index. The multidimensional feature tensor is reconstructed and labeled based on the dynamic weight distribution to obtain information labeled data.
[0014] A dual-path recommendation model is constructed based on the labeled information data. This model performs feature learning from both the user interest dimension and the information content dimension, fusing user behavior features and content semantic features from the labeled information data to output personalized recommendation results, including:
[0015] Based on the information-annotated data, user behavior data and content semantic data are mapped to a low-dimensional space through feature dimensionality reduction. The data in the low-dimensional space is used to construct a dual-path recommendation model. Based on the dual-path recommendation model, the influence of the upper-level node on the current node is calculated iteratively through gradient descent to obtain the influence strength between nodes.
[0016] Based on the influence intensity, adaptive clustering is performed on the user behavior data in the user interest dimension, and user behavior data with a cluster centrality higher than a preset influence threshold is taken as priority behavior data.
[0017] The content semantic data is semantically embedded to obtain content features. User behavior weights are calculated based on the influence intensity. The priority behavior data and the content features are weighted and combined according to the user behavior weights to generate personalized recommendation results.
[0018] Based on the personalized recommendation results and the information annotation data, an information analysis network is constructed using conditional random fields and probabilistic graphical models. The information propagation patterns in the personalized recommendation results are calculated using this information analysis network, including:
[0019] A conditional random field correlation mapping matrix is constructed for the user interaction behavior sequence and the item feature sequence. The value of each element in the correlation mapping matrix is set as the inner product operation result of the corresponding feature vector. The conditional transition probability of the item feature sequence relative to the user interaction behavior sequence is calculated based on the inner product operation result.
[0020] A hierarchical dynamic attention network is constructed based on the conditional transition probability, and the propagation intensity is assigned to the neighbor set of the recommended item node. Based on the propagation intensity assignment, an exponential weight function is used to perform nonlinear transformation and robustness enhancement on the conditional transition probability. Based on the result of the nonlinear transformation, a set of propagation relationship edges between nodes is constructed.
[0021] The conditional transition probability and the propagation relationship edge set are combined by weighting to construct an information analysis network. The propagation paths in the information analysis network are scored, and the information propagation rules in the personalized recommendation results are characterized based on the score calculation results.
[0022] A hierarchical dynamic attention network is constructed based on the conditional transition probabilities, and propagation intensity is assigned to the neighbor set of the recommended item node. Based on the propagation intensity assignment, an exponential weight function is used to perform a nonlinear transformation and robustness enhancement on the conditional transition probabilities. Based on the result of the nonlinear transformation, a set of propagation relationship edges between nodes is constructed, including:
[0023] A hierarchical dynamic attention network is constructed using the conditional transition probability as the baseline probability. In the hierarchical dynamic attention network, a robustness enhancement is achieved by using a positive sample propagation path, weak negative samples with randomly replaced nodes, and strong negative samples generated adversarially, resulting in the optimized conditional transition probability.
[0024] The propagation intensity is assigned to the neighbor set based on the optimized conditional transition probability, and the propagation intensity is obtained by adaptively fusing features at different levels through the hierarchical dynamic attention network.
[0025] An exponential weighting function is used to perform a nonlinear transformation on the propagation intensity, and an adaptive threshold is set based on the node feature vector and global context information to add random noise to the nonlinearly transformed propagation intensity.
[0026] Consistency regularization and adversarial training are performed on the propagation intensity with added random noise to obtain the final nonlinear transformation result. Based on the nonlinear transformation result, a set of propagation relationship edges between the recommended item node and its neighboring nodes is constructed.
[0027] Based on the aforementioned information dissemination patterns, an information evolution path diagram is generated. This diagram is then used to dynamically track the personalized recommendation results, generating a tracking and analysis report that includes the dissemination links and evolutionary trends.
[0028] The propagation data of personalized recommendation results during the propagation process is obtained, and the propagation data is subjected to hierarchical feature extraction using a deep feature extraction network to obtain a fused feature matrix;
[0029] The propagation intensity value is calculated based on the fusion feature matrix. The propagation intensity value is combined with the attention weight of the propagation node. The attention weight is determined by the temporal progression relationship of the historical propagation influence of the node. An information evolution path graph is constructed by the propagation intensity value and the attention weight.
[0030] The propagation nodes in the information evolution path graph are dynamically tracked, the tracking frequency is adaptively adjusted based on changes in the propagation environment, the propagation state changes of the nodes are calculated based on the propagation intensity value and the attention weight, and the edge weights between the propagation nodes in the information evolution path graph are updated based on the propagation state changes.
[0031] Multi-level analysis is performed on the updated edge weights to identify propagation association patterns between different levels and to mine the propagation links of the personalized recommendation results.
[0032] The analysis results of the propagation link are mapped based on the knowledge graph of the propagation domain. The propagation trend perception of the personalized recommendation results is predicted through graph reasoning, and a tracking analysis report containing the propagation link and an interpretable propagation evolution trend is generated.
[0033] Dynamically track the propagation nodes in the information evolution path graph, adaptively adjust the tracking frequency based on changes in the propagation environment, and calculate the propagation state changes of the nodes according to the propagation intensity value and the attention weight, including:
[0034] The propagation nodes in the information evolution path graph are dynamically tracked, the propagation probability between the propagation node and its neighboring nodes is calculated, the temporal characteristics of the propagation node are modeled based on the propagation probability to obtain the propagation probability evolution sequence, and the seepage entropy index of the propagation node is calculated based on the propagation probability evolution sequence.
[0035] The tracking frequency is adaptively adjusted based on changes in the propagation environment. The seepage entropy index is used as an environmental state feature, the rate of change of the environmental state feature is calculated, the calculation period is adjusted according to the rate of change, the calculation period is used as the tracking frequency, and the propagation state change of the node is calculated according to the propagation intensity value and the attention weight.
[0036] A second aspect of the present invention provides a system for intelligently searching and displaying important information on the Internet, comprising:
[0037] The first unit is used to acquire Internet information, perform semantic segmentation and noise removal on the Internet information, generate information feature vectors, calculate information entropy and information gain ratio on the information feature vectors, determine information weight distribution based on the information entropy and information gain ratio, and obtain information labeling data based on the information weight distribution.
[0038] The second unit is used to construct a dual-path recommendation model based on the information annotation data. The dual-path recommendation model performs feature learning from the user interest dimension and the information content dimension, respectively, and integrates the user behavior features and content semantic features in the information annotation data to output personalized recommendation results.
[0039] The third unit is used to construct an information analysis network based on the personalized recommendation results and the information annotation data, using conditional random fields and probabilistic graphical models, and to calculate the information propagation rules in the personalized recommendation results based on the information analysis network.
[0040] The fourth unit is used to generate an information evolution path diagram based on the information dissemination pattern, use the information evolution path diagram to dynamically track the personalized recommendation results, and generate a tracking analysis report containing the dissemination links and evolution trends.
[0041] A third aspect of the present invention provides an electronic device, comprising:
[0042] processor;
[0043] Memory used to store processor-executable instructions;
[0044] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0045] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0046] The beneficial effects of this application are as follows:
[0047] This invention generates high-quality information annotation data by semantically segmenting, removing noise, and determining information weight distribution of Internet information, which effectively improves the accuracy and efficiency of information processing and enables important information to be accurately identified and extracted.
[0048] By constructing a dual-path recommendation model that integrates user interest and information content dimensions, personalized information recommendation is achieved, meeting the differentiated needs of different users, while improving user experience and information acquisition efficiency, effectively solving the limitations of traditional recommendation systems with single-dimensional recommendations.
[0049] This invention utilizes conditional random fields and probabilistic graphical models to construct an information analysis network and generate an information evolution path diagram. It can dynamically track the laws and trends of information dissemination and provide users with comprehensive analysis of information dissemination links and development trends. This enhances the depth and breadth of information mining and is of great significance for the prediction, early warning, and decision support of important information. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating the method for intelligently searching and displaying important information on the Internet according to an embodiment of the present invention;
[0051] Figure 2 This is a flowchart illustrating the intelligent recommendation process based on multi-dimensional analysis, as described in an embodiment of the present invention.
[0052] Figure 3 This is an architecture diagram of an intelligent tracking system for network information propagation according to an embodiment of the present invention. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0055] Figure 1 This is a flowchart illustrating the method for intelligently searching and displaying important information on the Internet according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0056] Acquire Internet information, perform semantic segmentation and noise removal on the Internet information, generate information feature vectors, calculate information entropy value and information gain ratio on the information feature vectors, determine information weight distribution based on the information entropy value and the information gain ratio, and obtain information labeling data based on the information weight distribution;
[0057] A dual-path recommendation model is constructed based on the information-annotated data. The dual-path recommendation model performs feature learning from the user interest dimension and the information content dimension, respectively, and integrates the user behavior features and content semantic features in the information-annotated data to output personalized recommendation results.
[0058] Based on the personalized recommendation results and the information annotation data, an information analysis network is constructed using conditional random fields and probabilistic graphical models, and the information propagation rules in the personalized recommendation results are calculated based on the information analysis network.
[0059] Based on the aforementioned information dissemination patterns, an information evolution path diagram is generated. The personalized recommendation results are then dynamically tracked using this information evolution path diagram, generating a tracking analysis report that includes the dissemination links and evolution trends.
[0060] In one optional implementation, the information entropy value and information gain ratio are calculated for the information feature vector, the information weight distribution is determined based on the information entropy value and the information gain ratio, and the information annotation data is obtained based on the information weight distribution, including:
[0061] A multidimensional feature tensor is constructed from the information feature vector, and the information entropy value of each dimension is calculated based on the multidimensional feature tensor. Dynamic weights are calculated based on the information entropy values of each dimension, and the dynamic weight distribution is obtained through iterative optimization with information classification accuracy as the evaluation index. The multidimensional feature tensor is reconstructed and labeled based on the dynamic weight distribution to obtain information labeled data.
[0062] Obtain raw information data, which can be multimedia formats such as text, images, audio, or video. For example, for product review information on an e-commerce platform, raw data includes review text, rating, purchase time, user ID, and other information.
[0063] The acquired raw data undergoes preprocessing, which includes data cleaning, format conversion, and feature extraction. For text-based data, cleaning removes special characters, emojis, and redundant spaces; format conversion converts the text to a unified encoding; and feature extraction extracts features such as keywords, sentiment, and text length. For example, for a product review like "This phone takes great photos, but the battery life is just average," the extracted features include the keywords "phone," "photography," "battery," and "battery life," as well as the sentiment of "positive + negative."
[0064] After preprocessing, the processed data is converted into an information feature vector. The feature vector is a multi-dimensional array, with each dimension representing a feature. For example, for the above product reviews, a vector [0.2, 0.3, 0.1, 0.05, 0.7, 128] containing features such as word frequency statistics, sentiment score, and review length can be constructed. The first four values represent the normalized word frequency of the keywords, 0.7 represents the sentiment score (range 0-1), and 128 represents the character length.
[0065] Based on the constructed information feature vectors, a multidimensional feature tensor can be further constructed. A multidimensional feature tensor is a high-dimensional data structure composed of multiple feature vectors, which can more comprehensively express the relationships between information features. Taking product review analysis as an example, a three-dimensional tensor can be constructed. The first dimension represents different review samples, the second dimension represents different feature types (such as word frequency, sentiment, and length), and the third dimension represents specific feature values. For example, for 1000 review samples, 10 feature types are extracted from each review, and each type contains 20 specific features, resulting in a 1000×10×20 three-dimensional tensor.
[0066] The information entropy value of each dimension is calculated based on the constructed multidimensional feature tensor. Information entropy is used to measure the uncertainty of data; the higher the entropy value, the greater the uncertainty of the data. For discrete features, the data needs to be discretized first, and then the probability distribution of each value is calculated. For example, the comment length is discretized into four intervals: 0-50 characters, 51-100 characters, 101-200 characters, and more than 200 characters. The sample size ratio of each interval is calculated to be [0.15, 0.45, 0.3, 0.1]. Based on this, the information entropy value of this dimension is calculated to be 1.74. This process is repeated for each dimension in the tensor, and finally an array representing the information entropy of each dimension is obtained, such as [1.92, 1.74, 1.56, 2.01, 1.83].
[0067] Based on the calculated information entropy, the information gain ratio is further calculated. The information gain ratio takes into account the entropy value of the feature itself, which can avoid bias towards features with more values. For each feature dimension, the information gain of the feature on the classification result is calculated, and then divided by the entropy value of the feature itself to obtain the information gain ratio. For example, for the feature "comment length", assuming that its information gain on comment sentiment classification is 0.15 and its own entropy value is 1.74, then its information gain ratio is 0.15 / 1.74=0.086. Through similar calculations, an array of information gain ratios for all feature dimensions can be obtained, such as [0.102, 0.086, 0.134, 0.075, 0.118].
[0068] Based on the calculated information entropy values and information gain ratios of each dimension, dynamic weights are calculated using an iterative optimization approach, with information classification accuracy as the evaluation metric. Initially, each feature dimension is assigned an equal weight, such as 0.2 for all five dimensions. During iteration, the weight allocation is adjusted according to the information entropy value and information gain ratio of each dimension, with dimensions having lower information entropy values and higher information gain ratios receiving greater weights. For example, after the first iteration, the weights are adjusted to [0.22, 0.18, 0.25, 0.15, 0.2].
[0069] After each weight adjustment, the test data is classified using the new weight distribution, and the classification accuracy is calculated. The accuracy of different iterations is compared, and the weight distribution with the highest accuracy is selected as the final dynamic weight distribution. For example, after 10 iterations, the optimal weight distribution is [0.23, 0.17, 0.27, 0.13, 0.2], corresponding to a classification accuracy of 89.5%.
[0070] Based on the obtained dynamic weight distribution, the multidimensional feature tensor is reconstructed and labeled. The reconstruction process involves multiplying the original feature tensor by the weight distribution to obtain a weighted feature tensor. For example, for the feature vector [0.2, 0.3, 0.1, 0.05, 0.7], multiplying it by the weight distribution [0.23, 0.17, 0.27, 0.13, 0.2] yields the weighted feature vector [0.046, 0.051, 0.027, 0.0065, 0.14].
[0071] The labeling process determines the category label of information based on weighted feature vectors. Taking product review sentiment analysis as an example, rules can be set: a weighted sentiment score greater than 0.1 indicates a positive review, less than -0.1 indicates a negative review, and the rest are neutral reviews. Additionally, an importance score can be calculated using the weighted feature vectors to filter key information. For example, for a weighted feature vector [0.046, 0.051, 0.027, 0.0065, 0.14], the calculated importance score is 0.27, thus labeling the review as an "important positive review".
[0072] Through the above process, the final labeled data includes the original information, feature vectors, weight distribution, and labeling results. This labeled data can be used in subsequent applications such as information filtering, recommendation systems, or data analysis to improve the accuracy and efficiency of information processing.
[0073] In one optional implementation, a dual-path recommendation model is constructed based on the information-annotated data. The dual-path recommendation model performs feature learning from both the user interest dimension and the information content dimension, fusing user behavior features and content semantic features from the information-annotated data to output personalized recommendation results, including:
[0074] Based on the information-annotated data, user behavior data and content semantic data are mapped to a low-dimensional space through feature dimensionality reduction. The data in the low-dimensional space is used to construct a dual-path recommendation model. Based on the dual-path recommendation model, the influence of the upper-level node on the current node is calculated iteratively through gradient descent to obtain the influence strength between nodes.
[0075] Based on the influence intensity, adaptive clustering is performed on the user behavior data in the user interest dimension, and user behavior data with a cluster centrality higher than a preset influence threshold is taken as priority behavior data.
[0076] The content semantic data is semantically embedded to obtain content features. User behavior weights are calculated based on the influence intensity. The priority behavior data and the content features are weighted and combined according to the user behavior weights to generate personalized recommendation results.
[0077] like Figure 2 As shown, the method includes:
[0078] Principal component analysis (PCA) was used to map high-dimensional user behavior data and content semantic data to a unified low-dimensional feature space for feature reduction of the labeled data. User behavior data includes multi-dimensional information such as click records, browsing duration, purchase history, rating behavior, and sharing / forwarding, with an original feature dimension of 1024. PCA extracted the first 128 principal components, achieving a cumulative contribution rate of 95.2%, effectively preserving the main information features of the original data. For the behavior vector [45, 320, 8] containing user click counts, dwell time, and interaction depth, standardization yielded [0.75, 0.64, 0.53], which was then mapped to a 128-dimensional feature vector through PCA transformation. Content semantic data includes multi-modal information such as text content, image features, audio features, and tag information, with an original dimension of 2048. The same PCA method was used to reduce the dimension to 128, ensuring that user behavior data and content semantic data are processed in the same dimensional space.
[0079] The low-dimensional spatial data construction employs a dual-encoder architecture, establishing two parallel processing branches: a user interest path and a content semantic path. The user interest path encoder receives dimensionality-reduced user behavior features and performs feature extraction and representation learning through a three-layer fully connected neural network. The first layer has 256 neurons using the ReLU activation function, the second layer has 128 neurons, and the third layer outputs a 64-dimensional user interest representation vector. The content semantic path encoder receives dimensionality-reduced content features and performs feature learning using the same network structure, outputting a 64-dimensional content semantic representation vector. The two encoders share some network parameters, enhancing the consistency and complementarity of feature representations while maintaining their respective characteristics.
[0080] The dual-path recommendation model is constructed by fusing the outputs of the user interest path and the content semantic path. The model employs an attention mechanism to dynamically weight and combine the output features of the two paths, with the attention weights adaptively adjusted according to the characteristics of the current recommendation task. The 64-dimensional vector output from the user interest path and the 64-dimensional vector output from the content semantic path are concatenated to form a 128-dimensional fused feature vector. This fused feature vector is input to the final recommendation prediction layer, which uses a two-layer fully connected network: the first layer has 128 neurons, and the second layer outputs a recommendation score. During model training, the parameters of both paths are optimized simultaneously, and the network weights are updated using a backpropagation algorithm.
[0081] Gradient descent iterative calculation uses the backpropagation algorithm to calculate the influence of upper-layer nodes on the current node in the network. The calculation process proceeds layer by layer from the output layer to the input layer, and the gradient value of each node reflects its influence on the final recommendation result. For the first neuron in the recommendation prediction layer, its gradient value is 0.23, indicating the degree of influence of this neuron on the recommendation score. The gradient value of the 15th neuron in the fusion feature layer is 0.18, and the gradient value of the 8th neuron in the user interest path is 0.31. The gradient calculation adopts the chain rule, decomposing the gradient of the composite function into the product of the gradients of each layer. The learning rate is set to 0.001 in the iterative process. After 1000 iterations of training, the network converges to a stable state, and the influence strength of each layer node tends to stabilize.
[0082] The calculation of the influence strength between nodes is based on gradient propagation path analysis to determine the influence relationship of each node on other nodes. The influence strength equals the gradient value of the source node multiplied by the connection weight, and then multiplied by the sensitivity coefficient of the target node. The influence strength of the 8th neuron in the user interest path on the 15th neuron in the fusion layer is 0.31 multiplied by the connection weight 0.75 multiplied by the sensitivity coefficient 0.82, which equals 0.19. The sensitivity coefficient is calculated by applying a small perturbation to the input of the target node and observing the magnitude of the output change. The influence strength matrix records the influence relationship between all node pairs, with the matrix dimension being the product of the total number of nodes and the matrix elements being the corresponding influence strength values.
[0083] Adaptive clustering based on user interest dimensions groups user behavior data according to similarity based on influence strength. The clustering algorithm employs a density-based clustering method, using influence strength as a weighting factor for distance metrics. The weighted distance between user behavior data points is equal to the Euclidean distance multiplied by the inverse of influence strength; data points with greater influence strength are closer together and more likely to cluster in the same category. The clustering process sets a minimum cluster size of 5 and a search radius of 0.3, resulting in 12 user behavior clusters. For each cluster, a centrality metric is calculated. Centrality is equal to the inverse of the average distance from all data points within the cluster to the cluster center; higher centrality indicates better clustering quality.
[0084] Cluster centrality is calculated using a combined evaluation method of intra-cluster variance and inter-cluster distance. Intra-cluster variance reflects the tightness of the clusters, calculated as the average of the sum of squared distances between all data points within a cluster and the cluster center. The first cluster contains 8 user behavior data points with an intra-cluster variance of 0.15, and the second cluster contains 12 data points with an intra-cluster variance of 0.22. Inter-cluster distance measures the separation between different clusters, calculated as the minimum distance between the cluster centers. The centrality index is obtained by dividing the inter-cluster distance by the intra-cluster variance. The centrality of the first cluster is 1.8 divided by 0.15, which equals 12.0, and the centrality of the second cluster is 1.5 divided by 0.22, which equals 6.8.
[0085] The preset influence threshold was determined based on statistical analysis of historical clustering results and recommendation performance. By analyzing recommendation accuracy, recall, and user satisfaction under different threshold settings, the threshold with the best overall performance was selected as the preset value. Statistical analysis shows that when the influence threshold is set to 8.0, the recommendation accuracy reaches 87.3%, the recall rate is 82.1%, and the user satisfaction score is 4.2. Setting the threshold too high will result in insufficient priority behavior data, affecting recommendation coverage; setting the threshold too low will introduce noisy data, reducing recommendation accuracy. With the preset influence threshold set to 8.0, clusters with centrality higher than this threshold are selected as the source of priority behavior data.
[0086] The selection of priority behavior data involves extracting representative user behavior samples from clusters with centrality exceeding a preset influence threshold. The selection process prioritizes data points near the cluster centers, as these points best represent the typical characteristics of that type of user behavior. For the first cluster with a centrality of 12.0, the five data points closest to the cluster center are selected as priority behavior data. For the third cluster with a centrality of 10.5, the seven data points closest to the cluster center are selected. The priority behavior data contains a total of 23 high-quality user behavior samples, covering users' main interests and behavioral patterns.
[0087] Semantic embedding of content semantic data employs a pre-trained language model for deep semantic understanding of text content. This pre-trained model, trained on a large-scale text corpus, possesses powerful semantic representation capabilities. The text content undergoes preprocessing steps such as word segmentation, part-of-speech tagging, and syntactic analysis to convert it into an input format that the model can process. For a product description text containing 120 words, a 768-dimensional semantic vector is obtained after encoding by the pre-trained model. The semantic embedding process preserves the contextual information and semantic relationships of the text, and the generated vector accurately reflects the semantic features of the content. When multimodal content includes images and audio, convolutional neural networks and recurrent neural networks are used for feature extraction, respectively, and these features are concatenated with the text semantic vector to form comprehensive content features.
[0088] Content feature generation utilizes a multilayer perceptron (MLP) to further abstract and optimize the semantic embedding results. The MLP contains two hidden layers: a first layer with 256 neurons and a second layer with 128 neurons. The output layer generates a 64-dimensional content feature vector. Supervised learning is employed during network training, using manually labeled content as the supervision signal. The content feature vectors undergo L2 normalization to ensure a vector length of 1, facilitating subsequent similarity calculations and feature fusion. For movie recommendation scenarios, the content feature vectors for action movies are [0.23, 0.15, 0.31, ...], and those for romance movies are [0.18, 0.42, 0.09, ...]. The cosine similarity between the vectors reflects the semantic relevance of the content.
[0089] The calculation of user behavior weights is based on a comprehensive evaluation of the influence intensity and the importance score of the user behavior. The importance score is calculated through three dimensions: frequency, intensity, and timeliness of user behavior. Behavior frequency reflects the user's sustained attention to a certain type of content, with a frequency weight of 0.4. Behavior intensity measures the depth and level of user engagement, with an intensity weight of 0.35. Timeliness assesses the proximity of the behavior to the present time, with a timeliness weight of 0.25. For example, a user clicked on an action movie 15 times, with an intensity score of 0.7 and a timeliness score of 0.9. The importance score is calculated as 15 multiplied by 0.4, plus 0.7 multiplied by 0.35, plus 0.9 multiplied by 0.25, equaling 6.47. The user behavior weight equals the importance score multiplied by the corresponding influence intensity. When the influence intensity is 0.19, the final weight is 6.47 multiplied by 0.19, equaling 1.23.
[0090] The weighted combination process linearly combines priority behavior data and content features according to the calculated user behavior weights. A weighted average is used to ensure balanced fusion of features from different sources. For a set of 23 priority behavior data points, each data point corresponds to a 64-dimensional feature vector and a scalar weight value. The feature vector of the first priority behavior data point is [0.31, 0.22, 0.45, ...] with a weight of 1.23, and the feature vector of the second data point is [0.28, 0.35, 0.41, ...] with a weight of 1.15. The weighted average of all priority behavior data is then fused with the content feature vector. The fusion weights are dynamically adjusted according to the current recommendation scenario, with user behavior weight accounting for 60% and content feature weight accounting for 40%.
[0091] Personalized recommendations are generated based on similarity calculations and ranking of weighted, fused feature vectors. The cosine similarity between the fused feature vector and the feature vectors of candidate recommended items is calculated; higher similarity indicates a better match. The candidate item pool contains 1000 items, each corresponding to a 64-dimensional feature vector. Similarity calculation results are sorted from highest to lowest, and the top 20 items with the highest similarity are selected as recommendation candidates. The recommendations are further adjusted for diversity and filtered for novelty to avoid overly simplistic or repetitive results. The final personalized recommendation output includes 10 recommended items, each accompanied by a recommendation rating and a reason for the recommendation, providing users with a personalized recommendation service.
[0092] In one optional implementation, based on the personalized recommendation results and the information annotation data, an information analysis network is constructed using conditional random fields and probabilistic graphical models. The information analysis network is then used to calculate the information propagation patterns in the personalized recommendation results, including:
[0093] A conditional random field correlation mapping matrix is constructed for the user interaction behavior sequence and the item feature sequence. The value of each element in the correlation mapping matrix is set as the inner product operation result of the corresponding feature vector. The conditional transition probability of the item feature sequence relative to the user interaction behavior sequence is calculated based on the inner product operation result.
[0094] A hierarchical dynamic attention network is constructed based on the conditional transition probability, and the propagation intensity is assigned to the neighbor set of the recommended item node. Based on the propagation intensity assignment, an exponential weight function is used to perform nonlinear transformation and robustness enhancement on the conditional transition probability. Based on the result of the nonlinear transformation, a set of propagation relationship edges between nodes is constructed.
[0095] The conditional transition probability and the propagation relationship edge set are combined by weighting to construct an information analysis network. The propagation paths in the information analysis network are scored, and the information propagation rules in the personalized recommendation results are characterized based on the score calculation results.
[0096] User interaction behavior sequences are extracted from user feedback data of personalized recommendation results, including time-series records of various user interactions such as clicks, browsing, favorites, sharing, and comments. Each interaction behavior is represented by a vector, containing information such as behavior type, behavior intensity, behavior duration, and behavior occurrence time. Behavior type is encoded as a one-hot vector: click behavior is encoded as [1,0,0,0,0], browsing behavior as [0,1,0,0,0], favorite behavior as [0,0,1,0,0], sharing behavior as [0,0,0,1,0], and comment behavior as [0,0,0,0,1]. Behavior intensity is normalized to a range of 0 to 1, with click behavior intensity set to 0.3, browsing behavior intensity set to 0.5, favorite behavior intensity set to 0.7, sharing behavior intensity set to 0.8, and comment behavior intensity set to 0.9. Behavior duration is recorded in seconds, and the numerical range is compressed through logarithmic transformation. Behavior occurrence time is represented by a timestamp, converted into a periodic feature vector using time encoding technology.
[0097] The item feature sequence is constructed based on the multidimensional attribute information of recommended items, including four aspects: content features, statistical features, relational features, and dynamic features. Content features include static attributes such as item category, tags, and descriptive text, encoded into a 128-dimensional dense vector using a pre-trained word vector model. Statistical features include quantitative indicators such as click-through rate, conversion rate, user rating, and number of reviews, standardized to form a 16-dimensional feature vector. Relational features reflect the similarity and correlation between items and other items, calculated using a collaborative filtering algorithm to obtain a 32-dimensional relational vector. Dynamic features capture the popularity changes and trends of items over different time periods, generating a 24-dimensional temporal feature vector using a sliding window statistical method. The item feature sequence concatenates these four types of features to form a 200-dimensional comprehensive feature vector, which serves as the observed variable for a conditional random field.
[0098] The Conditional Random Field (CRF) association mapping matrix establishes a correspondence between user interaction behavior sequences and item feature sequences. The number of rows in the matrix equals the length of the user interaction behavior sequence, the number of columns equals the length of the item feature sequence, and the matrix dimension is the length of the behavior sequence multiplied by the dimension of the feature vector. For a user sequence containing 15 interaction behaviors and a 200-dimensional item feature vector, the association mapping matrix dimension is 15 multiplied by 200. The value of each element in the matrix is obtained by performing an inner product operation between the corresponding user behavior feature vector and the item feature vector. The feature vector of the user's third interaction behavior is [0.3, 0.6, 0.8, 0.2, 0.9], and the corresponding value of the item feature vector is [0.5, 0.7, 0.4, 0.8, 0.6]. The inner product result is 0.3 multiplied by 0.5 plus 0.6 multiplied by 0.7 plus 0.8 multiplied by 0.4 plus 0.2 multiplied by 0.8 plus 0.9 multiplied by 0.6, which equals 1.33.
[0099] The inner product calculation process employs vectorization to improve computational efficiency. Matrix multiplication is performed between the user interaction sequence matrix and the item feature sequence matrix to directly obtain the complete association mapping matrix. Batch processing is used to process multiple user interaction sequences simultaneously, reducing redundant computation overhead. The inner product result is then processed by an activation function, using a positive value function to compress the numerical range to between -1 and +1, enhancing the non-linearity of feature representation. For an inner product value of 1.33, positive activation yields 0.87; for an inner product value of -0.5, activation yields -0.46. The activated values serve as the basic weights for connections between nodes in the conditional random field.
[0100] The conditional transition probability calculation is based on probability distribution modeling using the inner product operation results in the association mapping matrix. For each feature vector in the item feature sequence, its conditional probability distribution relative to all behaviors in the user interaction behavior sequence is calculated. The probability calculation uses the softmax function to normalize the inner product operation results, ensuring that the probability distribution satisfies the normalization constraint. The inner product value of the item feature vector relative to the first user interaction behavior is 0.8, the inner product value relative to the second interaction behavior is 1.2, and the inner product value relative to the third interaction behavior is 0.6. After softmax normalization, the probability distribution is [0.28, 0.42, 0.30]. The rows of the conditional transition probability matrix correspond to user interaction behaviors, the columns correspond to item features, and the matrix element values are the corresponding conditional probabilities.
[0101] The hierarchical dynamic attention network employs a multi-layer neural network architecture to deeply model conditional transition probabilities. The network comprises three layers of attention mechanisms: the first layer handles local interaction patterns, the second layer handles mid-range dependencies, and the third layer handles global semantic information. Each attention mechanism includes three components: a query matrix, a key matrix, and a value matrix. Feature weight distributions are obtained through self-attention computation. The first layer's attention weight distribution is [0.15, 0.25, 0.35, 0.25], the second layer's is [0.20, 0.30, 0.20, 0.30], and the third layer's is [0.25, 0.25, 0.25, 0.25]. The multi-layer attention output is processed through residual connections and layer normalization to avoid gradient vanishing and overfitting problems.
[0102] The identification of the neighbor set for recommended item nodes is based on the connections in the item relationship graph. The neighbor set includes other item nodes that are directly related to the currently recommended item. These connections are established through various methods such as user interaction, content similarity, and collaborative filtering. For movie recommendation scenarios, the neighbor set for action movie nodes includes other movies by the same director, movies with the same actors, movies of the same genre, and movies watched by the same users. The size of the neighbor set is controlled between 20 and 50 nodes to ensure a balance between computational efficiency and effectiveness in propagation intensity allocation. Neighbor nodes are sorted according to their association strength, and the top 30 nodes with the strongest association are selected as the effective neighbor set.
[0103] The propagation intensity allocation is calculated based on the output of a hierarchical dynamic attention network and the importance scores of neighboring nodes. The weight distribution of the attention network output is weighted and combined with the importance scores of neighboring nodes to obtain the propagation intensity value for each neighboring node. The importance score is calculated using a combination of the PageRank algorithm, degree centrality, and betweenness centrality. The first neighboring node has an attention weight of 0.15, an importance score of 0.8, and a propagation intensity of 0.15 multiplied by 0.8, which equals 0.12. The second neighboring node has an attention weight of 0.25, an importance score of 0.6, and a propagation intensity of 0.15. The propagation intensity allocation results are normalized to ensure that the sum of the propagation intensities of all neighboring nodes equals 1.0.
[0104] The exponential weighting function applies a non-linear transformation to the conditional transition probabilities, enhancing the weighting ability of high-probability values. The exponential function uses a base-2 exponentiation, where the exponent is the conditional transition probability multiplied by a scaling factor of 3.0. After exponential transformation, an item with a conditional transition probability of 0.42 becomes 2^1.26, equal to 2.39, while an item with a conditional transition probability of 0.28 becomes 2^0.84, equal to 1.79. This non-linear transformation widens the gap between different probability values, giving high-probability items a more significant weighting advantage. The transformation result is normalized to maintain the validity of the probability distribution, avoiding numerical overflow and computational instability.
[0105] Robustness enhancement improves the stability of conditional transition probabilities through adversarial training and noise injection techniques. The adversarial training process generates adversarial examples to test attacks on the network, identifying vulnerabilities and providing targeted reinforcement. Noise injection adds Gaussian noise to the conditional transition probability calculation process, with the noise intensity set to 5% of the original probability value. For example, adding Gaussian noise with a standard deviation of 0.021 to a probability value of 0.42 yields a perturbed probability value of 0.435 or 0.408. The robustness enhancement training process iterates for 3000 epochs, evaluating the model's resistance to interference every 100 epochs to gradually improve the network's fault tolerance to abnormal inputs.
[0106] The propagation relationship edge set is constructed based on the conditional transition probabilities after nonlinear transformation to determine the connections between nodes. A probability threshold of 0.15 is set; node pairs with conditional transition probabilities exceeding the threshold are established as propagation relationship edges. The edge weight is set to the corresponding nonlinear transformation probability value, and the edge direction represents the directionality of information propagation. The propagation relationship edges are represented using a directed graph structure, supporting complex propagation path modeling and analysis. For a recommendation network containing 100 nodes, approximately 300 propagation relationship edges are retained after threshold filtering, forming a sparse but efficient propagation network structure. The edge set is stored in an adjacency list format, supporting efficient graph traversal and path search operations.
[0107] The information analysis network is constructed by weighted combination of conditional transition probabilities and propagation relationship edge sets. Conditional transition probabilities, as node-level features, are weighted at 0.6, reflecting the inherent propagation ability of each node. Propagation relationship edge sets, as edge-level features, are weighted at 0.4, representing the strength of propagation relationships between nodes. The weighted combination process employs a linear fusion approach, preserving the semantic meaning of the original features while enhancing their expressive power. The fused network contains both node and edge features, supporting multi-granularity propagation analysis and prediction. The network structure is represented using a graph neural network, with node embedding dimensions set to 256 and edge embedding dimensions set to 128.
[0108] The propagation path score calculation quantifies and evaluates all propagation paths in the information analysis network. The score calculation comprehensively considers four factors: path length, path weight, node importance, and edge strength. Path length has a weight of 0.2, with shorter paths receiving higher scores. Path weight accounts for 0.3, calculated by accumulating the weights of all edges along the path. Node importance has a weight of 0.3, calculated based on the centrality index of nodes traversed by the path. Edge strength has a weight of 0.2, reflecting the reliability of connections along the path. For a propagation path of length 3, the accumulated edge weight is 2.1, the average node importance is 0.7, the average edge strength is 0.8, and the overall score is 0.2 x 0.33 + 0.3 x 2.1 + 0.3 x 0.7 + 0.2 x 0.8 = 1.07.
[0109] The representation of information propagation patterns is achieved through statistical analysis and pattern recognition of propagation path scores. Cluster analysis groups propagation paths with similar scores into the same category, identifying different propagation patterns. High-scoring paths represent strong propagation patterns, medium-scoring paths represent general propagation patterns, and low-scoring paths represent weak propagation patterns. Propagation pattern features include four dimensions: propagation speed, propagation range, propagation depth, and propagation stability. Propagation speed is measured by the reciprocal of the path length; propagation range is assessed by the number of nodes covered by the path; propagation depth is calculated by the maximum number of hops; and propagation stability is derived through variance analysis of path scores. The information propagation patterns of personalized recommendation results are ultimately represented as a four-dimensional feature vector, used to guide subsequent recommendation optimization and propagation prediction tasks.
[0110] In one optional implementation, a hierarchical dynamic attention network is constructed based on the conditional transition probabilities, and propagation intensity is assigned to the neighbor set of the recommended item node. Based on the propagation intensity assignment, an exponential weighting function is used to perform a nonlinear transformation and robustness enhancement on the conditional transition probabilities. Based on the result of the nonlinear transformation, a set of propagation relationship edges between nodes is constructed, including:
[0111] A hierarchical dynamic attention network is constructed using the conditional transition probability as the baseline probability. In the hierarchical dynamic attention network, a robustness enhancement is achieved by using a positive sample propagation path, weak negative samples with randomly replaced nodes, and strong negative samples generated adversarially, resulting in the optimized conditional transition probability.
[0112] The propagation intensity is assigned to the neighbor set based on the optimized conditional transition probability, and the propagation intensity is obtained by adaptively fusing features at different levels through the hierarchical dynamic attention network.
[0113] An exponential weighting function is used to perform a nonlinear transformation on the propagation intensity, and an adaptive threshold is set based on the node feature vector and global context information to add random noise to the nonlinearly transformed propagation intensity.
[0114] Consistency regularization and adversarial training are performed on the propagation intensity with added random noise to obtain the final nonlinear transformation result. Based on the nonlinear transformation result, a set of propagation relationship edges between the recommended item node and its neighboring nodes is constructed.
[0115] Conditional transition probabilities serve as the baseline probabilities for constructing the hierarchical dynamic attention network's architecture. These probabilities reflect the likelihood of a recommended item node propagating information to its neighboring nodes, and are calculated using historical user behavior data and item attribute similarity. For movie recommendation scenarios, the conditional transition probability from an action movie node to a science fiction movie node is 0.75, to a romance movie node is 0.32, and to a documentary node is 0.18. The hierarchical dynamic attention network employs a multi-layer neural network structure, including an input layer, hidden layers, and an output layer. The input layer receives the conditional transition probability matrix and node feature vectors; the hidden layer uses an attention mechanism to weightedly fuse features from different levels; and the output layer generates the optimized conditional transition probabilities.
[0116] Positive propagation paths are constructed based on effective propagation links in real user behavior sequences. Continuous item access sequences are extracted from user browsing history, and causally related item pairs are identified as positive propagation paths. For example, a user's behavior sequence of immediately searching for related science fiction movies after watching an action movie is labeled as a positive path, with a path weight set to 1.0. During positive path training, the network learns feature representations of real propagation patterns, strengthening its ability to identify effective propagation relationships. Positive paths typically account for 60% to 70% of the total training samples to ensure the network can fully learn correct propagation patterns.
[0117] Weak negative sample generation creates erroneous propagation paths by randomly replacing nodes. A node is randomly selected from the positive sample propagation path and replaced with another node whose attributes are unrelated, forming a weak negative sample. For example, in the positive sample path from action movie to science fiction movie, the science fiction movie is replaced with a cooking tutorial video, creating a weak negative sample path from action movie to cooking tutorial. The label for weak negative samples is set to 0.0, and the path weight is negative, guiding the network to learn and identify illogical propagation relationships. The random replacement process is kept under control; the similarity threshold between the replaced node and the original node is set below 0.3 to ensure a clear negative sample contrast.
[0118] Strong negative examples are generated through adversarial generative networks (GANs) to create deceptive error propagation paths. The adversarial generator receives feature representations of positive example paths and generates propagation paths that are superficially similar to positive examples but semantically incorrect. The generator learns the statistical features of positive examples to create seemingly reasonable but actually flawed propagation relationships. The strong negative example corresponding to the positive example path from action movies to science fiction movies is the strong negative example from action movies to suspense movies; the two are similar in some feature dimensions but have different propagation logics. The generation process of strong negative examples involves multiple rounds of adversarial training, with the generator continuously improving its quality and the discriminator continuously improving its recognition capabilities.
[0119] The robustness enhancement training process uses positive, weak, and strong negative samples simultaneously to train the hierarchical dynamic attention network. In each training batch, positive samples comprise 50%, weak negative samples 30%, and strong negative samples 20%. The loss function comprehensively considers the fitting error of positive samples, the rejection loss of weak negative samples, and the adversarial loss of strong negative samples. During training, the network learns to distinguish between genuine and spurious propagation relationships, improving its robustness against noisy data and anomalous patterns. The number of training iterations is set to 5000 epochs, and an exponential decay strategy is adopted for the learning rate, with an initial learning rate of 0.001, decreasing by a factor of 0.1 every 1000 epochs.
[0120] The optimized conditional transition probabilities are calculated through forward propagation using a hierarchical dynamic attention network. The network output probability values are normalized to ensure that the sum of the conditional transition probabilities of all neighboring nodes equals 1.0. Compared to the original probabilities, the optimized conditional transition probabilities have better discriminative power and stability, and can more accurately reflect the true propagation relationships between nodes. After optimization, the conditional transition probability from the action movie node to the science fiction movie node is adjusted to 0.82, the probability from the romance movie node is reduced to 0.25, and the probability from the documentary node is adjusted to 0.15.
[0121] The neighbor set propagation strength assignment is calculated based on the optimized conditional transition probability. For each neighbor node of a recommended item node, the propagation strength equals the conditional transition probability multiplied by the importance weight of the neighbor node. The importance weight of the neighbor node is determined by a weighted combination of the node's degree centrality, betweenness centrality, and eigenvector centrality. The degree centrality weight is set to 0.4, reflecting the number of direct connections of the node. The betweenness centrality weight is set to 0.35, measuring the node's bridging role in the propagation path. The eigenvector centrality weight is set to 0.25, evaluating the node's global influence.
[0122] A hierarchical dynamic attention network adaptively fuses features from different layers to generate the final propagation strength. The first layer of the attention mechanism processes the basic features of nodes, including node type, attribute values, and historical activity. The second layer processes the relational features of nodes, including the distribution of neighboring nodes, edge weights, and local network structure. The third layer processes the global features of nodes, including their position in the entire network, their influence range, and propagation potential. Each layer outputs a feature vector, which is then weighted and fused to form a comprehensive feature representation. The adaptive weights are dynamically adjusted according to the characteristics of the current propagation task; the weights of the first and second layers are increased when local propagation is emphasized, while the weights of the third layer are increased when global propagation is emphasized.
[0123] The exponential weighting function performs a non-linear transformation on the propagation intensity, enhancing its ability to express differences in propagation intensity. The function uses an exponential form with the natural constant as the base, where the exponent is the propagation intensity value multiplied by a scaling factor. A scaling factor of 2.5 makes the differences between nodes with propagation intensities above 0.5 more pronounced after the exponential transformation. A neighbor node with a propagation intensity of 0.8 receives a value of 6.05 after the exponential transformation, a neighbor node with a propagation intensity of 0.6 receives a value of 4.48, and a neighbor node with a propagation intensity of 0.4 receives a value of 2.72. This non-linear transformation results in a greater weight gain for nodes with high propagation intensities and a relatively reduced weight for nodes with low propagation intensities.
[0124] The adaptive threshold setting is dynamically determined based on node feature vectors and global context information. The node feature vectors contain the node's content, structural, and behavioral features, which are reduced to 128 dimensions using principal component analysis. The global context information includes the overall characteristics of the current propagation network, such as network density, average clustering coefficient, and network diameter. During the adaptive threshold calculation, the node feature vectors and global context vectors are multiplied by a dot product, and then mapped to a value between 0 and 1 using the sigmoid function. The threshold is used to filter out neighboring nodes with low propagation strength after nonlinear transformation; nodes with propagation strength below the threshold are excluded from the propagation relationship edge set.
[0125] The random noise addition process introduces a controlled perturbation to the propagation intensity after nonlinear transformation, improving the network's generalization ability. The random noise is generated using a Gaussian distribution with a mean of 0 and a standard deviation of 0.05. The noise addition ratio is controlled within 5% of the propagation intensity to avoid excessive interference with the original propagation relationship. For a node with a propagation intensity of 6.05, Gaussian noise with a standard deviation of 0.3 is added, ultimately adjusting the propagation intensity to 5.87 or 6.23. The noise addition process is performed during both the training and inference phases. During training, it helps the network learn robust feature representations, and during inference, it increases the diversity of prediction results.
[0126] Consistency regularization ensures that the propagation intensity distribution remains relatively stable before and after noise is added. The distance between the original propagation intensity distribution and the propagation intensity distribution after noise perturbation is calculated, and the KL divergence is used to measure the distribution difference. When the KL divergence exceeds a preset threshold of 0.1, the noise intensity is adjusted or noise is regenerated. The consistency regularization loss function guides the network to enhance robustness while maintaining the validity of the propagation relationship, avoiding significant impacts of noise on propagation quality. The regularization weight is set to 0.02, which accounts for a small proportion of the total loss function but plays an important constraining role.
[0127] The adversarial training process constructs adversarial examples to train the propagation intensity allocation network for both attack and defense. The adversarial example generator adds carefully designed perturbations to the original propagation intensity, attempting to deceive the network into making incorrect propagation relationship judgments. Perturbation generation employs a gradient ascent method, adjusting the propagation intensity value along the gradient direction of the loss function, with the perturbation amplitude controlled within 10% of the original value. The defense network learns more stable feature representations by recognizing and resisting adversarial perturbations. The adversarial training alternates between attack and defense processes, with the attack intensity gradually increasing and the defense capability correspondingly improving.
[0128] The final nonlinear transformation result integrates the outputs of noise addition, consistency regularization, and adversarial training. The transformation preserves the original ranking relationship of propagation intensity while enhancing the expressive power and anti-interference ability of numerical differences. The propagation intensity after the complete processing has better discriminative power, stability, and reliability, accurately reflecting the strength of propagation relationships between nodes. The numerical distribution of the processed propagation intensity is more reasonable, the difference between high-intensity and low-intensity nodes is more obvious, and the distribution of medium-intensity nodes is smoother.
[0129] The propagation relationship edge set is constructed based on the final nonlinear transformation result to determine the connection relationship between the recommended item node and its neighbor nodes. A propagation strength threshold of 3.0 is set; neighbor nodes exceeding the threshold establish propagation relationship edges with the recommended item node. The edge weight is set to the corresponding nonlinear transformation propagation strength value, and the edge direction is from the recommended item node to the neighbor node. For a recommended item with 50 neighbor nodes, after threshold filtering, 18 neighbor nodes with high propagation strength are retained, forming 18 propagation relationship edges. The edge set is stored using a sparse matrix, reducing storage space and computational complexity, and supporting efficient graph operations and propagation simulation.
[0130] In one optional implementation, an information evolution path diagram is generated based on the information propagation pattern. This information evolution path diagram is then used to dynamically track the personalized recommendation results, generating a tracking analysis report that includes propagation links and evolution trends.
[0131] The propagation data of personalized recommendation results during the propagation process is obtained, and the propagation data is subjected to hierarchical feature extraction using a deep feature extraction network to obtain a fused feature matrix;
[0132] The propagation intensity value is calculated based on the fusion feature matrix. The propagation intensity value is combined with the attention weight of the propagation node. The attention weight is determined by the temporal progression relationship of the historical propagation influence of the node. An information evolution path graph is constructed by the propagation intensity value and the attention weight.
[0133] The propagation nodes in the information evolution path graph are dynamically tracked, the tracking frequency is adaptively adjusted based on changes in the propagation environment, the propagation state changes of the nodes are calculated based on the propagation intensity value and the attention weight, and the edge weights between the propagation nodes in the information evolution path graph are updated based on the propagation state changes.
[0134] Multi-level analysis is performed on the updated edge weights to identify propagation association patterns between different levels and to mine the propagation links of the personalized recommendation results.
[0135] The analysis results of the propagation link are mapped based on the knowledge graph of the propagation domain. The propagation trend perception of the personalized recommendation results is predicted through graph reasoning, and a tracking analysis report containing the propagation link and an interpretable propagation evolution trend is generated.
[0136] like Figure 3 As shown, the method includes:
[0137] The acquisition of personalized recommendation results propagation data employs a distributed data collection architecture, deploying multiple data collection nodes to monitor the propagation behavior of recommended content across different platforms in real time. These nodes acquire user interaction data such as clicks, shares, comments, and reposts through application programming interfaces (APIs), while simultaneously recording key information such as propagation timestamps, propagation path identifiers, and user identifiers. The collected raw data encompasses four dimensions: user behavior characteristics, content characteristics, time characteristics, and social relationship characteristics. Taking video recommendations as an example, the collected data includes a video viewing duration of 120 seconds, 85 likes, 12 shares, 23 comments, a propagation depth of 3 layers, and a propagation breadth of 156 nodes.
[0138] The deep feature extraction network employs a multi-layer convolutional neural network architecture to perform hierarchical feature extraction on the propagation data. The first convolutional layer extracts basic features of the propagation data, including user behavior frequency features, content attribute features, and time series features. The second convolutional layer combines the basic features to extract interaction features and association features. The third convolutional layer performs high-order feature abstraction to generate semantic-level propagation features. Batch normalization and activation functions are applied after each convolutional operation to prevent gradient vanishing and overfitting problems. During feature extraction, the feature dimension is set to 512 dimensions, resulting in a 128-dimensional fused feature vector after three convolutional layers. The fused feature matrix arranges the feature vectors of all propagation nodes according to time order and propagation level, forming a feature matrix with a dimension equal to the number of nodes multiplied by 128.
[0139] The propagation intensity value is calculated based on a weighted aggregation of feature vectors in the fusion feature matrix. For each propagation node, the user activity feature, content quality feature, propagation speed feature, and influence range feature in its feature vector are weighted and combined. The user activity feature weight is set to 0.3, reflecting the degree of user participation in the propagation process. The content quality feature weight is set to 0.25, reflecting the attractiveness and value of the recommended content. The propagation speed feature weight is set to 0.25, measuring the timeliness of information propagation. The influence range feature weight is set to 0.2, assessing the size of the user group covered by the propagation. The propagation intensity value of each node is calculated through weighted aggregation, with a value ranging from 0 to 1, where a larger value indicates a stronger propagation intensity.
[0140] Attention weights are calculated based on the temporal progression of a node's historical propagation influence. For each propagation node, its historical propagation data over the past 30 days is collected, including indicators such as the number of times it participated in propagation, propagation effect, and propagation range. Historical data is weighted using a time decay function, with higher weights for recent data and lower weights for older data. The time decay factor is set to 0.9, meaning the weight is multiplied by 0.9 for each additional day. A basic influence score for each node is calculated based on the weighted historical data, and this score, combined with the node's positional importance in the current propagation network, yields the final attention weight. Positional importance is evaluated using network topology features such as the number of in-degrees and out-degrees of a node, the shortest path length between nodes, and the node clustering coefficient.
[0141] The information evolution path graph construction combines propagation intensity values and attention weights to form node weights. Each node weight equals the propagation intensity value multiplied by the attention weight, constructing a directed graph structure. Each node in the graph represents a propagation participant, and the node weight reflects the participant's importance in the propagation process. Edges in the graph represent the direction and path of information propagation, with edge weights initialized to the difference in propagation intensity. For a source node with a propagation intensity of 0.8 and a target node with a propagation intensity of 0.6, the edge weight is set to 0.2. The graph structure is represented using an adjacency matrix, where matrix elements are the weights of corresponding edges, and unconnected nodes have elements with a value of 0.
[0142] The dynamic tracking mechanism monitors propagation nodes in the information evolution path graph in real time. The tracking frequency adaptively adjusts according to changes in the propagation environment, including changes in propagation speed, node activity, and the impact of external events. When an increased propagation speed is detected, the tracking frequency increases from once per hour to once every 30 minutes. When node activity significantly increases, the tracking frequency for that node and its neighboring nodes increases to once every 15 minutes. When an external emergency occurs, the tracking frequency for the entire network increases to once every 5 minutes. During the tracking process, propagation behavior data of nodes is collected in real time, including newly added clicks, shares, comments, and other interactive behaviors.
[0143] The calculation of propagation status changes is based on real-time collected propagation behavior data and historical status data. For each node, the rate of change of activity, the rate of change of influence, and the rate of change of propagation range within the current time window are calculated. The rate of change of activity equals the activity in the current time window minus the activity in the previous time window, and then divided by the activity in the previous time window. The rate of change of influence and the rate of change of propagation range are calculated using the same method. Combining the three rate of change indicators, a weighted average is used to obtain the comprehensive propagation status change value of the node. The weights are allocated as follows: activity rate of change 0.4, influence rate of change 0.35, and propagation range rate of change 0.25.
[0144] Edge weight updates adjust the connections between nodes in the information evolution path graph based on propagation state change values. When both the source node and target node have positive propagation state changes, the weight of the corresponding edge is increased. The weight increase is equal to the average of the two node propagation state changes multiplied by an adjustment factor of 0.1. When either the source or target node has a negative propagation state change, the edge weight is decreased. The weight decrease is equal to the absolute value of the negative change multiplied by an adjustment factor of 0.15. After edge weight updates, normalization is required to ensure that all edge weights are between 0 and 1.
[0145] Multi-level edge weight analysis performs hierarchical clustering and pattern recognition on the updated edge weights, dividing the propagation network into three layers according to edge weight: strong connection layer, medium connection layer, and weak connection layer. The strong connection layer contains connections with edge weights greater than 0.7, representing tight propagation relationships. The medium connection layer contains connections with edge weights between 0.3 and 0.7, representing general propagation relationships. The weak connection layer contains connections with edge weights less than 0.3, representing loose propagation relationships. Connection patterns within each layer are analyzed to identify structural features such as propagation clusters, propagation bridges, and propagation islands. Propagation clusters refer to densely connected regions between nodes, propagation bridges are key edges connecting different clusters, and propagation islands are sparsely connected groups of nodes.
[0146] The propagation path mining method identifies the complete propagation path of personalized recommendation results based on multi-level analysis results. Starting from the information source node, a depth-first search is performed along the paths with larger edge weights to discover all propagation paths. The discovered propagation paths are sorted according to indicators such as path length, total path weight, and number of nodes covered by the path. The top 10 paths with the largest total weight are selected as primary propagation paths, and the next 20 paths with the second largest total weight are selected as secondary propagation paths. Primary propagation paths represent the mainstream direction of information propagation, while secondary propagation paths represent auxiliary directions. Detailed information such as the starting node, ending node, intermediate nodes, propagation time, and propagation effect of each propagation path is recorded.
[0147] The knowledge graph construction in the field of communication includes entity types such as communication subjects, communication content, communication channels, and communication effects, as well as the relationships between them. The communication subject entity includes information such as user attributes, social relationships, and behavioral preferences. The communication content entity includes information such as content type, content quality, and content popularity. The communication channel entity includes information such as platform characteristics, algorithm mechanisms, and user distribution. The communication effect entity includes information such as communication scope, communication depth, and communication speed. Relationships between entities include types such as influence relationships, promotion relationships, inhibition relationships, and competition relationships. The knowledge graph is stored using a graph database, supporting complex relationship queries and inference computations.
[0148] The propagation link mapping process matches the mined propagation links with entities and relationships in the knowledge graph. For each node in the propagation link, the corresponding propagation entity is searched in the knowledge graph to obtain its detailed attribute information. For each edge in the propagation link, the corresponding propagation relationship is searched in the knowledge graph to obtain its strength and type information. Through this mapping process, the originally simple propagation links are transformed into knowledge-based propagation paths containing rich semantic information. The mapping results include in-depth information such as the role positioning of the propagation entity, the motivational analysis of the propagation behavior, and the factors influencing the propagation effect.
[0149] Knowledge graph reasoning predicts propagation trends based on entity relationships and rule bases within a knowledge graph. The reasoning rules include propagation diffusion rules, propagation attenuation rules, and propagation blocking rules. Propagation diffusion rules describe the spread and expansion patterns of information under specific conditions; propagation attenuation rules describe the natural decline in the intensity of information propagation; and propagation blocking rules describe the conditions that lead to the interruption of information propagation. Based on the current propagation state and historical propagation patterns, the reasoning rules are applied to predict future propagation trends. The prediction results include indicators such as expected propagation range, peak propagation time, propagation duration, and degree of propagation impact. Prediction accuracy is evaluated and optimized by comparing with actual propagation results.
[0150] The tracking and analysis report generates integrated dissemination link analysis and dissemination evolution trend prediction results. The report includes dissemination link visualization charts, showcasing the complete path and key nodes of information dissemination. The dissemination evolution trend section provides a curve graph of dissemination intensity changing over time, a heat map of dissemination scope expansion, and a bar chart of changes in the influence of key nodes. The interpretability analysis section, based on knowledge graph mapping results, elaborates on the causes of dissemination phenomena, the influencing factors of dissemination effects, and the driving mechanisms of dissemination trends. The report uses a template-based generation method, supporting different output formats and personalized customization needs.
[0151] In one optional implementation, the propagation nodes in the information evolution path graph are dynamically tracked, and the tracking frequency is adaptively adjusted based on changes in the propagation environment. The calculation of the propagation state changes of the nodes based on the propagation intensity value and the attention weight includes:
[0152] The propagation nodes in the information evolution path graph are dynamically tracked, the propagation probability between the propagation node and its neighboring nodes is calculated, the temporal characteristics of the propagation node are modeled based on the propagation probability to obtain the propagation probability evolution sequence, and the seepage entropy index of the propagation node is calculated based on the propagation probability evolution sequence.
[0153] The tracking frequency is adaptively adjusted based on changes in the propagation environment. The seepage entropy index is used as an environmental state feature, the rate of change of the environmental state feature is calculated, the calculation period is adjusted according to the rate of change, the calculation period is used as the tracking frequency, and the propagation state change of the node is calculated according to the propagation intensity value and the attention weight.
[0154] When dynamically tracking propagation nodes in the information evolution path graph, the connection relationships between nodes are constructed. Assuming the information evolution path graph G contains a node set V and an edge set E, for any node v∈V, its set of adjacent nodes is N(v). The system calculates the propagation probability p(v,u) by analyzing the historical interaction data between node v and its adjacent nodes u∈N(v). In the specific implementation, the propagation probability calculation considers factors such as the frequency of historical information flow between nodes, content similarity, and user interaction behavior. For example, if the historical interaction count between node A and node B is 50, with information flowing from A to B 35 times, the initial estimated propagation probability is 0.7. Furthermore, considering the content similarity factor of 0.85 and the user response speed factor of 0.92, the final determined propagation probability is 0.7×0.85×0.92≈0.55.
[0155] Based on the calculated propagation probabilities, the temporal characteristics of propagation nodes are modeled. By calculating the propagation probabilities within multiple time windows t1, t2, ..., tn, a propagation probability evolution sequence P(v,t)={p(v,t1),p(v,t2),...,p(v,tn)} is formed. Taking a certain social network node as an example, its propagation probabilities in five consecutive time windows are 0.23, 0.27, 0.45, 0.78, and 0.62, respectively, constituting a temporal sequence reflecting the changes in the node's propagation capability.
[0156] The percolation entropy index H(v) of a propagating node is calculated based on the propagation probability evolution sequence. The percolation entropy index reflects the degree of uncertainty in the propagation of information by a node in a network. In the calculation, the propagation probability sequence is divided into multiple intervals, and the frequency of probability values in each interval is counted to calculate the entropy value of the probability distribution. For example, if the probability range [0,1] is divided into five equal-length intervals, the frequencies of a node's propagation probability distribution within 50 time windows in these intervals are 5, 8, 12, 20, and 5, respectively. Based on this frequency distribution, the calculated percolation entropy is 2.05, indicating that the node's propagation behavior has a moderate degree of uncertainty.
[0157] In the stage of adaptively adjusting the tracking frequency based on changes in the propagation environment, the seepage entropy index is used as the environmental state characteristic S(t). By comparing the environmental state characteristics of adjacent time windows, the rate of change of the environmental state r(t) = |S(t) - S(t-1)| / S(t-1) is calculated. For example, if the seepage entropy values of a node in two adjacent evaluation periods are 1.85 and 2.12 respectively, then the rate of change of the environment is |2.12-1.85| / 1.85≈0.146.
[0158] The calculation period T is dynamically adjusted based on the rate of change of environmental conditions. When the rate of change is high, the calculation period is shortened to increase the tracking frequency; when the rate of change is low, the calculation period is extended to reduce computational overhead. Specific adjustment strategies can be implemented using thresholds for tiered adjustments. For example, when the rate of change exceeds 0.2, the calculation period is shortened to 0.6 times its original value; when the rate of change is below 0.05, the calculation period is extended to 1.5 times its original value; and when the rate of change is between 0.05 and 0.2, the calculation period remains unchanged. Taking an information dissemination scenario as an example, the initial calculation period was 30 minutes, which was adjusted three times to 18 minutes, 18 minutes, and 27 minutes respectively.
[0159] The change in a node's propagation state is calculated based on its propagation intensity value and attention weight. The propagation intensity value I(v) reflects node v's information propagation capability and can be calculated by considering factors such as the node's out-degree, influence range, and propagation speed. The attention weight A(v) represents the degree of attention information recipients pay to the content published by node v and can be calculated using indicators such as user dwell time, interaction behavior, and forwarding rate. The change in a node's propagation state ΔS(v) is determined by both the propagation intensity value and the attention weight, and is calculated as a weighted combination of the two values. For example, if a node has a propagation intensity value of 0.78 and an attention weight of 0.65, and the weighting coefficients are set to 0.6 and 0.4 respectively, then the change in propagation state is approximately 0.78 × 0.6 + 0.65 × 0.4 ≈ 0.728.
[0160] In practical applications, a threshold is set for the propagation state change value to classify it. For example, if the threshold is set to 0.7, a propagation state change value greater than or equal to 0.7 is judged as a high-activity state, and a propagation state change value less than 0.7 is judged as a low-activity state. Through this classification, the system can quickly identify active propagation nodes in the network and adjust the monitoring strategy accordingly.
[0161] The aforementioned method achieves efficient monitoring and analysis of information evolution paths by dynamically tracking propagation nodes, adaptively adjusting the tracking frequency, and calculating changes in node propagation states. This method has significant application value in scenarios such as large-scale social network information dissemination and marketing campaign effectiveness evaluation, helping relevant personnel quickly identify key propagation nodes and formulate corresponding intervention strategies.
[0162] A second aspect of the present invention provides a system for intelligently searching and displaying important information on the Internet, comprising:
[0163] The first unit is used to acquire Internet information, perform semantic segmentation and noise removal on the Internet information, generate information feature vectors, calculate information entropy and information gain ratio on the information feature vectors, determine information weight distribution based on the information entropy and information gain ratio, and obtain information labeling data based on the information weight distribution.
[0164] The second unit is used to construct a dual-path recommendation model based on the information annotation data. The dual-path recommendation model performs feature learning from the user interest dimension and the information content dimension, respectively, and integrates the user behavior features and content semantic features in the information annotation data to output personalized recommendation results.
[0165] The third unit is used to construct an information analysis network based on the personalized recommendation results and the information annotation data, using conditional random fields and probabilistic graphical models, and to calculate the information propagation rules in the personalized recommendation results based on the information analysis network.
[0166] The fourth unit is used to generate an information evolution path diagram based on the information dissemination pattern, use the information evolution path diagram to dynamically track the personalized recommendation results, and generate a tracking analysis report containing the dissemination links and evolution trends.
[0167] A third aspect of the present invention provides an electronic device, comprising:
[0168] processor;
[0169] Memory used to store processor-executable instructions;
[0170] The processor is configured to invoke instructions stored in the memory to execute the aforementioned method.
[0171] A fourth aspect of the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.
[0172] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligently searching and displaying important information on the Internet, characterized in that, include: Acquire Internet information, perform semantic segmentation and noise removal on the Internet information, generate information feature vectors, calculate information entropy value and information gain ratio on the information feature vectors, determine information weight distribution based on the information entropy value and the information gain ratio, and obtain information labeling data based on the information weight distribution; A dual-path recommendation model is constructed based on the information-annotated data. The dual-path recommendation model performs feature learning from the user interest dimension and the information content dimension, respectively, and integrates the user behavior features and content semantic features in the information-annotated data to output personalized recommendation results. Based on the personalized recommendation results and the information annotation data, an information analysis network is constructed using conditional random fields and probabilistic graphical models, and the information propagation rules in the personalized recommendation results are calculated based on the information analysis network. Based on the aforementioned information dissemination patterns, an information evolution path diagram is generated. The personalized recommendation results are then dynamically tracked using this information evolution path diagram, generating a tracking analysis report that includes the dissemination links and evolution trends.
2. The method according to claim 1, characterized in that, The information entropy and information gain ratio are calculated for the information feature vector. The information weight distribution is determined based on the information entropy and information gain ratio. Information annotation data is obtained based on the information weight distribution, including: A multidimensional feature tensor is constructed from the information feature vector, and the information entropy value of each dimension is calculated based on the multidimensional feature tensor. Dynamic weights are calculated based on the information entropy values of each dimension, and the dynamic weight distribution is obtained through iterative optimization with information classification accuracy as the evaluation index. The multidimensional feature tensor is reconstructed and labeled based on the dynamic weight distribution to obtain information labeled data.
3. The method according to claim 1, characterized in that, A dual-path recommendation model is constructed based on the labeled information data. This model performs feature learning from both the user interest dimension and the information content dimension, fusing user behavior features and content semantic features from the labeled information data to output personalized recommendation results, including: Based on the information-annotated data, user behavior data and content semantic data are mapped to a low-dimensional space through feature dimensionality reduction. The data in the low-dimensional space is used to construct a dual-path recommendation model. Based on the dual-path recommendation model, the influence of the upper-level node on the current node is calculated iteratively through gradient descent to obtain the influence strength between nodes. Based on the influence intensity, adaptive clustering is performed on the user behavior data in the user interest dimension, and user behavior data with a cluster centrality higher than a preset influence threshold is taken as priority behavior data. The content semantic data is semantically embedded to obtain content features. User behavior weights are calculated based on the influence intensity. The priority behavior data and the content features are weighted and combined according to the user behavior weights to generate personalized recommendation results.
4. The method according to claim 1, characterized in that, Based on the personalized recommendation results and the information annotation data, an information analysis network is constructed using conditional random fields and probabilistic graphical models. The information propagation patterns in the personalized recommendation results are calculated using this information analysis network, including: A conditional random field correlation mapping matrix is constructed for the user interaction behavior sequence and the item feature sequence. The value of each element in the correlation mapping matrix is set as the inner product operation result of the corresponding feature vector. The conditional transition probability of the item feature sequence relative to the user interaction behavior sequence is calculated based on the inner product operation result. A hierarchical dynamic attention network is constructed based on the conditional transition probability, and the propagation intensity is assigned to the neighbor set of the recommended item node. Based on the propagation intensity assignment, an exponential weight function is used to perform nonlinear transformation and robustness enhancement on the conditional transition probability. Based on the result of the nonlinear transformation, a set of propagation relationship edges between nodes is constructed. The conditional transition probability and the propagation relationship edge set are combined by weighting to construct an information analysis network. The propagation paths in the information analysis network are scored, and the information propagation rules in the personalized recommendation results are characterized based on the score calculation results.
5. The method according to claim 4, characterized in that, A hierarchical dynamic attention network is constructed based on the conditional transition probabilities, and propagation intensity is assigned to the neighbor set of the recommended item node. Based on the propagation intensity assignment, an exponential weight function is used to perform a nonlinear transformation and robustness enhancement on the conditional transition probabilities. Based on the result of the nonlinear transformation, a set of propagation relationship edges between nodes is constructed, including: A hierarchical dynamic attention network is constructed using the conditional transition probability as the baseline probability. In the hierarchical dynamic attention network, a robustness enhancement is achieved by using a positive sample propagation path, weak negative samples with randomly replaced nodes, and strong negative samples generated adversarially, resulting in the optimized conditional transition probability. The propagation intensity is assigned to the neighbor set based on the optimized conditional transition probability, and the propagation intensity is obtained by adaptively fusing features at different levels through the hierarchical dynamic attention network. An exponential weighting function is used to perform a nonlinear transformation on the propagation intensity, and an adaptive threshold is set based on the node feature vector and global context information to add random noise to the nonlinearly transformed propagation intensity. Consistency regularization and adversarial training are performed on the propagation intensity with added random noise to obtain the final nonlinear transformation result. Based on the nonlinear transformation result, a set of propagation relationship edges between the recommended item node and its neighboring nodes is constructed.
6. The method according to claim 1, characterized in that, Based on the aforementioned information dissemination patterns, an information evolution path diagram is generated. This diagram is then used to dynamically track the personalized recommendation results, generating a tracking and analysis report that includes the dissemination links and evolutionary trends. The propagation data of personalized recommendation results during the propagation process is obtained, and the propagation data is subjected to hierarchical feature extraction using a deep feature extraction network to obtain a fused feature matrix; The propagation intensity value is calculated based on the fusion feature matrix. The propagation intensity value is combined with the attention weight of the propagation node. The attention weight is determined by the temporal progression relationship of the historical propagation influence of the node. An information evolution path graph is constructed by the propagation intensity value and the attention weight. The propagation nodes in the information evolution path graph are dynamically tracked, the tracking frequency is adaptively adjusted based on changes in the propagation environment, the propagation state changes of the nodes are calculated based on the propagation intensity value and the attention weight, and the edge weights between the propagation nodes in the information evolution path graph are updated based on the propagation state changes. Multi-level analysis is performed on the updated edge weights to identify propagation association patterns between different levels and to mine the propagation links of the personalized recommendation results. The analysis results of the propagation link are mapped based on the knowledge graph of the propagation domain. The propagation trend perception of the personalized recommendation results is predicted through graph reasoning, and a tracking analysis report containing the propagation link and an interpretable propagation evolution trend is generated.
7. The method according to claim 6, characterized in that, Dynamically track the propagation nodes in the information evolution path graph, adaptively adjust the tracking frequency based on changes in the propagation environment, and calculate the propagation state changes of the nodes according to the propagation intensity value and the attention weight, including: The propagation nodes in the information evolution path graph are dynamically tracked, the propagation probability between the propagation node and its neighboring nodes is calculated, the temporal characteristics of the propagation node are modeled based on the propagation probability to obtain the propagation probability evolution sequence, and the seepage entropy index of the propagation node is calculated based on the propagation probability evolution sequence. The tracking frequency is adaptively adjusted based on changes in the propagation environment. The seepage entropy index is used as an environmental state feature, the rate of change of the environmental state feature is calculated, the calculation period is adjusted according to the rate of change, the calculation period is used as the tracking frequency, and the propagation state change of the node is calculated according to the propagation intensity value and the attention weight.
8. A system for intelligently searching and displaying key information on the Internet, used to implement the method of any one of claims 1-7, characterized in that, include: The first unit is used to acquire Internet information, perform semantic segmentation and noise removal on the Internet information, generate information feature vectors, calculate information entropy and information gain ratio on the information feature vectors, determine information weight distribution based on the information entropy and information gain ratio, and obtain information labeling data based on the information weight distribution. The second unit is used to construct a dual-path recommendation model based on the information annotation data. The dual-path recommendation model performs feature learning from the user interest dimension and the information content dimension, respectively, and integrates the user behavior features and content semantic features in the information annotation data to output personalized recommendation results. The third unit is used to construct an information analysis network based on the personalized recommendation results and the information annotation data, using conditional random fields and probabilistic graphical models, and to calculate the information propagation rules in the personalized recommendation results based on the information analysis network. The fourth unit is used to generate an information evolution path diagram based on the information dissemination pattern, use the information evolution path diagram to dynamically track the personalized recommendation results, and generate a tracking analysis report containing the dissemination links and evolution trends.
9. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 7.