Network information trend prediction method and system based on deep learning
Through deep learning combined with evolutionary algorithms and reinforcement learning methods, neural network structures are dynamically reconstructed, neuron pruning and online learning are performed, which solves the adaptability and resource efficiency problems of network information trend prediction, and improves the accuracy and real-time response capabilities of turning point prediction.
Patent Information
- Application Number
- CN202510692475.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-05-26
- Filing Date
- 2025-05-27
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing network information trend prediction methods are difficult to adapt to the diversified and dynamic network information environment, the prediction accuracy is reduced, the real-time response needs cannot be met, the calculation complexity is high, and it is difficult to deploy efficiently in resource-constrained environments. The lack of feature extraction and enhancement mechanisms for turning points, resulting in insufficient warning time.
Deep learning-based methods are adopted, combining evolutionary algorithms and reinforcement learning to automatically search the optimal neural network structure, dynamically reconstruct the model, perform importance-aware neuron pruning, and optimize the model through knowledge transfer and online learning to extract and enhance turning point features to prevent catastrophic forgetting.
The model is highly adaptable to diversified information types, responds to changes in the network environment in real time, reduces computing complexity and memory usage, and improves the accuracy and advance amount of turning point prediction, providing more accurate and timely decision-making support.
Smart Images

Figure CN120524979A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence and information dissemination analysis, and more specifically, to a network information trend prediction method and system based on deep learning. Background Art
[0002] With the rapid development of the internet and social media, online information is experiencing explosive growth and complex, ever-changing dissemination characteristics. Accurately predicting online information dissemination trends, especially key turning points, is crucial for decision-making by governments, businesses, and social organizations. Currently, online information trend prediction primarily relies on statistical models and machine learning methods, but faces the following technical challenges: Existing prediction methods often use fixed neural network architectures, which are designed to optimize for specific information types and are difficult to adapt to the diversity and dynamics of network information. When faced with different types of information propagation patterns, prediction accuracy decreases; The network information environment is highly dynamic. Once trained, traditional models retain a fixed structure and cannot be dynamically adjusted to environmental changes. When the information environment undergoes sudden changes, model performance often plummets, requiring time-consuming retraining to recover, making it impossible to meet real-time response requirements. High-precision prediction models typically have high computational complexity and a large number of parameters, making them difficult to deploy efficiently in resource-constrained environments. While model simplification can reduce resource consumption, the simplified models significantly reduce accuracy for complex tasks such as turning point prediction, making it difficult to balance prediction accuracy and resource efficiency. The turning point features in network information dissemination are often overwhelmed by a large amount of ordinary time series data. Existing methods lack a feature extraction and enhancement mechanism specifically for turning points, resulting in low turning point prediction accuracy, insufficient warning time, and difficulty in providing sufficient preparation time for decision-making. The network information environment is constantly changing, and new data is constantly being generated. However, traditional models have difficulty effectively utilizing new data for continuous learning after initial training, and are prone to the problem of "catastrophic forgetting," whereby the memory of previously learned knowledge is lost when learning new data, resulting in a decline in prediction performance over time. Therefore, a network information trend prediction method is needed that can adapt to the diversity and dynamics of network information, is resource-efficient, and has continuous learning capabilities. In particular, it is necessary to improve the prediction accuracy and lead time of key turning points to provide more accurate and timely decision support for various application scenarios. Summary of the Invention
[0003] The present invention provides a network information trend prediction method and system based on deep learning, which solves the technical problems in related technologies such as reduced prediction accuracy, inability to meet real-time response requirements, difficulty in balancing prediction accuracy and resource efficiency, and difficulty in providing sufficient preparation time for decision-making.
[0004] The present invention provides a network information trend prediction method based on deep learning, comprising the following steps: Based on historical network information dissemination data, the optimal neural network structure is automatically searched by combining evolutionary algorithms and reinforcement learning; Based on the optimal neural network structure, by monitoring the changes in the network information environment, the dynamic reconstruction of the model structure is triggered when changes in the information propagation pattern are detected; For the reconstructed neural network, adaptive pruning of neurons based on importance perception is performed according to the resource constraints of the deployment environment. Extract and enhance turning point features in network information propagation, and transfer the knowledge of large, high-precision models obtained through search to lightweight models after pruning; The lightweight model after knowledge migration is deployed, and it is subjected to online learning and continuous optimization, using the experience replay mechanism to prevent catastrophic forgetting.
[0005] In a preferred embodiment, the step of automatically searching for the optimal neural network structure by combining evolutionary algorithms and reinforcement learning includes: Constructing a neural architecture search space, including network layer types, connection methods, activation functions, and attention mechanisms; Define a multi-objective evaluation function that takes into account prediction accuracy, model complexity, and inference time; Randomly generate initial candidate architectures to form an initial population; Calculate the multi-objective evaluation function value for each candidate architecture; Select the best architecture as the parent based on the evaluation function value; Apply crossover and mutation operations to the selected parent architecture to generate new candidate architectures; Use reinforcement learning to optimize the mutation strategy; repeat the evaluation, selection, evolution, and optimization steps until convergence, and select the architecture with the highest evaluation function value in the final population as the optimal architecture.
[0006] In a preferred embodiment, the step of triggering dynamic reconstruction of the model structure upon detecting a change in the information propagation mode includes: Build an information environment change detection module to monitor the rate of change of information flow, topic distribution deviation, sentiment polarity changes, and changes in user interaction patterns; Pre-build and maintain a library of candidate architectures, each optimized for different types of network information; Create a performance index table for each architecture in the candidate architecture library, recording its prediction performance on different information categories; When a change in the information environment is detected, the current information flow characteristics are analyzed and the category to which it belongs is identified; Select the best architecture in the category from the candidate library; Migrate key parameters of the current model to the new architecture; Quickly fine-tune new models using recently collected data.
[0007] In a preferred embodiment, the step of performing importance-aware neuron adaptive pruning includes: Obtain resource constraints for the deployment environment, including computing power, memory limitations, energy requirements, and latency requirements. Determine the target ratio of model compression based on resource constraint parameters; Calculate the importance score for each neuron in the neural network, including activation importance, gradient importance, and turning point sensitivity; Define resource-accuracy balance optimization function to guide the pruning process; The neuron importance sorting, pruning, evaluation, and decision steps are iteratively performed, gradually removing neurons with low importance until the target compression ratio is reached or the performance degradation exceeds a threshold.
[0008] In a preferred embodiment, the steps of online learning and continuous optimization include: Build an incremental learning module and configure immediate and long-term buffers to store recently collected data and historical key data; Design a priority sampling strategy to prioritize turning point samples, high loss samples, and low-frequency pattern samples; Adopting elastic weight merging algorithm for incremental updates to balance new task learning and old knowledge retention; Configure the concept drift detection component to determine data distribution changes through distribution shift measurement and performance monitoring; Dynamically adjust the learning rate based on data characteristics and model performance; Select representative historical samples from the long-term buffer to construct mixed training batches; Set different loss weights for new data and historical samples to prevent catastrophic forgetting.
[0009] In a preferred embodiment, the step of extracting and enhancing turning point features in network information propagation includes: Decompose the original time series data into trend terms, cycle terms and residual terms; Calculate the first and second derivatives of time series data to capture the rate of change and acceleration of the data; Calculate Z scores to identify data points that deviate significantly from the normal range; Calculate cumulative distribution characteristics to capture long-term accumulation effects; The original features are combined with the enhanced features to form an enhanced feature vector.
[0010] In a preferred embodiment, the step of automatically searching for the optimal neural network structure by combining evolutionary algorithms and reinforcement learning further includes: Construct a multi-level architecture search space, including the overall architecture model at the macro level and the specific module design at the micro level; At the macro level, different types of information processing modules are defined, including temporal feature extraction module, attention enhancement module, multimodal fusion module and prediction output module; At the micro level, define optional internal structure and parameter configuration for each module; Adopting a hierarchical search strategy, first determine the macro-architecture and then optimize the micro-structure of each module; Introducing an adaptive resource constraint mechanism to dynamically adjust the search space boundaries based on the target deployment environment; Design dedicated evaluation indicators for network information dissemination characteristics, including turning point prediction accuracy, advance warning time, and environmental adaptability score; During the evolution process, a reinforcement learning agent is used to dynamically adjust the mutation probability and crossover strategy to optimize the search efficiency based on the historical search trajectory.
[0011] In a preferred embodiment, the step of extracting and enhancing turning point features in network information dissemination further includes: Build a multi-scale time window analysis module to simultaneously capture short-term fluctuations, medium-term changes, and long-term trends; Design an adaptive threshold detection algorithm to dynamically adjust the turning point determination criteria based on the propagation characteristics of different types of network information; Introducing network topology feature analysis to calculate the structural entropy, centrality distribution, and community evolution indicators of the information dissemination network; Construct a user influence weighted model to weight information dissemination characteristics according to the influence of user nodes; Design a self-learning mechanism for feature importance, and automatically discover the most predictive turning point features in different scenarios through meta-learning methods; Implement feature temporal correlation analysis to capture temporal dependencies and interaction patterns between features; A turning point feature enhancement network is constructed, and the feature representation of turning point samples is enhanced through contrastive learning methods to improve the discrimination in the prediction process.
[0012] In a preferred embodiment, the steps of online learning and continuous optimization further include: Set the sample retention strategy to prioritize turning point samples, high loss samples, and low-frequency pattern samples; Configure the sample attenuation factor to control the weight of historical samples to decrease over time; Use elastic weight merging algorithm for incremental updates to balance new knowledge learning with old knowledge retention; Parameter importance assessment based on Fisher information matrix estimation; Set the incremental update frequency and batch size to properly allocate the ratio of new data to historical samples; Configure a concept drift detection mechanism to determine data distribution changes through divergence calculation and performance monitoring; Set the continuous monitoring step count and benchmark distribution update period; Implement an adaptive learning rate adjustment mechanism to dynamically adjust the learning rate based on gradient changes and performance.
[0013] In a preferred embodiment, a network information trend prediction system based on deep learning is used to execute a network information trend prediction method based on deep learning, including: The neural architecture automatic search module is used to automatically search for the optimal neural network structure based on historical network information propagation data by combining evolutionary algorithms and reinforcement learning; The environment perception and model reconstruction module is used to monitor changes in the network information environment and trigger model structure reconstruction, while also adaptively pruning the neural network based on the resource constraints of the deployment environment; The turning point feature enhancement module is used to extract and enhance the turning point features in network information propagation and capture the key change nodes of information propagation; The knowledge transfer and distillation module is used to transfer the knowledge of large, high-precision models to lightweight models to achieve model compression and performance preservation; The online learning and continuous optimization module is used to continuously learn from new data and optimize model performance after deployment. At the same time, it prevents catastrophic forgetting through experience replay and achieves accurate prediction of turning points and trends in network information dissemination.
[0014] The beneficial effects of the present invention are: The present invention uses an evolutionary reinforcement learning architecture search mechanism to automatically find the optimal neural network architecture for different types of network information, greatly improving the model's adaptability to diverse information types.
[0015] The event-triggered model structure dynamic reconstruction mechanism constructed by the present invention can monitor changes in the network information environment in real time, and automatically select the architecture or combined architecture that best suits the current environment from the candidate architecture library when major changes are detected, thereby enabling the model to quickly adapt to environmental changes.
[0016] The adaptive model compression technology adopted in this invention, by combining technologies such as neuron importance evaluation, structured pruning and knowledge distillation, can significantly reduce the model's computational complexity and memory usage without significantly reducing the prediction accuracy.
[0017] The turning point feature enhancement module specially designed in the present invention can effectively extract and amplify turning point signals in network information flows through technologies such as multi-scale time series comparison, attention mechanism and feature importance weighting. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a flow chart of a network information trend prediction method based on deep learning of the present invention; Figure 2 It is a pie chart of the importance weights of different information environment change indicators of the present invention; Figure 3 is a bar graph comparing the prediction performance before and after the turning point feature enhancement of the present invention; Figure 4 It is a radar chart showing the performance changes of different types of models during the online learning process of the present invention. DETAILED DESCRIPTION
[0019] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0020] At least one embodiment of the present invention discloses a network information trend prediction method based on deep learning, such as Figure 1 As shown, the following steps are included: Step 1: Based on historical network information propagation data, the optimal neural network structure is automatically searched using a combination of evolutionary algorithms and reinforcement learning. The specific steps include: Step 1.1, neural architecture search space construction; First, a neural architecture search space is constructed to define the set of searchable network components.
[0021] The search space includes: Network layer types: convolutional layer, recurrent layer, attention layer, residual connection, etc. Connection mode: serial connection, jump connection, parallel connection, etc. Activation function: ReLU, Sigmoid, Tanh, LeakyReLU, etc. Attention mechanism: self-attention, cross-attention, temporal attention, etc.; Number of floors range: 2 to 20; The number of hidden units ranges from 16 to 1024.
[0022] The cable space can be expressed as: ; in, represents the cable space; Represents a set of network layer types; Represents a set of connection methods; represents the set of activation functions; represents the set of attention mechanisms; Indicates the range of layers; Indicates the range of the number of hidden units.
[0023] Step 1.2, multi-objective evaluation function definition; Define a multi-objective evaluation function that takes into account three dimensions: prediction accuracy, model complexity, and inference time: ; in, Represents the neural network model to be evaluated; Representation Model The comprehensive evaluation score of the model, the higher the score, the better the model; Representation Model The prediction accuracy of Representation Model complexity; Representation Model The reasoning time of the model, the larger the value, the slower the model reasoning speed; 、 、 They represent the weight coefficients of prediction accuracy, model complexity and inference time respectively.
[0024] Step 1.3, architecture search algorithm for evolutionary reinforcement learning; This implementation uses a method combined with evolutionary reinforcement learning to perform architecture search: Initialization: Randomly generated candidate architectures, forming the initial population ,in, Represents the initial population, which is the set of all candidate architectures; 、 、 Respectively represent 、 、 candidate neural network architectures; Represents the number of candidate neural network architectures.
[0025] Evaluation: For each candidate architecture , calculate the multi-objective evaluation function value on the validation set ; Selection: Based on the evaluation function value, the tournament selection method is used to select An excellent architecture as the parent; Evolutionary operation: Apply crossover and mutation operations to the selected parent architecture to generate new candidate architectures: Crossover: Randomly select components from each layer of the two parent architectures to form a child architecture; Mutation: by probability Randomly changing certain components of the architecture; Reinforcement learning optimization: Use reinforcement learning to provide guidance for mutation operations and optimize mutation strategies based on historical search experience: ; in, Indicates the current architecture status Select mutation operation The probability of Represents the state description of the current neural network architecture, Indicates possible mutation operations, represents the set of all possible mutation operations, Represents the state-action value function, which means that in state Execute the following operation The expected cumulative reward of represents the natural exponential function, and the denominator represents the normalization factor to ensure that the sum of all operation probabilities is 1.
[0026] Updated by TD (temporal difference) learning: ; in, Indicates that the status Take action The value function of Represents the first learning rate, which controls the step size of Q value update; Indicates the immediate reward obtained by the current mutation operation; Indicates execution of an action The new state to which it is transferred; New state All possible moves The maximum Q value in represents the estimate of the maximum possible reward in the future; is the timing difference error.
[0027] Iteration: Repeat the steps until the preset number of iterations is reached or the evaluation function value converges; Output: The architecture with the highest evaluation function value in the final population is selected as the optimal architecture.
[0028] Step 2: Based on the optimal neural network structure, by monitoring the changes in the network information environment, the model structure is dynamically reconstructed when changes in the information propagation pattern are detected; The specific steps include: Step 2.1, information environment change detection; Build an information environment change detection module to monitor changes in the network information environment through the following indicators: Information flow rate of change: ; in, Indicates time Relative to the previous time window The rate of change of information flow; Indicates time Information flow; Indicates the previous time window Information flow; Indicates the time window size; when A large value indicates a surge in information flow, which may indicate the occurrence of a major event or the formation of a hot topic.
[0029] Topic distribution deviation: , use KL divergence to measure the current time Topic distribution With the previous time window Topic distribution The difference between When the value is large, it indicates that the topic distribution has changed significantly, and there may be new hot topics or a significant change in the popularity of existing topics.
[0030] Emotional polarity changes: ; in, Indicates time Relative to the previous time window The absolute value of the change in emotional polarity; Indicates time The overall sentiment polarity value of Indicates the previous time window The overall sentiment polarity value of Larger values indicate a significant change in the emotional tone of online information, potentially signaling a turning point in public sentiment.
[0031] Changes in user interaction patterns: Using graph structure similarity Measuring the current user engagement network User interaction network with the previous time window Structural changes between .
[0032] When any of the above indicators exceeds the preset threshold, or the weighted combination value of the indicators exceeds the comprehensive threshold, an event detection signal is triggered: ; in, Indicates time Whether to trigger the event detection signal, 1 means trigger, 0 means not trigger; 、 、 、 The weight coefficients representing the information flow change rate, topic distribution deviation, sentiment polarity change, and user interaction pattern change respectively; Indicates time The rate of change of information flow; Indicates the current time KL divergence with the topic distribution of the previous time window; Indicates time The amount of change in emotional polarity; Indicates the graph structure similarity between the current and previous time windows of the user interaction network; Represents the comprehensive threshold, which determines the sensitivity of trigger event detection. The smaller the value, the more sensitive it is.
[0033] Step 2.2, candidate architecture library management; In an embodiment of the present application, a candidate architecture library is pre-built and maintained ,in, represents a set of candidate neural network architectures; 、 、 Respectively represent the 1st, 2nd, A neural network architecture; Indicates the total number of architectures in the candidate architecture library; “Different types of online information” refers to various types of online information data such as social media information flows, news dissemination, and public events; "Good performance" refers to achieving a high level in indicators such as accuracy, recall, prediction lead time, and computational efficiency.
[0034] The candidate architecture library is formed and updated in the following ways: Initialization: Use the neural architecture search algorithm to search for the optimal architecture on different types of historical network information data to form an initial candidate library; Performance index: for each candidate architecture Create a performance index table , record the architecture in different information categories Prediction performance on Dynamic update: Regularly use newly collected data to evaluate the performance of each candidate architecture and update the performance index table; when an architecture with significantly degraded performance is found, use neural architecture search to generate a new architecture to replace it.
[0035] Step 2.3, optimal architecture selection and structure reconstruction; When a change in the information environment is detected ( ), the model structure reconstruction process is triggered: Information category identification: Analyze the characteristics of the current information flow and identify its category ; Architecture matching: Based on the performance index table, select the architecture with the best performance in this category: ; in, represents the optimal neural network architecture selected; Indicates taking the maximum value of the following expression variable; Represents a library of candidate architectures The candidate neural network architectures; represents a set of candidate neural network architectures; Representation Architecture In the Information category Performance index values on ; Indicates the currently identified network information category.
[0036] Parameter migration: Migrate the key parameters of the current model to the new architecture, preserving the model's memory of historical data: ; in, Represents the newly selected optimal neural network architecture The set of parameters shared with the current architecture; Represents the set of parameters shared between the current neural network architecture and the new architecture; Indicates parameter migration operation; "Shared parameters" refers to network layer parameters with similar functions or the same structure in two different architectures, such as embedding layer weights, initial convolutional layer or feature extraction layer parameters, etc.
[0037] Fast fine-tuning: Use recently collected data to fine-tune the new model for a short period of time to make it better fit the current environment: ; in, Represents the optimal architecture The set of model parameters; Indicates parameter update operation; represents the second learning rate; Indicates about parameters The gradient operator of represents the loss function; Represents the most recently collected data set, which contains the latest network information samples and their corresponding labels.
[0038] The entire reconstruction process is completed within 30 seconds, ensuring that the model can quickly restore its predictive capabilities after changes in the information environment.
[0039] Step 3: For the reconstructed neural network, adaptively prune the neurons based on importance perception according to the resource constraints of the deployment environment. The specific steps include: Step 3.1, resource constraint awareness; First, obtain the resource constraint parameters of the deployment environment, including: Computing power: , which indicates the maximum computing power of the target device (in floating-point operations per second); Memory Limits: , indicating the available memory size; Energy consumption requirements: , indicating the maximum energy consumption level allowed; Latency requirements: , which represents the maximum allowed inference latency.
[0040] Based on the above constraint parameters, determine the target ratio of model compression : ; in, Indicates the target ratio of model compression; Indicates the maximum computing capability of the target device; Represents the computational requirements of the original model; Indicates the available memory size; Indicates taking the minimum value in the brackets; Represents the memory requirements of the original model; Indicates the maximum allowed energy consumption level; represents the energy consumption requirement of the original model; Indicates the maximum allowed inference delay; Represents the inference latency of the original model.
[0041] Step 3.2, neuron importance evaluation; For each neuron in the neural network , calculate its importance score .
[0042] This implementation adopts a weighted combination of three importance assessment criteria: Importance of activation: Calculate the average activation amplitude of neurons in the training set: ; Gradient importance: Calculate the gradient of the loss function with respect to the neuron output: ; Turning point sensitivity: Specifically evaluates the contribution of a neuron to turning point predictions: ; in, Represents neurons The activation importance score of Represents neurons The gradient importance score of ; Represents neurons Turning point sensitivity score; Represents the entire training sample set; is the loss function; Represents the partial derivative of the loss function with respect to the neuron activation value; Indicates the size of the sample set containing the turning point; represents an input sample in the training set; Represents a sample set containing turning points; Represents the turning point prediction output; It's a neuron For input samples The activation value of Represents the partial derivative of the turning point prediction output with respect to the neuron activation value.
[0043] The comprehensive importance score is calculated as: ; in, Represents neurons The comprehensive importance score of Indicates the importance of neuron activation; represents the gradient importance of the neuron; represents the turning point sensitivity of the neuron; 、 、 Weight coefficients for activation importance, gradient importance, and turning point sensitivity, respectively.
[0044] Step 3.3, resource accuracy balance optimization; Define the resource precision balance optimization function to guide the pruning process: ; in, Represents the value of the resource precision balance optimization function; represents the pruned model; Indicates the prediction accuracy of the model; Indicates the computational complexity of the model; represents the energy consumption of the model; is the balance parameter.
[0045] Step 3.4, iterative pruning algorithm; This implementation uses an iterative pruning algorithm to gradually remove neurons of low importance: Initialization: Set the original model as the current model ; Importance sorting: Calculate the importance scores of all neurons in the current model and sort them from low to high; Pruning: Remove the least important neurons to obtain a pruned model ; Evaluation: Calculate the resource-accuracy balance optimization function value of the pruned model ; decision making: if And the compression ratio has not yet reached the target ratio , then accept the pruning result and update the current model , and return the importance ranking to continue iterating; Otherwise, terminate the iteration and return the current model As the final pruning result.
[0046] In some implementations, the pruning process can employ the following variations: Layer-by-layer pruning: pruning in layers, first pruning the deep layers and then the shallow layers, or using a pruning order from shallow to deep; Structured pruning: not only prunes individual neurons, but also prunes entire structural units such as convolution kernels, attention heads or layers; Soft pruning: Instead of directly removing neurons, some weights are brought close to zero through regularization; Dynamic growth and pruning: Dynamically remove unimportant connections during training, while growing new connections to continuously evolve the network structure.
[0047] Step 4: Extract and enhance the turning point features in network information propagation, and transfer the knowledge of the large high-precision model obtained by the search to the pruned lightweight model; The specific steps include: Step 4.1, turning point feature extraction; Construct a turning point feature extraction module to enhance the features of the original network information time series data and highlight key patterns that may indicate turning points: Time series decomposition: ; in, Indicates at a point in time The original time series data; Represents the long-term trend term of the time series; Represents the periodic fluctuation term of the time series, capturing the cyclic change pattern of the data; The residual term represents the time series, which contains random fluctuations and changes that are difficult to explain by trends or cycles.
[0048] Rate of change calculation: ; in, Represents data at a point in time The instantaneous change rate of Indicates the speed of change of data change rate; Indicates at a point in time The original time series data value of Indicates at a point in time The original time series data value of Indicates at a point in time The first derivative value of ; Indicates the time interval between adjacent time points.
[0049] Mutation Detection: Calculating Z-Scores , identify data points that deviate significantly from the normal range: ; in, Indicates a time point Z score; Indicates a time point The original time series data value of Indicates a time point The average value of all data points in the sliding window centered at Indicates a time point The standard deviation of all data points in the sliding window centered on The size of the sliding window determines the range of data points used to calculate the mean and standard deviation. The window size is usually selected to reflect the normal fluctuation range.
[0050] Cumulative Characteristics: Calculates cumulative distribution characteristics , capturing the long-term accumulation effect: ; in, Indicates at a point in time The cumulative distribution eigenvalue of ; Indicates at a point in time The original time series data value of From the initial time point To the current time point The sum operation; Represents the time decay function, as the time interval decreases with the increase of is the attenuation coefficient.
[0051] Feature fusion: Combine the original features with the enhanced features to form an enhanced feature vector : ; in, Indicates at a point in time The enhanced feature vector of Indicates at a point in time The original time series data; represents the trend term obtained from the time series decomposition; represents the periodic term obtained from the time series decomposition; represents the residual term obtained from the time series decomposition; Represents data at a point in time rate of change; Represents data at a point in time rate of change; Indicates a time point Z score; Indicates at a point in time The cumulative distribution eigenvalue of is used to capture the long-term accumulation effect.
[0052] After feature enhancement, the characteristic pattern of the turning point is significantly strengthened, which helps the model to more accurately identify and predict the turning point.
[0053] Step 4.2, knowledge distillation model construction; Apply knowledge distillation techniques to transfer the knowledge of a large, high-precision teacher model (the original or unpruned model) to a pruned, lightweight student model: Teacher model preparation: Select a high-precision model as the teacher model ,This model performs well in predictive accuracy, but may have high computational complexity; Student model preparation: Use a lightweight model that has undergone adaptive pruning as the student model ,This model has low computational resource requirements, but the prediction accuracy may be reduced; Distillation loss function configuration: According to the embodiment of this application, define the comprehensive loss function , including hard target loss and soft target loss: ; in, Represents the total loss function of knowledge distillation; is the loss between the student model prediction and the true label; is the loss between the student model and the teacher model soft label; represents the true label of the training data; Represents the student model's response to input The predicted output of Indicates the teacher model's temperature parameter Lower pair input Softened output; Indicates the temperature parameter of the student model Lower pair input Softened output; Is the balance coefficient, which controls the relative importance of hard loss and soft loss, and its value range is ; is the temperature parameter.
[0054] when When it is close to 1, the model focuses more on matching the true label (hard loss) and ignores the knowledge of the teacher model; when When it is close to 0, the model focuses more on learning the knowledge of the teacher model (soft loss) and less on the true label; the typical setting is 0.5, which means that hard loss and soft loss are equally important; Higher A value of (e.g., 8 or 10) produces a smoother probability distribution, allowing information from non-maximum probability categories to be passed to the student model, helping to convey nuances and uncertainty information from the teacher model. The lower Smaller values (close to 1) will preserve the sharp distribution of the teacher model output, mainly transferring knowledge of high-confidence predictions.
[0055] Turning point sensitive knowledge transfer: During the knowledge distillation process, special attention is paid to the knowledge transfer of turning point samples, and the weight of turning point samples is increased: ; in, Represents the total loss function of knowledge distillation; represents the knowledge distillation loss on common samples; represents the knowledge distillation loss on the turning point sample; is the weight coefficient of the turning point sample.
[0056] Distillation training: Use a dataset rich in turning points to train the student model, minimizing the distillation loss: ; in, represents the optimal parameters of the student model after knowledge distillation training; Represents finding the student model parameters that minimize the following expression ; Represents the total loss function for knowledge distillation.
[0057] Step 5: Deploy the lightweight model after knowledge transfer and perform online learning and continuous optimization on it, using the experience replay mechanism to prevent catastrophic forgetting; The specific steps include: Step 5.1, incremental learning module construction; Build incremental learning modules that enable the model to learn from new data without retraining: Data buffer configuration: Create two levels of data buffer: Instant buffer: stores recently collected data for real-time incremental updates; Long-term buffer: Stores historical key data, especially samples containing turning points, to prevent catastrophic forgetting.
[0058] Priority sampling strategy: Assign priorities to samples in the buffer, keeping them first: Turning point sample: contains data on turning points in the spread of online information; High loss samples: samples with large model prediction errors; Low-frequency mode samples: samples representing rare propagation modes.
[0059] Incremental Update Algorithm: It should be understood that this embodiment uses the Elastic Weight Consolidation (EWC) algorithm to balance new task learning and old knowledge retention: ; in, Represents the total loss function of the elastic weight merging algorithm; is the prediction loss of the model on newly collected data; Represents the complete set of parameters of the model; is the first parameter in the current model parameter set. Parameters are being optimized; is the first parameter values, representing old knowledge; is the Fisher information matrix diagonal elements, quantizing the parameters The importance of the old task performance, the larger the value, the more important the parameter is to the old task; is the regularization coefficient, which controls the balance between learning new and old tasks. A larger value indicates a greater emphasis on retaining old knowledge. Indicates the sum of all model parameters; is a regularization term that prevents catastrophic forgetting by penalizing large changes in important parameters.
[0060] Progressive Networks: Add a new network branch for each new task while keeping the network structure of the old task unchanged and transferring knowledge through lateral connections: ; in, It is Task No. Layer hidden state; It is Task No. Layer hidden state; It is Task No. The weight matrix of the layer; It is from Tasks to The lateral connection weights of each task; It is Task No. Layer hidden states, representing features of early tasks; represents the weighted sum of features from all previous tasks; is the activation function.
[0061] Knowledge distillation assisted incremental learning: When learning a new task, the old model is used as a teacher model to guide the new model to retain the old knowledge: ; in, Represents the total loss function of knowledge distillation; is the prediction loss of the model on newly collected data; Represents the complete set of parameters of the model; is the distillation weight coefficient; Represents the output probability distribution of the old model (teacher model) for the input data; Represents the output probability distribution of the new model (student model) for the same input data; It is the KL divergence, which is used to measure the difference between two probability distributions. The smaller the value, the closer the two distributions are.
[0062] Step 5.2, concept drift detection; Configure the concept drift detection component to monitor changes in data distribution: Distribution Shift Metrics: Calculates the current data distribution With the benchmark distribution JS divergence between: ; in, Indicates the JS divergence between the current data distribution and the benchmark distribution; Represents the probability distribution of data collected in the current time window; represents the probability distribution of the benchmark data; Represents the average distribution of the current distribution and the benchmark distribution; represents the KL divergence.
[0063] Performance monitoring: Track the performance trend of the model on the validation set: ; in, Indicates a time point The change in accuracy at Indicates the current time point The prediction accuracy of Indicates the previous time point The prediction accuracy of Indicates the time interval parameter.
[0064] Drift detection judgment: When the distribution deviation metric exceeds the threshold Or the performance continues to degrade beyond the threshold The concept drift alert is triggered when: ; in, Indicates at a point in time A binary indicator function indicating whether concept drift is detected, where a value of 1 indicates that drift is detected and a value of 0 indicates that drift is not detected; Indicates the JS divergence between the current data distribution and the benchmark distribution; Represents the probability distribution of data collected in the current time window; represents the probability distribution of the benchmark data; is the threshold parameter for the distribution change; Indicates at a point in time The change in model accuracy at ; is the threshold parameter for performance degradation; It is the step number parameter of continuous detection, which requires continuous performance. A drift alarm is triggered only when the performance decreases over several time steps, which is used to filter out short-term fluctuations. “consecutivesteps” represents consecutive time steps, ensuring that only continuous performance degradation is regarded as a true concept drift signal.
[0065] Gradual drift detection: Use time series analysis techniques to detect slow concept drift, such as fitting performance trends using linear regression or ARIMA models and detecting significant downward slopes.
[0066] Drift detection based on ensemble: maintain multiple models trained at different times. When the prediction divergence between models increases significantly, it is considered a concept drift signal: ; in, Indicates at a point in time Diversity measure of model ensemble; represents the total number of models in the ensemble; represents the number of all possible model pairs; Represents all different model pairs Perform the summation, Make sure not to compare a model with itself; and Represents the first and Different models; Indicates at a point in time The collected input data samples; and Represent the model and Input The prediction results; Represents a metric function used to measure the difference between two prediction results.
[0067] Step 5.3, adaptive learning rate adjustment; Dynamically adjust the learning rate based on data characteristics and model performance: Based on gradient changes: Calculate the gradient change and determine the stability of parameter updates: ; in, Represents the exponential moving average of the gradient change at the current time step t; Represents the exponential moving average of the gradient change at the previous time step t-1; is the attenuation coefficient; Represents the loss function at the current time step t Model parameters gradient; Represents the gradient of the loss function of the previous time step t-1 with respect to the model parameters; Represents the square norm of the gradient difference between the current time step and the previous time step; is the weight coefficient of the new gradient change information.
[0068] Based on performance changes: Adjust the learning rate based on changes in model performance: ; in, Indicates the current time step The learning rate; Represents the previous time step The learning rate; Indicates the current time step Model performance metrics; Represents the previous time step Model performance metrics; represents the coefficient of increase of learning rate, and ; represents the reduction coefficient of the learning rate, and ; The threshold for judging the significance of performance change is used to determine whether the performance change is large enough to trigger learning rate adjustment.
[0069] Comprehensive learning rate calculation: ; in, Represents the learning rate finally applied to the model update; Represents the learning rate initially determined after performance evaluation at the current time step; Represents the exponential moving average of the gradient change at the current time step; Represents the square root of the gradient change; It is a small constant added to prevent the denominator from being zero and to ensure the stability of numerical calculations.
[0070] Step 5.4, experience replay mechanism; Configure the experience replay component to prevent catastrophic forgetting: Sample selection: Select representative historical samples from the long-term buffer zone, especially turning point samples; Mixed batch construction: Mix newly collected data with historical samples to construct training batches: ; in, Indicates that at time step Constructed mixed training batches; 、 、 Represents the newly collected 、 、 data samples; 、 、 They represent the first 、 、 A historical sample; Represents the union operation of sets; Indicates the number of new data samples; Indicates the number of samples drawn from historical data.
[0071] Balanced loss function: Set different loss weights for new data and historical samples: ; in, Represents the total loss function used in model training; Represents the loss function value calculated based on the newly collected data; Represents the loss function value calculated based on historical data sampled from the experience replay buffer; is the weight coefficient of historical sample loss.
[0072] In some implementations, the experience replay mechanism may employ the following enhancements: Core set extraction: Extract a representative core sample set from historical data, maximizing sample diversity and coverage while using a smaller buffer to store more representative samples: ; in, Indicates the selected representative core sample set; Indicates returning the parameter that maximizes the following expression, that is, finding the core set that maximizes the objective function ; express Is the complete dataset A subset of Represents the core set The size of , that is, only select samples; Represents the complete dataset Remove coreset The remaining samples; Indicates that in the remaining sample set Find the minimum value in Indicates that in the selected core set Find the minimum value in Representation sample and samples The distance measurement function between them is used to calculate the similarity or difference between samples.
[0073] Generative replay: Use a generative model (such as a variational autoencoder or a generative adversarial network) to learn the distribution of historical data and then generate synthetic samples for replay, reducing the storage requirements of the original historical data: ; in, Represents the generated synthetic samples, which are used to replace the original historical data for experience playback; represents the trained generator network responsible for mapping random noise into meaningful data samples; Represents the random noise vector input to the generator as a random seed for the generation process; Represents a standard normal distribution with a mean of 0 and a covariance matrix of the identity matrix I, which is random noise The sampling distribution of A mathematical symbol indicating "subjection" to a certain distribution, indicating that the noise vector is randomly sampled from a standard normal distribution.
[0074] Dynamic replay weighting: Dynamically adjust the weight of replay samples based on their scarcity, importance, and age: ; in, Representation sample The total weight value of determines the probability of the sample being selected during the playback process; Representation sample The importance score is usually based on the influence of the sample on the model performance, which can be quantified by the gradient magnitude or loss value. Representation sample The rarity score of , which measures the rarity of the sample in the feature space, is usually determined by calculating the average distance or density estimation with other samples; Representation sample The timeliness score evaluates the newness of the sample and is usually calculated based on the difference between the sample's timestamp and the current time; 、 、 The weight coefficients represent importance, rarity, and timeliness respectively.
[0075] The technical effects of this embodiment are as follows: According to the embodiments of the present application, the provided network information trend prediction method based on deep learning has the following significant technical effects: Significantly improved prediction performance: Through multi-objective optimization of neural architecture automatic search and turning point feature enhancement technology, the method of this embodiment has achieved significant breakthroughs in network information turning point prediction: The accuracy of turning point prediction increased by 28.3%, far exceeding the traditional fixed-architecture model; The forecast lead time increased by 15.7%, providing more sufficient preparation time for emergency response; It has strong adaptability to different types of network information and can automatically discover the optimal model structure for different information types.
[0076] Rapid adaptation to environmental changes: The event-triggered dynamic reconstruction technology of the model structure enables the method provided by this application to perform well in the face of changes in the network information environment: Complete model reconstruction within 30 seconds after the information environment changes and restore predictive capabilities; No need for complete retraining, the model's memory of historical data is retained through parameter migration; It has strong adaptability to changes in the propagation pattern of emergency information and its prediction performance remains stable.
[0077] Resource consumption is greatly reduced: By leveraging importance-aware adaptive neuron pruning and knowledge distillation techniques, the method provided in this application significantly reduces resource consumption while maintaining prediction accuracy: The number of model parameters is reduced by 32.5%, significantly reducing storage requirements; Computing resource requirements decreased by 41.2%, reducing processing latency; Energy consumption is reduced by 35.8%, extending device battery life; Ability to flexibly adjust model size based on resource constraints of different deployment environments.
[0078] Continuous learning and adaptation: According to one embodiment of the present application, the method has the ability to continuously evolve through online learning and continuous optimization technology: Continuously learn from new data, and predictive performance improves over time; Effectively prevent catastrophic forgetting and preserve the memory of historical communication patterns; Ability to detect and adapt to concept drift in network information distribution; Model performance steadily improves with extended usage, eliminating the need for frequent offline retraining.
[0079] Wide applicability: The method of this application has broad application prospects and is applicable to various network information trend prediction scenarios: Public opinion monitoring and early warning of public emergencies; Analysis of the dissemination trend of social hot topics; Prediction of changes in business brand reputation; Evaluation of the effectiveness of policy interpretation information dissemination; Suitable for a variety of deployment environments, from high-performance servers to resource-constrained edge devices.
[0080] Figures 2 to 4 The importance weights of different information environment change indicators; comparison of prediction performance before and after turning point feature enhancement; performance changes of different types of models during online learning.
[0081] It can be seen that the method provided in this application has significant advantages in prediction performance, environmental adaptability, resource efficiency and continuous learning ability, and can effectively solve the core technical challenges faced by existing network information trend prediction methods.
[0082] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A network information trend prediction method based on deep learning, characterized in that: The following steps are involved: Based on historical network information dissemination data, the optimal neural network structure is automatically searched by combining evolutionary algorithms and reinforcement learning; Based on the optimal neural network structure, by monitoring the changes in the network information environment, the dynamic reconstruction of the model structure is triggered when changes in the information propagation pattern are detected; For the reconstructed neural network, adaptive pruning of neurons based on importance perception is performed according to the resource constraints of the deployment environment. Extract and enhance turning point features in network information propagation, and transfer the knowledge of large, high-precision models obtained through search to lightweight models after pruning; The lightweight model after knowledge migration is deployed, and it is subjected to online learning and continuous optimization, using the experience replay mechanism to prevent catastrophic forgetting.
2. A network information trend prediction method based on deep learning according to claim 1, characterized in that: The step of automatically searching for the optimal neural network structure by combining evolutionary algorithm and reinforcement learning includes: Constructing a neural architecture search space, including network layer types, connection methods, activation functions, and attention mechanisms; Define a multi-objective evaluation function that takes into account prediction accuracy, model complexity, and inference time; Randomly generate initial candidate architectures to form an initial population; Calculate the multi-objective evaluation function value for each candidate architecture; Select the best architecture as the parent based on the evaluation function value; Apply crossover and mutation operations to the selected parent architecture to generate new candidate architectures; Use reinforcement learning to optimize the mutation strategy; repeat the evaluation, selection, evolution, and optimization steps until convergence, and select the architecture with the highest evaluation function value in the final population as the optimal architecture.
3. The method for predicting network information trends based on deep learning according to claim 1, characterized in that: The step of triggering dynamic reconstruction of the model structure when a change in the information propagation mode is detected includes: Build an information environment change detection module to monitor the rate of change of information flow, topic distribution deviation, sentiment polarity changes, and changes in user interaction patterns; Pre-build and maintain a library of candidate architectures, each optimized for different types of network information; Create a performance index table for each architecture in the candidate architecture library, recording its prediction performance on different information categories; When a change in the information environment is detected, the current information flow characteristics are analyzed and the category to which it belongs is identified; Select the best architecture in the category from the candidate library; Migrate key parameters of the current model to the new architecture; Quickly fine-tune new models using recently collected data.
4. The method for predicting network information trends based on deep learning according to claim 1, characterized in that: The step of performing importance-aware neuron adaptive pruning includes: Obtain resource constraints for the deployment environment, including computing power, memory limitations, energy requirements, and latency requirements. Determine the target ratio of model compression based on resource constraint parameters; Calculate the importance score for each neuron in the neural network, including activation importance, gradient importance, and turning point sensitivity; Define resource-accuracy balance optimization function to guide the pruning process; The neuron importance sorting, pruning, evaluation, and decision steps are iteratively performed, gradually removing neurons with low importance until the target compression ratio is reached or the performance degradation exceeds a threshold.
5. The method for predicting network information trends based on deep learning according to claim 1, characterized in that: The steps for online learning and continuous optimization include: Build an incremental learning module and configure immediate and long-term buffers to store recently collected data and historical key data; Design a priority sampling strategy to prioritize turning point samples, high loss samples, and low-frequency pattern samples; Adopting elastic weight merging algorithm for incremental updates to balance new task learning and old knowledge retention; Configure the concept drift detection component to determine data distribution changes through distribution shift measurement and performance monitoring; Dynamically adjust the learning rate based on data characteristics and model performance; Select representative historical samples from the long-term buffer to construct mixed training batches; Set different loss weights for new data and historical samples to prevent catastrophic forgetting.
6. The method for predicting network information trends based on deep learning according to claim 1, characterized in that: The step of extracting and enhancing turning point features in network information dissemination includes: Decompose the original time series data into trend terms, cycle terms and residual terms; Calculate the first and second derivatives of time series data to capture the rate of change and acceleration of the data; Calculate Z scores to identify data points that deviate significantly from the normal range; Calculate cumulative distribution characteristics to capture long-term accumulation effects; The original features are combined with the enhanced features to form an enhanced feature vector.
7. The method for predicting network information trends based on deep learning according to claim 1, characterized in that: The step of automatically searching for the optimal neural network structure by combining evolutionary algorithms and reinforcement learning also includes: Construct a multi-level architecture search space, including the overall architecture model at the macro level and the specific module design at the micro level; At the macro level, different types of information processing modules are defined, including temporal feature extraction module, attention enhancement module, multimodal fusion module and prediction output module; At the micro level, define optional internal structure and parameter configuration for each module; Adopting a hierarchical search strategy, first determine the macro-architecture and then optimize the micro-structure of each module; Introducing an adaptive resource constraint mechanism to dynamically adjust the search space boundaries based on the target deployment environment; Design dedicated evaluation indicators for network information dissemination characteristics, including turning point prediction accuracy, advance warning time, and environmental adaptability score; During the evolution process, a reinforcement learning agent is used to dynamically adjust the mutation probability and crossover strategy to optimize the search efficiency based on the historical search trajectory.
8. The method for predicting network information trends based on deep learning according to claim 1, characterized in that: The step of extracting and enhancing turning point features in network information dissemination also includes: Build a multi-scale time window analysis module to simultaneously capture short-term fluctuations, medium-term changes, and long-term trends; Design an adaptive threshold detection algorithm to dynamically adjust the turning point determination criteria based on the propagation characteristics of different types of network information; Introducing network topology feature analysis to calculate the structural entropy, centrality distribution, and community evolution indicators of the information dissemination network; Construct a user influence weighted model to weight information dissemination characteristics according to the influence of user nodes; Design a self-learning mechanism for feature importance, and automatically discover the most predictive turning point features in different scenarios through meta-learning methods; Implement feature temporal correlation analysis to capture temporal dependencies and interaction patterns between features; A turning point feature enhancement network is constructed, and the feature representation of turning point samples is enhanced through contrastive learning methods to improve the discrimination in the prediction process.
9. The method for predicting network information trends based on deep learning according to claim 1, characterized in that: Online learning and continuous optimization steps also include: Set the sample retention strategy to prioritize turning point samples, high loss samples, and low-frequency pattern samples; Configure the sample attenuation factor to control the weight of historical samples to decrease over time; Use elastic weight merging algorithm for incremental updates to balance new knowledge learning with old knowledge retention; Parameter importance assessment based on Fisher information matrix estimation; Set the incremental update frequency and batch size to properly allocate the ratio of new data to historical samples; Configure a concept drift detection mechanism to determine data distribution changes through divergence calculation and performance monitoring; Set the continuous monitoring step count and benchmark distribution update period; Implement an adaptive learning rate adjustment mechanism to dynamically adjust the learning rate based on gradient changes and performance.
10. A network information trend prediction system based on deep learning, used to execute a network information trend prediction method based on deep learning according to any one of claims 1 to 9, characterized in that: include: The neural architecture automatic search module is used to automatically search for the optimal neural network structure based on historical network information propagation data by combining evolutionary algorithms and reinforcement learning; The environment perception and model reconstruction module is used to monitor changes in the network information environment and trigger model structure reconstruction, while also adaptively pruning the neural network based on the resource constraints of the deployment environment; The turning point feature enhancement module is used to extract and enhance the turning point features in network information propagation and capture the key change nodes of information propagation; The knowledge transfer and distillation module is used to transfer the knowledge of large, high-precision models to lightweight models to achieve model compression and performance preservation; The online learning and continuous optimization module is used to continuously learn from new data and optimize model performance after deployment. At the same time, it prevents catastrophic forgetting through experience replay and achieves accurate prediction of turning points and trends in network information dissemination.
Citation Information
Patent Citations
Self-distillation training method and scalable dynamic prediction method of convolutional neural network
CN110472730A
Compression method and device of neural network model, image processing method and processor
CN111275190A
Neural network model searching method and device, image processing method and processor
CN111340219A
Attention mechanism-based information propagation evolution trend prediction method
CN112182423A
Automatic pruning method and platform for convolutional neural network general compression architecture
CN112396181A
Cited By
SiC groove etching optimization method and system based on deep learning
CN120744411A
SiC trench etching optimization method and system based on deep learning
CN120744411B
Robust target recognition method and device based on noise analysis and storage medium thereof
CN121093126A
A robust target recognition method based on noise analysis, device and storage medium thereof
CN121093126B
Incremental learning driven security equipment self-optimization method and system
CN121119313A