Sensor network intention reasoning method based on time sequence knowledge graph

By combining the construction of a temporal knowledge graph with reinforcement learning algorithms, the dynamic problem of intent reasoning in sensor networks is solved, enabling intelligent management and efficient resource allocation of sensor networks, and improving the accuracy of intent analysis and the system's predictive capabilities.

CN120952176APending Publication Date: 2025-11-14XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511065818.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Sensor network management systems cannot effectively utilize dynamic information, resulting in insufficient accuracy and real-time performance of intent reasoning, inefficient resource allocation, and impact on user service quality.

Method used

An intent reasoning method based on temporal knowledge graphs is adopted. By collecting and preprocessing feature data of sensor nodes, a temporal knowledge graph is constructed, and an intent reasoning model is designed in combination with reinforcement learning algorithms to realize intelligent management of sensor networks.

Benefits of technology

It improves the accuracy of intent analysis and the system's predictive capabilities, enhances the network's adaptability and resource allocation efficiency in dynamic environments, and realizes intelligent intent reasoning and decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952176A_ABST
    Figure CN120952176A_ABST
Patent Text Reader

Abstract

The invention discloses a sensor network intention reasoning method based on a time sequence knowledge graph. The method comprises the following steps: S1, preprocessing acquired feature data and time sequence data; s2, constructing a time sequence knowledge graph as a data set; s3, constructing a sensor network intention reasoning model based on a reinforcement learning algorithm; s4, dividing the constructed data set into a training set and a test set, inputting data of the training set into the constructed intention reasoning model of the sensor network to obtain a predicted value, and adjusting parameters of the intention reasoning model of the sensor network through training until the obtained predicted value is closest to an actual value; and S5, collecting sensor node time sequence characteristic data in real time, inputting the data into the trained and optimized sensor network intention reasoning model, and performing intention reasoning prediction. According to the method, the accuracy and the real-time performance of intention recognition in the sensor network are improved, intelligent intention reasoning and decision support are realized, and sensor network resources are efficiently managed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of knowledge reasoning technology of knowledge graphs, and specifically relates to a sensor network intent reasoning method based on temporal knowledge graphs. Background Technology

[0002] Internet of Things (IoT) devices and applications process data extracted from Wireless Sensor Networks (WSN) devices and transmit the results to remote locations for further analysis and decision-making. A WSN network is a specific type of distributed network system. Its terminals consist of a series of spatially dispersed nodes capable of sensing the real-world environment and possessing network communication capabilities that can be adjusted according to specific application scenarios. Through interconnected wireless network communication, they collaboratively monitor the development of target events or collect information about targets, achieving the goal of securely and efficiently acquiring valuable information.

[0003] Despite the broad application prospects of sensor networks, their management and maintenance face numerous challenges, primarily including the following: Firstly, sensor networks are densely populated with a vast number of nodes distributed across a wide geographical area, and may encounter problems such as node failure and signal interference. Secondly, sensor networks are highly dynamic; the sensors and observed objects within the network may be mobile, resulting in a highly dynamic topology and information transmission paths. Traditional network management systems struggle to adapt to the highly elastic and dynamic business needs of complex sensor networks, failing to acquire real-time dynamic data of all network resources and inefficiently allocate resources, thereby reducing the user's quality of service experience.

[0004] Intent-driven Networks (IDNs) offer an effective solution to the aforementioned problems. IDNs automatically manage and optimize networks by translating high-level user intents into specific network configurations and operations, thus achieving automated and closed-loop optimization of network services. IDNs hold immense potential in applications from users and devices to data centers and the cloud, offering a programmable and customizable automated network that integrates deep application intent mining, global network state awareness, and real-time network configuration optimization. Intent reasoning is a crucial task in sensor network management, enabling automated network configuration and optimization. Traditional intent reasoning methods often rely on static data, neglecting the dynamic and spatiotemporal attributes of real-world data, leading to poor prediction and recommendation performance. To improve the accuracy and relevance of intent reasoning, researchers have recently begun exploring methods based on temporal knowledge graphs.

[0005] Temporal Knowledge Graph (TKG) is a dynamic knowledge graph that expands each fact into a quadruple (subject, relation, tail entity, timestamp) by introducing timestamps to the triples of a Knowledge Graph (KG). This effectively manages dynamically evolving temporal knowledge and provides crucial support for applications tightly coupled with time. TKG reasoning focuses on inferring new facts from known facts, effectively predicting future events by mining information from historical temporal quadruples.

[0006] However, existing research on knowledge graphs focuses on static knowledge reasoning, neglecting the temporal information of knowledge graphs, making it impossible to dynamically update knowledge graphs and utilize dynamic information to realize the evolution of knowledge graphs over time. Summary of the Invention

[0007] To overcome the shortcomings of the existing technology, the present invention aims to provide a sensor network intent reasoning method based on temporal knowledge graph. By using dynamic modeling of temporal knowledge graph and intent-driven network management mechanism, the method improves the accuracy and real-time performance of intent recognition in sensor networks, realizes intelligent intent reasoning and decision support, and efficiently manages sensor network resources.

[0008] In this invention, for sensor networks:

[0009] Intent recognition refers to extracting the corresponding attribute requirements for sensor nodes, bandwidth requirements, latency requirements, etc., for the sensor network from the user's proposed task requirements, and mapping them into a computable and verifiable structured intent representation.

[0010] Intent reasoning refers to predicting unknown intents by leveraging the dynamic evolution capabilities of temporal knowledge graphs and combining temporal logic rules with reinforcement learning, based on identified structured intents.

[0011] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0012] A sensor network intent reasoning method based on temporal knowledge graph includes the following steps;

[0013] S1: Collect feature data of sensor member nodes in the sensor network, and synchronously collect time series data of sensor nodes. This data records the working status of the sensor at different time points and data collection timestamp information. Preprocess the feature data and time series data after collection.

[0014] S2: Using the preprocessed sensor node feature data and time series data, construct a time series knowledge graph as the dataset, with sensor nodes as entities in the graph and the relationships and time relationships between nodes as edges in the graph;

[0015] S3: Design reinforcement learning strategies and build a sensor network intent reasoning model based on reinforcement learning algorithms;

[0016] S4: Divide the constructed dataset into a training set and a test set, where the training set is used for model training and the test set is used for model evaluation. Input the training set data into the constructed sensor network intent reasoning model to obtain the predicted value. Adjust the parameters of the sensor network intent reasoning model through training until the predicted value is closest to the actual value.

[0017] S5: Real-time acquisition of time-series feature data from sensor nodes is input into the trained and optimized sensor network intent reasoning model for intent reasoning prediction.

[0018] Specifically, S1 is:

[0019] Step 1: Collect feature data from each member node in the sensor network. The feature data includes sensor type, sensor ID, data type, communication capability, and resolution.

[0020] Step two involves preprocessing the collected feature data and time series data. The preprocessing methods include data cleaning, data normalization, feature dimensionality reduction, and data labeling.

[0021] The data cleaning process specifically involves cleaning the collected feature data to remove duplicate, missing, or abnormal data; filling in missing data using interpolation methods (such as linear interpolation and time series interpolation); detecting and correcting abnormal data through statistical analysis or rule-based methods; detecting and removing duplicate node records from sensor node feature data and automatically correcting abnormal fields using linear interpolation; and deduplicating time series data by timestamp and retaining only the unique sampled frame.

[0022] The data normalization specifically involves normalizing different types of feature data to eliminate differences in dimensions and data ranges; and unifying the dimensions of sensor node feature data and time feature data.

[0023] Specifically, the feature dimensionality reduction involves performing dimensionality reduction processing on high-dimensional feature data according to the actual needs of the sensor network. Principal component analysis (PCA), linear discriminant analysis (LDA), or other dimensionality reduction algorithms are used to extract key features and reduce computational complexity. Dimensionality compression is performed on sensor feature data with node attribute dimensions > 20 to reduce storage and computational burden.

[0024] The data annotation specifically involves annotating the collected sensor node feature data according to the service requirements of the sensor network, marking its corresponding intent category or application scenario, and providing labeled samples for subsequent model training.

[0025] Specifically, S2 is:

[0026] The labeled sensor node feature data is represented in the form of a time-series knowledge graph quadruple, which includes a head entity, a tail entity, a relation, and a time series.

[0027] The head entity describes the sensor member nodes and their capability attributes, including observation and sensing capabilities, communication and transmission capabilities, storage and computing capabilities, environmental adaptability, and energy endurance.

[0028] The relationship is used to describe the relationship between the head entity and the tail entity;

[0029] The tail entity is used to describe the specific characteristic attributes of the relationship;

[0030] The time series is used to describe the changes in sensor node feature data over time, including the timestamp sequence of data acquisition and the numerical sequence of various capability attribute indicators at the corresponding time points.

[0031] The observation and sensing capability is used to describe the ability of a sensor to observe and sense a target object. The main indicators include observation range, observation resolution, observation accuracy, observation frequency, and resolution.

[0032] The communication transmission capability is used to describe the sensor's ability to complete network access, communication and data transmission. The main indicators include communication / interface protocol, real-time-non-real-time communication type, upload and download bandwidth, transmission rate and signaling latency.

[0033] The storage and computing capabilities are used to describe the sensor's performance attributes, with key indicators including the number of processing bits (16 / 32 bits), storage capacity, positioning function, and synchronous communication function.

[0034] The environmental adaptability is used to describe the degree to which a sensor can adapt to the operating environment, and the indicators include operating temperature range, humidity range, protection level, etc.

[0035] The S3 uses a reinforcement learning algorithm for inference and prediction. The algorithm can be divided into six parts: representation of the environment state, representation of the behavior action, design of the reward function, design of the policy network, design of the value network, and setting of the optimization function.

[0036] The representation of environmental state is used to characterize the external environment and internal state during the intent reasoning process of sensor network, integrating entity, relation, and time series information in temporal knowledge graph;

[0037] The representation of behavior and action defines the operations that an intelligent agent can perform in a specific environmental state;

[0038] The reward function is designed to evaluate the effect of an agent performing a certain action. Its output depends on the environmental state and the corresponding action. When an agent performs a specific action in a certain environmental state, the reward function will give a corresponding reward value based on the result of the action and affect the agent's choice of action. The policy network is designed to generate the agent's behavior policy, that is, to output the probability distribution of each action given the environmental state.

[0039] The training of the policy network depends on the feedback provided by the reward function. The higher the reward value of an action, the higher its probability of appearing in the policy network. At the same time, the output of the policy network is also directly affected by the environmental state.

[0040] Value networks are designed to evaluate the expected long-term cumulative reward that an agent can obtain under a specific environmental state or after performing a certain action.

[0041] The optimization function is set to maximize the cumulative reward obtained by the agent by adjusting the parameters of the policy network and the value network.

[0042] Among these six parts, environmental state and behavioral actions are the basic inputs, the reward function is the feedback bridge, the policy network and value network are the decision-making core, and the optimization function is the driver of parameter adjustment. Together, they enable the reinforcement learning algorithm to achieve accurate reasoning of the sensor network's intent.

[0043] Specifically, S3 is:

[0044] S101, the representation of the environmental state, let S represent the state space, and the current state of the environment is represented by a quintuple s. l =(e l ,t l ,e p ,t p ,r p ) indicates that (e l ,t l ) represents the node visited by the agent in the current step, (e p ,t p ,r p The target to be inferred and predicted is the entity node e; the reasoning process of the agent starts from the initial entity node e in this inference. p Initially, the initial state of the entire environment is represented as s. p =(ep ,t p ,e p ,t p ,r p );

[0045] S102, the representation of the action: In the TKG reasoning task, the action depends on the current environmental state of the agent. The agent selects the corresponding action based on the historical facts associated with the current entity. The selectable action is A. l ={(r l ,e l ,t l ),t l ≤t cur ,t l ≤t p}, where A l r represents the set of actions that the agent can choose at the current node. l Represents all entity relationships associated with the current node, e l t represents all entity nodes associated with the current node. l t represents the timestamp of the corresponding relationship and node interaction. cur The current inference time represents the upper limit of the time range for defining the action. p This indicates the upper limit of the historical window, limiting the time range within which the agent can refer to historical facts.

[0046] S103, the reward function is designed such that, in each round of training, the reward varies depending on the prediction result, and the reward function is designed as R = R intent +R concept , where R intent The reward for the prediction result is 1 if the agent reaches the correct target intention state at the end of the inference process; otherwise, it receives 0. concept The reward is for conceptual semantic information; when the agent searches for a state with an incorrect target intent, i.e., R... intent When the prediction is incorrect (e.g., 0), this soft reward is added to guide the model's learning.

[0047] R concept The network management concept is further subdivided, such as network load, latency, and bandwidth usage; the specific optimization reward function is as follows:

[0048] R context =R load +R delay +R bandwidth

[0049] in:

[0050] Rload A reward is given for network load when the predicted network load and the target load are within the same range.

[0051] R delay A reward is given for network latency when the predicted network latency is within the same range as the target latency.

[0052] R bandwidth A reward is given for bandwidth usage when the predicted bandwidth usage falls within the same range as the target bandwidth usage.

[0053] Specific reward calculation:

[0054]

[0055] Where I{·} is an indicator function, which takes the value 1 when the condition is true and 0 otherwise;

[0056] S104, the policy network is designed so that the agent selects appropriate actions based on the current environmental state and the policy network model, thereby changing its own environmental state; the policy network uses π θ (a l |s l )=P(a l |s l The model is represented by θ, where θ is the parameter of the entire model. The entire model includes the embedding representation of entity nodes and entity relationships, the encoding of inference paths, and the scoring of the agent's optional actions in the current environment state.

[0057] The entity nodes and the relationships between them are embedded and represented. For entity nodes, time information is fused with the feature information of the entity node. For entity node e at a certain moment... l The embedding result is represented as e i This represents the embedding representation of entity nodes, where Φ(Δt) represents the current time of the node and is calculated using the formula: Φ(Δt)=σ(ωΔt+b), where Δt=t p -t l , || indicates a concatenation operation on the embedded vectors;

[0058] The agent employs multi-hop reasoning to predict and reason about corresponding events. Therefore, each reasoning path needs to be encoded. This part uses a Gate Recurrent Unit (GRU) network to encode the paths during the agent's reasoning process. For the current environmental state of the agent, the historical reasoning path h is encoded. cur =((e) p ,t p),r1,(e1,t1),…,r l ,(e l ,t l The following formula can be used to calculate:

[0059]

[0060] For each state, we need to calculate its possible actions and the probability of state transitions. For state s l By weighting the target entity nodes and relationships selectable by the intelligent agent, and using multi-layer perception (MLP) to calculate the current state information, the expected optimal target node e is output. i and the optimal entity relationship r i The representation is as follows:

[0061]

[0062] In the formula, W1, W e W r For learnable variables;

[0063] Then, according to φ(a) l ,s l ) = 0.5* <e i ,e n >+0.5* <r i ,r n Calculate the similarity between the target entity nodes and relationships of candidate actions as the score φ(a) for the corresponding action. l ,s l The policy network calculates the probability of each action based on the action score, and then selects the corresponding entity relationship and target entity node according to the probability.

[0064] S105, the design of the value network, the value network is based on the current environmental state s l A score is calculated by using a multi-layer perception mechanism to determine the value of the environment in which the agent is currently located. Among them W v These are learnable parameters;

[0065] S106, the optimization function is set by using the Proximal Policy Optimization (PPO) algorithm to optimize the model parameters. Here, GRU, MLP, Φ, and the embedded learning representations of entities and relations are all optimizable parameter variables. The objective optimization function is expressed as:

[0066]

[0067] In the formula, ρ t Importance sampling, i.e., the ratio of the difference between the new strategy and the old strategy:

[0068]

[0069] `clip` is a truncation function. When the value of the importance sample is higher (lower) than the corresponding upper (lower) limit, that is, when a destructive large update occurs during model training, this function will limit the magnitude of the model parameter update to help the algorithm remain stable during the learning process and ensure the effectiveness of model learning.

[0070] The dominance function can be calculated using the following formula:

[0071]

[0072] In strategy Under the current state s t The reward is the reward obtained by selecting the corresponding action 'a'.

[0073] Specifically, S4 is:

[0074] Dataset partitioning: The constructed time-series knowledge dataset is divided into a training set and a test set in a 7:3 ratio. Stratified sampling is used during the partitioning to ensure that the types of sensor nodes and the time-series distribution characteristics in the training set and the test set are consistent with the original dataset, avoiding model training bias due to uneven data distribution. The training set is used for model training, and the test set is used for model evaluation.

[0075] Model training: Training set data is input into the constructed sensor network intent reasoning model in batches according to time sequence. For each batch of input data, the model generates an intent reasoning prediction value based on the current policy network. The deviation between the prediction value and the actual intent is calculated through the reward function, and the reward signal is fed back to the value network. The value network evaluates the long-term value of the action in the current state in combination with temporal features. The policy network adjusts the behavior probability distribution according to the reward value and the value evaluation result. The optimization function takes the loss of the policy network and the loss of the value network as input and updates the network parameters through backpropagation of the optimizer.

[0076] Model optimization: Repeat the above training process. Each iteration of the full training set is called an epoch. After each epoch, the model performance is evaluated using the validation set (10%-20% of the training set). If the prediction error on the validation set does not decrease for several consecutive epochs, an early stopping mechanism is triggered to avoid model overfitting. Finally, when the average error between the predicted and actual values ​​on the training set is lower than a preset threshold and the performance of the validation set is stable, training ends.

[0077] The S5 system collects real-time time-series feature data from sensor nodes and preprocesses it. The preprocessed real-time data is then converted into a time-series knowledge graph fragment consistent with the training data format and input into the trained and optimized sensor network intent reasoning model for intent reasoning prediction. The model's policy network generates an intent reasoning probability distribution based on the current input data and historical reasoning experience, selecting the result with the highest probability as the final predicted intent.

[0078] An electronic device, comprising:

[0079] Memory, used to store programs;

[0080] A processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method.

[0081] A storage medium storing a computer program that, when run on a processor, causes the processor to perform the method.

[0082] The beneficial effects of this invention are:

[0083] 1. Improve the accuracy of intent analysis. By synchronously collecting and preprocessing features and time series data through S1, constructing a time series knowledge graph to establish time correlations through S2, and combining the utilization of time series dynamic features by the reinforcement learning model in S3 and the training optimization in S4, the system can effectively distinguish semantically similar intents, filter out interference items, and improve the accuracy of intent analysis.

[0084] 2. Enhance the system's predictive capabilities. Based on the historical dynamic information provided by the S2 temporal knowledge graph, the reward function in S3 reinforcement learning combined with temporal trend feedback, and the inference model optimized by S4 training, the integration of historical and current data can more accurately understand and predict changes in network state and user intent. The integration of temporal data enhances the quality of network decision-making, enabling the network to make more reasonable decisions based on historical and current data.

[0085] 3. By using reinforcement learning algorithms, this invention can train the intent reasoning model more effectively, improving the efficiency of model training and the model's generalization ability.

[0086] 4. Enhanced system adaptability: In a dynamically changing network environment, this invention can quickly adapt to and respond to changes in network status, thereby improving the overall efficiency of the system. Attached Figure Description

[0087] Figure 1 This is a flowchart of an intent reasoning method for sensor networks based on temporal knowledge graphs.

[0088] Figure 2A flowchart for data preprocessing.

[0089] Figure 3 Build a flowchart for time-series knowledge graphs.

[0090] Figure 4 These are the capability attributes of sensor member nodes.

[0091] Figure 5 This is a diagram illustrating the overall framework of a temporal knowledge graph reasoning and prediction algorithm based on reinforcement learning.

[0092] Figure 6 This is a flowchart illustrating the algorithm steps of a temporal knowledge graph intent reasoning model based on reinforcement learning. Detailed Implementation

[0093] The present invention will now be described in further detail with reference to the accompanying drawings.

[0094] To address the problems existing in the prior art, this invention provides a sensor network intent reasoning method based on temporal knowledge graphs. The technical solution of this invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0095] To address the challenges of increasingly large-scale sensor networks, with a surge in node types, numbers, and service types, traditional network management methods struggle to efficiently allocate resources and suffer from high network configuration complexity. This invention employs knowledge graph technology, network data collection technology, and artificial intelligence technology to enable intent reasoning in sensor networks, thereby promoting intelligent network management.

[0096] This invention discloses a sensor network intent reasoning method based on temporal knowledge graphs. The technical solutions provided in the embodiments of this invention will be described below.

[0097] like Figure 1 As shown, the sensor network intent reasoning method based on temporal knowledge graph provided in this embodiment of the invention includes the following steps:

[0098] Step S1, taking a ground-based collaborative monitoring network as an example, the user collects and preprocesses the feature data of the sensor member nodes in the sensor network. The collected sensor member nodes include radar sensors, imaging sensors, multispectral sensors, and infrared sensors of UAVs, etc., and feature data such as battery level, monitoring distance, resolution, and band of the sensor nodes are collected. The collected feature data and its time-series feature data of the sensor member nodes are preprocessed, and the data preprocessing process is as follows: Figure 2 As shown, it includes four steps: data cleaning, data normalization, feature dimensionality reduction, and data labeling.

[0099] Data cleaning: The collected feature data is cleaned to remove duplicate, missing, or outlier data. For missing data, interpolation methods (such as linear interpolation and time series interpolation) are used to impute it; for outlier data, statistical analysis or rule-based methods are used for detection and correction.

[0100] Data normalization: Normalizing different types of feature data to eliminate differences in units and data ranges;

[0101] Feature dimensionality reduction: Based on the actual needs of sensor networks, high-dimensional feature data is reduced in dimensionality by using principal component analysis (PCA), linear discriminant analysis (LDA) or other dimensionality reduction algorithms to extract key features and reduce computational complexity.

[0102] Data annotation: Based on the business needs of the sensor network, the collected feature data is annotated to mark its corresponding intent category or application scenario, such as clear weather monitoring, continuous monitoring, etc., to provide labeled samples for subsequent model training.

[0103] Step S2 involves constructing a time-series knowledge graph as a dataset using the feature data and time-series data from collected radar imaging, multispectral, and infrared sensor nodes. Specifically, the capability attributes of each sensor node at the time of acquisition, such as battery level, monitoring distance, resolution, and band, are written into the knowledge graph in the form of quadruples, following the sequence shown below. Figure 3 The process of constructing a time-series knowledge graph is shown below, which involves constructing a sensor time-series knowledge graph.

[0104] Sensor node feature data is represented in the form of quadruples of a time-series knowledge graph, which includes head entity, tail entity, relation, and time series.

[0105] like Figure 4 As shown, the capabilities of sensor member nodes are divided into five categories: observation and sensing capabilities, communication and transmission capabilities, storage and computing capabilities, environmental adaptability, and energy endurance. Observation and sensing capabilities describe the sensor's ability to observe and sense target objects, with key indicators including observation range, observation resolution, observation accuracy, observation frequency, and resolution. Communication and transmission capabilities describe the sensor's ability to complete network access, communication, and data transmission, with key indicators including communication / interface protocols, real-time / non-real-time communication types, upload / download bandwidth, transmission rate, and signaling latency. Storage and computing capabilities describe the sensor's performance attributes, with key indicators including processing power (16 / 32 bits), storage capacity, positioning function, and synchronous communication function. Environmental adaptability describes the degree to which the sensor can adapt to the operating environment, with indicators including operating temperature range, humidity range, and protection level.

[0106] Step S3: Design a reinforcement learning strategy and construct a sensor network intent reasoning model based on the reinforcement learning algorithm. The overall framework of the algorithm is as follows: Figure 5 As shown;

[0107] The algorithm can be divided into six parts: representation of the environment state, representation of actions, design of the reward function, design of the policy network, design of the value network, and setting of the optimization function. The training steps of the algorithm are as follows: Figure 6 As shown;

[0108] S101, the representation of the environmental state, let S represent the state space, and the current state of the environment is represented by a quintuple s. l =(e l ,t l ,e p ,t p ,r p ) indicates that (e l ,t l ) represents the node visited by the agent in the current step, (e p ,t p ,r p The target to be inferred and predicted is the entity node e. The reasoning process of the agent starts from the initial entity node e in this inference. p Initially, the initial state of the entire environment can be represented as s. p =(e p ,t p ,e p ,t p ,r p )

[0109] S102, the representation of the action: In the TKG reasoning task, the action depends on the current environmental state of the agent. The agent selects the corresponding action based on the historical facts associated with the current entity. The selectable action is A. l ={(r l ,e l ,t l ),t'≤t cur ,t'≤t p}, where r' represents the current node e l All associated entity relationships, where e' represents all associated entity nodes.

[0110] S103, the reward function is designed such that, in each round of training, the reward varies depending on the prediction result, and the reward function is designed as R = R intent +R concept , where R intentThe reward for the prediction result is 1 if the agent reaches the correct target intention state at the end of the inference process; otherwise, it receives 0. concept The reward is for conceptual semantic information; when the agent searches for a state with an incorrect target intent, i.e., R... intent =0 When the prediction is incorrect, this soft reward is added to guide the model's learning. Taking the objective "confirm fire risk within 10 minutes" as an example, if the agent correctly determines the fire risk level within 10 minutes, it receives a terminal reward of 1; otherwise, it receives 0. Network load, latency, and bandwidth usage are respectively given as R. load R delay R bandwidth Soft rewards.

[0111] R concept It can be further subdivided based on different concepts in network management, such as network load, latency, and bandwidth usage. The specific optimization reward function is as follows:

[0112] R context =R load +R delay +R bandwidth

[0113] in:

[0114] R load A reward is given for network load when the predicted network load and the target load are within the same range.

[0115] R delay A reward is given for network latency when the predicted network latency is within the same range as the target latency.

[0116] R bandwidth A reward is given for bandwidth usage when the predicted bandwidth usage falls within the same range as the target bandwidth usage.

[0117] Specific reward calculation:

[0118]

[0119] Where I{·} is an indicator function, which takes the value 1 when the condition is true, and 0 otherwise.

[0120] S104, the policy network is designed so that the agent selects appropriate actions based on the current environmental state and the policy network model, thereby changing its own environmental state. The policy network can be represented by π. θ (a l |s l )=P(a l |s lThe model is represented by θ, where θ is the parameter of the entire model. The entire model includes the embedding representation between entity nodes and entity relationships, the encoding of inference paths, and the scoring of the agent's optional actions in the current environment state.

[0121] The entity nodes and the relationships between them are embedded and represented. For entity nodes, time information is fused with the feature information of the entity node. For entity node e at a certain moment... l The embedding result is represented as e i This represents the embedding representation of entity nodes, where Φ(Δt) represents the current time of the node and can be calculated using the formula: Φ(Δt) = σ(ωΔt + b), where Δt = t p -t l , || indicates a concatenation operation on the embedded vectors.

[0122] The agent employs multi-hop reasoning to predict and reason about corresponding events. Therefore, each reasoning path needs to be encoded. This part uses a Gate Recurrent Unit (GRU) network to encode the paths during the agent's reasoning process. For the current environmental state of the agent, the historical reasoning path h is encoded. cur =((e) p ,t p ),r1,(e1,t1),…,r l ,(e l ,t l The following formula can be used to calculate:

[0123]

[0124] For each state, we need to calculate its possible actions and the probability of state transitions. For state s l By weighting the target entity nodes and relationships selectable by the intelligent agent, and using multi-layer perception (MLP) to calculate the current state information, the expected optimal target node e is output. i and the optimal entity relationship r i The representation is as follows:

[0125]

[0126] In the formula, W1, W e W r These are learnable variables.

[0127] Then, according to φ(a) l ,s l ) = 0.5* <e i,e n >+0.5* <r i ,r n Calculate the similarity between the target entity nodes and relationships of candidate actions as the score φ(a) for the corresponding action. l ,s l The policy network can calculate the probability of each action based on the action score, and then select the corresponding entity relationship and target entity node according to the probability.

[0128] S105, the design of the value network, the value network is based on the current environmental state s l A score is calculated by using a multi-layer perception mechanism to determine the value of the environment in which the agent is currently located. Among them W v These are learnable parameters.

[0129] S106, the optimization function is set by using the Proximal Policy Optimization (PPO) algorithm to optimize the model parameters. Here, GRU, MLP, Φ, and the embedded learning representations of entities and relations are all optimizable parameter variables. The objective optimization function is expressed as:

[0130]

[0131] In the formula, ρ t Importance sampling, i.e., the ratio of the difference between the new strategy and the old strategy:

[0132]

[0133] `clip` is a truncation function. When the value of the importance sample is higher (lower) than the corresponding upper (lower) limit, that is, when a destructive large update occurs during model training, this function will limit the magnitude of the model parameter update to help the algorithm remain stable during the learning process and ensure the effectiveness of model learning.

[0134] The dominance function can be calculated using the following formula:

[0135]

[0136] In strategy Under the current state s t The reward is the reward obtained by selecting the corresponding action 'a'.

[0137] Step S4: Input a portion of the dataset into the constructed sensor network intent reasoning model to obtain predicted values. Adjust the model parameters through training until the predicted values ​​are closest to the actual values.

[0138] The model training algorithm first sets the number of training iterations (N), then randomly selects a batch of training data in each iteration, including entities, relations, start time, and end time. The algorithm sets the inference path length and, in each inference iteration, calculates all possible action spaces of the current entity before a given time, records the current state, and calculates the probability of each action based on that state. The agent selects an action based on these probabilities and calculates the corresponding value. After interacting with the environment, the algorithm updates the entity, time, and historical information. If the end of the inference path is reached, the algorithm calculates the reward for each step. Afterward, the algorithm stores the current state, the selected action and its probability, and the value of the current state. By calculating the reward and advantage value, the algorithm finally optimizes the model parameters using the Proximal Policy Optimization (PPO) algorithm until all training iterations are completed, returning the final optimized model parameters.

[0139] Step S5: Real-time acquisition of sensor time series feature data is input into the model for intent reasoning. For example, if the temperature sensor temperature continues to rise, it is determined to be a high fire risk. Imaging sensors are immediately dispatched for monitoring and inspection, and fire extinguishing bombs are dispatched to the monitoring coordinates to realize sensor network intent reasoning based on time-series knowledge graph.

[0140] The model was trained according to the method described in the above embodiments, and will not be repeated here.

[0141] This establishes a sensor network intent reasoning method based on temporal knowledge graphs.

[0142] Based on the methods described in the above embodiments, this invention provides an electronic device. The device may include a memory for storing a program and a processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor performs the methods described in the above embodiments.

[0143] Based on the methods in the above embodiments, this embodiment of the invention provides a storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0144] The processor in the embodiments of the present invention can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any other combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.

[0145] The method steps in the embodiments of the present invention can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0146] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented, in whole or in part, as a computer program product, the computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)). The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention, and within the spirit and principles of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A sensor network intent reasoning method based on temporal knowledge graph, characterized in that, Includes the following steps; S1: Collect feature data of sensor member nodes in the sensor network, synchronously collect time series data of sensor nodes, and perform preprocessing; S2: Using the preprocessed sensor node feature data and time series data, construct a time series knowledge graph as the dataset, with sensor nodes as entities in the graph and the relationships and time relationships between nodes as edges in the graph; S3: Design reinforcement learning strategies and build a sensor network intent reasoning model based on reinforcement learning algorithms; S4: Divide the constructed dataset into a training set and a test set. The training set is used for model training, and the test set is used for model evaluation. Input the training set data into the constructed sensor network intent reasoning model to obtain the predicted value. Adjust the parameters of the sensor network intent reasoning model through training until the predicted value is closest to the actual value. S5: Real-time acquisition of time-series feature data from sensor nodes is input into the trained and optimized sensor network intent reasoning model for intent reasoning prediction.

2. The sensor network intent reasoning method based on temporal knowledge graph according to claim 1, characterized in that, Specifically, S1 is: Step one: Collect feature data from each member node in the sensor network. The feature data includes sensor type, sensor ID, data type, communication capability, and resolution; Step 2 involves preprocessing the collected feature data and time series data. The preprocessing methods include data cleaning, data normalization, feature dimensionality reduction, and data labeling. The data cleaning process specifically involves cleaning the collected feature data to remove duplicate, missing, or abnormal data; filling in missing data using interpolation methods; detecting and correcting abnormal data through statistical analysis or rule-based methods; detecting and removing duplicate node records from sensor node feature data and automatically correcting abnormal fields using linear interpolation; and deduplicating time series data by timestamp and retaining only the unique sampled frame. The data normalization specifically involves normalizing different types of feature data to eliminate differences in dimensions and data ranges; and unifying the dimensions of sensor node feature data and time feature data. Specifically, the feature dimensionality reduction involves performing dimensionality reduction processing on high-dimensional feature data according to the actual needs of the sensor network. Principal component analysis, linear discriminant analysis, or other dimensionality reduction algorithms are used to extract key features and reduce computational complexity. Dimensionality compression is performed on sensor feature data with node attribute dimensions > 20 to reduce storage and computational burden. The data annotation specifically involves annotating the collected sensor node feature data according to the service requirements of the sensor network, marking its corresponding intent category or application scenario, and providing labeled samples for subsequent model training.

3. The sensor network intent reasoning method based on temporal knowledge graph according to claim 2, characterized in that, Specifically, S2 is: The labeled sensor node feature data is represented in the form of a time-series knowledge graph quadruple, which includes a head entity, a tail entity, a relation, and a time series. The head entity describes the sensor member nodes and their capability attributes, including observation and sensing capabilities, communication and transmission capabilities, storage and computing capabilities, environmental adaptability, and energy endurance. The relationship is used to describe the relationship between the head entity and the tail entity; The tail entity is used to describe the specific characteristic attributes of the relationship; The time series is used to describe the changes in sensor node feature data over time, including the timestamp sequence of data acquisition and the numerical sequence of various capability attribute indicators at the corresponding time points.

4. The sensor network intent reasoning method based on temporal knowledge graph according to claim 3, characterized in that, The observation and perception capability is used to describe the sensor's ability to observe and perceive target objects; The communication transmission capability is used to describe the sensor's ability to complete network access, communication, and data transmission. The storage and computing capabilities are used to describe the sensor's performance attributes; The environmental adaptability is used to describe the degree to which a sensor can adapt to its operating environment.

5. The sensor network intent reasoning method based on temporal knowledge graph according to claim 4, characterized in that, The S3 uses a reinforcement learning algorithm for inference and prediction. The algorithm consists of six parts: representation of the environment state, representation of the behavior action, design of the reward function, design of the policy network, design of the value network, and setting of the optimization function. The representation of environmental state is used to characterize the external environment and internal state during the intent reasoning process of sensor network, integrating entity, relation, and time series information in temporal knowledge graph; The representation of behavior and action defines the operations that an intelligent agent can perform in a specific environmental state; The reward function is designed to evaluate the effect of an agent performing a certain action. Its output depends on the environmental state and the corresponding action. When an agent performs a specific action in a certain environmental state, the reward function will give a corresponding reward value based on the result of the action and affect the agent's choice of action. The policy network is designed to generate the agent's behavior policy, that is, to output the probability distribution of each action given the environmental state. The training of the policy network depends on the feedback provided by the reward function. The higher the reward value of an action, the higher its probability of appearing in the policy network. At the same time, the output of the policy network is also directly affected by the environmental state. Value networks are designed to evaluate the expected long-term cumulative reward that an agent can obtain under a specific environmental state or after performing a certain action. The optimization function is set to maximize the cumulative reward obtained by the agent by adjusting the parameters of the policy network and the value network. Among these six parts, environmental state and behavioral actions are the basic inputs, the reward function is the feedback bridge, the policy network and value network are the decision-making core, and the optimization function is the driver of parameter adjustment. Together, they enable the reinforcement learning algorithm to achieve accurate reasoning about the intentions of the sensor network.

6. The sensor network intent reasoning method based on temporal knowledge graph according to claim 5, characterized in that, Specifically, S3 is: S101, the representation of the environmental state, let S represent the state space, and the current state of the environment is represented by a quintuple s. l =(e l ,t l ,e p ,t p ,r p ) indicates that (e l ,t l ) represents the node visited by the agent in the current step, (e p ,t p ,r p The target to be inferred and predicted is the entity node e; the agent's reasoning process begins from the initial entity node e in this inference. p Initially, the initial state of the entire environment is represented as s. p =(e p ,t p ,e p ,t p ,r p ); S102, the representation of the action: In the TKG reasoning task, the action depends on the current environmental state of the agent. The agent selects the corresponding action based on the historical facts associated with the current entity. The selected action is A. l ={(r l ,e l ,t l ),t l ≤t cur ,t l ≤t p }, where A l r represents the set of actions selected by the agent at the current node. l Represents all entity relationships associated with the current node, e l t represents all entity nodes associated with the current node. l t represents the timestamp of the corresponding relationship and node interaction. cur The current inference time represents the upper limit of the time range for defining the action. p This indicates the upper limit of the historical window, limiting the time range within which the agent can refer to historical facts. S103, the reward function is designed such that, in each round of training, the reward varies depending on the prediction result, and the reward function is designed as R = R intent +R concept , where R intent The reward for the prediction result is 1 if the agent reaches the correct target intention state at the end of the inference process; otherwise, it receives 0. concept The reward is for conceptual semantic information; when the agent searches for a state with an incorrect target intent, i.e., R... intent When the prediction is incorrect (e.g., 0), this soft reward is added to guide the model's learning. R concept Based on different concepts in network management, the specific optimized reward function is as follows: R context =R load +R delay +R bandwidth in: R load A reward is given for network load when the predicted network load and the target load are within the same range. R delay A reward is given for network latency when the predicted network latency is within the same range as the target latency. R bandwidth A reward is given for bandwidth usage when the predicted bandwidth usage is within the same range as the target bandwidth usage. Specific reward calculation: Where I{·} is an indicator function, which takes the value 1 when the condition is true and 0 otherwise; S104, the policy network is designed so that the agent selects appropriate actions based on the current environmental state and the policy network model, thereby changing its own environmental state; the policy network uses π θ (a l |s l )=P(a l |s l The model is represented by θ, where θ is the parameter of the entire model. The entire model includes the embedding representation of entity nodes and entity relationships, the encoding of inference paths, and the scoring of the agent's optional actions in the current environment state. The entity nodes and the relationships between them are embedded and represented. For entity nodes, time information is fused with the feature information of the entity node. For entity node e at a certain moment... l The embedding result is represented as e i This represents the embedding representation of entity nodes, where Φ(Δt) represents the current time of the node and is calculated using the formula: Φ(Δt)=σ(ωΔt+b), where Δt=t p -t l , || indicates a concatenation operation on the embedded vectors; The agent employs multi-hop reasoning to predict and reason about corresponding events, encoding each reasoning path. This part uses a gated recurrent unit network to encode the paths during the agent's reasoning process. For the current environmental state of the agent, the historical reasoning path h is encoded. cur =((e) p ,t p ),r1,(e1,t1),…,r l ,(e l ,t l The following formula is used for calculation: For each state, we need to calculate its possible actions and the probability of state transitions. For state s l By using a weighted intelligent agent to select target entity nodes and relationships, and by using a multi-layer perceptron to calculate the current state information, the expected optimal target node e is output. i and the optimal entity relationship r i The representation is as follows: In the formula, W1, W e W r These are learnable variables; Then, according to φ(a) l ,s l ) = 0.5* <e i ,e n >+0.5* <r i ,r n Calculate the similarity between the target entity nodes and relationships of candidate actions as the score φ(a) for the corresponding action. l ,s l The policy network calculates the probability of each action based on the action score, and then selects the corresponding entity relationship and target entity node according to the probability. S105, the design of the value network, the value network is based on the current environmental state s l A score is calculated by using a multi-layer perception mechanism to determine the value of the environment in which the agent is currently located. Among them W v These are learnable parameters; S106, the optimization function is set by using a proximal strategy optimization algorithm to optimize the model parameters. Here, GRU, MLP, Φ, and the embedded learning representations of entities and relations are all optimizable parameter variables. The objective optimization function is expressed as: In the formula, ρ t Importance sampling, i.e., the ratio of the difference between the new strategy and the old strategy: clip is the truncation function; The dominant function is calculated using the following formula: In strategy Under the current state s t The reward is the reward obtained by selecting the corresponding action 'a'.

7. The sensor network intent reasoning method based on temporal knowledge graph according to claim 6, characterized in that, Specifically, S4 is: Dataset partitioning: The constructed time-series knowledge dataset is divided into a training set and a test set. Stratified sampling is used to ensure that the types of sensor nodes and the time-series distribution characteristics in the training set and the test set are consistent with the original dataset. The training set is used for model training, and the test set is used for model evaluation. Model training: Training set data is input into the constructed sensor network intent reasoning model in batches according to time sequence. For each batch of input data, the model generates an intent reasoning prediction value based on the current policy network. The deviation between the prediction value and the actual intent is calculated through the reward function, and the reward signal is fed back to the value network. The value network evaluates the long-term value of the action in the current state in combination with temporal features. The policy network adjusts the behavior probability distribution according to the reward value and the value evaluation result. The optimization function takes the loss of the policy network and the loss of the value network as input and updates the network parameters through backpropagation of the optimizer. Model optimization: Repeat the above training process. Each iteration of the full training set is called an epoch. After each epoch, the model performance is evaluated using the validation set. If the prediction error on the validation set does not decrease for several consecutive epochs, an early stopping mechanism is triggered to avoid model overfitting. Finally, when the average error between the predicted and actual values ​​on the training set is lower than a preset threshold and the performance of the validation set is stable, the training ends.

8. The sensor network intent reasoning method based on temporal knowledge graph according to claim 7, characterized in that, After S5 collects the time-series feature data of sensor nodes in real time, it performs preprocessing and converts the preprocessed real-time data into a time-series knowledge graph fragment with the same format as the training data. This fragment is then input into the trained and optimized sensor network intent reasoning model for intent reasoning prediction. The model's policy network generates an intent reasoning probability distribution based on the current input data and historical reasoning experience, and selects the result with the highest probability as the final predicted intent.

9. An electronic device, characterized in that, include: Memory, used to store programs; A processor for executing a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method according to any one of claims 1-8.

10. A storage medium, characterized in that, The storage medium stores a computer program that, when run on a processor, causes the processor to perform the method described in any one of claims 1-8.