Intelligent inspection path planning method and system based on reinforcement learning
Through a reinforcement learning-based method, combined with multimodal perception and digital twin systems, self-organized group intelligence algorithms and multi-objective conflict coordinators are designed, which solves the problem of insufficient flexibility and robustness of intelligent patrol path planning in complex environments, and realizes efficient and intelligent path planning and resource management.
Patent Information
- Application Number
- CN202510261795.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent patrol path planning methods are difficult to adapt to complex and changeable environments, lack semantic understanding capabilities, and lack effective coordination mechanisms in multi-robot collaborative patrol tasks, resulting in insufficient flexibility and robustness of path planning.
Using reinforcement learning methods, a knowledge-driven scenario understanding model is constructed by obtaining multimodal perceptual information, and a dynamic interaction model between the agent and the environment is established by combining graph neural networks and recursive memory networks to form a cognitively enhanced digital twin system. Design self-organized group intelligence algorithms, multi-objective conflict coordinators and group behavior emergence prediction models, use the federated learning framework to generate highly adaptable group evolution strategies, and enhance the adaptability and robustness of the system through edge computing and knowledge distillation mechanisms.
It improves the inspection efficiency and intelligence level, optimizes resource allocation and risk control, enhances the adaptability and robustness of the system, and realizes closed-loop optimization from task execution to cognitive enhancement.
Smart Images

Figure CN119990496A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a path planning technology, and in particular to an intelligent inspection path planning method and system based on reinforcement learning. Background Art
[0002] Intelligent inspection robot technology plays an increasingly important role in ensuring infrastructure safety and improving production efficiency. Traditional inspection path planning methods are usually based on pre-set rules or simple optimization algorithms, which are difficult to adapt to complex and changing actual environments. With the development of artificial intelligence technology, intelligent inspection path planning methods based on reinforcement learning have gradually become a research hotspot. These methods autonomously optimize inspection paths through interactive learning between intelligent agents and the environment to improve inspection efficiency and quality.
[0003] Limited ability to perceive the environment: Most methods rely only on simple sensor data and have difficulty understanding complex environments at a semantic level, resulting in a lack of flexibility in path planning and a tendency to fall into local optimality.
[0004] Lack of group coordination mechanism: In multi-robot collaborative inspection tasks, existing methods usually lack an effective coordination mechanism, making it difficult to achieve information sharing and task coordination between robots, resulting in low overall inspection efficiency.
[0005] Insufficient adaptability: The actual inspection environment is complex and changeable. Existing methods are difficult to quickly adapt to new environmental changes, such as equipment failures and obstacles, resulting in insufficient robustness of path planning. Summary of the invention
[0006] The embodiments of the present invention provide a method and system for intelligent inspection path planning based on reinforcement learning, which can solve the problems in the prior art.
[0007] According to a first aspect of the embodiments of the present invention, Provides an intelligent inspection path planning method based on reinforcement learning, including: Acquire the multimodal perception information of the intelligent inspection robot, build a knowledge-driven scene understanding model based on the multimodal perception information, and realize semantic-level analysis of the inspection environment; establish a dynamic interaction model between the intelligent agent and the environment based on the graph neural network and the recursive memory network; integrate the scene understanding model and the dynamic interaction model into the cognitive enhanced digital twin system to form a bidirectional mapping architecture from physical space to cognitive space; Based on the cognitive enhanced digital twin system, a self-organizing swarm intelligence algorithm is designed, which integrates meta-heuristic search strategy and social learning mechanism; a multi-objective conflict coordinator is constructed, which balances inspection efficiency, energy consumption distribution and risk aversion through the Pareto optimization method; a swarm behavior emergence prediction model is designed, which is based on chaos theory and complex network analysis to achieve dynamic prediction of the collaborative mode of swarm intelligent agents; a federated learning framework is used to aggregate the experience knowledge of multiple intelligent agents under the premise of protecting privacy, and generate a swarm evolution strategy with adaptability higher than a preset adaptation threshold; The swarm evolution strategy is deployed to the edge computing network to establish a distributed decision-making system. A robustness evaluation mechanism with real-time feedback is constructed, which includes a fault warning module, an emergency response module and a strategy reconstruction module. An attention-based task allocation optimizer is designed, which can adaptively adjust the task allocation plan according to the specialized capabilities and current status of the intelligent agent. A multi-level knowledge distillation mechanism is established to compress the experiential knowledge gained during the swarm evolution process and migrate it to new scenarios. The decision-making strategy is continuously optimized through a meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement.
[0008] Based on the cognitive enhanced digital twin system, a self-organizing swarm intelligence algorithm is designed, which integrates meta-heuristic search strategy and social learning mechanism; a multi-objective conflict coordinator is constructed, which balances inspection efficiency, energy consumption distribution and risk avoidance through the Pareto optimization method, including: The cognitive enhanced digital twin system includes a knowledge reasoning module, which uses a graph neural network to establish a multi-agent interaction model; based on the multi-agent interaction model, a group behavior feature sequence and an environmental dynamic feature sequence are generated; According to the group behavior feature sequence and the environmental dynamic feature sequence, a self-organizing group intelligence algorithm is designed, and the self-organizing group intelligence algorithm integrates a meta-heuristic search strategy and a social learning mechanism; the meta-heuristic search strategy constructs a task decomposition model based on a dynamic memory structure, and the social learning mechanism constructs a knowledge transfer mechanism through a hierarchical attention network; based on the task decomposition model, a complex inspection task is divided into a sub-task sequence, and the group experience knowledge base is extracted using the knowledge transfer mechanism; the sub-task sequence and the group experience knowledge base are input into a reinforcement learning framework to generate an initial behavior decision plan; A multi-objective conflict coordinator is constructed to process the initial behavior decision plan. The multi-objective conflict coordinator uses a Pareto optimization method to dynamically weigh the inspection efficiency objective function, the energy consumption balance objective function and the risk control objective function.
[0009] A group behavior emergence prediction model is designed. The group behavior emergence prediction model is based on chaos theory and complex network analysis to achieve dynamic prediction of the collaborative mode of group intelligent agents. The federated learning framework is used to aggregate the experience knowledge of multiple intelligent agents under the premise of protecting privacy, and the group evolution strategy with adaptability higher than the preset adaptation threshold is generated, including: Constructing a group behavior emergence prediction model, wherein the group behavior emergence prediction model establishes an agent interaction network; extracting time-varying network features based on the agent interaction network, constructing a chaotic dynamics equation based on the time-varying network features, and calculating a chaotic feature index; combining the time-varying network features and the chaotic feature index to form a collaborative mode feature; Inputting the collaborative mode features into a recurrent neural network, the recurrent neural network predicting the evolution trend of the group intelligent body based on the historical collaborative sequence; calculating the group behavior emergence probability and the collaborative mode conversion probability according to the evolution trend; and generating a group behavior prediction result based on the group behavior emergence probability and the collaborative mode conversion probability; A federated learning framework is constructed based on the group behavior prediction results, and the federated learning framework trains an empirical model locally in an intelligent agent; differential privacy encryption is performed on the parameters of the empirical model, and the differential privacy encryption performs parameter perturbations based on local sensitivity and privacy budget; the encrypted model parameters are transmitted to a central server, and the central server assigns aggregation weights according to the knowledge quality and collaborative contribution of the intelligent agent, and uses a federated averaging algorithm to aggregate the encrypted parameters of multiple intelligent agents; The environmental fitness of the aggregated model is calculated, and the environmental fitness is combined with the group behavior prediction result to evaluate the strategy stability and knowledge coverage; when the environmental fitness exceeds a preset adaptation threshold, the group evolution strategy is decrypted and generated.
[0010] A federated learning framework is constructed based on the group behavior prediction result, and the federated learning framework trains an empirical model locally in the intelligent agent; differential privacy encryption is performed on the parameters of the empirical model, and the differential privacy encryption performs parameter perturbation based on local sensitivity and privacy budget; the encrypted model parameters are transmitted to a central server, and the central server allocates aggregation weights according to the knowledge quality and collaborative contribution of the intelligent agent, including: Building a federated learning framework based on the group behavior prediction results, wherein the group behavior prediction results include the behavior trajectory and decision-making mode of the agent; dividing the group behavior prediction results into a training data set, wherein the training data set includes a state vector and an action vector; configuring a local training environment for each agent in the federated learning framework; training an empirical model locally in the agent based on the training data set, wherein the empirical model outputs an action value evaluation result; Extract the network parameters of the empirical model and construct a parameter matrix; calculate the local sensitivity of the parameter matrix, where the local sensitivity characterizes the degree of response of the network parameters to data disturbance; determine the amount of noise injection according to the local sensitivity and a preset privacy budget; perform differential privacy encryption on the parameters of the empirical model, where the differential privacy encryption is achieved by injecting random noise that conforms to the Laplace distribution into the parameter matrix; Calculate the knowledge quality of the agent, the knowledge quality is evaluated based on the model prediction accuracy and knowledge coverage; count the collaborative contribution of the agent, the collaborative contribution is evaluated based on the training participation and sample quality; input the knowledge quality and the collaborative contribution into a weight distributor; the weight distributor outputs the aggregate weight of the agent in federated learning; transmit the encrypted model parameters and the aggregate weight to a central server, the central server receives the encrypted parameters and aggregate weight uploaded by multiple agents.
[0011] The swarm evolution strategy is deployed to the edge computing network to establish a distributed decision-making system; a robustness evaluation mechanism with real-time feedback is constructed, which includes a fault warning module, an emergency response module, and a strategy reconstruction module; a task allocation optimizer based on the attention mechanism is designed, which can adaptively adjust the task allocation scheme according to the specialized capabilities and current status of the intelligent agent, including: Deploy the swarm evolution strategy to edge nodes in an edge computing network; collect the computing load, storage capacity and network bandwidth of the edge nodes to generate a resource status matrix; evaluate the task processing capability score of the edge nodes according to the resource status matrix; convert the swarm evolution strategy into a local execution instruction of the edge node based on the task processing capability score; establish a collaborative decision-making channel between edge nodes according to the local execution instruction, and the collaborative decision-making channel is used to transmit node status information and task scheduling instructions; Continuously monitor the execution status of the edge node, calculate the deviation value between the real-time status and the expected status; calculate the probability of failure based on the deviation value and historical fault data; when the probability of failure exceeds the preset fault threshold, generate corresponding resource allocation instructions according to the degree of deviation; execute the resource allocation instructions to reallocate computing resources and network bandwidth; update the local execution instructions according to the reallocated resources; adjust the data transmission strategy of the collaborative decision-making channel based on the updated local execution instructions; Obtain the skill score and current load status of the agent in the edge node, and construct the agent capability vector; analyze the computational complexity and delay constraint of the task to be assigned, and generate a task requirement vector; input the agent capability vector and the task requirement vector into the attention calculation unit; the attention calculation unit outputs the fitness weight of the agent and the task; and generate a task allocation plan according to the fitness weight and the current resource status matrix; Deploy the task allocation scheme and monitor the task execution process, record the execution quality and resource utilization; input the execution quality and resource utilization into the attention calculation unit, dynamically adjust the fitness weight; adaptively adjust the task allocation scheme according to the adjusted fitness weight.
[0012] Continuously monitoring the execution status of the edge node, calculating the deviation value between the real-time status and the expected status; calculating the probability of failure based on the deviation value and historical failure data; when the probability of failure exceeds a preset failure threshold, generating a corresponding resource allocation instruction according to the degree of deviation; executing the resource allocation instruction to reallocate computing resources and network bandwidth includes: Continuously monitor the execution status of the edge node, collect performance indicators of the edge node, the performance indicators include memory usage, network throughput and task response time; obtain expected operating parameters of the edge node, the expected operating parameters are standard operating indicators set by the system; build a performance evaluation model, the performance evaluation model calculates the deviation value between the real-time state and the expected state based on the performance indicators and the expected operating parameters; The deviation value is input into a time series feature extractor, and the time series feature extractor generates a state deviation sequence; a fault record is read from a historical fault database, and the fault record contains state deviation data before the historical fault occurs; a fault prediction model is trained based on the state deviation sequence and the state deviation data, and the fault prediction model outputs a probability of fault occurrence; and the fault probability is compared with a preset fault threshold; When the probability of the failure occurring exceeds the preset failure threshold, the deviation value is input into the resource evaluator, and the resource evaluator calculates the resource demand according to the degree of deviation; generates a resource allocation instruction based on the resource demand, and the resource allocation instruction includes computing resources and network bandwidth; determines the resource allocation priority according to the changing trend of the performance indicator; executes the resource allocation instruction, reallocates the computing resources according to the resource allocation priority, and adjusts the network bandwidth.
[0013] Establish a multi-level knowledge distillation mechanism to compress and transfer the experiential knowledge gained during the group evolution process to new scenarios; continuously optimize decision-making strategies through the meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement, including: Establish a multi-level knowledge distillation mechanism, based on which the behavior sequence, state information and task completion status of the intelligent agents in the group evolution process are collected; the behavior sequence and the state information are input into the first layer of distillation network to extract temporal knowledge features, the temporal knowledge features and the task completion status are input into the second layer of distillation network to refine the evolution experience, and the evolution experience is compressed through the third layer of distillation network to obtain a knowledge compression representation; Analyze the feature distribution of the target scene, compare the knowledge compression representation with the feature distribution of the target scene, and calculate the knowledge migration loss; adaptively adjust the migration parameters based on the knowledge migration loss to achieve the progressive migration of the knowledge compression representation to the new scene, and output the scene-adaptive decision knowledge; Construct a meta-reinforcement learning framework, and input the scenario-adapted decision knowledge into the meta-reinforcement learning framework as prior information; the meta-reinforcement learning framework generates an action distribution based on the current state, and obtains immediate feedback by interacting with the environment; optimizes meta-strategy parameters based on the immediate feedback, continuously optimizes decision effects, and collects task completion data during execution; The task completion data is analyzed and evaluated to extract cognitive level improvement indicators; the cognitive level improvement indicators are fed back to each level of the multi-level knowledge distillation mechanism, and the distillation parameters and compression ratio are dynamically adjusted; the extraction and compression effects of group evolution knowledge are optimized through parameter adjustment, thereby improving the quality of knowledge transfer and forming a complete optimization closed loop from task execution to cognitive enhancement.
[0014] According to a second aspect of the embodiments of the present invention, Provides an intelligent inspection path planning system based on reinforcement learning, including: The first unit is used to obtain the multimodal perception information of the intelligent inspection robot, build a knowledge-driven scene understanding model based on the multimodal perception information, and realize the semantic level analysis of the inspection environment; based on the graph neural network and the recursive memory network, establish a dynamic interaction model between the intelligent agent and the environment; integrate the scene understanding model and the dynamic interaction model into the cognitive enhanced digital twin system to form a bidirectional mapping architecture from physical space to cognitive space; The second unit is used to design a self-organizing swarm intelligence algorithm based on the cognitive enhanced digital twin system, the self-organizing swarm intelligence algorithm integrates meta-heuristic search strategy and social learning mechanism; construct a multi-objective conflict coordinator, the multi-objective conflict coordinator balances inspection efficiency, energy consumption distribution and risk aversion through the Pareto optimization method; design a swarm behavior emergence prediction model, the swarm behavior emergence prediction model is based on chaos theory and complex network analysis to achieve dynamic prediction of the collaborative mode of swarm intelligent agents; use a federated learning framework to aggregate the experience knowledge of multiple intelligent agents under the premise of protecting privacy, and generate a swarm evolution strategy with adaptability higher than a preset adaptation threshold; The third unit is used to deploy the group evolution strategy to the edge computing network and establish a distributed decision-making system; construct a robustness evaluation mechanism with real-time feedback, which includes a fault warning module, an emergency response module and a strategy reconstruction module; design a task allocation optimizer based on the attention mechanism, which can adaptively adjust the task allocation plan according to the specialized capabilities and current status of the intelligent agent; establish a multi-level knowledge distillation mechanism to compress and migrate the experiential knowledge gained during the group evolution process to new scenarios; and continuously optimize the decision-making strategy through the meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement.
[0015] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0016] A fourth aspect of the embodiments of the present invention is: A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0017] The beneficial effects of this application are as follows: 1. Improve inspection efficiency and intelligence level: This method combines technologies such as multimodal perception, knowledge-driven scene understanding, and graph neural networks to achieve semantic-level analysis of the inspection environment and dynamic interaction between the intelligent agent and the environment, and builds a cognitive-enhanced digital twin system, which can more accurately understand the inspection tasks and environment, and perform more intelligent path planning, ultimately improving inspection efficiency and intelligence level.
[0018] 2. Optimize resource allocation and risk control: Based on the self-organizing swarm intelligence algorithm, multi-objective conflict coordinator and swarm behavior emergence prediction model, this method can optimize the energy consumption distribution of the intelligent inspection robot, effectively avoid risks, and balance inspection efficiency, energy consumption and risks under the Pareto optimal framework, so as to achieve optimal resource allocation and effective risk control.
[0019] 3. Enhance system adaptability and robustness: By utilizing federated learning, edge computing, real-time feedback robustness evaluation mechanism, task allocation optimizer and knowledge distillation mechanism, this method can realize the sharing and transfer of swarm intelligence experience knowledge, enhance the adaptability and robustness of the system, and continuously optimize the decision-making strategy through the meta-reinforcement learning framework, ultimately achieving closed-loop optimization from task execution to cognitive enhancement, and improving the overall performance and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 A schematic diagram of a process flow of an intelligent inspection path planning method based on reinforcement learning according to an embodiment of the present invention; Figure 2 It is a structural diagram of an intelligent inspection path planning system based on reinforcement learning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0022] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.
[0023] Figure 1 FIG. 1 is a flow chart of an intelligent inspection path planning method based on reinforcement learning according to an embodiment of the present invention. Figure 1 As shown, the method includes: S11. Obtain multimodal perception information of the intelligent inspection robot, build a knowledge-driven scene understanding model based on the multimodal perception information, and realize semantic-level analysis of the inspection environment; establish a dynamic interaction model between the intelligent agent and the environment based on the graph neural network and the recursive memory network; integrate the scene understanding model and the dynamic interaction model into the cognitive enhanced digital twin system to form a bidirectional mapping architecture from physical space to cognitive space; S12. Based on the cognitive enhanced digital twin system, a self-organizing swarm intelligence algorithm is designed, which integrates meta-heuristic search strategy and social learning mechanism; a multi-objective conflict coordinator is constructed, which balances inspection efficiency, energy consumption distribution and risk aversion through the Pareto optimization method; a swarm behavior emergence prediction model is designed, which is based on chaos theory and complex network analysis to achieve dynamic prediction of the collaborative mode of swarm intelligent agents; a federated learning framework is used to aggregate the experience knowledge of multiple intelligent agents under the premise of protecting privacy, and generate a swarm evolution strategy with adaptability higher than a preset adaptation threshold; S13. Deploy the swarm evolution strategy to the edge computing network and establish a distributed decision-making system; construct a robustness evaluation mechanism with real-time feedback, which includes a fault warning module, an emergency response module and a strategy reconstruction module; design a task allocation optimizer based on an attention mechanism, which can adaptively adjust the task allocation plan according to the specialized capabilities and current status of the intelligent agent; establish a multi-level knowledge distillation mechanism to compress and migrate the experiential knowledge gained during the swarm evolution process to new scenarios; continuously optimize the decision-making strategy through a meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement.
[0024] In an optional implementation, based on the cognitive enhanced digital twin system, a self-organizing swarm intelligence algorithm is designed, the self-organizing swarm intelligence algorithm integrates a meta-heuristic search strategy and a social learning mechanism; a multi-objective conflict coordinator is constructed, and the multi-objective conflict coordinator balances inspection efficiency, energy consumption distribution and risk avoidance through a Pareto optimization method, including: The cognitive enhanced digital twin system includes a knowledge reasoning module, which uses a graph neural network to establish a multi-agent interaction model; based on the multi-agent interaction model, a group behavior feature sequence and an environmental dynamic feature sequence are generated; According to the group behavior feature sequence and the environmental dynamic feature sequence, a self-organizing group intelligence algorithm is designed, and the self-organizing group intelligence algorithm integrates a meta-heuristic search strategy and a social learning mechanism; the meta-heuristic search strategy constructs a task decomposition model based on a dynamic memory structure, and the social learning mechanism constructs a knowledge transfer mechanism through a hierarchical attention network; based on the task decomposition model, a complex inspection task is divided into a sub-task sequence, and the group experience knowledge base is extracted using the knowledge transfer mechanism; the sub-task sequence and the group experience knowledge base are input into a reinforcement learning framework to generate an initial behavior decision plan; A multi-objective conflict coordinator is constructed to process the initial behavior decision plan. The multi-objective conflict coordinator uses a Pareto optimization method to dynamically weigh the inspection efficiency objective function, the energy consumption balance objective function and the risk control objective function.
[0025] First, a cognitive enhanced digital twin system is constructed. The core component of the system is the knowledge reasoning module, which uses graph neural networks to build a multi-agent interaction model. Specifically, each agent and its environment state are represented as nodes in the graph, and the interactions between agents, such as communication and collaboration, are represented as edges in the graph. Graph neural networks capture the complex relationships between agents and their interaction patterns with the environment by learning the features of nodes and edges. For example, in a patrol robot team, the position, power, sensor data, etc. of each robot are used as node features, and the communication information between robots is used as edge features. By training the graph neural network, the behavior patterns of the robot team and the impact of the environment can be learned. The knowledge reasoning module generates group behavior feature sequences and environmental dynamic feature sequences based on the multi-agent interaction model. The group behavior feature sequence describes the overall behavior of the robot team, such as movement trajectory, task allocation, etc. The environmental dynamic feature sequence describes changes in the environment, such as the appearance of obstacles, temperature changes, etc. Assuming there are three robots, the group behavior feature sequence can be "[robot 1 moves to area A, robot 2 moves to area B, robot 3 is on standby], [robot 1 and robot 2 collaboratively explore area C, robot 3 charges]". The environmental dynamic feature sequence can be “[an obstacle appears in area A], [the temperature in area C rises]”.
[0026] Next, a self-organizing swarm intelligence algorithm is designed. The algorithm combines a meta-heuristic search strategy and a social learning mechanism. The meta-heuristic search strategy builds a task decomposition model based on a dynamic memory structure. The dynamic memory structure stores the decomposition schemes of historical inspection tasks and the corresponding performance indicators. For new inspection tasks, the algorithm retrieves similar historical tasks from the dynamic memory structure according to the characteristics of the task and draws on their decomposition schemes. For example, the inspection task of a large factory can be decomposed into inspection tasks for multiple sub-areas. The social learning mechanism builds a knowledge transfer mechanism through a hierarchical attention network. The hierarchical attention network can capture the experience differences between different agents and selectively learn the successful experience of other agents. For example, an experienced robot can share its strategy for avoiding obstacles with a newly joined robot. After dividing the complex inspection task into a sub-task sequence, the knowledge transfer mechanism is used to extract the group experience knowledge base. The group experience knowledge base contains the successful experiences and failed lessons accumulated in historical inspection tasks. The sub-task sequence and the group experience knowledge base are input into the reinforcement learning framework to generate an initial behavior decision plan. The reinforcement learning framework continuously optimizes the behavior strategy of the agent through trial and error learning to maximize inspection efficiency, minimize energy consumption, and avoid risks. For example, robots can learn to choose the best path and inspection strategy in different environments through reinforcement learning.
[0027] Finally, a multi-objective conflict coordinator is constructed to handle the initial behavior decision plan. The multi-objective conflict coordinator uses the Pareto optimization method to dynamically weigh the inspection efficiency objective function, energy consumption balance objective function, and risk control objective function. The Pareto optimization method aims to find a set of optimal solutions. Without reducing the values of other objective functions, it is impossible to further improve the value of any objective function. For example, improving inspection efficiency may increase energy consumption, and reducing risks may reduce efficiency. The multi-objective conflict coordinator needs to find a balance point so that all three objective functions reach a good value. Assuming that the initial behavior decision plan is "[Robot 1 inspects area A, robot 2 inspects area B]", the multi-objective conflict coordinator may be adjusted to "[Robot 1 inspects area B first, then inspects area A, robot 2 charges and stands by]" to achieve better energy consumption balance.
[0028] The solution of this application can: Improve inspection efficiency: Through task decomposition and knowledge transfer, the intelligent agent can quickly learn the best inspection strategy, reduce redundant inspections, and improve overall inspection efficiency. Optimize energy consumption distribution: Through the multi-objective conflict coordinator, it is possible to balance inspection efficiency and energy consumption, avoid excessive energy consumption of individual intelligent agents, and extend the endurance of the entire team. Enhance risk avoidance ability: Through reinforcement learning and group experience knowledge base, the intelligent agent can learn how to avoid risks, such as avoiding obstacles and collisions, to improve inspection safety.
[0029] In an optional implementation, a group behavior emergence prediction model is designed, which is based on chaos theory and complex network analysis to achieve dynamic prediction of the collaborative mode of group intelligent agents; using a federated learning framework to aggregate the experience knowledge of multiple intelligent agents under the premise of protecting privacy, and generating a group evolution strategy with adaptability higher than a preset adaptation threshold includes: Constructing a group behavior emergence prediction model, wherein the group behavior emergence prediction model establishes an agent interaction network; extracting time-varying network features based on the agent interaction network, constructing a chaotic dynamics equation based on the time-varying network features, and calculating a chaotic feature index; combining the time-varying network features and the chaotic feature index to form a collaborative mode feature; Inputting the collaborative mode features into a recurrent neural network, the recurrent neural network predicting the evolution trend of the group intelligent body based on the historical collaborative sequence; calculating the group behavior emergence probability and the collaborative mode conversion probability according to the evolution trend; and generating a group behavior prediction result based on the group behavior emergence probability and the collaborative mode conversion probability; A federated learning framework is constructed based on the group behavior prediction results, and the federated learning framework trains an empirical model locally in an intelligent agent; differential privacy encryption is performed on the parameters of the empirical model, and the differential privacy encryption performs parameter perturbations based on local sensitivity and privacy budget; the encrypted model parameters are transmitted to a central server, and the central server assigns aggregation weights according to the knowledge quality and collaborative contribution of the intelligent agent, and uses a federated averaging algorithm to aggregate the encrypted parameters of multiple intelligent agents; The environmental fitness of the aggregated model is calculated, and the environmental fitness is combined with the group behavior prediction result to evaluate the strategy stability and knowledge coverage; when the environmental fitness exceeds a preset adaptation threshold, the group evolution strategy is decrypted and generated.
[0030] First, we build a group behavior emergence prediction model. We build an agent interaction network. For example, we can use a graph structure to represent the interaction relationship between agents. Nodes represent agents, edges represent the interaction behaviors between agents, and the weight of edges can represent the interaction intensity or frequency.
[0031] Next, we extract time-varying network features from the agent interaction network. For example, we can calculate the average degree, clustering coefficient, network diameter and other indicators of the network. These indicators change over time and reflect the dynamic evolution of the network structure. Assume that at time t1, the average degree of the network is 5, the clustering coefficient is 0.8, and the network diameter is 4; at time t2, these indicators become 6, 0.7, and 3 respectively.
[0032] Then, a chaotic dynamics equation is constructed based on the extracted time-varying network characteristics. For example, the Lorenz equation or other chaotic models can be used to describe the dynamic changes of the network, and the parameters of the equation can be adjusted according to the time-varying network characteristics. Assume that the initial values of the parameters x, y, z of the Lorenz equation are 1, 1, 1. After the evolution of time t1, the values of x, y, z are 2, 3, 4; after the evolution of time t2, the values of x, y, z are 5, 6, 7.
[0033] Calculate chaotic characteristic indicators. For example, you can calculate indicators such as the Lyapunov index and correlation dimension, which reflect the chaotic characteristics of the system. Assume that the calculated Lyapunov index is 0.5 and the correlation dimension is 2.5.
[0034] The time-varying network characteristics and chaotic characteristic indicators are combined to form the collaborative pattern characteristics. For example, the network average degree, clustering coefficient, network diameter, Lyapunov index, correlation dimension and other indicators can be combined into a feature vector. Assume that at time t1, the collaborative pattern feature vector is [5, 0.8, 4, 0.5, 2.5]; at time t2, the collaborative pattern feature vector is [6, 0.7, 3, 0.6, 2.6].
[0035] The collaborative pattern features are input into a recurrent neural network, such as a long short-term memory network (LSTM), to predict the evolution trend of the swarm agent based on the historical collaborative sequence. Assuming the collaborative pattern feature vectors at time t1 and t2 are input, the LSTM network predicts that the average degree of the network at time t3 is 7, the clustering coefficient is 0.6, and the network diameter is 2.
[0036] The probability of group behavior emergence and the probability of collaborative mode conversion are calculated based on the evolution trend. For example, the possibility of group behavior emergence and the probability of conversion from one collaborative mode to another can be judged based on the predicted changes in network characteristics. Assume that the probability of group behavior emergence at time t3 is predicted to be 0.8 and the probability of collaborative mode conversion is 0.6.
[0037] The group behavior prediction results are generated based on the group behavior emergence probability and the coordination mode conversion probability. For example, the predicted group behavior type, coordination mode and corresponding probability can be output.
[0038] Next, a federated learning framework is constructed based on the group behavior prediction results. The experience model is trained locally in the agent. For example, each agent can use a reinforcement learning algorithm to train a policy model based on its own interaction experience with the environment.
[0039] The parameters of the empirical model are differentially privately encrypted. For example, Laplace noise or Gaussian noise can be added to the model parameters according to the local sensitivity and privacy budget. Assuming the privacy budget is 0.1 and the local sensitivity is 1, the added noise follows a Laplace distribution with a mean of 0 and a standard deviation of 1 / 0.1.
[0040] The encrypted model parameters are transmitted to the central server. The central server assigns aggregate weights based on the knowledge quality and collaborative contribution of the agents. For example, weights can be assigned based on the historical performance or contribution of the agents. Suppose the weight of agent A is 0.3 and the weight of agent B is 0.7.
[0041] The central server aggregates the encryption parameters of multiple agents using a federated averaging algorithm, for example, taking a weighted average of the encryption parameters of all agents.
[0042] Calculate the environmental fitness of the aggregated model. For example, the strategy stability and knowledge coverage can be evaluated in combination with the group behavior prediction results.
[0043] When the environmental fitness exceeds the preset adaptation threshold, the population evolution strategy is decrypted and generated.
[0044] The solution of this application can: Improve prediction accuracy: Combining chaos theory and complex network analysis can more accurately predict the emergence pattern of group behavior and capture the complexity and nonlinear characteristics of group behavior. Protect privacy: Adopting the federated learning framework and differential privacy technology, it can aggregate the experience knowledge of multiple agents without leaking the information of individual agents and protect the privacy of agents. Enhance adaptability: Through continuous learning and evolution, the generated group evolution strategy can better adapt to the dynamically changing environment and improve the stability and knowledge coverage of the strategy.
[0045] In an optional implementation, a federated learning framework is constructed based on the group behavior prediction result, wherein the federated learning framework trains an empirical model locally in an intelligent agent; differential privacy encryption is performed on the parameters of the empirical model, wherein the differential privacy encryption performs parameter perturbations based on local sensitivity and privacy budget; and the encrypted model parameters are transmitted to a central server, wherein the central server allocates aggregation weights according to the knowledge quality and collaborative contribution of the intelligent agent, including: Building a federated learning framework based on the group behavior prediction results, wherein the group behavior prediction results include the behavior trajectory and decision-making mode of the agent; dividing the group behavior prediction results into a training data set, wherein the training data set includes a state vector and an action vector; configuring a local training environment for each agent in the federated learning framework; training an empirical model locally in the agent based on the training data set, wherein the empirical model outputs an action value evaluation result; Extract the network parameters of the empirical model and construct a parameter matrix; calculate the local sensitivity of the parameter matrix, where the local sensitivity characterizes the degree of response of the network parameters to data disturbance; determine the amount of noise injection according to the local sensitivity and a preset privacy budget; perform differential privacy encryption on the parameters of the empirical model, where the differential privacy encryption is achieved by injecting random noise that conforms to the Laplace distribution into the parameter matrix; Calculate the knowledge quality of the agent, the knowledge quality is evaluated based on the model prediction accuracy and knowledge coverage; count the collaborative contribution of the agent, the collaborative contribution is evaluated based on the training participation and sample quality; input the knowledge quality and the collaborative contribution into a weight distributor; the weight distributor outputs the aggregate weight of the agent in federated learning; transmit the encrypted model parameters and the aggregate weight to a central server, the central server receives the encrypted parameters and aggregate weight uploaded by multiple agents.
[0046] This method uses the group behavior prediction results to build a federated learning framework and trains the empirical model locally in the intelligent agent. In order to protect privacy, the model parameters are differentially encrypted, and then the encrypted parameters are transmitted to the central server for aggregation, and weights are assigned according to the contribution of the intelligent agent.
[0047] First, collect a large amount of group behavior data, which contains the behavior trajectory and decision-making patterns of the intelligent agents. For example, you can collect users' browsing history, purchase records, evaluation information, etc. on e-commerce platforms. These data will be used to train the prediction model.
[0048] Next, the collected group behavior data is divided into a training data set. The training data set contains state vectors and action vectors. The state vector describes the state of the environment in which the agent is located, such as the user's age, gender, region, interests and hobbies, etc. The action vector represents the action taken by the agent in a specific state, such as clicking on a product, adding it to the shopping cart, and finally purchasing it. For example, a state vector can be [25 years old, male, Beijing, technology enthusiast], and the corresponding action vector can be [click on the mobile phone product and add it to the shopping cart].
[0049] Then, a local training environment is configured for each agent in the federated learning framework. Each agent will obtain a portion of the training data and train the empirical model locally. The input of the empirical model is the state vector, and the output is the action value evaluation result. For example, the model can predict the probability of a user purchasing a certain product.
[0050] After training is completed, the network parameters of the empirical model are extracted and the parameter matrix is constructed. For example, the parameters of a simple linear model can be represented as a vector [w1, w2, w3, …], where each element represents the weight of a feature.
[0051] In order to protect privacy, it is necessary to calculate the local sensitivity of the parameter matrix. Local sensitivity characterizes the degree of response of network parameters to data perturbations. Simply put, it is to evaluate the impact of changing a data sample on the model parameters. For example, if adding or deleting a user's data causes a large change in the model parameters, it means that the local sensitivity of the model is high.
[0052] The amount of noise injection is determined based on the local sensitivity and the preset privacy budget. The privacy budget is a parameter that controls the degree of privacy leakage. The smaller the privacy budget, the stricter the privacy protection. The amount of noise injection is proportional to the local sensitivity and the privacy budget.
[0053] Perform differential privacy encryption on the parameters of the empirical model. Differential privacy encryption is achieved by injecting random noise that conforms to the Laplace distribution into the parameter matrix. The amplitude of the noise is determined by the amount of noise injection. For example, a random noise can be added to each element of the parameter vector [w1, w2, w3, …].
[0054] In order to fairly aggregate model parameters, it is necessary to calculate the knowledge quality and collaborative contribution of the agents. Knowledge quality is evaluated based on model prediction accuracy and knowledge coverage. For example, the accuracy of the model on the test set and the diversity of data types that the model can handle can be calculated. Collaborative contribution is evaluated based on training participation and sample quality. For example, the number of times the agent participates in training and the quality of the training data provided can be counted.
[0055] The knowledge quality and collaborative contribution are input into the weight allocator. The weight allocator outputs the aggregate weight of the agents in federated learning. For example, if an agent provides a high model accuracy and good data quality, it will get a higher weight. Suppose there are two agents, the first agent has a knowledge quality score of 0.8 and a collaborative contribution score of 0.9; the second agent has a knowledge quality score of 0.7 and a collaborative contribution score of 0.8. The weight allocator can calculate their weights based on these two scores, for example, the first agent has a weight of 0.6 and the second agent has a weight of 0.4.
[0056] Finally, the encrypted model parameters and aggregate weights are transmitted to the central server. The central server receives the encrypted parameters and aggregate weights uploaded by multiple agents, and performs weighted averaging based on the weights to obtain the final global model.
[0057] The solution of this application can: Enhanced privacy protection: Through differential privacy encryption technology, the privacy of local data of intelligent agents is effectively protected to prevent the leakage of sensitive information. Improved model performance: Using the federated learning framework, the local data of multiple intelligent agents can be integrated to train more accurate and robust prediction models. Promoted collaboration: Through the weight distribution mechanism, intelligent agents are encouraged to actively participate in training, improve data quality and model performance, and promote collaboration.
[0058] In an optional implementation, the swarm evolution strategy is deployed to an edge computing network to establish a distributed decision-making system; a robustness evaluation mechanism with real-time feedback is constructed, and the robustness evaluation mechanism includes a fault warning module, an emergency response module, and a strategy reconstruction module; a task allocation optimizer based on an attention mechanism is designed, and the task allocation optimizer can adaptively adjust the task allocation scheme according to the specialized capabilities and current status of the intelligent agent, including: Deploy the swarm evolution strategy to edge nodes in an edge computing network; collect the computing load, storage capacity and network bandwidth of the edge nodes to generate a resource status matrix; evaluate the task processing capability score of the edge nodes according to the resource status matrix; convert the swarm evolution strategy into a local execution instruction of the edge node based on the task processing capability score; establish a collaborative decision-making channel between edge nodes according to the local execution instruction, and the collaborative decision-making channel is used to transmit node status information and task scheduling instructions; Continuously monitor the execution status of the edge node, calculate the deviation value between the real-time status and the expected status; calculate the probability of failure based on the deviation value and historical fault data; when the probability of failure exceeds the preset fault threshold, generate corresponding resource allocation instructions according to the degree of deviation; execute the resource allocation instructions to reallocate computing resources and network bandwidth; update the local execution instructions according to the reallocated resources; adjust the data transmission strategy of the collaborative decision-making channel based on the updated local execution instructions; Obtain the skill score and current load status of the agent in the edge node, and construct the agent capability vector; analyze the computational complexity and delay constraint of the task to be assigned, and generate a task requirement vector; input the agent capability vector and the task requirement vector into the attention calculation unit; the attention calculation unit outputs the fitness weight of the agent and the task; and generate a task allocation plan according to the fitness weight and the current resource status matrix; Deploy the task allocation scheme and monitor the task execution process, record the execution quality and resource utilization; input the execution quality and resource utilization into the attention calculation unit, dynamically adjust the fitness weight; adaptively adjust the task allocation scheme according to the adjusted fitness weight.
[0059] Edge node preparation phase: First, select appropriate edge nodes in the edge computing network as the execution units of the swarm evolution strategy. In order to ensure the efficiency and stability of strategy execution, the selected edge nodes need to be initialized and configured, including installing the necessary software environment, configuring network parameters, and allocating storage space. For example, select 10 edge nodes, numbered Node1 to Node10, and each node is configured with a 2-core CPU, 4GB memory, and 100GB storage space.
[0060] Resource status matrix construction phase: In order to accurately evaluate the task processing capabilities of each edge node, it is necessary to collect information such as the node's computing load, storage capacity, and network bandwidth in real time. This information will be organized into a resource status matrix for subsequent decision-making. For example, at a certain moment, Node1's CPU load is 50%, the remaining storage space is 80GB, and the network bandwidth utilization is 30%. This information is recorded in the corresponding position of the resource status matrix.
[0061] Swarm evolution strategy conversion and deployment phase: Swarm evolution strategies usually exist in the form of algorithms, which need to be converted into local instructions that edge nodes can understand and execute. The conversion process needs to consider the hardware architecture and software environment of the node. For example, the swarm evolution strategy is converted into a Python script and distributed to each edge node.
[0062] Collaborative decision-making channel establishment phase: In order to achieve collaboration between edge nodes, a collaborative decision-making channel needs to be established to transmit node status information and task scheduling instructions. This channel can be implemented based on technologies such as message queues or distributed databases. For example, using Kafka message queues as collaborative decision-making channels, nodes exchange information by publishing and subscribing to messages.
[0063] Fault warning and emergency response stage: In order to improve the robustness of the system, it is necessary to establish a fault warning module to monitor the execution status of edge nodes in real time and calculate the probability of fault occurrence based on state deviation and historical fault data. When the fault probability exceeds the preset threshold, the emergency response module is triggered to execute the corresponding resource allocation instructions. For example, when the CPU load of Node3 continues to exceed 90% and the fault probability exceeds 80%, some computing tasks are migrated from Node3 to nodes with lower loads.
[0064] Strategy reconstruction stage: When a serious failure or environmental change occurs, the group evolution strategy needs to be reconstructed to adapt to the new situation. Strategy reconstruction can be based on expert experience or machine learning algorithms. For example, based on historical failure data and current environmental information, a reinforcement learning algorithm is used to optimize the parameters of the group evolution strategy.
[0065] Agent capability vector construction phase: In order to implement task allocation based on the attention mechanism, it is necessary to obtain the skill score and current load status of the agent in the edge node and construct the agent capability vector. For example, there are three agents on Node2, which are good at image recognition, natural language processing and data analysis respectively. Their skill scores are 0.9, 0.8 and 0.7 respectively, and their current load status are 40%, 60% and 20% respectively. This information is organized into a vector to represent the agent capability of Node2.
[0066] Task requirement vector generation phase: Analyze the computational complexity and delay constraints of the task to be assigned and generate a task requirement vector. For example, the computational complexity of an image recognition task is 0.8 and the delay constraint is 100ms. Organize this information into a vector to represent the requirements of the task.
[0067] Attention calculation and task allocation stage: The agent capability vector and task requirement vector are input into the attention calculation unit to calculate the fitness weight between the agent and the task. The task allocation scheme is generated based on the fitness weight and the current resource status matrix. For example, according to the calculation results, the fitness weight between Node2's image recognition agent and the image recognition task is 0.95, so the task is assigned to Node2.
[0068] Task allocation scheme adaptive adjustment stage: deploy the task allocation scheme and monitor the task execution process, record the execution quality and resource utilization. Feed this information back to the attention calculation unit, dynamically adjust the fitness weight, and adaptively adjust the task allocation scheme according to the adjusted weight. For example, if the image recognition agent of Node2 performs the task with high quality and reasonable resource utilization, then increase its fitness weight to make it easier to obtain similar tasks.
[0069] The solution of this application can: Improve the resource utilization efficiency of edge computing networks: By real-time monitoring of resource status and adaptively adjusting task allocation schemes, the computing resources of edge nodes can be maximized, resource waste can be avoided, and overall resource utilization efficiency can be improved. Enhance the robustness and fault tolerance of the system: Through fault warning, emergency response and strategy reconstruction mechanisms, it is possible to effectively respond to edge node failures and environmental changes, ensure the stable operation of the system, and enhance the robustness and fault tolerance of the system. Optimize task allocation and execution efficiency: The task allocation optimizer based on the attention mechanism can allocate tasks to the most suitable agent according to the agent's specialized capabilities and current status, thereby improving the efficiency and quality of task execution.
[0070] In an optional implementation, the execution state of the edge node is continuously monitored, and a deviation value between the real-time state and the expected state is calculated; the probability of a fault occurrence is calculated based on the deviation value and historical fault data; when the probability of a fault occurrence exceeds a preset fault threshold, a corresponding resource allocation instruction is generated according to the degree of deviation; and executing the resource allocation instruction to reallocate computing resources and network bandwidth includes: Continuously monitor the execution status of the edge node, collect performance indicators of the edge node, the performance indicators include memory usage, network throughput and task response time; obtain expected operating parameters of the edge node, the expected operating parameters are standard operating indicators set by the system; build a performance evaluation model, the performance evaluation model calculates the deviation value between the real-time state and the expected state based on the performance indicators and the expected operating parameters; The deviation value is input into a time series feature extractor, and the time series feature extractor generates a state deviation sequence; a fault record is read from a historical fault database, and the fault record contains state deviation data before the historical fault occurs; a fault prediction model is trained based on the state deviation sequence and the state deviation data, and the fault prediction model outputs a probability of fault occurrence; and the fault probability is compared with a preset fault threshold; When the probability of the failure occurring exceeds the preset failure threshold, the deviation value is input into the resource evaluator, and the resource evaluator calculates the resource demand according to the degree of deviation; generates a resource allocation instruction based on the resource demand, and the resource allocation instruction includes computing resources and network bandwidth; determines the resource allocation priority according to the changing trend of the performance indicator; executes the resource allocation instruction, reallocates the computing resources according to the resource allocation priority, and adjusts the network bandwidth.
[0071] First, continuously monitor the execution status of edge nodes. Collect multiple performance indicators of edge nodes, including memory usage, network throughput, and task response time. For example, collect performance indicators of edge node A every 5 seconds, and measure that the memory usage is 70%, the network throughput is 10Mbps, and the average task response time is 200ms. At the same time, obtain the standard operating indicators preset by the system, for example, the expected value of memory usage is 60%, the expected value of network throughput is 12Mbps, and the expected value of average task response time is 150ms.
[0072] Then, a performance evaluation model is constructed to calculate the deviation between the real-time status and the expected status. The collected performance indicators are compared with the preset standard operation indicators. For example, the memory usage deviation of edge node A is +10%, the network throughput deviation is -2Mbps, and the task average response time deviation is +50ms.
[0073] Next, the deviation value is input into the time series feature extractor. The time series feature extractor analyzes the deviation value sequence over a period of time. For example, the memory usage deviation value sequence of edge node A in the past minute [ +5%, +8%, +10%, +12%, +10%, +10%, +10%, +9%, +10%, +11%, +10%, +10% ] is used as input.
[0074] At the same time, fault records are read from the historical fault database. These records contain state deviation data before the historical fault occurred. For example, the database records the memory usage deviation value sequence of edge node B within 1 minute before a certain fault occurred in the past [ +2%, +5%, +10%, +15%, +20%, +25%, +30%, +35%, +40%, +42%, +45%, +48% ].
[0075] Based on the state deviation sequence and historical fault data, the fault prediction model is trained. The model learns the state deviation pattern before the historical fault occurred and predicts the probability of the fault occurring in the current state. For example, based on the current state deviation sequence and historical fault data of edge node A, the prediction model outputs a fault probability of 15%.
[0076] The probability of failure occurrence is compared with the preset failure threshold. The preset failure threshold is set to 20%, for example. In this example, since 15% is less than 20%, the system does not trigger resource allocation.
[0077] If the probability of a failure exceeds a preset failure threshold, such as 25%, the deviation value is input into the resource evaluator. The resource evaluator calculates the resource requirements based on the degree of deviation. For example, based on the deviation value of edge node A, the resource evaluator calculates that 2GB of memory and 1Mbps of network bandwidth are required.
[0078] Generate resource allocation instructions based on resource demand. The instructions include computing resources and network bandwidth. For example, the generated instructions are "add 2GB of memory and 1Mbps of network bandwidth to edge node A".
[0079] Determine resource allocation priorities based on the changing trends of performance indicators. For example, if memory usage continues to rise, memory allocation takes precedence over network bandwidth adjustment.
[0080] Finally, execute resource allocation instructions to reallocate computing resources according to resource allocation priorities and adjust network bandwidth.
[0081] The solution of this application can: Improve system stability: By predicting potential failures and adjusting resources in advance, you can effectively avoid system crashes or service interruptions caused by insufficient resources, thereby improving system stability and reliability. Optimize resource utilization: Dynamically allocate resources according to actual needs to avoid idle or wasted resources, improve resource utilization efficiency, and reduce operating costs. Ensure service quality: Through timely resource allocation, you can ensure that edge applications always have enough resources to meet service needs, ensure service quality, and improve user experience.
[0082] In an optional implementation, a multi-level knowledge distillation mechanism is established to compress and transfer the experiential knowledge gained during the group evolution process to new scenarios; the decision-making strategy is continuously optimized through the meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement, including: Establish a multi-level knowledge distillation mechanism, based on which the behavior sequence, state information and task completion status of the intelligent agents in the group evolution process are collected; the behavior sequence and the state information are input into the first layer of distillation network to extract temporal knowledge features, the temporal knowledge features and the task completion status are input into the second layer of distillation network to refine the evolution experience, and the evolution experience is compressed through the third layer of distillation network to obtain a knowledge compression representation; Analyze the feature distribution of the target scene, compare the knowledge compression representation with the feature distribution of the target scene, and calculate the knowledge migration loss; adaptively adjust the migration parameters based on the knowledge migration loss to achieve the progressive migration of the knowledge compression representation to the new scene, and output the scene-adaptive decision knowledge; Construct a meta-reinforcement learning framework, and input the scenario-adapted decision knowledge into the meta-reinforcement learning framework as prior information; the meta-reinforcement learning framework generates an action distribution based on the current state, and obtains immediate feedback by interacting with the environment; optimizes meta-strategy parameters based on the immediate feedback, continuously optimizes decision effects, and collects task completion data during execution; The task completion data is analyzed and evaluated to extract cognitive level improvement indicators; the cognitive level improvement indicators are fed back to each level of the multi-level knowledge distillation mechanism, and the distillation parameters and compression ratio are dynamically adjusted; the extraction and compression effects of group evolution knowledge are optimized through parameter adjustment, thereby improving the quality of knowledge transfer and forming a complete optimization closed loop from task execution to cognitive enhancement.
[0083] First, a multi-level knowledge distillation mechanism is constructed. The mechanism consists of three distillation network levels, which are used to extract temporal knowledge features, refine evolutionary experience, and compress knowledge representation.
[0084] Take a group robot navigation task as an example. A group of robots explore and learn in a complex maze environment, with the goal of finding the exit. During the evolution process, each robot's behavior sequence (e.g., moving forward, turning left, turning right), state information (e.g., current location, surrounding environment information), and task completion status (e.g., whether it reaches the exit, how much time it takes) are recorded.
[0085] The collected behavior sequences and state information are fed into the first layer of distillation network - for example, a recurrent neural network (RNN). RNN can capture the temporal dependencies in the behavior sequence and extract temporal knowledge features that represent the robot's behavior pattern. For example, the robot's combined action of turning left and then moving forward may represent a specific navigation strategy.
[0086] The extracted temporal knowledge features and task completion are fed into the second layer of distillation network - for example, a multi-layer perceptron (MLP). MLP associates temporal features with task results and extracts successful evolutionary experience. For example, in certain specific locations, it is easier to find an exit by executing the strategy of "turn left first and then go forward".
[0087] Finally, the refined evolutionary experience is input into the third layer of distillation network - for example, an autoencoder. The autoencoder compresses the high-dimensional evolutionary experience into a low-dimensional knowledge compression representation for easy storage and migration. For example, a complex navigation strategy is compressed into a concise vector representation. Suppose that after compression, a successful navigation strategy is represented as a vector [0.8, 0.2, 0.5].
[0088] Analyze the feature distribution of the target scene. Assume that the target scene is a new maze environment, whose features may include the size of the maze, the number and distribution of obstacles, etc. Represent the features of the target scene as a vector, such as [0.5, 0.7, 0.1].
[0089] The knowledge compression representation is compared with the feature distribution of the target scene to calculate the knowledge transfer loss. For example, the Euclidean distance between vectors can be used to measure the difference between the knowledge compression representation and the target scene features.
[0090] Adaptively adjust the migration parameters based on the knowledge migration loss. For example, if the migration loss is large, it means that the source scene and the target scene are very different, and the proportion of knowledge migration needs to be reduced to avoid negative migration. On the contrary, if the migration loss is small, the proportion of knowledge migration can be increased. Assuming that the calculated migration loss is 0.6, according to the preset rules, the migration parameter is set to 0.4, which means that 40% of the evolution experience is migrated to the new scene.
[0091] Construct a meta-reinforcement learning framework and input the scene-adapted decision knowledge as prior information. For example, the adjusted knowledge compression representation is input into the meta-reinforcement learning policy network to guide the initialization of the policy.
[0092] The meta-reinforcement learning framework generates an action distribution based on the current state. For example, in a new maze environment, the robot generates a probability distribution of different actions based on its current position and surrounding environment information.
[0093] Get instant feedback by interacting with the environment. For example, after the robot performs an action, it observes the changes in the environment and receives rewards or penalties.
[0094] Optimize meta-strategy parameters based on immediate feedback and continuously optimize decision-making results. For example, use the policy gradient method to update meta-strategy parameters so that the robot tends to choose more favorable actions.
[0095] During the execution process, the task completion data is collected. For example, whether the robot successfully reaches the exit, how long it takes, etc.
[0096] Analyze and evaluate the task completion data to extract cognitive level improvement indicators. For example, count the number of times the robot successfully completes tasks, the average time taken, and other indicators to evaluate whether the robot's cognitive level has been improved.
[0097] Feedback the cognitive level improvement index to each level of the multi-level knowledge distillation mechanism to dynamically adjust the distillation parameters and compression ratio. For example, if the cognitive level is significantly improved, it means that the knowledge transfer is effective, and the compression ratio can be increased to extract more refined knowledge.
[0098] The solution of this application can: Improve the efficiency of knowledge transfer: Through the multi-level knowledge distillation mechanism, the complex evolution experience is compressed into a concise knowledge representation, which effectively reduces the cost of knowledge transfer and improves the efficiency of transfer. Enhance scene adaptability: By analyzing the characteristic distribution of the target scene and adaptively adjusting the migration parameters, the gradual transfer of knowledge is achieved, and the adaptability of the model to new scenes is enhanced. Promote the improvement of cognitive ability: Through the meta-reinforcement learning framework, the transferred knowledge is combined with environmental feedback, and the decision-making strategy is continuously optimized, which ultimately achieves the improvement of the robot's cognitive ability, enabling it to better complete complex tasks.
[0099] Figure 2 FIG. 1 is a schematic diagram of the structure of an intelligent inspection path planning system based on reinforcement learning according to an embodiment of the present invention. Figure 2 As shown, the system comprises: The first unit is used to obtain the multimodal perception information of the intelligent inspection robot, build a knowledge-driven scene understanding model based on the multimodal perception information, and realize the semantic level analysis of the inspection environment; based on the graph neural network and the recursive memory network, establish a dynamic interaction model between the intelligent agent and the environment; integrate the scene understanding model and the dynamic interaction model into the cognitive enhanced digital twin system to form a bidirectional mapping architecture from physical space to cognitive space; The second unit is used to design a self-organizing swarm intelligence algorithm based on the cognitive enhanced digital twin system, the self-organizing swarm intelligence algorithm integrates meta-heuristic search strategy and social learning mechanism; construct a multi-objective conflict coordinator, the multi-objective conflict coordinator balances inspection efficiency, energy consumption distribution and risk aversion through the Pareto optimization method; design a swarm behavior emergence prediction model, the swarm behavior emergence prediction model is based on chaos theory and complex network analysis to achieve dynamic prediction of the collaborative mode of swarm intelligent agents; use a federated learning framework to aggregate the experience knowledge of multiple intelligent agents under the premise of protecting privacy, and generate a swarm evolution strategy with adaptability higher than a preset adaptation threshold; The third unit is used to deploy the group evolution strategy to the edge computing network and establish a distributed decision-making system; construct a robustness evaluation mechanism with real-time feedback, which includes a fault warning module, an emergency response module and a strategy reconstruction module; design a task allocation optimizer based on the attention mechanism, which can adaptively adjust the task allocation plan according to the specialized capabilities and current status of the intelligent agent; establish a multi-level knowledge distillation mechanism to compress and migrate the experiential knowledge gained during the group evolution process to new scenarios; and continuously optimize the decision-making strategy through the meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement.
[0100] According to a third aspect of the embodiments of the present invention, An electronic device is provided, comprising: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0101] A fourth aspect of the embodiments of the present invention is: A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0102] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.
[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent inspection path planning method based on reinforcement learning, characterized in that: include: Acquire the multimodal perception information of the intelligent inspection robot, build a knowledge-driven scene understanding model based on the multimodal perception information, and realize semantic-level analysis of the inspection environment; establish a dynamic interaction model between the intelligent agent and the environment based on the graph neural network and the recursive memory network; integrate the scene understanding model and the dynamic interaction model into the cognitive enhanced digital twin system to form a bidirectional mapping architecture from physical space to cognitive space; Based on the cognitive enhanced digital twin system, a self-organizing swarm intelligence algorithm is designed, wherein the self-organizing swarm intelligence algorithm integrates a meta-heuristic search strategy and a social learning mechanism; Constructing a multi-objective conflict coordinator, which balances inspection efficiency, energy consumption distribution and risk avoidance through a Pareto optimization method; Designing a group behavior emergence prediction model, which is based on chaos theory and complex network analysis to achieve dynamic prediction of the coordination mode of group intelligent agents; Using the federated learning framework, the experience knowledge of multiple agents is aggregated under the premise of protecting privacy, generating a group evolution strategy with adaptability higher than the preset adaptation threshold; Deploy the group evolution strategy to the edge computing network to establish a distributed decision-making system; Construct a robustness evaluation mechanism with real-time feedback, which includes a fault warning module, an emergency response module, and a strategy reconstruction module; design a task allocation optimizer based on an attention mechanism, which can adaptively adjust the task allocation scheme according to the specialized capabilities and current status of the intelligent agent; establish a multi-level knowledge distillation mechanism to compress the experiential knowledge gained during the group evolution process and transfer it to new scenarios; The decision-making strategy is continuously optimized through the meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement.
2. The method according to claim 1, characterized in that: Based on the cognitive enhanced digital twin system, a self-organizing swarm intelligence algorithm is designed, wherein the self-organizing swarm intelligence algorithm integrates a meta-heuristic search strategy and a social learning mechanism; Construct a multi-objective conflict coordinator, which balances inspection efficiency, energy consumption distribution and risk avoidance through the Pareto optimization method, including: The cognitive enhanced digital twin system includes a knowledge reasoning module, which uses a graph neural network to establish a multi-agent interaction model; based on the multi-agent interaction model, a group behavior feature sequence and an environmental dynamic feature sequence are generated; According to the group behavior feature sequence and the environmental dynamic feature sequence, a self-organizing group intelligence algorithm is designed, and the self-organizing group intelligence algorithm integrates a meta-heuristic search strategy and a social learning mechanism; the meta-heuristic search strategy constructs a task decomposition model based on a dynamic memory structure, and the social learning mechanism constructs a knowledge transfer mechanism through a hierarchical attention network; based on the task decomposition model, a complex inspection task is divided into a sub-task sequence, and the group experience knowledge base is extracted using the knowledge transfer mechanism; the sub-task sequence and the group experience knowledge base are input into a reinforcement learning framework to generate an initial behavior decision plan; A multi-objective conflict coordinator is constructed to process the initial behavior decision plan. The multi-objective conflict coordinator uses a Pareto optimization method to dynamically weigh the inspection efficiency objective function, the energy consumption balance objective function and the risk control objective function.
3. The method according to claim 1, characterized in that Designing a group behavior emergence prediction model, which is based on chaos theory and complex network analysis to achieve dynamic prediction of the coordination mode of group intelligent agents; Using the federated learning framework, the experience knowledge of multiple agents is aggregated under the premise of protecting privacy, and the group evolution strategy with adaptability higher than the preset adaptation threshold is generated, including: Constructing a group behavior emergence prediction model, wherein the group behavior emergence prediction model establishes an agent interaction network; extracting time-varying network features based on the agent interaction network, constructing a chaotic dynamics equation based on the time-varying network features, and calculating a chaotic feature index; combining the time-varying network features and the chaotic feature index to form a collaborative mode feature; Inputting the collaborative mode features into a recurrent neural network, the recurrent neural network predicting the evolution trend of the group intelligent body based on the historical collaborative sequence; calculating the group behavior emergence probability and the collaborative mode conversion probability according to the evolution trend; and generating a group behavior prediction result based on the group behavior emergence probability and the collaborative mode conversion probability; A federated learning framework is constructed based on the group behavior prediction results, and the federated learning framework trains an empirical model locally in an intelligent agent; differential privacy encryption is performed on the parameters of the empirical model, and the differential privacy encryption performs parameter perturbations based on local sensitivity and privacy budget; the encrypted model parameters are transmitted to a central server, and the central server assigns aggregation weights according to the knowledge quality and collaborative contribution of the intelligent agent, and uses a federated averaging algorithm to aggregate the encrypted parameters of multiple intelligent agents; The environmental fitness of the aggregated model is calculated, and the environmental fitness is combined with the group behavior prediction result to evaluate the strategy stability and knowledge coverage; when the environmental fitness exceeds a preset adaptation threshold, the group evolution strategy is decrypted and generated.
4. The method according to claim 3, characterized in that: A federated learning framework is constructed based on the group behavior prediction result, wherein the federated learning framework trains an empirical model locally in the intelligent agent; differential privacy encryption is performed on the parameters of the empirical model, wherein the differential privacy encryption performs parameter perturbations based on local sensitivity and privacy budget; The encrypted model parameters are transmitted to the central server, and the central server assigns aggregation weights according to the knowledge quality and collaborative contribution of the agent, including: Building a federated learning framework based on the group behavior prediction results, wherein the group behavior prediction results include the behavior trajectory and decision-making mode of the agent; dividing the group behavior prediction results into a training data set, wherein the training data set includes a state vector and an action vector; configuring a local training environment for each agent in the federated learning framework; training an empirical model locally in the agent based on the training data set, wherein the empirical model outputs an action value evaluation result; Extract the network parameters of the empirical model and construct a parameter matrix; calculate the local sensitivity of the parameter matrix, where the local sensitivity characterizes the degree of response of the network parameters to data disturbance; determine the amount of noise injection according to the local sensitivity and a preset privacy budget; perform differential privacy encryption on the parameters of the empirical model, where the differential privacy encryption is achieved by injecting random noise that conforms to the Laplace distribution into the parameter matrix; Calculate the knowledge quality of the agent, the knowledge quality is evaluated based on the model prediction accuracy and knowledge coverage; count the collaborative contribution of the agent, the collaborative contribution is evaluated based on the training participation and sample quality; input the knowledge quality and the collaborative contribution into a weight distributor; the weight distributor outputs the aggregate weight of the agent in federated learning; transmit the encrypted model parameters and the aggregate weight to a central server, the central server receives the encrypted parameters and aggregate weight uploaded by multiple agents.
5. The method according to claim 1, characterized in that Deploy the swarm evolution strategy to the edge computing network to establish a distributed decision-making system; construct a robustness evaluation mechanism with real-time feedback, which includes a fault warning module, an emergency response module, and a strategy reconstruction module; Designing a task allocation optimizer based on the attention mechanism, which can adaptively adjust the task allocation scheme according to the agent's specialized capabilities and current state, including: Deploy the swarm evolution strategy to edge nodes in an edge computing network; collect the computing load, storage capacity and network bandwidth of the edge nodes to generate a resource status matrix; evaluate the task processing capability score of the edge nodes according to the resource status matrix; convert the swarm evolution strategy into a local execution instruction of the edge node based on the task processing capability score; establish a collaborative decision-making channel between edge nodes according to the local execution instruction, and the collaborative decision-making channel is used to transmit node status information and task scheduling instructions; Continuously monitor the execution status of the edge node, calculate the deviation value between the real-time status and the expected status; calculate the probability of failure based on the deviation value and historical fault data; when the probability of failure exceeds the preset fault threshold, generate corresponding resource allocation instructions according to the degree of deviation; execute the resource allocation instructions to reallocate computing resources and network bandwidth; update the local execution instructions according to the reallocated resources; adjust the data transmission strategy of the collaborative decision-making channel based on the updated local execution instructions; Obtain the skill score and current load status of the agent in the edge node, and construct the agent capability vector; analyze the computational complexity and delay constraint of the task to be assigned, and generate a task requirement vector; input the agent capability vector and the task requirement vector into the attention calculation unit; the attention calculation unit outputs the fitness weight of the agent and the task; and generate a task allocation plan according to the fitness weight and the current resource status matrix; Deploy the task allocation scheme and monitor the task execution process, record the execution quality and resource utilization; input the execution quality and resource utilization into the attention calculation unit, dynamically adjust the fitness weight; adaptively adjust the task allocation scheme according to the adjusted fitness weight.
6. The method according to claim 5, characterized in that Continuously monitor the execution status of the edge node, calculate the deviation value between the real-time status and the expected status; calculate the probability of failure based on the deviation value and historical failure data; When the probability of the failure occurring exceeds a preset failure threshold, a corresponding resource allocation instruction is generated according to the degree of deviation; Executing the resource allocation instruction to reallocate computing resources and network bandwidth includes: Continuously monitor the execution status of the edge node, collect performance indicators of the edge node, the performance indicators include memory usage, network throughput and task response time; obtain expected operating parameters of the edge node, the expected operating parameters are standard operating indicators set by the system; build a performance evaluation model, the performance evaluation model calculates the deviation value between the real-time state and the expected state based on the performance indicators and the expected operating parameters; The deviation value is input into a time series feature extractor, and the time series feature extractor generates a state deviation sequence; a fault record is read from a historical fault database, and the fault record contains state deviation data before the historical fault occurs; a fault prediction model is trained based on the state deviation sequence and the state deviation data, and the fault prediction model outputs a probability of fault occurrence; and the fault probability is compared with a preset fault threshold; When the probability of the failure occurring exceeds the preset failure threshold, the deviation value is input into the resource evaluator, and the resource evaluator calculates the resource demand according to the degree of deviation; generates a resource allocation instruction based on the resource demand, and the resource allocation instruction includes computing resources and network bandwidth; determines the resource allocation priority according to the changing trend of the performance indicator; executes the resource allocation instruction, reallocates the computing resources according to the resource allocation priority, and adjusts the network bandwidth.
7. The method according to claim 1, characterized in that Establish a multi-level knowledge distillation mechanism to compress the experiential knowledge gained during group evolution and transfer it to new scenarios; Continuously optimize decision strategies through the meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement, including: Establish a multi-level knowledge distillation mechanism, based on which the behavior sequence, state information and task completion status of the intelligent agents in the group evolution process are collected; the behavior sequence and the state information are input into the first layer of distillation network to extract temporal knowledge features, the temporal knowledge features and the task completion status are input into the second layer of distillation network to refine the evolution experience, and the evolution experience is compressed through the third layer of distillation network to obtain knowledge compression representation; Analyze the feature distribution of the target scene, compare the knowledge compression representation with the feature distribution of the target scene, and calculate the knowledge migration loss; adaptively adjust the migration parameters based on the knowledge migration loss to achieve the progressive migration of the knowledge compression representation to the new scene, and output the scene-adaptive decision knowledge; Construct a meta-reinforcement learning framework, and input the scenario-adapted decision knowledge into the meta-reinforcement learning framework as prior information; the meta-reinforcement learning framework generates an action distribution based on the current state, and obtains immediate feedback by interacting with the environment; optimizes meta-strategy parameters based on the immediate feedback, continuously optimizes decision effects, and collects task completion data during execution; The task completion data is analyzed and evaluated to extract cognitive level improvement indicators; the cognitive level improvement indicators are fed back to each level of the multi-level knowledge distillation mechanism, and the distillation parameters and compression ratio are dynamically adjusted; the extraction and compression effects of group evolution knowledge are optimized through parameter adjustment, thereby improving the quality of knowledge transfer and forming a complete optimization closed loop from task execution to cognitive enhancement.
8. An intelligent inspection path planning system based on reinforcement learning, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the multimodal perception information of the intelligent inspection robot, build a knowledge-driven scene understanding model based on the multimodal perception information, and realize the semantic level analysis of the inspection environment; based on the graph neural network and the recursive memory network, establish a dynamic interaction model between the intelligent agent and the environment; integrate the scene understanding model and the dynamic interaction model into the cognitive enhanced digital twin system to form a bidirectional mapping architecture from physical space to cognitive space; The second unit is used to design a self-organizing swarm intelligence algorithm based on the cognitive enhanced digital twin system, wherein the self-organizing swarm intelligence algorithm integrates a meta-heuristic search strategy and a social learning mechanism; Constructing a multi-objective conflict coordinator, which balances inspection efficiency, energy consumption distribution and risk avoidance through a Pareto optimization method; Designing a group behavior emergence prediction model, which is based on chaos theory and complex network analysis to achieve dynamic prediction of the coordination mode of group intelligent agents; Using the federated learning framework, the experience knowledge of multiple agents is aggregated under the premise of protecting privacy, generating a group evolution strategy with adaptability higher than the preset adaptation threshold; The third unit is used to deploy the group evolution strategy to the edge computing network to establish a distributed decision-making system; Construct a robustness evaluation mechanism with real-time feedback, which includes a fault warning module, an emergency response module, and a strategy reconstruction module; design a task allocation optimizer based on an attention mechanism, which can adaptively adjust the task allocation scheme according to the specialized capabilities and current status of the intelligent agent; establish a multi-level knowledge distillation mechanism to compress the experiential knowledge gained during the group evolution process and transfer it to new scenarios; The decision-making strategy is continuously optimized through the meta-reinforcement learning framework to achieve closed-loop optimization from task execution to cognitive enhancement.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Intelligent agent construction method and system based on agentive workflow
CN120144577A
Federal learning-based roadbed compaction quality evaluation system and method thereof
CN120146710A
Unmanned aerial vehicle inspection fault intelligent identification system and method
CN120298938A
Virtual workplace training task allocation method and system based on multi-agent collaboration
CN120355204A
Unmanned cluster cooperative combat digital twin deduction optimization method under complex meteorological conditions
CN120597701A