A multi-resource task planning method and system based on reinforcement learning and performance evaluation

By adopting a multi-resource task planning method based on reinforcement learning and performance evaluation, the real-time responsiveness and practicality of the task resource management system are solved, and efficient collaborative task planning for complex objectives is achieved.

CN116541797BActive Publication Date: 2025-12-23SICHUAN JIUZHOU ELECTRIC GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310344511.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2025-12-23
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

Existing task resource management and task planning systems have limited applicability, cannot respond to many real-time factors in a timely manner, and are relatively lacking in practicality.

Method used

A multi-resource task planning method based on reinforcement learning and performance evaluation is used to solve resources and plan collaboratively through digital modeling and simulation, target behavior threat library, task resource library and collaborative strategy library. Feedback evaluation is carried out by combining performance indicators, performance indicators, contribution degree, activity degree and other indicators to optimize collaborative strategy solution, generate the best matching collaborative planning strategy and execute target task planning.

Benefits of technology

It achieves a collaborative planning strategy for the efficient, scientific and rational use of task resources, and executes target task planning through network access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0004159170000000011
    Figure HDA0004159170000000011
  • Figure HDA0004159170000000012
    Figure HDA0004159170000000012
  • Figure HDA0004159170000000021
    Figure HDA0004159170000000021
Patent Text Reader

Abstract

The application discloses a multi-resource task planning method and system based on reinforcement learning and performance evaluation, and the target task is subjected to resource calculation and cooperative strategy calculation based on a target behavior threat library, a task resource library and a cooperative strategy library which have been constructed to obtain a cooperative planning strategy; different means task resource data are accessed through a network, and the data are collected and analyzed through middleware to form a multi-means task resource library; and the application forms a heterogeneous task resource cooperative behavior strategy according to existing task resource occupation conditions and target behavior threat analysis calculation, and supports multiple task resources to cooperatively complete the same target task.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of task planning, in particular to a multi-resource task planning method and system based on reinforcement learning and performance evaluation. BACKGROUND

[0002] In actual use, there are different types of task resources according to signal types, such as visible light, spectrum, microwave, infrared, electromagnetic induction, etc. They have different functions and use ranges on different platforms (space-based, sea-based, land-based), and perform targeted tasks. Due to different means, there are certain differences in the use method and equipment management system of the resource equipment, and when they are used together to perform collaborative tasks, it is necessary to plan and schedule appropriately and reasonably according to some rules or algorithms, so that the whole multi-means system can obtain target information more efficiently and accurately.

[0003] With the development and progress of computer science and technology such as artificial intelligence, big data analysis and prediction, knowledge graph (expert system) decision, data chain, information grid, distributed technology and multi-agent technology, many organizations have begun to study neural networks, expert systems, multi-agent technology and other technologies, and based on these technologies, they have started to establish multi-task resource management and task planning systems. There is a maximum detection probability management method based on annealing algorithm management resource expectation constraint condition, which greatly improves the monitoring probability of the target using dynamic programming method, but the multi-target detection efficiency is poor; the dynamic alliance method based on multi-agent theory manages multiple targets according to priority, organizes task resources according to alliance and competition management method, and can better and more scientifically allocate resources. These methods have a certain scientific theoretical basis, but they have great requirements for target modeling and threat modeling, and cannot respond to many real-time factors in time. The target data used for training is mostly simulated modeling, and the practicality is relatively lacking. SUMMARY

[0004] The technical problem to be solved by the present application is that the existing task resource management and task planning system has a single application scope and cannot respond to many real-time factors in time, and the practicability is relatively lacking.

[0005] The present application is implemented by the following technical solutions:

[0006] The present application provides a multi-resource task planning method based on reinforcement learning and performance evaluation, comprising the following steps:

[0007] Step 1: Obtain the target task and perform digital modeling simulation.

[0008] Step 2: Based on the target behavior threat library, task resource library and collaborative strategy library that have been constructed, the target task is solved for resources and collaborative strategy to obtain a collaborative planning strategy.

[0009] Step 3: Execute the target task according to the collaborative planning strategy: in the process of executing the target task, real-time resource scheduling is performed based on the target real-time behavior, task result feedback and real-time resource performance evaluation results; at the same time, the target real-time behavior is reinforced to optimize the target behavior threat library, the task result feedback is reinforced to optimize the collaborative planning strategy, and the real-time resource performance evaluation is reinforced to optimize the resource performance evaluation index library.

[0010] The working principle of the present application is as follows:

[0011] The single means task resource is difficult to complete the detection task of the composite target, and often needs the cooperation of multiple means task resources to monitor the composite target, which requires a high degree of information sharing. The present scheme is based on the target behavior threat library, the task resource library and the cooperative strategy library to solve the target task and the cooperative strategy, access different means task resource data through the network, and analyze the data through middleware to form a multi-means task resource library. According to the existing task resource occupation situation, the heterogeneous task resource cooperative behavior strategy is formed according to the target behavior threat analysis, and multiple task resources are supported to cooperatively complete the same target task.

[0012] Further optimization scheme is that the target behavior threat library is used for describing the task, behavior and threat degree of the target after entering the area;

[0013] The task resource library is used for describing the function index, hardware device capability index, physical parameter and means index matched with the target possessed by the task resource;

[0014] The cooperative strategy library is used for describing the use setting instruction, power setting instruction and cycle setting instruction of the task resource at different target distances, different target loads, different target behaviors and different threat levels;

[0015] The resource efficiency evaluation index library is used for describing the performance index and efficiency index of the task resource to the target or load, and describing the contribution index, activity index, availability index and credibility index of the task resource in the cooperative process.

[0016] Further optimization scheme is that step two includes the following substeps:

[0017] S21, correlation analysis based on the task resource library: the optimal matching strategy is generated by fusion analysis of the feedback of the availability, credibility, contribution and activity of the multi-source task resource in the task resource library, and the target resource action optimal strategy atlas is established;

[0018] S22, based on the target resource action optimal strategy atlas, the cooperative strategy is solved, and the resources required by the target task and the instructions required by the resources in the task are calculated.

[0019] And according to the instruction feedback result, the cooperative planning strategy is optimized, the target behavior modeling process is repeated, the target model library is widened, so that the resource solving and cooperative strategy selection can be more accurate and efficient in subsequent task planning.

[0020] Further optimization scheme is that the real-time resource efficiency evaluation method includes:

[0021] T1, establish a resource performance evaluation index factor set;

[0022] T2, divide the resource performance evaluation index factors into performance indicators, efficiency indicators and hardware parameters; take information collection, information identification and information fusion in the performance indicators as feedback factors; take availability, reliability, activity and participation in the efficiency indicators as calculation factors;

[0023] T3, according to the resource performance evaluation index factor set, establish a target adaptation degree, load adaptation degree and collaborative contribution degree evaluation set;

[0024] T4, calculate the evaluation matrix according to the fuzzy comprehensive evaluation algorithm, and combine the performance evaluation weight factor to obtain the comprehensive evaluation score.

[0025] Further optimization scheme is that each task resource is configured with a task execution agent, and the task execution agent converts the collaborative planning strategy into a control instruction of the task resource according to the task time and target detection information, and issues the control instruction to each task resource for execution; at the same time, the task execution agent also collects the detection information of the task resource and the state information of the task resource. Since the planning process involves the collaborative work of multiple task resources, the system task resource needs to make real-time adjustments according to time or target information, such as a node coordinating the work of each means task resource at the same time, the system calculation amount will be extremely heavy, and the stability of the system will be reduced. Therefore, through the distributed agent architecture, a task execution agent is created for each task resource.

[0026] The scheme also provides a multi-resource task planning system based on reinforcement learning and performance evaluation, which is used to realize the multi-resource task planning method based on reinforcement learning and performance evaluation of the above scheme, comprising:

[0027] The acquisition module is used to acquire the target task and perform digital modeling simulation;

[0028] The calculation module is used to calculate the target task based on the constructed target behavior threat library, task resource library and collaborative strategy library to obtain the collaborative planning strategy;

[0029] The execution module is used to execute the target task according to the collaborative planning strategy: in the process of executing the target task, real-time resource scheduling is performed based on the target real-time behavior, task result feedback and real-time resource performance evaluation result; at the same time, the target real-time behavior is reinforced to optimize the target behavior threat library, the task result feedback is reinforced to optimize the collaborative planning strategy, and the real-time resource performance evaluation is reinforced to optimize the resource performance evaluation index library.

[0030] Further optimization scheme is that the multi-resource task planning system comprises:

[0031] A data management subsystem is configured to manage target behavior threat data, task resource data, target coordination strategy data, coordination strategy data, and resource performance evaluation index data used for task planning, and to provide data basis for resource rapid calculation, target behavior prediction, target threat analysis, coordination strategy calculation, and performance evaluation in a task process.

[0032] A task planning and execution subsystem is configured to perform rapid planning of task resources based on target task information, to control task start, task pause, task end, and task re-planning according to a state of task execution and feedback information, to support human-in-the-loop task resource task control, to support target data simulation, and to support sensor data simulation.

[0033] A task resource coordination system is configured to evaluate a target based on a task result, to construct an index system for task result evaluation, to calculate a task evaluation result based on task data and the index system for task result evaluation, and to visually display the evaluation result in a form of a graph or a table.

[0034] A service and feedback subsystem is configured to provide resource calculation services for resource selection based on target behavior and threat, to provide coordination strategy calculation services for a task execution process, to provide instruction distribution control services for task process control and task resource task control, to provide reinforcement learning services for task resource target adaptation, and to provide feedback evolution services for coordination strategies.

[0035] A data communication service subsystem is configured to send task planning instructions to specific task resources for execution, to receive data of the task resources, and to communicate with external systems as an interface of the system and the external systems, to receive external information data, and to send data of the system to the external systems.

[0036] A further optimization scheme is that the data communication service subsystem includes multiple communication interfaces, and is compatible with input and output of different interface data, and can guarantee parallel access and parallel execution of instructions calculated by a coordination strategy of multiple task resources in a target task process.

[0037] A further optimization scheme is that the data communication service subsystem provides a unified clock service.

[0038] A further optimization scheme is that a system is designed based on a centralized management and distributed execution idea, and in a process of effectively allocating heterogeneous multi-means tasks, complexity of multi-task resource management is reduced, and system scalability is well satisfied. According to demand analysis and design constraints of the system, a system architecture of the multi-resource task planning system is logically divided into a user layer, an application function layer, a service layer, and an interface layer.

[0039] User layer, used as a user interaction means with the multi-resource task planning system, to establish data editing positions, task planning positions, task evaluation positions, and complete the business work of each position;

[0040] Application layer, used for centralized processing of data and business involved in the target task execution process; receiving business work instructions from the user layer, converting them into system service instructions, and sending them to the service layer for execution; also used for receiving data pushed by the service layer and visualizing the situation data;

[0041] Service layer, used to provide resource solving services, collaborative strategy resource allocation task decomposition, reinforcement learning services, performance evaluation services, and data subscription and distribution services;

[0042] Interface layer, used to establish a connection between the multi-resource task planning system and the task resource hardware devices, drive the task resource hardware devices to execute instructions, collect state information and detection results of the task resource hardware devices, fuse the detection results, and feed back to the service layer and the application layer.

[0043] The multi-resource task planning system meets the rapid planning and execution of task resource detection, data collection, data fusion positioning, ensures the target matching of task planning and the efficiency of task execution, and effectively supports the identification and generation of task target trajectories; the multi-resource task planning system supports reinforcement (deep) learning through feedback in the task planning process, task execution process, and simulation tasks, and strengthens the adaptability between different task resources and targets and target loads, to meet the efficient calculation of task resources for different task targets, different target behaviors, and different threat levels.

[0044] The system uses reinforcement learning method to predict target behavior and threat according to the target library, calculate multi-means task resource specific parameters, deploy resources, and issue corresponding instructions; the task resources execute the instructions and provide effectiveness feedback, perform resource-target performance evaluation, evaluate target individual fitness and contribution, use optimal selection method to adjust resource occupation and resource parameters in real time for target individuals, and form an efficient behavior strategy for a certain type of target; the system can establish a task when a certain type of target performs different tasks, predict target tasks and behaviors, and if the current collaborative planning strategy cannot meet the task, a new collaborative strategy is selected and new resources and instructions are recalculated, and the performance is evaluated through real-time feedback of the task resources, the collaborative strategy is diversified adjusted according to the evaluation results, a new efficient collaborative planning strategy is generated, the target modeling and collaborative strategy modeling process is repeated, the target and collaborative strategy model library is expanded, and the target and collaborative strategy are reinforced.

[0045] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0046] The application provides a multi-resource task planning method based on reinforcement learning and performance evaluation, and the method comprises the following steps: performing resource calculation and cooperative strategy calculation on a target task based on a constructed target behavior threat library, a task resource library and a cooperative strategy library to obtain a cooperative planning strategy; accessing different means task resource data through a network, and collecting and analyzing the data through middleware to form a multi-means task resource library; and according to an existing task resource occupation condition, forming a heterogeneous task resource cooperative behavior strategy according to target behavior threat analysis calculation, and supporting multiple task resources to cooperatively complete a same target task.

[0047] The application provides a multi-resource task planning system based on reinforcement learning and performance evaluation, and the system is based on reinforcement learning (task resource-target-performance evaluation feedback modeling) and big data analysis (target behavior threat modeling analysis) technologies, and key nodes in an algorithm are modeled and fused with input and output in actual use, and more scientific results and services are obtained according to algorithm training results, and performance evaluation is performed in real time according to detection feedback in a task, and a database and a cooperative planning strategy are optimized, and ideas and research directions are provided for construction of a multi-target multi-resource cooperative management system. BRIEF DESCRIPTION OF DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the example embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor. In the drawings:

[0049] Figure 1 It is a multi-resource task planning principle diagram based on reinforcement learning and performance evaluation;

[0050] Figure 2 It is a multi-resource task planning system structure diagram based on reinforcement learning and performance evaluation;

[0051] Figure 3 It is a multi-resource task planning system logic architecture diagram based on reinforcement learning and performance evaluation;

[0052] Figure 4 It is a multi-resource task planning system working principle diagram A based on reinforcement learning and performance evaluation;

[0053] Figure 5 It is a multi-resource task planning system working principle diagram B based on reinforcement learning and performance evaluation;

[0054] Figure 6This is a schematic diagram of a multi-resource task planning process based on reinforcement learning and performance evaluation. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.

[0056] Example 1

[0057] like Figure 1 and Figure 6 As shown, this embodiment provides a multi-resource task planning method based on reinforcement learning and performance evaluation, including the following steps:

[0058] Step 1: Obtain the target task and perform digital modeling and simulation;

[0059] Step 2: Based on the established target behavior threat library, task resource library, and collaboration strategy library, perform resource calculation and collaboration strategy calculation on the target task to obtain the collaboration planning strategy;

[0060] Step 3: Execute the target task according to the collaborative planning strategy: During the execution of the target task, real-time resource scheduling is carried out based on the target's real-time behavior, task result feedback, and real-time resource efficiency evaluation results; at the same time, the target's real-time behavior is reinforced to optimize the target behavior threat database, the task result feedback is reinforced to optimize the collaborative planning strategy, and the real-time resource efficiency evaluation is reinforced to optimize the resource efficiency evaluation index database.

[0061] The target behavior threat database is used to describe the tasks, behaviors, and threat levels of targets after entering the area;

[0062] The task resource library is used to describe the functional indicators, hardware equipment capability indicators, physical parameters, and means indicators that match the target of the task resources.

[0063] The collaborative strategy library is used to describe the usage setting instructions, power setting instructions, and cycle setting instructions for task resources under different target distances, different target payloads, different target behaviors, and different threat levels.

[0064] The resource performance evaluation index library is used to describe the performance and effectiveness of task resources to the target or payload, as well as the contribution, activity, availability and reliability of task resources in the collaborative process.

[0065] Step two includes the following sub-steps:

[0066] S21, correlation analysis based on the task resource library: through the fusion analysis of the feedback of the availability, reliability, contribution and activity of the multi-source task resources in the task resource library, the optimal matching strategy is generated, and the target resource action optimal strategy atlas is established;

[0067] S22, based on the target resource action optimal strategy atlas, the collaborative strategy is calculated, and the resources required by the target task and the instructions required by the resources in the task are calculated.

[0068] According to the instruction feedback result, the collaborative strategy is optimized, the target behavior modeling process is repeated, the target model library is widened, and the subsequent task planning can be more accurate and efficient in resource solving and collaborative strategy selection.

[0069] The multi-resource task planning method based on reinforcement learning and performance evaluation has the characteristics that the real-time resource performance evaluation method comprises:

[0070] T1, a set of resource performance evaluation index factors is established;

[0071] T2, the resource performance evaluation index factors are divided into performance index, performance index and hardware parameter; the information collection, information identification and information fusion in the performance index are taken as feedback factors; the availability, reliability, activity and participation in the performance index are taken as calculation factors;

[0072] T3, according to the set of resource performance evaluation index factors, a set of target adaptation degree, load adaptation degree and collaborative contribution degree evaluation is established;

[0073] T4, according to the fuzzy comprehensive evaluation algorithm, the evaluation matrix is calculated, and the comprehensive evaluation score is obtained by combining the performance evaluation weight factor.

[0074] Embodiment 2

[0075] The embodiment provides a multi-resource task planning system based on reinforcement learning and performance evaluation, which is used for realizing the multi-resource task planning method based on reinforcement learning and performance evaluation of the previous embodiment, and comprises:

[0076] The acquisition module is used for acquiring the target task and performing digital modeling simulation;

[0077] The solving module is used for resource solving and collaborative strategy solving of the target task based on the target behavior threat library, the task resource library and the collaborative strategy library which have been constructed to obtain the collaborative planning strategy;

[0078] The execution module is configured to execute the target task according to the cooperative planning strategy, and in the process of executing the target task, real-time resource scheduling is performed based on the target real-time behavior, task result feedback and real-time resource performance evaluation result; meanwhile, the target real-time behavior is reinforced to optimize the target behavior threat library, the task result feedback is reinforced to optimize the cooperative planning strategy, and the real-time resource performance evaluation is reinforced to optimize the resource performance evaluation index library.

[0079] As shown in Figure 2 The multi-resource task planning system comprises:

[0080] The data management subsystem is configured to manage target behavior threat data, task resource data, target cooperative strategy data, cooperative strategy data and resource performance evaluation index data used for task planning, and provide a data basis for resource rapid calculation, target behavior prediction, target threat analysis, cooperative strategy calculation and performance evaluation in a task process.

[0081] The task planning and execution subsystem is configured to perform rapid task resource planning based on target task information, control task start, task pause, task end and task re-planning according to a state of task execution and feedback information, support human-in-the-loop task resource task control, support target data simulation, and support sensor data simulation.

[0082] The task resource cooperative system is configured to construct an index system for task result evaluation according to a task result evaluation target, calculate a task evaluation result according to task data and the index system for task result evaluation, and visually display the evaluation result in the form of a graph or a table.

[0083] The service and feedback subsystem is configured to provide resource calculation services for resource selection according to target behavior and threats, provide cooperative strategy calculation services for a task execution process, provide instruction distribution control services for task process control and task resource task control, provide reinforcement learning services for task resource target adaptation, and provide feedback evolution services for cooperative strategies.

[0084] The data communication service subsystem is configured to issue task planning instructions to specific task resources for execution, receive data of the task resources, and serve as an interface of the system and an external system to receive external information data and send system data to the external system.

[0085] The data communication service subsystem comprises a plurality of communication interfaces compatible with input and output of different interface data, and can guarantee parallel access and parallel execution of instructions calculated by the cooperative strategy of multiple task resources in the process of executing the target task.

[0086] The data communication service subsystem provides a unified clock service. In order to ensure time synchronization of all task resources, the communication service provides a unified clock service. In communication with the task resources, the task resource communication service is not directly connected with the task resources, but is connected through individual task resource agents. The task resource agents are customized according to the interface parameters of the task resources, can convert task resource data and publish it to the system, and can also convert user instructions into control commands of the task resources and send them to the task resources. All task resource agents work in parallel under the scheduling of the task resource service, improve the execution efficiency of the system, and at the same time, when new task resources need to be expanded, new agent nodes can be added, meeting the scalability of the system.

[0087] As shown in Figure 3 , the system architecture of the multi-resource task planning system includes:

[0088] A user layer, which is used as an interactive means between users and the multi-resource task planning system, establishes data editing positions, task planning positions, and task evaluation positions, and completes business work of each position;

[0089] An application layer, which is used to centrally process data and business involved in the execution process of target tasks, receives business work instructions from the user layer, converts them into system service instructions, and sends them to the service layer for execution, and is also used to receive data pushed by the service layer and visually display the situation data;

[0090] A service layer, which is used to provide resource solving services, collaborative strategy resource allocation task decomposition, reinforcement learning services, effectiveness evaluation services, and data subscription and distribution services;

[0091] An interface layer, which is used to establish a connection between the multi-resource task planning system and task resource hardware devices, drive the task resource hardware devices to execute instructions, collect state information and detection results of the task resource hardware devices, fuse the detection results, and feed back to the service layer and the application layer.

[0092] As shown in Figure 4 and Figure 5As shown, each task resource is configured with a task execution agent, which converts the cooperative planning strategy into control instructions of the task resource according to the task time and target detection information, and issues the control instructions to each task resource for execution; meanwhile, the task execution agent also collects the detection information of the task resource and the state information of the task resource. According to the target behavior threat modeling, the task resource capability model, and the cooperative strategy modeling, the target threat assessment, the target intention prediction, the target action direction prediction and assessment are performed, and through the resource solving service and the cooperative strategy solving service, the task resources participating in the task and the parameter instructions are calculated for the task planning, and are sent to the subscription distribution service and distributed to each task resource for execution to perform target detection and situation awareness; the resource agent receives the task, calculates according to the cooperative planning strategy, and drives the task resource to execute the task in real time according to the time or detection result, and collects the task resource state and resource detection fusion result and sends them to the task front end to provide data basis for the task resource efficiency evaluation; meanwhile, according to the detection information, the situation information, and the task resource instruction feedback information, the cooperative strategy is calculated, the real-time task resource scheduling and parameter instructions are calculated to match the current situation, and the target is tracked to achieve the final target of target track detection and generation.

[0093] The system uses a method based on reinforcement learning, uses real-time target tasks, target simulation data, target real track record playback, and other target data for driving training, uses the target detection result as an internal reinforcement signal, uses the time difference prediction method TD algorithm for learning for the efficiency evaluation, and performs genetic operation for the cooperative instruction solving, uses the internal reinforcement signal (detection result modeling) as the fitness function of the cooperative instruction, solves the behavior reaction of the task resource for the target and the detection result, that is, the specific instruction (action reinforcement signal) of the task resource, and then issues the instruction to the task resource through the agent service, and obtains the feedback (external reinforcement signal) of the task resource for this behavior, which can effectively train the single-means task resource, can effectively detect the adaptability of various means task resources to a certain type of target and corresponding load, and can train the efficiency evaluation index of each means task resource to the target and load, calculate the adaptability of each task resource to the target, and establish a task resource cooperative strategy model to provide basis and service for the task planning resource solution and the task process control algorithm.

[0094] The multi-resource task planning system is a complete digital modeling, strategy analysis, task planning, task resource control, instruction feedback, and evaluation evolution system. In the early stage, a target digital database, a target behavior database, a threat analysis database, a task resource capability database, a collaborative strategy database, and an efficiency evaluation index database are constructed. In subsequent task planning, task decomposition is reasonably performed through multi-target priority evaluation (threat degree), and then scientific and reasonable task resource matching planning is performed through target behavior prediction. Real-time resource scheduling is performed based on real-time target behavior or resource detection results feedback and real-time resource efficiency evaluation results. Task sensor instructions are calculated according to the collaborative strategy, and on the basis of ensuring the completion of the task, each database is perfected. Thus, the multi-means system can achieve a global optimal solution in the next task planning, more scientifically and reasonably use task resources, and more efficiently and accurately perform multi-target collaborative detection tasks.

[0095] The above specific embodiments further illustrate the purposes, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A multi-resource task planning method based on reinforcement learning and performance evaluation, characterized in that, The method comprises the steps of: Step 1: obtaining a target task and performing digital modeling simulation; Step 2: performing resource calculation and collaborative strategy calculation on the target task based on the constructed target behavior threat library, task resource library and collaborative strategy library to obtain a collaborative planning strategy; Step 3: executing the target task according to the collaborative planning strategy: in the process of executing the target task, real-time resource scheduling is performed based on target real-time behavior, task result feedback and real-time resource efficiency evaluation results; at the same time, the target real-time behavior is reinforced to optimize the target behavior threat library, the task result feedback is reinforced to optimize the collaborative planning strategy, and the real-time resource efficiency evaluation is reinforced to optimize the resource efficiency evaluation index library; The target behavior threat library is used to describe the task, behavior and threat degree of the target after entering the area; The task resource library is used to describe the function index, hardware device capability index, physical parameter and means index matched with the target possessed by the task resource; The collaborative strategy library is used to describe the use setting instruction, power setting instruction and cycle setting instruction of the task resource under different target distances, different target loads, different target behaviors and different threat levels; The resource efficiency evaluation index library is used to describe the performance index and efficiency index of the task resource to the target or load, and describe the contribution index, activity index, availability index and credibility index of the task resource in the collaborative process; Step 2 comprises the following substeps: S21, performing correlation analysis based on the task resource library: the optimal matching strategy is generated by fusion analysis on the feedback of the availability, credibility, contribution and activity of the multi-source task resource in the task resource library, and the target resource action optimal strategy atlas is established; S22, taking the target resource action optimal strategy atlas as the blueprint, the resource required by the target task and the instruction required by the resource in the task are calculated; The method for real-time resource efficiency evaluation comprises: T1, establishing a resource efficiency evaluation index factor set; T2, dividing the resource efficiency evaluation index factors into performance index, efficiency index and hardware parameter; taking information collection, information identification and information fusion in the performance index as feedback factors; taking availability, credibility, activity and participation in the efficiency index as calculation factors; T3, establishing a target adaptation degree, load adaptation degree and collaborative contribution degree evaluation set according to the resource efficiency evaluation index factor set; T4, calculating an evaluation matrix according to a fuzzy comprehensive evaluation algorithm, and combining an efficiency evaluation weight factor to obtain a comprehensive evaluation score. 2.The method of claim 1, wherein, Each task resource is configured with a task execution agent, and the task execution agent converts the collaborative planning strategy into control instructions of the task resource according to the task time and target detection information, and issues the control instructions to each task resource for execution; at the same time, the task execution agent also collects detection information of the task resource and state information of the task resource.

3. A multi-resource task planning system based on reinforcement learning and performance evaluation, characterized in that, The method for realizing the multi-resource task planning based on reinforcement learning and efficiency evaluation in claim 1 or 2 comprises: An acquisition module is configured to obtain a target task and perform digital modeling simulation; The solving module is configured to perform resource solving and cooperative strategy solving on the target task based on the constructed target behavior threat library, the task resource library and the cooperative strategy library to obtain a cooperative planning strategy. The execution module is configured to execute the target task according to the cooperative planning strategy, perform real-time resource scheduling based on the target real-time behavior, the task result feedback and the real-time resource performance evaluation result in the process of executing the target task, and perform reinforcement learning on the target real-time behavior to optimize the target behavior threat library, reinforcement learning on the task result feedback to optimize the cooperative planning strategy, and reinforcement learning on the real-time resource performance evaluation to optimize the resource performance evaluation index library.

4. The multi-resource task planning system based on reinforcement learning and performance evaluation of claim 3, wherein, The system comprises: The data management subsystem is configured to manage target behavior threat data, task resource data, target cooperative strategy data, cooperative strategy data and resource performance evaluation index data used for task planning. The task planning and execution subsystem is configured to perform rapid task resource planning based on target task information, control task start, task pause, task end and task re-planning according to the state of task execution and feedback intelligence. The task resource cooperation system is configured to construct an index system for task result evaluation according to the task result evaluation target, calculate the task evaluation result according to the task data and the index system for task result evaluation, and visually display the evaluation result in the form of a graph or a table. The service and feedback subsystem is configured to provide resource solving services for resource selection according to the target behavior and threat, and provide cooperative strategy solving services for the task execution process. The data communication service subsystem is configured to send task planning instructions to specific task resources for execution, receive data of the task resources, and serve as an interface for communication with external systems.

5. The multi-resource task planning system based on reinforcement learning and performance evaluation according to claim 4, characterized in that, The data communication service subsystem comprises a plurality of communication interfaces compatible with input and output of different interface data, and can guarantee parallel access of multiple task resources in the process of executing the target task.

6. The multi-resource task planning system based on reinforcement learning and performance evaluation according to claim 5, characterized in that, The data communication service subsystem provides unified clock services.

7. The multi-resource task planning system based on reinforcement learning and performance evaluation of claim 3, wherein, The system architecture of the multi-resource task planning system comprises: The user layer is configured to serve as an interactive means between users and the multi-resource task planning system, establish data editing positions, task planning positions and task evaluation positions, and complete business work of the positions. The application layer is configured to centrally process data and business involved in the process of executing the target task, receive business work instructions of the user layer, convert the business work instructions into instructions of system services, send the instructions to the service layer for execution, and receive data pushed by the service layer and visually display the situation data. The service layer is configured to provide resource solving services, cooperative strategy solving services, resource allocation and task decomposition services, reinforcement learning services, performance evaluation services, and data subscription and distribution services. The interface layer is configured to establish a connection between the multi-resource task planning system and task resource hardware devices, drive the task resource hardware devices to execute instructions, collect state information and detection results of the task resource hardware devices, fuse the detection results, and feed back the detection results to the service layer and the application layer.

Citation Information

Patent Citations

  • Intelligent game confrontation deduction system and method for electronic reconnaissance

    CN115796042A

  • KR20220026449A