Reinforcement Learning-Based Maintenance Plan Generation System, Method, and Program Product
By dividing into multiple modal layers in the heterogeneous knowledge graph and adopting reinforcement learning methods, the problem that secondary agents can only retrieve one voice information in the prior art is solved, and rapid retrieval and comprehensive evaluation of multiple modal maintenance solutions are achieved.
Patent Information
- Application Number
- CN202510368659.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-27
AI Technical Summary
In the prior art, secondary agents can only retrieve information in one voice in the database, resulting in poor comprehensiveness and the inability to quickly retrieve maintenance solutions of multiple modalities.
A maintenance solution generation system based on reinforcement learning is adopted. By dividing it into multiple modal layers in the heterogeneous knowledge graph, each sub-agent group is searched in one modal layer, and the results are evaluated and reward allocation are performed through the main agent to optimize the search control strategy of each sub-agent.
It improves the comprehensiveness and speed of searches, can quickly retrieve maintenance solutions of multiple modalities, and enhances the adaptability and efficiency of the system.
Smart Images

Figure CN119884396B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a maintenance plan generation system, method and program product based on reinforcement learning, belonging to the technical field of artificial intelligence. Background Art
[0002] The fault diagnosis and maintenance strategies of complex equipment usually require highly customized designs, which rely heavily on expert knowledge, making them expensive and difficult to scale or standardize.
[0003] To solve these problems, a Chinese patent application with publication number CN118798859A discloses a maintenance plan generation system, method and program product based on a large language model. In the system, a task allocator is used to allocate maintenance problems to K secondary agents; a relevance calculator is used to calculate the relevance between the maintenance plans retrieved by each secondary agent and the allocated maintenance problems; a first comparator is used to compare the relevance with a threshold value and provide the maintenance modification plans with relevance exceeding the threshold value to a trainer; the trainer calculates a first time variance target value using the set of maintenance plans with relevance exceeding the threshold value, the set of predicted retrieval strategies and the set of rewards, and broadcasts it to the K secondary agents; among the secondary agents, a retrieval strategy network retrieves and generates maintenance plans from a database according to the maintenance problems allocated by the task allocator, and reports the retrieval strategy and reward to a primary agent.
[0004] However, in the technical solution disclosed in this patent application, all secondary agents can only retrieve information in one voice in the database, and the comprehensiveness is poor. Summary of the Invention
[0005] To overcome the disadvantages in the prior art, the invention aims to provide a maintenance plan generation system, method and program product based on reinforcement learning, which can retrieve information in multiple voices and has high comprehensiveness; when each sub-agent performs a retrieval action, it needs to consider the retrieval actions, rewards and retrieval results of other sub-agents for self-retrieval, and can quickly retrieve maintenance plans in multiple modalities.
[0006] To achieve the above invention purpose, the present invention provides a maintenance plan generation system based on reinforcement learning, which includes a heterogeneous knowledge graph in P layers of modalities, a primary agent A m and N sub-agents A n |n = 1, …, N, where P is greater than or equal to 3, , wherein, the N sub-agents A n |n = 1, …, N are divided into P sub-agent groups, and each sub-agent group includes one or more sub-agents; the sub-agents in the P sub-agent groups respectively retrieve according to the primary agent A mThe P modal information of the issued maintenance plan query request label performs retrieval actions in different modal layers of the heterogeneous knowledge graph and reports its retrieval results to the main agent A m , the main agent A m is responsible for receiving and processing the retrieval results of each sub-agent from P sub-agent groups, respectively evaluating the matching degrees between the retrieval results of each sub-agent in the P sub-agent groups and the P modalities of the maintenance plan query request label, as well as the matching degree between the combined retrieval results of the optimal retrieval results of each sub-agent group and the maintenance plan query request label, calculating the reward values assigned to each sub-agent in each sub-agent group, and broadcasting the reward values, retrieval actions, and retrieval results of each sub-agent; each sub-agent in each sub-agent group optimizes its retrieval control strategy according to the reward values, retrieval actions, and retrieval results of all sub-agents received
[0007] To achieve the above invention purpose, the present invention also provides a maintenance plan generation method based on reinforcement learning, which includes the following steps
[0008] Step 1: Divide N sub-agents A n |n = 1,…, N into P sub-agent groups
[0009] Step 2: The sub-agents in each sub-agent group respectively perform retrieval actions in different modal layers of the heterogeneous knowledge graph according to the P modal information of the maintenance plan query request label issued by the main agent A m and report their retrieval results to the main agent A m ;
[0010] Step 3: The main agent A m is responsible for receiving and processing the retrieval results of each sub-agent from P sub-agent groups, respectively evaluating the matching degrees between the retrieval results of each sub-agent in the P sub-agent groups and the P modalities of the maintenance plan query request label, as well as the matching degree between the combined retrieval results of the optimal retrieval results of each sub-agent group and the maintenance plan query request label, calculating the reward values assigned to each sub-agent, and broadcasting the reward values, retrieval actions, and retrieval results of each sub-agent
[0011] Step 4: Each sub-agent in each sub-agent group optimizes its retrieval control strategy according to the reward values, retrieval actions, and retrieval results of all sub-agents received
[0012] To achieve the above invention purpose, the present invention also provides a computer program product, which includes a computer program that implements the above-mentioned maintenance plan generation method based on reinforcement learning when executed by a processor
[0013] Compared with the prior art, the reinforcement learning-based maintenance plan generation system, method, and program product provided by the present invention capture implicit relational knowledge from the maintenance practices of experienced technicians, integrate multi-modal data into a unified heterogeneous knowledge graph, divide the heterogeneous knowledge graph into multiple modal layers, enable one or more sub-agent groups to perform retrieval in the knowledge graph layers of one modality respectively, and then synthesize the retrievals of the sub-agent groups, greatly improving comprehensiveness. Moreover, when each sub-agent performs a retrieval action, it needs to consider the retrieval actions, rewards, and retrieval results of all sub-agents for self-retrieval to optimize its own retrieval control strategy and learn from the preferred retrieval results, greatly improving the retrieval speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a block diagram of the composition of the reinforcement learning-based maintenance plan generation system provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] In order to make the technical problems, technical solutions, and beneficial effects to be solved by the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0016] Figure 1 is a block diagram of the composition of the reinforcement learning-based maintenance plan generation system provided by the present invention. As Figure 1 shown, the reinforcement learning-based maintenance plan generation system provided by the present invention includes a heterogeneous knowledge graph with P layers of modalities, a main agent A m and N sub-agents A n |n = 1, …, N, where P is greater than or equal to 3. , wherein, the N sub-agents A n |n = 1, …, N are divided into P sub-agent groups, and each sub-agent group includes one or more sub-agents; the sub-agents of the P sub-agent groups respectively perform retrieval actions in different modal layers of the heterogeneous knowledge graph according to the P types of modal information of the maintenance plan query request tags issued by the main agent A m and report their retrieval results to the main agent A m , the main agent A mResponsible for receiving and processing the retrieval results of each sub-agent from P sub-agent groups, respectively evaluating the matching degrees between the retrieval results of each sub-agent in the P sub-agent groups and the P modalities of the maintenance plan query request label, as well as the matching degree between the combined retrieval result of the optimal retrieval results of each sub-agent group and the maintenance plan query request label, calculating the reward values assigned to the sub-agents in each sub-agent group, and broadcasting the reward values, retrieval actions, and retrieval results of each sub-agent; each sub-agent in each sub-agent group optimizes its retrieval control strategy according to the reward values, retrieval actions, and retrieval results of all sub-agents received. The modalities are, for example, text, image, formula, code, etc.
[0017] In the present invention, the sub-agents in the P sub-agent groups respectively perform retrieval actions in the first modality layer,..., the p-th layer,..., the P-th modality layer of the heterogeneous knowledge graph according to the P modality information of the maintenance plan query request label issued by the main agent A m and obtain retrieval results.
[0018] The present invention also provides a maintenance plan generation method based on reinforcement learning, which includes the following steps:
[0019] Step 1: Divide N sub-agents A n |n = 1,…, N into P sub-agent groups, such as sub-agent group C1,..., sub-agent group C p ,..., sub-agent group C P ;
[0020] Step 2: The sub-agents in each sub-agent group respectively perform retrieval actions in different modality layers of the heterogeneous knowledge graph according to the P modality information of the maintenance plan query request label issued by the main agent A m and report their retrieval results to the main agent A m ;
[0021] Step 3: The main agent A m is responsible for receiving and processing the retrieval results of each sub-agent from the P sub-agent groups, respectively evaluating the matching degrees between the retrieval results of each sub-agent in the P sub-agent groups and the P modalities of the maintenance plan query request label, as well as the matching degree between the combined retrieval result of the optimal retrieval results of each sub-agent group and the maintenance plan query request label, calculating the reward values assigned to each sub-agent, and broadcasting the reward values, retrieval actions, and retrieval results of each sub-agent;
[0022] Step 4: Each sub-agent in each sub-agent group optimizes its retrieval control strategy according to the reward values, retrieval actions, and retrieval results of all sub-agents received.
[0023] In the present invention, the construction of the heterogeneous knowledge graph includes the following steps:
[0024] S1-01: Construct knowledge graphs for M different domains respectively, and use the self-training method based on GCN to obtain the semantic vectors of the labels of each knowledge node in the knowledge graph of each domain, where M is a positive integer greater than or equal to 3;
[0025] S1-02: Associate and merge the knowledge graphs from M domains to create a heterogeneous knowledge graph;
[0026] S1-03: Stratify the heterogeneous knowledge graph by modality into P layers.
[0027] In the present invention, when N = P, the retrieval control strategies of each sub-agent in each sub-agent group are optimized according to the reward values, retrieval actions and retrieval results of all sub-agents received, including the following process:
[0028] S2-01: Master agent A m Assign the modalities of the maintenance plan query request labels to N sub-agents A n |n = 1,…,N. The sub-agent A n Performs a retrieval action at time t in the corresponding modality layer of the heterogeneous knowledge graph according to the different modalities of the assigned maintenance plan query request labels , and obtains the local retrieval result at time t ;
[0029] S2-02: Calculate the state-action function of sub-agent A n according to the following formula:
[0030] ,
[0031] In the formula, π n is the current retrieval control strategy of agent A n ; is the local retrieval result of sub-agent A n at time t, is the retrieval action of sub-agent A n at time t; is for sub-agent A n When performing the retrieval action at time t , the probability of transferring from the local retrieval result at time t being to the local retrieval result at the next time t+1 ; is for sub-agent A n When performing the retrieval action at time t to obtain the reward of master agent A m ; is the state value function of the local retrieval result ; represents sub-agent A nAll observation spaces; For sub-agent A n The improved retrieval control strategy; r is the number of training times;
[0032] S2-03: Sub-agent A n Calculate the state value function value of the local retrieval result at time t according to the following formula: when:
[0033] ,
[0034] In the formula, represents the detection action retrieval adopted by sub-agent A n when the detection state is for the control strategy of; ; represents all retrieval actions of sub-agent A n ;
[0035] S2-04: Sub-agent A n Judge whether it is less than or equal to the threshold , if so, execute step S2-09, if not, execute step S2-05;
[0036] S2-05: , and recalculate the state-action function of sub-agent A n ; ;
[0037] S2-06: Master agent A m Calculate the global state-action function according to the following formula:
[0038] ,
[0039] In the formula, ; , ;
[0040] S2-07: Sub-agent A n Improve its retrieval control strategy according to the following formula:
[0041] ;
[0042] S2-08: Let r←r + 1, and return to S2-03;
[0043] S2-09: Output the optimal retrieval control strategy of sub-agent A n ; .
[0044] Technical solutions that include some steps of the above technical solutions provided by the present invention are also within the scope of disclosure of the present invention.
[0045] In the present invention, the retrieval action of the P sub-agent groups in the corresponding modal layer of the heterogeneous knowledge graph according to the maintenance plan query request tags assigned by the main agent includes:
[0046] S2-10: Sub-agent A n According to the optimal retrieval control strategy In the corresponding state knowledge layer of the heterogeneous knowledge graph Perform a retrieval, Is the node set of the n-th modal layer of the heterogeneous knowledge graph, Is the edge set of the n-th modal layer of the heterogeneous knowledge graph. The retrieval process includes:
[0047] 2-10-1: Select the central node of the n-th modal layer of the heterogeneous knowledge graph As the retrieval starting point , Record it in the set And report the maintenance plan represented by the starting point as a retrieval report to the main agent A m , the main agent A m Evaluate the matching degree between the retrieval result of the sub-agent A n And the n-th modal of the maintenance plan query request tag;
[0048] S2-10-2: Sub-agent A n Access the next node according to the following formula:
[0049] ,
[0050] In the formula, Represents the distance;
[0051] S2-10-3: Record In the set , and record The maintenance plan represented by the starting point as a retrieval report to the main agent A m, The main agent A m Combine the retrieval results of all sub-agents A n |n = 1,…, N to form a combined retrieval result and calculate the matching degree between the combined retrieval result and the maintenance plan query request tag. Calculate the rewards of each sub-agent A n According to the matching degree between the retrieval result of the sub-agent A n And the n-th modal of the maintenance plan query request tag and the matching degree between the combined retrieval result and the maintenance plan query request tag;
[0052] Repeat steps S2-10-2 to S2-10-3 to obtain a series of rewards, sort all the rewards, and select the maintenance plan represented by the node with the highest reward as the retrieval result of this retrieval action.
[0053] In the present invention, according to the sub-agent A n The rewards of each sub-agent A n are calculated and the rewards are distributed according to the following formula:
[0054] ,
[0055] wherein, is the reward obtained by the main agent A n when performing a retrieval action at time t and transferring from the local retrieval result at time t to the local retrieval result at the next time t+1; m is the reward obtained by the main agent A when the sub-agent A k performs a retrieval action at time t and transferring from the local retrieval result at time t to the local retrieval result at the next time t+1; m is the contribution factor of the sub-agent A to the sub-agent A n to the sub-agent A k , and and .
[0056] In the present invention, when, each sub-agent of each sub-agent group optimizes its retrieval control strategy according to the reward values, retrieval actions, and retrieval results of all received sub-agents, including the following process:
[0057] S3-01: Group the N sub-agents A n |n=1,…,N into P groups of sub-agents, and each group of sub-agent groups includes one or more sub-agents.
[0058] S3-02: The main agent A m assigns the modalities of the maintenance plan query request tags to the P sub-agent groups C p |p=1,…,P. Each sub-agent A p of the sub-agent group C pq performs a retrieval action at time t in the corresponding modality layer of the heterogeneous knowledge graph according to the assigned modality of the maintenance plan query request tag. , Obtain the local retrieval result at time t , and send the local retrieval result to the main agent; q = 1, …, Q, where Q is a positive integer greater than or equal to 1. In the present invention, for heterogeneous knowledge graphs of different modal layers, different numbers of sub-agents can be set as needed;
[0059] S3-03: The main agent A m matches the Q local retrieval results of the sub-agent group C p with the p-th modality of the maintenance plan query request label respectively, filters out the local retrieval result with the best matching degree , and records the corresponding retrieval action and the corresponding sub-agent A ; p* ;
[0060] S3-04: The main agent A m calculates the state-action function of the sub-agent group C according to the following formula p :
[0061] ,
[0062] In the formula, π p* is the current retrieval control strategy of the sub-agent A p* ; is the local retrieval result of the sub-agent A p* at time t, is the retrieval action corresponding to the local retrieval result ; is the probability of transferring to the local retrieval result at the next time t + 1 when the retrieval action is executed at time t and the local retrieval result is ; ; is the reward obtained by the main agent A p* when the sub-agent A executes the retrieval action at time t; m ; is the state value function of the local retrieval result p* of the sub-agent A ; represents all the observation spaces of the sub-agent A p* ; r is the number of training times;
[0063] S3-05: The sub-agent A pq calculates the state value function value when the local retrieval result at time t is according to the following formula:
[0064]
[0065] In the formula, represents sub-agent A pq performs a retrieval action according to the retrieval control policy π pq to obtain a retrieval result ; ; is the improved retrieval control policy for sub-agent A pq ; represents sub-agent A pq all retrieval actions; is the gap factor between sub-agent A pq and the optimal sub-agent A p* ;
[0066] S3-06: Judge whether it is less than or equal to the threshold . If so, execute step S3-11. If not, execute step S3-07;
[0067] S3-07: , and calculate the state-action function of sub-agent A pq :
[0068] ,
[0069] In the formula, π pq is the current retrieval control policy of agent A pq ; is the optimal local retrieval result of sub-agent A pq at time t, is the retrieval action corresponding to the local retrieval result ; is the probability of transferring to the local retrieval result at the next time t+1 when the optimal local retrieval result is when the retrieval action is executed at time t; ; is the reward obtained by sub-agent A pq when the retrieval action is executed at time t to obtain the master agent A m ; is the state value function of the optimal local retrieval result ; represents all the observation spaces of sub-agent A pq ;
[0070] S3-08: Calculate the global state-action function according to the following formula:
[0071] ,
[0072] In the formula, ; , ;
[0073] S3-09: Improve sub-agent A according to the following formula pq Retrieval control strategy:
[0074] ;
[0075] S3-10: Let r ← r + 1, and return to S3-03;
[0076] S3-11: Output the optimal retrieval control strategy of sub-agent A pq .
[0077] Technical solutions that include some steps of the above technical solutions provided by the present disclosure are also within the scope of disclosure of the present invention.
[0078] The present invention integrates multi-modal data into a unified heterogeneous knowledge graph, divides the heterogeneous knowledge graph into multiple modal layers, enables an agent group composed of one or more sub-agents to perform retrieval in a heterogeneous knowledge graph layer of one modality respectively, and the main agent combines the retrieval results of the sub-agents in multiple groups. The combined retrieval results greatly improve the comprehensiveness. Moreover, when each sub-agent executes the retrieval action, it optimizes its own retrieval control strategy according to the retrieval actions, rewards, and retrieval results of all sub-agents, which greatly improves the retrieval speed.
[0079] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined. "Several" means one or more unless otherwise specifically defined.
[0080] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.
Claims
1. A maintenance plan generation system based on reinforcement learning, characterized in that: It includes a heterogeneous knowledge graph with P layers of modalities, a master agent A m and N subagents A n |n=1,…,N, said P is greater than or equal to 3, , where N subagents A n |n=1,…,N is divided into P sub-agent groups, each sub-agent group includes one or more sub-agents; the sub-agents of the P sub-agent groups are respectively based on the main agent A m The P modal information of the maintenance solution query request label sent by the agent performs a retrieval action in the corresponding modal layer of the heterogeneous knowledge graph and reports its retrieval results to the main agent A. m , master agent A m Responsible for receiving and processing the retrieval results of each subagent from P subagent groups, evaluating the matching degree between the retrieval results of each subagent in the P subagent groups and the P modes of the maintenance solution query request label, and the matching degree between the joint retrieval results of the optimal retrieval results of each subagent group and the maintenance solution query request label, calculating the reward value assigned to the subagent of each subagent group and broadcasting the reward value, retrieval action and retrieval result of each subagent; each subagent in each subagent group optimizes its retrieval control strategy according to the reward values, retrieval actions and retrieval results received from all subagents.
2. The maintenance plan generation system based on reinforcement learning according to claim 1 is characterized in that: When N=P, each subagent in each subagent group optimizes its retrieval control strategy according to the reward values, retrieval actions and retrieval results received from all subagents, including the following steps: Subagent A is calculated according to the following formula n The state-action function: , In the formula, π n Agent A n Current retrieval control strategy; Is Subagent A n The local search results at time t, Is Subagent A n The retrieval action at time t; For subagent A n Perform a retrieval action at time t When the local search results at time t Transfer to the local search results at the next time t+1 probability; For subagent A n Perform a search action at time t Get the main agent A m Rewards To adopt the retrieval control strategy π n Local search results obtained by searching The state value function of Represents subagent A n All observation spaces; For subagent A n Improved retrieval control strategy; r is the number of training times; The global state-action function is calculated according to the following formula: , In the formula, , , , Subagent A n Improve its retrieval control strategy according to the following formula: , where A n Refers to Subagent A n The collection of all actions.
3. The maintenance plan generation system based on reinforcement learning according to claim 1, characterized in that: When each subagent in each subagent group optimizes its retrieval control strategy according to the reward values, retrieval actions and retrieval results received from all subagents, the process includes the following: N subagents A n |n=1,…,N groups, divided into P groups of sub-agents; Master Agent A m Give P sub-agent groups C p |p=1,…,P assigns the maintenance plan query request label mode, sub-agent group C p Each subagent A pq According to the modality of the assigned maintenance solution query request label, a retrieval action is performed in the corresponding modality layer of the heterogeneous knowledge graph at time t , get the local search results at time t ; q=1,…,Q, Q is a positive integer greater than or equal to 1; The global state-action function is calculated according to the following formula: , In the formula, , ; π pq Agent A pq Current control strategies; Is Subagent A p* Local search results at time t; Improve subagent A according to the following formula pq Retrieval control strategy: , In the formula, A pq Refers to Subagent A pq The collection of all actions.
4. A maintenance plan generation method based on reinforcement learning, characterized in that: The steps include: Step 1: Set up N subagents A n |n=1,…,N is divided into P sub-agent groups; Step 2: Each sub-agent group’s sub-agents are assigned to the master agent A. m The P modal information of the maintenance solution query request label sent by the agent performs a retrieval action in the corresponding modal layer of the heterogeneous knowledge graph and reports its retrieval results to the main agent A. m ; Step 3: Master Agent A m Responsible for receiving and processing the search results from each subagent of the P subagent groups, evaluating the matching degree between the search results of each subagent of the P subagent groups and the P modes of the maintenance solution query request label, and the matching degree between the joint search results of the optimal search results of each subagent group and the maintenance solution query request label, calculating the reward value assigned to each subagent and broadcasting the reward value, search action and search result of each subagent; Step 4: Each subagent in each subagent group optimizes its retrieval control strategy according to the reward values, retrieval actions and retrieval results received from all subagents.
5. The maintenance plan generation method based on reinforcement learning according to claim 4 is characterized in that: The construction of heterogeneous knowledge graph includes the following steps: S1-01: construct knowledge graphs for M different fields respectively, and use a GCN-based self-training method to obtain the semantic vector of the label of each knowledge node of the knowledge graph of each field, where M is a positive integer greater than or equal to 3; S1-02: Associate and merge knowledge graphs from M domains to create a heterogeneous knowledge graph; S1-03: Layer the heterogeneous knowledge graph by modality into P layers.
6. The maintenance plan generation method based on reinforcement learning according to claim 4 is characterized in that: When N=P, each subagent in each subagent group optimizes its retrieval control strategy according to the reward values, retrieval actions and retrieval results received from all subagents, including the following process: Subagent A is calculated according to the following formula n The state-action function: , In the formula, π n Is Subagent A n Current control strategies; Is Subagent A n The local search results at time t, Is Subagent A n The retrieval action at time t; For subagent A n Perform a retrieval action at time t When , the local search result at time t is Transfer to the local search results at the next time t+1 probability; For subagent A n Perform a search action at time t Get the main agent A m Rewards To adopt the retrieval control strategy π n Local search results obtained by searching The state value function of Represents subagent A n All observation spaces; For subagent A n Improved retrieval control strategy; r is the number of training times; The global state-action function is calculated according to the following formula: , In the formula, , ; Subagent A n Improve its retrieval control strategy according to the following formula: , where A n Refers to the total action space of the subagent.
7. The maintenance plan generation method based on reinforcement learning according to claim 6 is characterized in that: The P sub-agent groups perform retrieval actions in the corresponding modal layer of the heterogeneous knowledge graph according to the modality of the maintenance plan query request label assigned by the master agent, including: Subagent A n According to the improved retrieval control strategy In the corresponding modality layer of heterogeneous knowledge graph Search in is the node set of the nth modal layer of the heterogeneous knowledge graph, is the edge set of the nth modal layer of the heterogeneous knowledge graph. The retrieval process includes: 2-10-1: Select the central node of the nth modal layer of the heterogeneous knowledge graph As the starting point of J search , Record collection In the process, the maintenance plan of the starting point representative is reported to the main agent A as a search report. m , master agent A m Evaluate Subagent A n The matching degree between the retrieval result and the nth mode of the maintenance solution query request label; S2-10-2: Subagent A n Access the next node according to the following formula: , In the formula, Indicates distance; S2-10-3: Record collection , and The maintenance plan of the starting point representative is reported to the main agent A as a search report m, Master Agent A m Unite all subagents A n |n=1,…,N The search results form a joint search result and calculate the matching degree between the joint search result and the maintenance solution query request label. n The matching degree between the search result and the maintenance solution query request label nth mode and the matching degree between the joint search result and the maintenance solution query request label are calculated for each sub-agent A. n Rewards; Repeat steps S2-10-2 to S2-10-3 to obtain a series of rewards, sort all rewards, and select the maintenance plan represented by the node with the highest reward as the retrieval result of this retrieval action.
8. The maintenance plan generation method based on reinforcement learning according to claim 7 is characterized in that: According to Subagent A n The matching degree between the search result and the maintenance solution query request label nth mode and the matching degree between the joint search result and the maintenance solution query request label are calculated for each sub-agent A. n The rewards include the following distribution of rewards: , In the formula, For subagent A n Perform a retrieval action at time t When the local search results at time t Transfer to the local search results at the next time t+1 Get the main agent A m Rewards; For subagent A k Perform a search action at time t , based on the local search results at time t Transfer to the local search results at the next time t+1 Get the main agent A m Rewards, For subagent A n For Subagent A k The contribution factor of and .
9. The maintenance plan generation method based on reinforcement learning according to claim 4 is characterized in that: When each subagent in each subagent group optimizes its retrieval control strategy according to the reward values, retrieval actions and retrieval results received from all subagents, the process includes the following: N subagents A n |n=1,…,N groups, divided into P groups of sub-agents; Master Agent A m Give P sub-agent groups C p |p=1,…,P assigns the maintenance plan query request label mode, sub-agent group C p Each subagent A pq According to the modality of the assigned maintenance solution query request label, a retrieval action is performed in the corresponding modality layer of the heterogeneous knowledge graph at time t , get the local search results at time t ; q=1,…,Q, Q is a positive integer greater than or equal to 1; The global state-action function is calculated according to the following formula: , In the formula, ; , , π pq Is Subagent A pq Current control strategies; Is Subagent A p* Local search results at time t; Improve subagent A according to the following formula pq Retrieval control strategy: , where A pq Refers to Subagent A pq The collection of all actions.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the maintenance plan generation method based on reinforcement learning according to any one of claims 4 to 9 is implemented.
Citation Information
Patent Citations
Maintenance scheme generation system and method based on large language model, and program product
CN118798859A
Recommendation method and system based on multi-level comparative learning and multi-modal knowledge graph
CN116091152A
Dynamic hierarchical multi-agent control method based on reinforcement learning
CN119337962A