Multi-agent cooperative vehicle chassis maintenance method, device, equipment and medium
Through a multi-agent collaborative vehicle chassis maintenance method, the target maintenance strategy is determined using local and global evaluation values, which solves the problem of low maintenance efficiency of special vehicle chassis systems and achieves overall efficiency and reliability improvement of the system.
Patent Information
- Application Number
- CN202411753218.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-12-02
AI Technical Summary
In the existing technology, the maintenance method of special vehicle chassis system assumes that the components are independent, resulting in low maintenance efficiency and difficulty in determining a reasonable maintenance strategy.
A multi-agent collaborative vehicle chassis maintenance method is adopted. By obtaining the local collaborative maintenance evaluation value and global maintenance evaluation value of the agent system, the target maintenance strategy for each vehicle chassis component is determined, and collaborative reinforcement learning is used for collaborative decision-making.
It improves the overall maintenance efficiency of special vehicle chassis systems, enhances the reliability and flexibility of the system, and achieves global optimality.
Smart Images

Figure CN119941216B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of chassis system maintenance, and in particular, to a multi-agent collaborative vehicle chassis maintenance method, device, equipment and medium. BACKGROUND
[0002] Maintenance optimization of special vehicle chassis systems is a hot research topic in the field of reliability. Maintenance theory generally believes that there are various correlations between these components, making it difficult to determine a reasonable maintenance strategy. Research on maintenance optimization of special vehicle chassis systems can help enterprises save costs and improve system reliability.
[0003] In related technologies, common chassis system maintenance optimization methods are mainly implemented based on fixed threshold strategies. In the implementation process, a set of preventive maintenance thresholds is usually set to simplify the maintenance optimization problem. For example, ant colony algorithm and simulated annealing algorithm are used to optimize the maintenance thresholds of different components in the system.
[0004] However, in existing methods, maintenance is performed by assuming that each component is independent of the others, resulting in low overall system maintenance efficiency. SUMMARY
[0005] The embodiments described herein provide a multi-agent collaborative vehicle chassis maintenance method, device, equipment and medium, which overcome the above problems.
[0006] In a first aspect, according to the content of the present disclosure, a multi-agent collaborative vehicle chassis maintenance method is provided, comprising:
[0007] An agent composition system corresponding to a special vehicle chassis system is obtained, the agent composition system being used to describe an intelligent running system composed of a plurality of vehicle chassis components, each vehicle chassis component corresponding to a first agent, and each two first agents being connected by a second agent.
[0008] In the current running state of the agent composition system, a local collaborative maintenance evaluation value corresponding to each first agent is obtained, the local collaborative maintenance evaluation value corresponding to each first agent being a system reward value obtained after each first agent executes a predetermined maintenance strategy on the corresponding vehicle chassis component.
[0009] Based on the local collaborative maintenance evaluation value corresponding to each first agent, a global maintenance evaluation value corresponding to the agent composition system in the current running state is determined.
[0010] determine a target maintenance strategy corresponding to each vehicle chassis component based on the global maintenance evaluation value corresponding to the current running state of the agent composition system, so that each first agent performs component maintenance on the corresponding vehicle chassis component based on the target maintenance strategy.
[0011] In a second aspect, the present disclosure provides a multi-agent collaborative vehicle chassis maintenance device, comprising:
[0012] A first acquisition module is configured to acquire an agent composition system corresponding to a special vehicle chassis system, the agent composition system being configured to describe an intelligent running system composed of a plurality of vehicle chassis components, each vehicle chassis component corresponding to a first agent, and each two first agents being connected by a second agent.
[0013] A second acquisition module is configured to acquire a local collaborative maintenance evaluation value corresponding to each first agent when the special vehicle is in a current running state of the agent composition system, the local collaborative maintenance evaluation value corresponding to each first agent being a system reward value obtained after each first agent executes a preset maintenance strategy on the corresponding vehicle chassis component.
[0014] A first determination module is configured to determine a global maintenance evaluation value corresponding to the current running state of the agent composition system based on the local collaborative maintenance evaluation value corresponding to each first agent.
[0015] A second determination module is configured to determine a target maintenance strategy corresponding to each vehicle chassis component based on the global maintenance evaluation value corresponding to the current running state of the agent composition system, so that each first agent performs component maintenance on the corresponding vehicle chassis component based on the target maintenance strategy.
[0016] In a third aspect, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the multi-agent collaborative vehicle chassis maintenance method in any one of the above embodiments.
[0017] In a fourth aspect, a computer readable storage medium is provided, the computer readable storage medium storing a computer program, and the computer program being executed by a processor to implement the steps of the multi-agent collaborative vehicle chassis maintenance method in any one of the above embodiments.
[0018] The multi-agent cooperative vehicle chassis maintenance method provided by the embodiment of the application obtains an agent composition system corresponding to a special vehicle chassis system, the agent composition system is used for describing an intelligent operation system composed of a plurality of vehicle chassis components, each vehicle chassis component corresponds to a first agent, and each two first agents are connected in association by a second agent; in a current operation state of the special vehicle in the agent composition system, a local cooperative maintenance evaluation value corresponding to each first agent is obtained, the local cooperative maintenance evaluation value corresponding to the first agent is a system reward value obtained after each first agent respectively executes a preset maintenance strategy on the corresponding vehicle chassis component; based on the local cooperative maintenance evaluation value corresponding to each first agent, a global maintenance evaluation value of the agent composition system corresponding to the current operation state is determined; based on the global maintenance evaluation value of the agent composition system corresponding to the current operation state, a target maintenance strategy corresponding to each vehicle chassis component is determined, so that each first agent performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy. In this way, by connecting a plurality of vehicle chassis components with each other, the cooperative maintenance processing is performed according to the agent corresponding to each vehicle chassis component, so as to facilitate effectively improving the overall maintenance efficiency of the vehicle chassis system.
[0019] The above description is only a summary of the technical solutions of the embodiments of the application. In order to more clearly understand the technical means of the embodiments of the application, the embodiments of the application can be implemented according to the content of the description, and in order to make the above and other purposes, characteristics and advantages of the embodiments of the application more obvious and easy to understand, the following specific embodiments of the application are described. BRIEF DESCRIPTION OF DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. It should be noted that the drawings described below only relate to some embodiments of the present disclosure, but not limit the present disclosure, wherein:
[0021] Figure 1 is a flowchart of a multi-agent cooperative vehicle chassis maintenance method provided by the present disclosure.
[0022] Figure 2A is a structural schematic diagram of an agent composition system provided by the present disclosure.
[0023] Figure 2B is a structural schematic diagram of a series production system with buffer inventory provided by the present disclosure.
[0024] Figure 3 is a structural schematic diagram of a multi-agent cooperative vehicle chassis maintenance device provided by the present disclosure.
[0025] Figure 4 is a structural schematic diagram of a computer device provided by the present disclosure.
[0026] It should be noted that the elements in the drawings are schematic and not drawn to scale. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be described clearly and completely below in conjunction with the drawings. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative effort also belong to the scope of protection of the present disclosure.
[0028] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this present subject matter belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the specification and relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. As used herein, the statement that two or more parts are "connected" or "coupled" together refer to an indirect or direct connection or coupling.
[0029] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase "in an embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. A person of ordinary skill in the art will readily recognize from the disclosure herein, given the total volume of this application that the embodiments described herein can be combined with one another in various ways.
[0030] The term "and / or", merely an associative relationship of the associated objects, means that there can be three relationships, for example, A and / or B, can represent: there is A, there are A and B, and there is B. In addition, the character " / " herein generally represents that the front and rear associated objects are a "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or part of a component) from another component (or another part of a component).
[0031] In the description of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more (including two), and similarly, "a plurality of groups" means two or more groups (including two groups).
[0032] The embodiment solves the complex state and action space of the chassis system of the special vehicle by introducing coordinated reinforcement learning, and realizes coordinated decision between components, thereby breaking through the limitations of traditional maintenance strategies and having high flexibility, expansibility and global optimality.
[0033] In order to enable personnel in the technical field to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings.
[0034] Figure 1 is a flowchart of a vehicle chassis maintenance method provided by the embodiment of the present disclosure, as shown in Figure 1 The specific process of the vehicle chassis maintenance method of the multi-agent collaboration includes:
[0035] S110, acquiring an agent composition system corresponding to a chassis system of a special vehicle, the agent composition system being used for describing an intelligent running system composed of a plurality of vehicle chassis components, and each vehicle chassis component corresponding to a first agent.
[0036] The chassis system of the special vehicle can be composed of a drive train, a running train and a steering train. The drive train is mainly composed of a differential, a main reducer, a universal transmission device, a transmission and a clutch and other vehicle chassis components. The running train is mainly composed of a suspension, an axle and a wheel and a frame and other vehicle chassis components. The steering system is composed of a steering column, a steering shaft and a steering wheel and other vehicle chassis components.
[0037] Each two first agents are associated and connected by a second agent, so as to connect two first agents by one second agent.
[0038] Each vehicle chassis component can be equivalent to a first agent, and each vehicle chassis component corresponds to an independent first agent, and the first agent is responsible for state monitoring and maintenance decision of the vehicle chassis component. For example, a component (such as a differential of a drive system) of a chassis of a special vehicle can have a dedicated first agent to process its running state and maintenance decision.
[0039] Each two vehicle chassis components are connected by a second agent, and the second agent can be a node connecting the components in the graph structure, and is used for processing the coordinated decision between the two vehicle chassis components. In coordinated reinforcement learning (CRL), a second agent can need to make joint decision on the states of two or more components, and consider the state information of the two components to select a coordinated action.
[0040] On the basis of traditional single-agent reinforcement learning, a coordination graph is introduced, and each first agent interacts and shares information through a coordination mechanism, thereby overcoming the problem that a single agent cannot handle large-scale state and action spaces.
[0041] A structural diagram of the agent composition system is shown in FIG. 1. Figure 2A Figure 2A A coordination graph of a series system composed of 12 vehicle chassis components is shown in FIG. 2. Each square represents a vehicle chassis component, corresponding to a first agent, and a circle represents a second agent, used to connect two adjacent vehicle chassis components.
[0042] In S120, a local coordination maintenance evaluation value corresponding to each first agent is obtained under the current running state of the special vehicle in the agent composition system.
[0043] The local coordination maintenance evaluation value corresponding to the first agent is a system reward value obtained after each first agent executes a preset maintenance strategy for the corresponding vehicle chassis component. That is, the impact incentive of the entire vehicle chassis system after the first agent corresponding to each vehicle chassis component executes a preset maintenance strategy for the component.
[0044] In some embodiments, each vehicle chassis component has different preset maintenance strategies, and the preset maintenance strategy is determined by the second agent associated with the first agent corresponding to the vehicle chassis component based on the component state of the vehicle chassis component and the component state of the adjacent chassis component; the preset maintenance strategy includes at least one candidate maintenance strategy.
[0045] The component state can include a health state, an operating state, a maintenance requirement, and an inventory state. The health state can be used to reflect the degree of degradation of the component, such as discrete states of "brand new", "mild degradation", "moderate degradation", "severe degradation", "failure", and "invalidation". The operating state can include "normal operation", "maintenance", "failure stop", "starvation state" (lack of resources to support operation), and "blocking state" (unable to operate normally due to insufficient downstream resources). The maintenance requirement is used to describe the type of maintenance required by the component, such as whether the preventive maintenance threshold has been reached or after-service maintenance is required. The inventory state can be, for example, the storage amount of the component (such as the storage amount of the buffer inventory).
[0046] The second agent connecting two first agents can determine the corresponding preset maintenance strategy of the corresponding vehicle chassis component by obtaining the component state of the component from the first agent. The preset maintenance strategy can include preventive maintenance, after-maintenance, and continue running. The preventive maintenance is to maintain the component before it fails completely to prolong its service life; the after-maintenance is to maintain the component after it fails to restore its function; and the continue running is to keep the component in the current state without any maintenance operation. In addition, a new preset maintenance strategy can be dynamically adjusted according to the component state and system demand, such as degraded use and adjustment of component load.
[0047] For example, if the second agent determines that the component state of a vehicle chassis component is "brand new" or "slight degradation", the second agent can determine that the corresponding preset maintenance strategy of the component is "continue running"; if the second agent determines that the component state of a vehicle chassis component is "moderate degradation" or "severe degradation" but has not failed, the second agent can determine that the corresponding preset maintenance strategy of the component is "preventive maintenance" to prevent failure; if the second agent determines that the component state of a vehicle chassis component is "failure", the second agent can determine that the corresponding preset maintenance strategy of the component is "after-maintenance"; and if the second agent determines that the component state of a vehicle chassis component is "blocked state" or "starvation state", the second agent can determine that the corresponding preset maintenance strategy of the component is "continue running" or "preventive maintenance" to avoid system downtime or prolong the life of the component.
[0048] It should be noted that in the collaborative learning environment, the second agent can also consider the component state of the adjacent component. For example, the component state of a certain component can affect the decision of the adjacent component (for example, if a key component is close to failure, preventive maintenance can be preferentially selected on the adjacent component).
[0049] The second agent can select the preset maintenance strategy of the two connected components according to the component states of the two connected components. The second agent connecting component m and component n is represented as Ag n,m , the state and action of Ag n,m are represented as X n,m =[X n ,X m ] and A n,m =[A n ,A m ] respectively.
[0050] The preset maintenance strategy of a component can be made by two or more agents, and the component states of the component can be shared between the agents. The best maintenance strategy is determined by combining the decision results of the strategies of different agents. Through the cooperation of multiple agents, it is ensured that the decision for a single component not only optimizes its own operating state, but also takes into account the relevance to other components and the overall benefit of the system. Figure 2A The parameter table of the middle agent is shown in Table 1.
[0051] Table 1 Agent parameter table
[0052]
[0053]
[0054] In the current running state of the special vehicle in the agent composition system, the local collaborative maintenance evaluation value corresponding to each first agent is obtained, including:
[0055] In the current running state of the special vehicle in the agent composition system, each first agent is controlled to execute each candidate maintenance strategy in turn for the corresponding vehicle chassis component, system running data obtained after each first agent executes all the corresponding candidate maintenance strategies is obtained, and the local collaborative maintenance evaluation value corresponding to each first agent is determined based on the system running data obtained after each first agent executes all the corresponding candidate maintenance strategies.
[0056] The system running data can be used to describe the current running parameters of the vehicle chassis system, such as the steering shaft angle and the steering wheel angle.
[0057] Each first agent is controlled to execute each candidate maintenance strategy in turn for the corresponding vehicle chassis component, and system running data obtained after each first agent executes all the corresponding candidate maintenance strategies is obtained, for example.
[0058] The agent composition system includes three first agents, namely agent 11, agent 12 and agent 13, the preset maintenance strategy of the vehicle chassis component corresponding to the agent 11 is strategy A1, the preset maintenance strategy of the vehicle chassis component corresponding to the agent 12 is strategy A2 and A3, and the preset maintenance strategy of the vehicle chassis component corresponding to the agent 13 is strategy A4, each first agent is controlled to execute each candidate maintenance strategy in turn for the corresponding vehicle chassis component, the execution sequence of the strategy can be obtained as {A1-A2-A4} and {A1-A3-A4}, and the system running data obtained after each first agent executes all the corresponding candidate maintenance strategies can include the system running data corresponding to the strategy {A1-A2-A4} and the system running data corresponding to the strategy {A1-A3-A4}.
[0059] In some embodiments, based on the system running data obtained after each first agent executes all the corresponding candidate maintenance strategies, the local collaborative maintenance evaluation value corresponding to each first agent is determined, including:
[0060] determine a system operation risk value of the overall operation system of the special vehicle based on the system operation data obtained after each first agent executes all the candidate maintenance strategies corresponding thereto; and determine a local collaborative maintenance evaluation value corresponding to each first agent based on the system operation risk value of the overall operation system of the special vehicle.
[0061] The system operation risk value of the overall operation system of the special vehicle can be used to describe a corresponding component-related cost, and the component-related cost can include a maintenance cost, an allocated detection cost and an allocated downtime cost. The maintenance cost can be a preventive maintenance cost, and the downtime cost can be a cost paid per unit time when the system is in a fault state.
[0062] In some embodiments, the local collaborative maintenance evaluation value corresponding to each first agent is determined based on the system operation risk value of the overall operation system of the special vehicle, including:
[0063] The local collaborative maintenance evaluation value corresponding to each first agent is determined based on the system operation risk value of the overall operation system of the special vehicle, and a collaborative influence value between each first agent and an adjacent agent.
[0064] For example, the local reward value C n,m As shown in the following formula (1).
[0065] C n,m = C n / |C(n)|+C m / |C(m)| (1)
[0066] In formula (1), c n represents a cost related to component n, i.e., a component-related cost corresponding to component n; c m represents a cost related to component m, i.e., a component-related cost corresponding to component m; represents the number of second agents connected to component n; represents the number of second agents connected to component m.
[0067] c n As shown in the following formula (2).
[0068]
[0069] In formula (2), C PR,n represents a preventive maintenance cost of component n, A nrepresents the execution action of component n, PR represents that the candidate maintenance strategy is "preventive maintenance", I(A n represents that the current execution action corresponds to the strategy of "preventive maintenance"; C INS represents the detection cost of component n; X represents the current state of the system, X' represents the next state of the system, C F represents the cost paid per unit time when the system is in a fault state; C ST represents the setting cost, N represents the total number of components of the series system, DN represents that the candidate maintenance strategy is "no maintenance", I(A n represents that the current execution action corresponds to the strategy of "no maintenance".
[0070] The global reward value corresponding to the first agent can be the system operation risk value of the overall operation system of the special vehicle, and the sum of the collaborative influence values between the first agent and the adjacent agents.
[0071] The local reward value corresponding to the first agent is set based on the immediate influence of the maintenance decision of the agent on the system. For example, when the first agent performs "preventive maintenance", if the occurrence of potential failure is successfully avoided, a corresponding reward will be obtained; and after the "after-maintenance" is performed to restore the normal operation of the component, a certain reward will also be obtained. At the same time, the cost of the maintenance process and the downtime loss caused by the maintenance and the like will be regarded as cost, and thus the corresponding scores will be deducted. These reward and deduction mechanisms directly reflect the good and bad of the first agent's own decision, so that the first agent can adjust its own decision strategy according to the immediate feedback. In the collaborative decision of the agent, the local reward influences the evaluation of the first agent on different actions, and when each first agent calculates the local reward value, the local reward brought by the own decision will be considered in the range, so as to more accurately evaluate the immediate effect of taking a certain action in the current state.
[0072] The first agent can also obtain an additional reward, i.e. a collaborative influence value, according to the collaborative effect between the adjacent components (or adjacent agents). The collaborative reward considers the overall operation state of the global system, for example, the health state of a component can affect the working efficiency of the adjacent components, and a successful collaborative maintenance decision (such as keeping the system stable operation under the cooperation of multiple key components) can obtain a reward for the improvement of the global performance, and such a reward encourages the agent to consider the reliability and long-term benefit of the overall system when making a decision.
[0073] The reward function is defined as the benefit of the system in a unit of time when the system state is X and the action is A, as shown in the following formula (3).
[0074] R(X,A)=R r (X,A)-C P (X,A)-C M (X,A)-CB (X,A) (3)
[0075] In formula (3), R r (X,A) is the revenue of the system producing products, which can be expressed as shown in formula (4).
[0076] R r (X,A) = v N,norm I(A N = DN and v N = v N,norm ) (4)
[0077] C P (X,A) is the total operating cost of each production unit, which can be expressed as shown in formula (5).
[0078]
[0079] C M (X,A) is the total maintenance cost of each production unit, which can be expressed as shown in formula (6).
[0080]
[0081] In formula (6), ω is true, and the function I(ω) = 1; ω is false, and the function I(ω) = 0.
[0082] C B (X,A) is the total holding cost of buffer inventory, which can be expressed as shown in formula (7).
[0083]
[0084] It is assumed that the degradation of each production unit is independent of each other, so the system transition probability can be simplified as shown in formula (8).
[0085]
[0086] When A n = DN, according to whether M n is in a runnable state, the state transition probability is shown in formula (9) or formula (10).
[0087]
[0088] When A n = PM, the state transition probability of M n is shown in formula (11) or formula (12).
[0089] P(S n ′ = 1 | S n = s n ) = pn,PM (11)
[0090] P(S n ′=D n |S n =s n )=1-p n,PM (12)
[0091] When A n =CM, the state transition probability of M n is shown in formula (13) or formula (14).
[0092] P(S′ n =1 S n =S n )=P n,CM (13)
[0093] P(S′ n =D n |S n =s n )=1-p n,CM (14)
[0094] S130, determining a global maintenance evaluation value corresponding to the current running state of the system composed of the agents based on the local collaborative maintenance evaluation values corresponding to each first agent.
[0095] Wherein, the global maintenance evaluation value corresponding to the current running state of the system composed of the agents can be the sum of the local collaborative maintenance evaluation values corresponding to the plurality of first agents.
[0096] For example, the local Q function of the agent Ag n,m is represented as Q n,m (X n,m ,A n,m ), and the global Q function is the sum of the local Q functions. The local Q function defines the "local Q value" of each first agent under the current state-action pair, which can be used to evaluate the value of taking a certain action in the current state, i.e. the local collaborative maintenance evaluation value.
[0097] The global Q function integrates the local Q values of multiple agents to form the global Q function of the system, to measure the impact of different decisions on the system as a whole, as shown in formula (15) below.
[0098]
[0099] In formula (15), represents the set of component pairs connected by the agents in the collaborative graph. The local Q function can be updated according to the sampled quadruple , and the update formula is shown in formula (16).
[0100]
[0101] In formula (16), η n,m represents the learning rate of the agent Ag n,m , and γ represents the discount factor.
[0102] The local optimal action is part of the global optimal action A'* and A′ * is determined by all agents together. Since the action space increases exponentially with the number of components, the selection of the global optimal action is also affected by the curse of dimensionality, that is, in a high-dimensional state or action space, the computational complexity and the amount of data required increase exponentially with the increase of dimension, resulting in the phenomenon that algorithms are difficult to handle large-scale problems. To solve this problem, a variable elimination algorithm is used to find the global optimal action A′ * .
[0103] This embodiment takes a series system composed of three components as an example to illustrate the implementation process of the variable elimination method. The global Q function is decomposed as shown in formula (17).
[0104] Q(X,A)=Q 1,2 (X 1,2 ,A1,A2)+Q 2,3 (X 2,3 ,A2,A3)+Q 3,1 (X 3,1 ,A3,A1) (17)
[0105] The minimum value can be converted to formula (19) shown by introducing formula (18).
[0106]
[0107] So far, the action variable A1 is eliminated. Similarly, the action variable A2 can be eliminated by introducing formula (20).
[0108]
[0109] Because A3 is the last remaining variable in the global Q function, the optimal value of A3 is shown in formula (21).
[0110]
[0111] Subsequently, the optimal values of variables A2 and A1 can be determined recursively as shown in formula (22) and formula (23).
[0112]
[0113] Finally, the global optimal action is obtained The order of variable elimination is variable, and the order will affect the computational complexity in the action selection process.
[0114] This embodiment uses the Q-value maximization principle, and the agent selects the optimal action (maintenance decision) at each time, realizes the dynamic update of the strategy. At each decision time, the agent selects the optimal action based on the Q-value maximization principle. For example, in a series system composed of 12 components, each agent calculates the local Q-value according to the state of two connected components, and interacts with other agents through the coordination graph mechanism, and then selects the action that can maximize the overall benefit of the system based on the global Q-value. Action selection: including "preventive maintenance", "after maintenance", "continue running", etc., the agent determines which action can bring the maximum long-term benefit in the current state according to the calculated value, and makes a decision. By decomposing the global reward into the local reward of each agent, while ensuring the unity of the global goal (such as system stability and high yield) and local decision.
[0115] The global reward takes into account the synergy effect between agents and the overall operation state of the system. For example, when multiple key components maintain the stable operation of the system through collaborative maintenance decisions, all agents participating in the collaboration can obtain a reward for improving the global performance. This reward mechanism encourages agents to consider the reliability and long-term benefits of the overall system when making decisions, and avoids pursuing only local maximum benefits at the expense of the overall performance of the system.
[0116] In some embodiments, based on the local collaborative maintenance evaluation value corresponding to each first agent, the global maintenance evaluation value of the system composed of agents corresponding to the current operating state is determined, including:
[0117] Obtaining the system set weight corresponding to each first agent; based on the system set weight corresponding to each first agent and the local collaborative maintenance evaluation value corresponding to each first agent, determining the global maintenance evaluation value of the system composed of agents corresponding to the current operating state.
[0118] Wherein, different system set weights can be set for each first agent, and then the weight summation processing is performed, so as to effectively determine the global maintenance evaluation value of the system composed of agents corresponding to the current operating state.
[0119] S140, based on the global maintenance evaluation value of the system composed of agents corresponding to the current operating state, determining the target maintenance strategy corresponding to each vehicle chassis component.
[0120] The target maintenance strategy corresponding to each vehicle chassis component is determined based on the global maintenance evaluation value of the agent composition system corresponding to the current running state, so that each first agent performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
[0121] In some embodiments, determining the target maintenance strategy corresponding to each vehicle chassis component based on the global maintenance evaluation value of the agent composition system corresponding to the current running state comprises:
[0122] When the global maintenance evaluation value of the agent composition system corresponding to the current running state is obtained, the candidate maintenance strategy executed by each first agent; the candidate maintenance strategy executed by the first agent is determined as the target maintenance strategy corresponding to the corresponding vehicle chassis component.
[0123] In combination with the above example, if it is determined that the candidate maintenance strategy executed by agent 11 is strategy A1, the candidate maintenance strategy executed by agent 12 is strategy A2, and the candidate maintenance strategy executed by agent 13 is strategy A4 when the global maintenance evaluation value of the agent composition system corresponding to the current running state is obtained, then the target maintenance strategy corresponding to the vehicle chassis component of agent 11 is strategy A1, the target maintenance strategy corresponding to the vehicle chassis component of agent 12 is strategy A2, and the target maintenance strategy corresponding to the vehicle chassis component of agent 13 is strategy A4. Thus, the target maintenance strategy suitable for global maximum benefit of each vehicle chassis component is effectively determined.
[0124] In this embodiment, the agent composition system corresponding to the special vehicle chassis system is obtained, the agent composition system is used to describe an intelligent running system composed of a plurality of vehicle chassis components, each vehicle chassis component corresponds to a first agent, and each two first agents are connected by a second agent. In the current running state of the agent composition system of the special vehicle, the local collaborative maintenance evaluation value corresponding to each first agent is obtained, and the local collaborative maintenance evaluation value corresponding to the first agent is the system reward value obtained after each first agent executes a predetermined maintenance strategy on the corresponding vehicle chassis component. The global maintenance evaluation value of the agent composition system corresponding to the current running state is determined based on the local collaborative maintenance evaluation value corresponding to each first agent. The target maintenance strategy corresponding to each vehicle chassis component is determined based on the global maintenance evaluation value of the agent composition system corresponding to the current running state, so that each first agent performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy. In this way, by connecting a plurality of vehicle chassis components, collaborative maintenance processing is performed according to the agent corresponding to each vehicle chassis component, which facilitates effective improvement of the overall maintenance efficiency of the vehicle chassis system.
[0125] This embodiment constructs the basic framework for the entire special vehicle chassis system by initializing the agent composition system, assigning an agent to each production unit and buffer inventory in the system, and defining the system's state space, action space, and reward function. This provides information such as the state space, action space, and reward function for the subsequent reinforcement learning process. The initialized system state space, action space, and reward function provide basic data for subsequent agent decision-making, policy updates, and training. In particular, these definitions are used for input and decision support in the "Agent Collaborative Decision-Making" and "Policy Update and Iteration" stages.
[0126] The construction of the system (including the design of production units and buffer stocks) reduces the coupling between production units through the introduction of buffer stocks, allowing the system to continue to partially operate when a unit is repaired or fails. The inventory mechanism avoids system downtime caused by single point failure, thereby improving the overall reliability of the system.
[0127] Each production unit and buffer stock can be modeled through a Markov decision process, which clearly defines the system state (e.g., health status, inventory level) and action (such as repair or continue operation). Through state modeling, a standardized feedback mechanism can be provided to the intelligent agent, thereby achieving dynamic decision-making.
[0128] For example, the MDP (Markov decision process) model is used to define the system state, action, transition probability and immediate reward, and the maintenance problem of the entire system is transformed into a reinforcement learning problem, which is solved using a collaborative reinforcement learning algorithm.
[0129] Serial production system with buffer stock Figure 2B As shown in the figure, the squares represent production units, and the circles represent buffer stock. Buffer stock is used to temporarily store products from upstream production units before supplying them to downstream units. The addition of buffer stock also ensures that the production system does not suffer downtime due to maintenance or failure of a single production unit. This type of system is widely used in real-world production, such as tire production systems and tile manufacturing systems.
[0130] Figure 2B In, M n (n∈1,2,…,N) is the production unit, B n (n∈1,2,…,N-1) is the buffer stock. n The inventory at time t is denoted as K n (t)=0,1,...,N n,B , where N n,B For buffer stock The inventory capacity of each production unit is assumed to be a discrete-time, discrete-state Markov process.n The degradation amount can be expressed as a discrete state {1,2,…,D n}, where 1 represents a brand new state, D n Indicates a fault state. If the production unit M n The upstream buffer stock is 0, then M n In a hungry state; if M n The downstream buffer inventory reaches the upper limit, then M n In particular, M1 will not starve, M N Will not be blocked. n The productivity is expressed as v n+1 When production unit M n When in failure, starvation, obstruction or maintenance, v n =0; at other times, v n =v n,norm ; Assumption The inventory at the current moment is K n , then its inventory at the next moment is K′ n As shown in formula (24).
[0131] K n ′=K n +v n -v n+1 (twenty four)
[0132] Assume M n The degradation process is affected by its production rate v n The impact of M n When working normally, its state transfer matrix is P n,norm ; in M n When in a starvation or blocking state, its state transition matrix is P n,idle In this embodiment, only two maintenance activities are considered: preventive maintenance and post-maintenance. Both maintenance activities can restore the production unit to a brand new state. Assuming that the maintenance time follows a geometric distribution, in one unit of time, the production unit M n The completion probabilities of preventive maintenance and post-maintenance are p n,PM 、p n,CM , and p n,PM <p n,CM The corresponding preventive maintenance and post-maintenance costs are C n,PM and C n,CM , and C n,PM <C n,CM . M per unit time n The operating cost is C n,op Buffer stock The cost of holding a component per unit time is C n,HR r The objective of maintenance optimization is to maximize the average expected reward per unit of time.
[0133] The MDP model can be represented by a four-tuple <ψ,ζ,ξ,s>, where ψ represents the system state set, ζ represents the system action set, ξ represents the system transition probability model, and s represents the system immediate reward. The decision maker performs an action A∈ζ according to the current system state X∈ψ, obtains an immediate reward and changes to the next state x′ according to the state transition probability P(X′|X,A)∈ξ.
[0134] The embodiment adds a D n +1 state to the production unit to represent the preventive maintenance state of the production unit. Therefore, the state of the M n DP model can be represented as S n ∈{1,2,…,D n ,D n +1}, and the system state can be represented as X=[S1,S2,…,S N ,K1,K2,…,K N-1 ].
[0135] After initializing the system model, the state information of each production unit (such as the degradation state of the component, the storage amount of the buffer inventory, etc.) can be obtained in real time and input into the MDP model, and the next state of the system is output. The next state refers to the subsequent state set to which the system changes from the current state after the agent performs an action, which reflects the dynamic change of the system over time and is the key basis for the agent to make decisions and learn.
[0136] After the agent selects an action according to the current state, the next state generated is observed, and the immediate reward (such as the revenue of the system producing products, the operating cost of each production unit, the maintenance cost, the holding cost of the buffer inventory, etc.) is combined to evaluate the action. For example, if the agent performs a maintenance action, the next state of the system shows that the component failure rate is reduced and the production efficiency is improved, and a higher immediate reward is obtained, the agent will consider that the action is beneficial, and will be more inclined to choose the action in the future similar situation.
[0137] Each agent makes maintenance decisions according to the collaborative reinforcement learning algorithm and the reward shaping mechanism, combined with the state information of itself and adjacent components. The decision of the agent includes “preventive maintenance”, “after-the-fact maintenance” or “continue running” and the like.
[0138] The self-state information refers to the instant state of the specific component managed by the agent, which can reflect the wear, aging or damage degree of the component itself, the running state ("normal running", "maintenance", "fault shutdown", "starvation state", "blocking state", etc.), indicating the current working condition of the component, the maintenance requirement (whether the preventive maintenance threshold is reached or after-service maintenance is needed), used to determine whether the component needs to be maintained, and the inventory of the component (such as the storage capacity of the buffer inventory) for the system involving inventory. For example, in the drive train of the chassis system, the current state of a certain gear component can be "mild degradation" and in "normal running" state, but has approached the preventive maintenance threshold, which needs to be considered by the agent in decision-making.
[0139] The state information of the adjacent component refers to the state of the component directly connected in the system structure or closely related in function to the current component managed by the agent. In the chassis system of the special vehicle, there are physical connections and functional cooperation relationships between different components, such as the gears, shafts and other components in the drive train working together, and the state change of a component can affect the performance and running state of the adjacent component. For example, in the drive train, the wear degree (health state) of a certain driving gear can affect the stress condition of the driven gear meshing with it, and further affect the running state and life of the driven gear.
[0140] In some embodiments, the agents correspond to the vehicle chassis components included in the system, and the upstream production unit and the downstream production unit correspond to the buffer inventory for transferring the stored production products of the upstream production unit to the downstream production unit.
[0141] The method of the embodiment further includes:
[0142] If the buffer inventory does not meet the preset inventory, the maintenance level of the vehicle chassis component corresponding to the upstream production unit is improved; if the buffer inventory meets the preset inventory, the maintenance level of the vehicle chassis component corresponding to the downstream production unit is improved.
[0143] The buffer inventory serves as an intermediate link between the upstream and downstream production units, making the collaborative learning among the agents more efficient. The state (such as the inventory) of the buffer inventory can affect the maintenance priority of the adjacent production unit, and the state of each production unit and the buffer inventory in the system can change in real time, allowing dynamic adjustment of the maintenance strategy.
[0144] For example, when the buffer inventory is below a certain threshold, the agent may prioritize repairing the upstream production unit to restore supply; when the buffer inventory is above a certain threshold, the work of the downstream production unit is adjusted. The setting of the buffer inventory not only reduces the impact of downtime on revenue, but also optimizes maintenance decisions by diversifying risks. For example, when a production unit is being repaired, the buffer inventory can temporarily support system operation, reducing economic losses.
[0145] The embodiment uses a multi-agent reinforcement learning algorithm to optimize the maintenance strategy of a complex special vehicle chassis system, breaking through the limitations of traditional single-agent methods in handling large-scale systems, and significantly improving decision-making efficiency and strategy flexibility. Through the design of the collaborative graph mechanism, each agent can collaborate according to the state of the adjacent components, ensuring globally optimal maintenance decisions, rather than just local optima for each agent. The collaboration between agents enhances the overall maintenance optimization capability of the system. Through reward shaping technology, the global reward of the system is reasonably decomposed into the local reward of each agent, combining global goals with local decision optimization to improve learning efficiency and global optimality of the strategy. This mechanism ensures that each agent considers global impact when making local decisions, ensuring the maximization of the overall revenue of the system. Combining deep Q-learning with multi-agent collaborative learning solves the convergence problem of traditional Q-learning in large-scale state space. Through the introduction of a deep neural network, the system can efficiently approximate complex state-action value functions, handling large-scale state and action spaces to ensure efficient learning and decision-making processes. Agents can dynamically choose "preventive maintenance" or "after-the-fact maintenance" strategies based on the real-time state of the production unit and the overall operation of the system, ensuring system reliability and economy.
[0146] Figure 3 A structural schematic diagram of a multi-agent collaborative vehicle chassis maintenance device is provided for the embodiment, which can include:
[0147] The first acquisition module 310 is configured to acquire an agent composition system corresponding to the special vehicle chassis system, the agent composition system being used to describe an intelligent running system composed of a plurality of vehicle chassis components, each vehicle chassis component corresponding to a first agent, and each two first agents being connected by a second agent.
[0148] The second acquisition module 320 is configured to acquire a local collaborative maintenance evaluation value corresponding to each first agent when the special vehicle is in a current running state of the agent composition system, the local collaborative maintenance evaluation value corresponding to the first agent being a system reward value obtained after each first agent executes a preset maintenance strategy on the corresponding vehicle chassis component.
[0149] The first determining module 330 is configured to determine a global maintenance evaluation value of the agent group system corresponding to the current running state based on the local collaborative maintenance evaluation value of each first agent.
[0150] The second determining module 340 is configured to determine a target maintenance strategy corresponding to each vehicle chassis component based on the global maintenance evaluation value of the agent group system corresponding to the current running state, so that each first agent performs component maintenance on the corresponding vehicle chassis component based on the target maintenance strategy.
[0151] In this embodiment, optionally, each vehicle chassis component corresponds to a different preset maintenance strategy, the preset maintenance strategy being a second agent associated with the first agent corresponding to the vehicle chassis component and being determined based on the component state of the vehicle chassis component and the component state of an adjacent chassis component; the preset maintenance strategy includes at least one candidate maintenance strategy.
[0152] The second obtaining module 320 includes a control unit, an obtaining unit and a determining unit.
[0153] The control unit is configured to control each first agent to sequentially execute each candidate maintenance strategy for the corresponding vehicle chassis component under the current running state of the agent group system.
[0154] The obtaining unit is configured to obtain system running data obtained after each first agent executes all the corresponding candidate maintenance strategies.
[0155] The determining unit is configured to determine the local collaborative maintenance evaluation value of each first agent based on the system running data obtained after each first agent executes all the corresponding candidate maintenance strategies.
[0156] In this embodiment, optionally, the determining unit is specifically configured to:
[0157] determine a system running risk value of the overall running system of the special vehicle based on the system running data obtained after each first agent executes all the corresponding candidate maintenance strategies; and determine the local collaborative maintenance evaluation value of each first agent based on the system running risk value of the overall running system of the special vehicle.
[0158] In this embodiment, optionally, the determining unit is specifically configured to:
[0159] determine the local reward value corresponding to each first agent based on the system operation risk value of the overall operation system of the special vehicle; determine the global reward value corresponding to each first agent based on the system operation risk value of the overall operation system of the special vehicle, and the collaborative influence value between each first agent and the adjacent agent; and determine the local collaborative maintenance evaluation value corresponding to each first agent based on the local reward value and the global reward value corresponding to each first agent.
[0160] In the embodiment, the second determination module 340 is specifically configured to:
[0161] When the global maintenance evaluation value corresponding to the current running state of the agent system is acquired, the candidate maintenance strategy executed by each first agent is determined; and the candidate maintenance strategy executed by the first agent is determined as the target maintenance strategy corresponding to the vehicle chassis component.
[0162] In the embodiment, the first determination module 330 is specifically configured to:
[0163] acquire the system set weight corresponding to each first agent; and determine the global maintenance evaluation value corresponding to the current running state of the agent system based on the system set weight corresponding to each first agent and the local collaborative maintenance evaluation value corresponding to each first agent.
[0164] In the embodiment, the vehicle chassis components included in the agent system correspond to the upstream production unit and the downstream production unit respectively, the upstream production unit corresponds to the buffer inventory, and the buffer inventory is used to transfer the production products of the upstream production unit stored to the downstream production unit.
[0165] The method further includes a promotion module.
[0166] The promotion module is configured to promote the maintenance level of the vehicle chassis component corresponding to the upstream production unit if the buffer inventory does not meet the preset inventory amount, and promote the maintenance level of the vehicle chassis component corresponding to the downstream production unit if the buffer inventory meets the preset inventory amount.
[0167] The multi-agent collaborative vehicle chassis maintenance device provided by the present disclosure can execute the method embodiments, and the specific implementation principles and technical effects can be referred to the method embodiments, which will not be described here in detail.
[0168] The embodiment of the present disclosure further provides a computer device. For details, please refer to Figure 4 , Figure 4 The embodiment of the present disclosure further provides a computer device. For details, please refer to
[0169] The computer device includes a memory 410 and a processor 420 that are interconnected and communicate with each other via a system bus. It should be noted that the figure only shows a computer device with a memory 410 and a processor 420, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.
[0170] Computer devices can be desktop computers, laptops, PDAs, cloud servers, etc. Computer devices can interact with users through keyboards, mice, remote controls, touchpads, or voice control devices.
[0171] The memory 410 includes at least one type of readable storage medium, including non-volatile memory or volatile memory, for example, flash memory, a hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. The RAM can include static RAM or dynamic RAM. In some embodiments, the memory 410 can be an internal storage unit of the computer device, for example, a hard disk or a memory of the computer device. In other embodiments, the memory 410 can also be an external storage device of the computer device, for example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, or a flash card, etc. equipped on the computer device. Of course, the memory 410 can include both an internal storage unit and an external storage device of the computer device. In the present embodiment, the memory 410 is generally used to store an operating system and various application software installed on the computer device, for example, program codes of the above-described method, etc. In addition, the memory 410 can also be used to temporarily store various data that has been output or will be output.
[0172] The processor 420 is generally used to perform the overall operation of the computer device. In the present embodiment, the memory 410 is used to store program codes or instructions, which include computer operation instructions, and the processor 420 is used to execute the program codes or instructions stored in the memory 410 or process data, for example, run the program codes of the above-described method.
[0173] In this document, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus system can be a system of address, data, control, and the like. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0174] Another embodiment of the present application also provides a computer readable medium, which can be a computer readable signal medium or a computer readable medium. The processor in the computer reads the computer readable program code stored in the computer readable medium, so that the processor can perform the function actions specified in each step or combination of steps in the above method; and generates a device that implements the function actions specified in each block or combination of blocks in the block diagram.
[0175] The computer readable medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any appropriate combination of the foregoing, for storing program codes or instructions, which include computer operation instructions, and the processor for executing the program codes or instructions of the above method stored in the memory.
[0176] The definition of the memory and the processor can refer to the description of the foregoing computer device embodiment, which will not be repeated here.
[0177] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other ways. For example, the device embodiment described above is only schematic, for example, the division of the module or unit is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between the units or devices, which can be electrical, mechanical or other forms.
[0178] The function units or modules in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software function unit.
[0179] If the integrated unit is implemented in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or say the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.
[0180] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of both hardware and software, and any combination thereof. In the device claim enumerating several means, several of these means can be embodied by one and the same item of hardware. The mere fact that certain measures are recited in mutually different claims does not indicate that a combination of these measures cannot be used to advantage. The use of relative terms such as "first", "second" and "third", etc. does not connote any prioritization, but such terms are used to distinguish a certain feature from another feature with the same name. The steps of the methods described in the above embodiments should not be understood as necessarily limited in their sequence, except when this is explicitly specified.
[0181] The above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; even though the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A multi-agent collaborative vehicle chassis maintenance method, characterized in that: include: Obtain an agent composition system corresponding to a special vehicle chassis system, the agent composition system being used to describe an intelligent operating system composed of multiple vehicle chassis components, wherein each vehicle chassis component corresponds to a first agent, and every two first agents are associated and connected by a second agent; each vehicle chassis component corresponds to a different preset maintenance strategy, wherein the preset maintenance strategy is determined by a second agent associated with the first agent corresponding to the vehicle chassis component based on the component status of the vehicle chassis component and the component status of adjacent chassis components; The preset maintenance strategy includes: at least one candidate maintenance strategy; When the special vehicle is in the current operating state of the agent system, obtaining a local collaborative maintenance evaluation value corresponding to each of the first agents, where the local collaborative maintenance evaluation value corresponding to the first agents is a system reward value obtained after each of the first agents executes a preset maintenance strategy for a corresponding vehicle chassis component; Determining a global maintenance evaluation value of the system of agents corresponding to a current operating state based on a local collaborative maintenance evaluation value corresponding to each of the first agents; Based on the global maintenance evaluation value of the intelligent agent composition system corresponding to the current operating state, the target maintenance strategy corresponding to each of the vehicle chassis components is determined, so that each of the first intelligent agents performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
2. The method according to claim 1, characterized in that The obtaining of a local collaborative maintenance evaluation value corresponding to each first agent when the special vehicle is in the current operating state of the agent composition system includes: The special vehicle controls each of the first agents to sequentially execute each candidate maintenance strategy for the corresponding vehicle chassis component under the current operating state of the agent composition system; Obtaining system operation data obtained after each of the first intelligent agents executes all corresponding candidate maintenance strategies; Based on the system operation data obtained after each of the first intelligent agents executes all corresponding candidate maintenance strategies, a local collaborative maintenance evaluation value corresponding to each of the first intelligent agents is determined.
3. The method according to claim 2, characterized in that Determining a local collaborative maintenance evaluation value corresponding to each first agent based on system operation data obtained after each first agent executes all corresponding candidate maintenance strategies includes: determining a system operation risk value of the overall operation system of the special vehicle based on system operation data obtained after each of the first agents executes all corresponding candidate maintenance strategies; Based on the system operation risk value of the overall operation system of the special vehicle, a local collaborative maintenance assessment value corresponding to each of the first intelligent agents is determined.
4. The method according to claim 3, characterized in that The determining of the local collaborative maintenance assessment value corresponding to each first agent based on the system operation risk value of the overall operation system of the special vehicle includes: Determining a local reward value corresponding to each of the first intelligent agents based on a system operation risk value of the overall operation system of the special vehicle; Determining a global reward value corresponding to each of the first intelligent agents based on a system operation risk value of the overall operation system of the special vehicle and a collaborative influence value between each of the first intelligent agents and adjacent intelligent agents; Based on the local reward value and the global reward value corresponding to each of the first agents, a local collaborative maintenance evaluation value corresponding to each of the first agents is determined.
5. The method according to claim 1, wherein The determining of a target maintenance strategy corresponding to each vehicle chassis component based on a global maintenance evaluation value of the intelligent agent composition system corresponding to a current operating state includes: obtaining a candidate maintenance strategy executed by each of the first agents when the global maintenance evaluation value of the agent-composed system corresponds to the current operating state; The candidate maintenance strategy executed by the first agent is determined to be a target maintenance strategy corresponding to the vehicle chassis component.
6. The method according to claim 1, characterized in that The determining, based on the local collaborative maintenance evaluation value corresponding to each of the first agents, a global maintenance evaluation value corresponding to the current operating state of the agent composition system includes: Obtaining a system-set weight corresponding to each of the first intelligent agents; Based on the system setting weight corresponding to each of the first agents and the local collaborative maintenance evaluation value corresponding to each of the first agents, a global maintenance evaluation value of the agent composition system corresponding to the current operating state is determined.
7. The method according to claim 1, characterized in that The vehicle chassis components included in the intelligent agent composition system correspond to upstream production units and downstream production units respectively. The upstream production units correspond to buffer stocks, and the buffer stocks are used to transfer the stored products produced by the upstream production units to the downstream production units. The method further comprises: If the buffer stock does not meet the preset stock level, the maintenance level of the vehicle chassis component corresponding to the upstream production unit is increased; If the buffer stock meets the preset stock amount, the maintenance level of the vehicle chassis component corresponding to the downstream production unit is increased.
8. A multi-agent collaborative vehicle chassis maintenance device, characterized in that: include: A first acquisition module is configured to acquire an agent composition system corresponding to a special vehicle chassis system, wherein the agent composition system is configured to describe an intelligent operation system composed of a plurality of vehicle chassis components, wherein each vehicle chassis component corresponds to a first agent, and each two first agents are associated with each other by a second agent; and each vehicle chassis component corresponds to a different preset maintenance strategy, wherein the preset maintenance strategy is determined by a second agent associated with the first agent corresponding to the vehicle chassis component based on the component status of the vehicle chassis component and the component status of adjacent chassis components. The preset maintenance strategy includes: at least one candidate maintenance strategy; a second acquisition module configured to acquire, when the special vehicle is in a current operating state of the agent system, a local collaborative maintenance evaluation value corresponding to each of the first agents, the local collaborative maintenance evaluation value corresponding to the first agents being a system reward value obtained after each of the first agents executes a preset maintenance strategy for a corresponding vehicle chassis component; a first determining module configured to determine a global maintenance evaluation value of the agent system corresponding to a current operating state based on a local collaborative maintenance evaluation value corresponding to each of the first agents; The second determination module is used to determine the target maintenance strategy corresponding to each of the vehicle chassis components based on the global maintenance evaluation value of the intelligent agent composition system corresponding to the current operating state, so that each of the first intelligent agents performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, a multi-agent collaborative vehicle chassis maintenance method as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-agent collaborative vehicle chassis maintenance method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
A multi-agent based bi-directional joint scheduling strategy decision method for maintenance resources
CN109190995A
Intelligent maintenance decision-making method and system for underwater production control system
CN118096121A