Multi-agent cooperative vehicle chassis maintenance method, device, equipment and medium
Through the multi-agent collaborative method, coordinated maintenance decisions for special vehicle chassis systems are solved, and the maintenance efficiency problem caused by the assumption that components are independent in the existing technology is solved, and a more efficient maintenance strategy is achieved.
Patent Information
- Application Number
- CN202411753218.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-12-02
AI Technical Summary
The existing chassis system maintenance methods assume that the components are independent of each other, resulting in inefficient overall maintenance of the system.
By introducing a multi-agent collaboration method, the agent composition system of the special vehicle chassis system is obtained. Each vehicle chassis component corresponds to a first agent, and the second agent is associated and connected between the two first agents to coordinate maintenance decisions.
It effectively improves the overall maintenance efficiency of the vehicle chassis system and achieves a more flexible, scalable and global optimal maintenance strategy.
Smart Images

Figure CN119941216A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the technical field of chassis system maintenance, and in particular, to a vehicle chassis maintenance method, device, equipment and medium suitable for multi-agent collaboration. Background Art
[0002] The maintenance optimization of special vehicle chassis systems is a hot topic in the field of reliability research. Maintenance theory generally believes that there are multiple correlations between these components, making it difficult to determine a reasonable maintenance strategy. Research on the maintenance optimization of special vehicle chassis systems can help companies save costs and improve system reliability.
[0003] In the related art, common chassis system maintenance optimization methods are mainly implemented based on fixed threshold value strategies. During the implementation process, a set of preventive maintenance threshold values are usually set to simplify the maintenance optimization problem. For example, ant colony algorithms and simulated annealing algorithms are used to optimize the maintenance threshold values of different components in the system.
[0004] However, in the existing approach, maintenance is performed on the assumption that each component is independent of each other, resulting in low overall maintenance efficiency of the system. Summary of the invention
[0005] The embodiments described herein provide a multi-agent collaborative vehicle chassis maintenance method, apparatus, device, and medium that overcome the above-mentioned problems.
[0006] In a first aspect, according to the present disclosure, a multi-agent collaborative vehicle chassis maintenance method is provided, comprising:
[0007] Acquire an agent composition system corresponding to a special vehicle chassis system, wherein the agent composition system is used to describe an intelligent operation system composed of a plurality of vehicle chassis components, each of the vehicle chassis components corresponds to a first agent, and every two of the first agents are associated and connected by a second agent;
[0008] When the special vehicle is in the current running state of the agent system, a local collaborative maintenance evaluation value corresponding to each of the first agents is obtained, where the local collaborative maintenance evaluation value corresponding to the first agent is a system reward value obtained after each of the first agents respectively executes a preset maintenance strategy for the corresponding vehicle chassis component;
[0009] Determine, based on the local collaborative maintenance evaluation value corresponding to each of the first agents, a global maintenance evaluation value of the agent composition system corresponding to the current operating state;
[0010] Based on the global maintenance evaluation value of the agent composition system corresponding to the current operating state, the target maintenance strategy corresponding to each of the vehicle chassis components is determined, so that each of the first agents performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
[0011] In a second aspect, according to the content of the present disclosure, a multi-agent collaborative vehicle chassis maintenance device is provided, comprising:
[0012] A first acquisition module is used to acquire an agent composition system corresponding to a special vehicle chassis system, wherein the agent composition system is used to describe an intelligent operation system composed of a plurality of vehicle chassis components, each of the vehicle chassis components corresponds to a first agent, and every two of the first agents are associated and connected by a second agent;
[0013] A second acquisition module is used to acquire a local collaborative maintenance evaluation value corresponding to each of the first agents when the special vehicle is in the current operating state of the agent composition system, wherein the local collaborative maintenance evaluation value corresponding to the first agent is a system reward value obtained after each of the first agents respectively executes a preset maintenance strategy for the corresponding vehicle chassis component;
[0014] A first determination module, configured to determine a global maintenance evaluation value of the agent composition system corresponding to a current operating state based on a local collaborative maintenance evaluation value corresponding to each of the first agents;
[0015] The second determination module is used to determine the target maintenance strategy corresponding to each of the vehicle chassis components based on the global maintenance evaluation value of the intelligent agent composition system corresponding to the current operating state, so that each of the first intelligent agents performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
[0016] In a third aspect, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the multi-agent collaborative vehicle chassis maintenance method in any of the above embodiments are implemented.
[0017] In a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the multi-agent collaborative vehicle chassis maintenance method in any of the above embodiments are implemented.
[0018] The multi-agent collaborative vehicle chassis maintenance method provided in the embodiment of the present application obtains an agent composition system corresponding to the special vehicle chassis system, the agent composition system is used to describe the intelligent operation system composed of multiple vehicle chassis components, each vehicle chassis component corresponds to a first agent, and each two first agents are associated and connected by a second agent; when the special vehicle is in the current operation state of the agent composition system, the local collaborative maintenance evaluation value corresponding to each first agent is obtained, and the local collaborative maintenance evaluation value corresponding to the first agent is the system reward value obtained after each first agent executes the preset maintenance strategy for the corresponding vehicle chassis component; based on the local collaborative maintenance evaluation value corresponding to each first agent, the global maintenance evaluation value corresponding to the agent composition system under the current operation state is determined; based on the global maintenance evaluation value corresponding to the agent composition system under the current operation state, the target maintenance strategy corresponding to each vehicle chassis component is determined, so that each first agent performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy. In this way, by connecting multiple vehicle chassis components to each other, collaborative maintenance processing is performed according to the agent corresponding to each vehicle chassis component, so as to effectively improve the overall maintenance efficiency of the vehicle chassis system.
[0019] The above description is only an overview of the technical solution of the embodiment of the present application. In order to more clearly understand the technical means of the embodiment of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiment of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly described below. It should be noted that the drawings described below only relate to some embodiments of the present disclosure, but are not intended to limit the present disclosure, wherein:
[0021] Figure 1 It is a flow chart of a multi-agent collaborative vehicle chassis maintenance method provided by the present invention.
[0022] Figure 2A It is a structural diagram of an intelligent agent composition system provided by the present invention.
[0023] Figure 2B It is a structural schematic diagram of a series production system with buffer inventory provided by the present invention.
[0024] Figure 3 It is a structural schematic diagram of a multi-agent collaborative vehicle chassis maintenance device provided by the present invention.
[0025] Figure 4 It is a structural schematic diagram of a computer device provided by the present disclosure.
[0026] It should be noted that the elements in the drawings are schematic and not drawn to scale. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the embodiments of the present disclosure clearer, the technical solution of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative work also fall within the scope of protection of the present disclosure.
[0028] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by a person skilled in the art to which the subject matter of the present disclosure belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the specification and the relevant art, and will not be interpreted in an idealized or overly formal form unless otherwise explicitly defined herein. As used herein, a statement that two or more parts are "connected" or "coupled" together shall mean that the parts are joined together directly or through one or more intermediate components.
[0029] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase "embodiments" in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0030] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists, A and B exist at the same time, and B exists. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship. Terms such as "first" and "second" are only used to distinguish one component (or a part of a component) from another component (or another part of a component).
[0031] In the description of the present application, unless otherwise specified, "plurality" means more than two (including two), and similarly, "multiple groups" means more than two groups (including two).
[0032] This embodiment effectively solves the complex state and action space of the special vehicle chassis system by introducing collaborative reinforcement learning, and realizes collaborative decision-making among various components, breaking through the limitations of traditional maintenance strategies and having high flexibility, scalability and global optimality.
[0033] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.
[0034] Figure 1 is a flow chart of a multi-agent collaborative vehicle chassis maintenance method provided by an embodiment of the present disclosure, such as Figure 1 As shown in FIG. 1 , the specific process of the multi-agent collaborative vehicle chassis maintenance method includes:
[0035] S110. Obtain an intelligent agent composition system corresponding to the special vehicle chassis system, where the intelligent agent composition system is used to describe an intelligent operation system composed of multiple vehicle chassis components, and each vehicle chassis component corresponds to a first intelligent agent.
[0036] Among them, the chassis system of special vehicles can be composed of transmission system, running system and steering system. The transmission system is mainly composed of vehicle chassis components such as differential, main reducer, universal transmission device, transmission and clutch. The running system is mainly composed of vehicle chassis components such as suspension, axle wheels and frame. The steering system is composed of vehicle chassis components such as steering column, steering shaft and steering wheel.
[0037] Every two first agents are associated and connected by a second agent, so that two first agents are connected through one second agent.
[0038] Each vehicle chassis component can be equivalent to a first agent, and each vehicle chassis component corresponds to an independent first agent, which is responsible for the status monitoring and maintenance decision of the vehicle chassis component. For example, a component of a special vehicle chassis (such as the differential of the transmission system) can have an exclusive first agent to handle its operating status and maintenance decision.
[0039] Every two vehicle chassis components are connected by a second agent, which can be used as a node connecting components in the graph structure to handle the collaborative decision-making between two vehicle chassis components. In Coordinated Reinforcement Learning (CRL), a second agent may need to make joint decisions on the states of two or more components and select coordinated actions by considering the state information of the two components.
[0040] Based on traditional single-agent reinforcement learning, a coordination graph is introduced. The first agents interact and share information through a coordination mechanism, overcoming the problem that a single agent cannot handle large-scale state and action spaces.
[0041] The structural diagram of the intelligent agent system is as follows: Figure 2A shown. Figure 2A The collaborative graph of a series system consisting of 12 vehicle chassis components is shown in Figure 1. Each square represents a vehicle chassis component, corresponding to a first agent, and a circle represents a second agent, which is used to connect two adjacent vehicle chassis components.
[0042] S120. When the special vehicle is in the current operating state of the intelligent agent system, obtain a local collaborative maintenance evaluation value corresponding to each first intelligent agent.
[0043] The local collaborative maintenance evaluation value corresponding to the first agent is the system reward value obtained after each first agent executes the preset maintenance strategy for the corresponding vehicle chassis component. That is, after the first agent corresponding to all vehicle chassis components executes a preset maintenance strategy for the component, the impact incentive on the entire vehicle chassis system.
[0044] In some embodiments, each vehicle chassis component corresponds to a different preset maintenance strategy, and the preset maintenance strategy is determined by the second intelligent agent associated with the first intelligent agent corresponding to the vehicle chassis component based on the component status of the vehicle chassis component and the component status of adjacent chassis components; the preset maintenance strategy includes: at least one candidate maintenance strategy.
[0045] Among them, the component status may include: health status, operating status, maintenance requirements, and inventory status. The health status can be used to reflect the degree of degradation of the component, such as discrete states such as "brand new", "mildly degraded", "moderately degraded", "severely degraded", "faulty", and "failed"; the operating status may include "normal operation", "under maintenance", "fault shutdown", "starvation state" (no material or lack of resources to support operation) and "blocked state" (unable to operate normally due to insufficient downstream resources); maintenance requirements are used to describe the type of maintenance required for the component, such as whether the preventive maintenance threshold has been reached or post-maintenance is required; the inventory status may be the inventory of the component (such as the storage volume of the buffer inventory).
[0046] The second agent connecting the two first agents can determine the corresponding preset maintenance strategy by obtaining the component status of the corresponding vehicle chassis component from the first agent. The preset maintenance strategy may include: preventive maintenance, post-maintenance, and continued operation. Preventive maintenance is to repair the component before it fails completely to extend its service life; post-maintenance is to repair the component after it fails to restore its function; continued operation is to keep the component in the current state without any maintenance operation. In addition, the new preset maintenance strategy can be dynamically adjusted according to the component status and system requirements, such as downgrading use, adjusting component load, etc.
[0047] For example, if the second intelligent agent determines that the component state of a vehicle chassis component is "brand new" or "slightly degraded", it can determine that the preset maintenance strategy corresponding to the component is "continue to operate"; if it is determined that the component state of a vehicle chassis component is "moderately degraded" or "severely degraded" but has not yet failed, it can determine that the preset maintenance strategy corresponding to the component is "preventive maintenance" to prevent failure; if it is determined that the component state of a vehicle chassis component is "failed", it can determine that the preset maintenance strategy corresponding to the component is "post-maintenance"; if it is determined that the component state of a vehicle chassis component is "blocked state" or "starved state", it can determine that the preset maintenance strategy corresponding to the component is "continue to operate" or "preventive maintenance" to avoid system downtime or extend component life.
[0048] It should be noted that in a collaborative learning environment, the second agent can also consider the component status of adjacent components. For example, the component status of a certain component may affect the decision of adjacent components (e.g., if a critical component is close to failure, preventive maintenance can be prioritized on adjacent components).
[0049] The second agent can select the preset maintenance strategy of the two connected components according to the component states of the two connected components. The second agent connecting component m and component n is represented as Ag n,m , Ag n,m The state and action are represented as X n,m =[X n ,X m ] and A n,m =[A n ,A m ].
[0050] The preset maintenance strategy for a component can be made by two or more intelligent agents. The component status of the component can be shared among the intelligent agents, and the optimal maintenance strategy can be jointly determined by combining the strategic decision results of different intelligent agents. Through the collaboration of multiple intelligent agents, it is ensured that the decision on a single component not only optimizes its own operating status, but also takes into account the correlation with other components and the overall benefits of the system. Figure 2A The parameter table of the agent is shown in Table 1.
[0051] Table 1 Agent parameters
[0052]
[0053]
[0054] When the special vehicle is in the current operating state of the agent composition system, a local collaborative maintenance evaluation value corresponding to each first agent is obtained, including:
[0055] When the special vehicle is in the current operating state of the agent composition system, each first agent is controlled to execute each candidate maintenance strategy in turn for the corresponding vehicle chassis component; the system operation data obtained after each first agent executes all the corresponding candidate maintenance strategies is obtained; based on the system operation data obtained after each first agent executes all the corresponding candidate maintenance strategies, the local collaborative maintenance evaluation value corresponding to each first agent is determined.
[0056] Among them, the system operation data can be used to describe the current operating parameters of the vehicle chassis system, such as the steering shaft angle and the steering wheel angle.
[0057] Each first agent is controlled to execute each candidate maintenance strategy in turn for the corresponding vehicle chassis component, and the system operation data obtained after each first agent executes all corresponding candidate maintenance strategies is obtained. An example is as follows.
[0058] The intelligent agent composition system includes three first intelligent agents, namely intelligent agent 11, intelligent agent 12 and intelligent agent 13. The preset maintenance strategy of the vehicle chassis component corresponding to intelligent agent 11 is strategy A1, the preset maintenance strategy of the vehicle chassis component corresponding to intelligent agent 12 is strategies A2 and A3, and the preset maintenance strategy of the vehicle chassis component corresponding to intelligent agent 13 is strategy A4. Each first intelligent agent is controlled to execute each candidate maintenance strategy in turn for the corresponding vehicle chassis component, and the strategy execution sequence can be obtained as: {A1-A2-A4} and {A1-A3-A4}. Then, the system operation data obtained after each first intelligent agent executes all the corresponding candidate maintenance strategies may include the system operation data corresponding to strategy {A1-A2-A4} and the system operation data corresponding to strategy {A1-A3-A4}.
[0059] In some embodiments, determining the local collaborative maintenance evaluation value corresponding to each first agent based on the system operation data obtained after each first agent executes all corresponding candidate maintenance strategies includes:
[0060] Based on the system operation data obtained after each first intelligent agent executes all corresponding candidate maintenance strategies, the system operation risk value of the overall operation system of the special vehicle is determined; based on the system operation risk value of the overall operation system of the special vehicle, the local collaborative maintenance assessment value corresponding to each first intelligent agent is determined.
[0061] Among them, the system operation risk value of the overall operation system of the special vehicle can be used to describe the corresponding component-related costs, which may include: maintenance costs, allocated inspection costs, and allocated downtime costs. Maintenance costs include preventive maintenance costs, and downtime costs include costs per unit time when the system is in a fault state.
[0062] In some embodiments, determining the local collaborative maintenance assessment value corresponding to each first agent based on the system operation risk value of the overall operation system of the special vehicle includes:
[0063] Based on the system operation risk value of the overall operation system of the special vehicle, the local reward value corresponding to each first intelligent agent is determined; based on the system operation risk value of the overall operation system of the special vehicle and the collaborative influence value between each first intelligent agent and the adjacent intelligent agent, the global reward value corresponding to each first intelligent agent is determined; based on the local reward value and the global reward value corresponding to each first intelligent agent, the local collaborative maintenance evaluation value corresponding to each first intelligent agent is determined.
[0064] For example, the local reward value C corresponding to the first agent n,m As shown in the following formula (1).
[0065] C n,m =C n / |C(n)|+C m / |C(m)| (1)
[0066] In formula (1), c n represents the cost associated with component n, that is, the component-related cost corresponding to component n; c m represents the cost associated with component m, that is, the component-related cost corresponding to component m; represents the number of second agents connected to component n; Represents the number of second agents connected to component m.
[0067] c n As shown in the following formula (2).
[0068]
[0069] In formula (2), C PR,n represents the preventive maintenance cost of component n, A nrepresents the execution action of component n, PR represents the candidate maintenance strategy is “preventive maintenance”, I(A n =PR) indicates that the strategy corresponding to the current execution action is "preventive maintenance"; C INS represents the detection cost of component n; X represents the current state of the system, X' represents the next state of the system, C F It represents the cost per unit time when the system is in a faulty state; C ST represents the setup cost, N represents the total number of components in the series system, DN represents the candidate maintenance strategy is “no maintenance”, I(A n =DN) indicates that the policy corresponding to the current execution action is "no maintenance".
[0070] The global reward value corresponding to the first intelligent agent may be the system operation risk value of the overall operation system of the special vehicle, and the sum of the synergistic influence values between the first intelligent agent and the adjacent intelligent agents.
[0071] The local reward value corresponding to the first agent is set based on the immediate impact of the agent's own maintenance decision on the system. For example, when the first agent performs "preventive maintenance", if it successfully avoids the occurrence of potential failures, it will receive a corresponding benefit reward; and after performing "post-maintenance" to restore the normal operation of the component, it will also receive a certain reward. At the same time, the expenses incurred during the maintenance process and the downtime losses caused by maintenance will be regarded as costs, and the corresponding points will be deducted. These reward and deduction mechanisms directly reflect the quality of the first agent's own decision, enabling the first agent to adjust its decision-making strategy based on immediate feedback. In the collaborative decision-making of agents, local rewards affect the first agent's evaluation of different actions. When each first agent calculates the local reward value, it will take into account the local rewards brought by its own decision, so as to more accurately evaluate the immediate effect of taking a certain action in the current state.
[0072] The first agent can also obtain additional rewards based on the synergy effect with adjacent components (or adjacent agents), namely the synergy impact value. The synergy reward takes into account the overall operating status of the global system. For example, the health status of a component may affect the working efficiency of adjacent components. Successful collaborative maintenance decisions (such as maintaining stable operation of the system under the cooperation of multiple key components) can obtain rewards for improving global performance. This type of reward encourages agents to take into account the reliability and long-term benefits of the overall system when making decisions.
[0073] The reward function is defined as the system's gain in one unit of time when the system state is X and the action is A, as shown in the following formula (3).
[0074] R(X,A)=R r (X,A)-C P (X,A)-C M (X,A)-CB (X,A) (3)
[0075] In formula (3), R r (X, A) is the revenue of the system’s products, which can be expressed as shown in formula (4).
[0076] R r (X,A)=v N,norm I(A N =DN and v N =v N,norm ) (4)
[0077] C P (X, A) is the sum of the operating costs of each production unit, expressed as shown in formula (5).
[0078]
[0079] C M (X, A) is the sum of the maintenance costs of each production unit, expressed as shown in formula (6).
[0080]
[0081] In formula (6), when ω is true, function I(ω)=1; when ω is false, function I(ω)=0.
[0082] C B (X, A) is the total holding cost of the buffer stock, expressed as shown in formula (7).
[0083]
[0084] Assuming that the degradation of each production unit is independent of each other, the system transition probability can be simplified as shown in formula (8).
[0085]
[0086] When A n = DN, according to M n Whether it is in an operational state, the state transition probability is shown in formula (9) or formula (10).
[0087]
[0088] When A n =PM, M n The state transition probability is shown in formula (11) or formula (12).
[0089] P(S n ′=1|S n =s n )=pn,PM (11)
[0090] P(S n ′=D n +1|S n =s n )=1-p n,PM (12)
[0091] When A n =CM, M n The state transition probability is shown in formula (13) or formula (14).
[0092] P(S′ n =1 S n =S n )=P n,CM (13)
[0093] P(S′ n =D n |S n =s n )=1-p n,CM (14)
[0094] S130. Determine a global maintenance evaluation value of the system of agents under a current operating state based on the local collaborative maintenance evaluation value corresponding to each first agent.
[0095] The global maintenance evaluation value of the agent composition system corresponding to the current operating state may be the sum of the local collaborative maintenance evaluation values corresponding to the multiple first agents.
[0096] For example, the agent Ag n,m The local Q function is expressed as Q n,m (X n,m ,A n,m ), the global Q function is the sum of the local Q functions. The local Q function defines the “local Q value” of each first agent under the current state-action pair, which can be used to evaluate the value of taking a certain action in the current state, that is, the local collaborative maintenance evaluation value.
[0097] The global Q function integrates the local Q values of multiple agents to form the global Q function of the system to measure the impact of different decisions on the overall system, as shown in the following formula (15).
[0098]
[0099] In formula (15), represents the set of component pairs connected by the agents in the collaborative graph. The local Q function can be based on the sampled quadruple To update, the update formula is shown in formula (16).
[0100]
[0101] In formula (16), η n,m Represents the agent Ag n,m The learning rate is , and γ represents the discount factor.
[0102] Locally optimal action is part of the global optimal action A'*, A' * is jointly determined by all agents. Since the action space grows exponentially with the number of components, the selection of the global optimal action will also be affected by the "curse of dimensionality", that is, in high-dimensional state or action space, as the dimension increases, the computational complexity and the amount of data required grow exponentially, making it difficult for the algorithm to handle large-scale problems. To solve this problem, a variable elimination algorithm is used to find the global optimal action A′ * .
[0103] This embodiment takes a series system composed of three components as an example to illustrate the implementation process of the variable elimination method. The global Q function is decomposed as shown in formula (17).
[0104] Q(X,A)=Q 1,2 (X 1,2 ,A1,A2)+Q 2,3 (X 2,3 ,A2,A3)+Q 3,1 (X 3,1 ,A3,A1) (17)
[0105] The minimum value can be converted into formula (19) by introducing formula (18).
[0106]
[0107] At this point, the action variable A1 is eliminated. Similarly, the action variable A2 can be eliminated by introducing formula (20).
[0108]
[0109] Since A3 is the last variable remaining in the global Q function, the optimal value of A3 is shown in formula (21).
[0110]
[0111] Subsequently, the optimal values of variables A2 and A1 can be recursively determined as shown in formula (22) and formula (23).
[0112]
[0113] Finally, the global optimal action is obtained The order of variable elimination is variable, and different orders will affect the computational complexity of the action selection process.
[0114] This embodiment uses the principle of Q-value maximization, and the agent selects the optimal action (maintenance decision) at each moment to achieve dynamic updating of the strategy. At each decision moment, the agent uses the principle of Q-value maximization to select the optimal action. For example, in a series system consisting of 12 components, each agent calculates the local Q-value based on the status of two connected components, and exchanges information with other agents through the collaborative graph mechanism, and then selects the action that maximizes the overall benefit of the system based on the global Q-value. Action selection: including "preventive maintenance", "post-maintenance", "continued operation", etc. The agent determines which action can bring the greatest long-term benefit in the current state based on the calculated value, and then makes a decision. By decomposing the global reward into the local reward of each agent, the unity of global goals (such as system stability and high returns) and local decisions is ensured at the same time.
[0115] Global rewards take into account the synergy between agents and the overall operating status of the system. For example, when multiple key components maintain stable system operation through collaborative maintenance decisions, all agents participating in the collaboration can receive rewards for improving global performance. This reward mechanism encourages agents to consider the reliability and long-term benefits of the overall system when making decisions, and avoid pursuing only the maximization of their own local interests at the expense of the overall performance of the system.
[0116] In some embodiments, determining the global maintenance evaluation value of the agent composition system corresponding to the current operating state based on the local collaborative maintenance evaluation value corresponding to each first agent includes:
[0117] Obtain the system setting weight corresponding to each first agent; based on the system setting weight corresponding to each first agent and the local collaborative maintenance evaluation value corresponding to each first agent, determine the global maintenance evaluation value of the agent composition system corresponding to the current operating state.
[0118] Among them, different system setting weights can be set for each first intelligent agent respectively, and then the weights are summed up to effectively determine the global maintenance evaluation value of the intelligent agent composition system corresponding to the current operating state.
[0119] S140. Determine a target maintenance strategy corresponding to each vehicle chassis component based on the global maintenance evaluation value of the agent composition system corresponding to the current operating state.
[0120] Among them, based on the global maintenance evaluation value of the intelligent agent composition system corresponding to the current operating state, the target maintenance strategy corresponding to each vehicle chassis component is determined, so that each first intelligent agent performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
[0121] In some embodiments, determining a target maintenance strategy corresponding to each vehicle chassis component based on a global maintenance evaluation value of the agent composition system corresponding to the current operating state includes:
[0122] The candidate maintenance strategy executed by each first agent is obtained when the global maintenance evaluation value of the agent composition system corresponding to the current operating state; and the candidate maintenance strategy executed by the first agent is determined to be the target maintenance strategy corresponding to the corresponding vehicle chassis component.
[0123] Combined with the above example, if it is determined that the global maintenance evaluation value of the agent composition system under the current operating state, the candidate maintenance strategy executed by agent 11 is strategy A1, the candidate maintenance strategy executed by agent 12 is strategy A2, and the candidate maintenance strategy executed by agent 13 is strategy A4, then the target maintenance strategy of the vehicle chassis component corresponding to agent 11 is determined to be strategy A1, the target maintenance strategy of the vehicle chassis component corresponding to agent 12 is strategy A2, and the target maintenance strategy of the vehicle chassis component corresponding to agent 13 is strategy A4. Thus, the target maintenance strategy for each vehicle chassis component suitable for maximizing global benefits is effectively determined.
[0124] In this embodiment, an intelligent agent composition system corresponding to the special vehicle chassis system is obtained, and the intelligent agent composition system is used to describe an intelligent operation system composed of multiple vehicle chassis components, each vehicle chassis component corresponds to a first intelligent agent, and each two first intelligent agents are associated and connected by a second intelligent agent; when the special vehicle is in the current operation state of the intelligent agent composition system, the local collaborative maintenance evaluation value corresponding to each first intelligent agent is obtained, and the local collaborative maintenance evaluation value corresponding to the first intelligent agent is the system reward value obtained after each first intelligent agent executes the preset maintenance strategy for the corresponding vehicle chassis component; based on the local collaborative maintenance evaluation value corresponding to each first intelligent agent, the global maintenance evaluation value corresponding to the intelligent agent composition system in the current operation state is determined; based on the global maintenance evaluation value corresponding to the intelligent agent composition system in the current operation state, the target maintenance strategy corresponding to each vehicle chassis component is determined, so that each first intelligent agent performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy. In this way, by connecting multiple vehicle chassis components to each other, collaborative maintenance processing is performed according to the intelligent agent corresponding to each vehicle chassis component, so as to effectively improve the overall maintenance efficiency of the vehicle chassis system.
[0125] This embodiment can construct the basic framework of the entire special vehicle chassis system by initializing the system of intelligent agents, assigning intelligent agents to each production unit and buffer stock in the system, and defining the system's state space, action space, and reward function, thereby providing information such as state space, action space, and reward function for the subsequent reinforcement learning process. The initialized system state space, action space, and reward function provide basic data for subsequent intelligent agent decision-making, strategy updating, and training. In particular, these definitions will be used for input and decision support in the "collaborative decision-making of each intelligent agent" and "strategy updating and iteration" stages.
[0126] The construction of the system (including the design of production units and buffer stocks) reduces the coupling between production units through the introduction of buffer stocks, allowing the system to continue to partially operate when a unit is repaired or fails. The inventory mechanism avoids system downtime caused by single point failures, thereby improving the overall reliability of the system.
[0127] Each production unit and buffer stock can be modeled through a Markov decision process, clearly defining the system state (e.g., health status, inventory level) and action (such as repair or continue operation). Through state modeling, a standardized feedback mechanism can be provided to the intelligent agent, thereby achieving dynamic decision-making.
[0128] For example, the MDP (Markov decision process) model is used to define the system state, action, transition probability and immediate reward, and the maintenance problem of the entire system is transformed into a reinforcement learning problem, which is solved using a collaborative reinforcement learning algorithm.
[0129] Serial production system with buffer stock Figure 2B As shown in the figure, the square represents the production unit and the circle represents the buffer stock. The buffer stock is used to temporarily store the products of the upstream production unit and provide them to the downstream production unit. The addition of buffer stock also ensures that the production system will not be shut down due to maintenance or failure of a single production unit. Such systems are widely used in actual production, such as tire production systems and tile manufacturing systems.
[0130] Figure 2B In, M n (n∈1,2,…,N) is the production unit, B n (n∈1,2,…,N-1) is the buffer stock. n The inventory at time t is denoted by K n (t)=0,1,...,N n,B , where N n,B For buffer stock The inventory capacity of each production unit is assumed to be subject to a discrete-time, discrete-state Markov process.n The degradation amount can be expressed as a discrete state {1,2,…,D n}, where 1 represents a brand new state, D n Indicates a fault state. If the production unit M n The upstream buffer stock is 0, then M n In a hungry state; if M n The downstream buffer inventory reaches the upper limit, then M n In particular, M1 will not starve, M N Will not be blocked. n The productivity is denoted as v n+1 When production unit M n When in failure, starvation, obstruction or maintenance, v n =0; at other times, v n =v n,norm Assumptions The inventory at the current moment is K n , then its inventory at the next moment is K′ n As shown in formula (24).
[0131] K n ′=K n +v n -v n+1 (twenty four)
[0132] Assume M n The degradation process is affected by its production rate v n The impact of M n When working normally, its state transfer matrix is P n,norm ; in M n When in a starvation or blocking state, its state transition matrix is P n,idle In this example, only two maintenance activities are considered: preventive maintenance and post-maintenance. Both maintenance activities can restore the production unit to a brand new state. Assuming that the maintenance time follows a geometric distribution, in a unit of time, the production unit M n The completion probabilities of preventive maintenance and post-maintenance are p n,PM 、p n,CM , and p n,PM <p n,CM The corresponding preventive maintenance and post-maintenance costs are C n,PM and C n,CM , and C n,PM <C n,CM . M per unit time n The operating cost is C n,op ; Buffer stock The cost of holding a component per unit time is C n,HThe system generates revenue R for each component produced. r ,The goal of maintenance optimization is to maximize the average expected ,revenue per unit time.
[0133] The MDP model can be represented by a four-tuple <ψ,ζ,ξ,s>, where ψ represents the system state set, ζ represents the system action set, ξ represents the system transition probability model, and s represents the system instant reward. The decision maker performs action A∈ζ based on the current system state X∈ψ and obtains an instant reward. And according to the state transition probability P(X′|X,A)∈ξ, it transitions to the next state x′.
[0134] This embodiment adds a D n +1 status is used to indicate that the production unit is in preventive maintenance status. n The state is represented by S n ∈{1,2,…,D n ,D n +1}, and the system state can be expressed as X = [S1, S2, ..., S N ,K1,K2,…,K N-1 ].
[0135] After initializing the system model, the status information of each production unit (such as the degradation status of components, the storage capacity of buffer inventory, etc.) can be obtained in real time and input into the MDP model to output the next state of the system. The next state refers to the set of subsequent states to which the system transfers from the current state after the agent performs a certain action. It reflects the dynamic changes of the system over time and is the key basis for the agent to make decisions and learn.
[0136] After the agent selects an action based on the current state, it observes the next state that is generated, and can evaluate the quality of the action by combining the immediate rewards (such as the revenue of the system's production products, the operating costs of each production unit, the maintenance costs, the buffer inventory holding costs, etc.). For example, if the agent performs a maintenance action, the next state of the system shows that the component failure rate is reduced and the production efficiency is improved, and a higher immediate reward is obtained, the agent will think that the action is beneficial and will be more inclined to choose this action in similar situations in the future.
[0137] Each agent makes maintenance decisions based on the collaborative reinforcement learning algorithm and reward shaping mechanism, combining the status information of itself and adjacent components. The agent's decisions include "preventive maintenance", "post-maintenance" or "continue operation".
[0138] Self-state information refers to the instantaneous state of the specific component that the agent is responsible for managing, which can reflect the degree of wear, aging or damage of the component itself; the operating state ("normal operation", "maintenance", "fault shutdown", "starvation state", "blocked state", etc.) indicates the current working condition of the component; the maintenance demand (whether the preventive maintenance threshold has been reached or post-maintenance is required) is used to determine whether the component needs to be repaired; for systems involving inventory, the inventory of components (such as the storage volume of buffer inventory) is also one of the component state information. For example, in the transmission system of the chassis system, the current state of a gear component may be "slightly degraded" and in a "normal operation" state, but it is close to the preventive maintenance threshold. The agent to which it belongs needs to consider this state information when making decisions; in the transmission system, the current state of a gear component may be "slightly degraded" and in a "normal operation" state, but it is close to the preventive maintenance threshold. The agent to which it belongs needs to consider this state information when making decisions.
[0139] The state information of adjacent components refers to the state of components that are directly connected to the current component that the agent is responsible for in the system structure or closely related in function. In the chassis system of special vehicles, there are physical connections and functional collaborations between different components. For example, the various gears, shafts and other components in the transmission system work together, and the state change of one component may affect the performance and operating status of adjacent components. For example, in the transmission system, the degree of wear (health status) of a certain driving gear will affect the force of the driven gear that meshes with it, and then affect the operating status and life of the driven gear.
[0140] In some embodiments, the vehicle chassis components included in the intelligent agent composition system correspond to upstream production units and downstream production units, respectively, and the upstream production unit corresponds to a buffer inventory, and the buffer inventory is used to transfer the production products of the upstream production unit to the downstream production unit;
[0141] The method of this embodiment also includes:
[0142] If the buffer stock does not meet the preset inventory level, the maintenance level of the vehicle chassis components corresponding to the upstream production unit will be increased; if the buffer stock meets the preset inventory level, the maintenance level of the vehicle chassis components corresponding to the downstream production unit will be increased.
[0143] Among them, the buffer stock serves as an intermediate link between upstream and downstream production units, making collaborative learning between agents more efficient. The status of the buffer stock (such as inventory) will affect the maintenance priority of adjacent production units. The status of each production unit and buffer stock in the system can change in real time, allowing dynamic adjustment of maintenance strategies.
[0144] For example, when the buffer inventory is below a certain threshold, the agent may prioritize repairing the upstream production unit to restore supply; when the buffer inventory is above a certain threshold, the work of the downstream production unit is adjusted. The setting of the buffer inventory not only reduces the impact of downtime on revenue, but also optimizes maintenance decisions by dispersing risks. For example, when a production unit is under maintenance, the buffer inventory can temporarily support the system operation and reduce economic losses.
[0145] This embodiment uses a multi-agent reinforcement learning algorithm to optimize the maintenance strategy of a complex special vehicle chassis system, breaking through the limitations of traditional single-agent methods in dealing with large-scale systems, and significantly improving decision-making efficiency and strategy flexibility. Through the design of a collaborative graph mechanism, each agent can collaborate according to the state of adjacent components to ensure the global optimal maintenance decision, rather than just being limited to the local optimality of each agent. The collaboration between agents enhances the overall maintenance optimization capability of the system. Through reward shaping technology, the global reward of the system is reasonably decomposed into the local reward of each agent, and the global optimality of the learning efficiency and strategy is improved by combining global goals with local decision optimization. This mechanism ensures that each agent can also consider the global impact when making local decisions, ensuring that the overall benefits of the system are maximized. Combining deep Q learning with multi-agent collaborative learning solves the convergence problem of traditional Q learning in large-scale state space. Through the introduction of deep neural networks, the system can efficiently approximate complex state-action value functions, handle large-scale state and action spaces, and ensure efficient learning and decision-making processes. The intelligent agent can dynamically select "preventive maintenance" or "post-maintenance" strategies according to the real-time status of the production unit and the overall operation of the system to ensure the reliability and economy of the system's operation.
[0146] Figure 3 The schematic diagram of the structure of a multi-agent collaborative vehicle chassis maintenance device provided in this embodiment, the multi-agent collaborative vehicle chassis maintenance device may include:
[0147] The first acquisition module 310 is used to obtain the intelligent agent composition system corresponding to the special vehicle chassis system. The intelligent agent composition system is used to describe the intelligent operation system composed of multiple vehicle chassis components. Each vehicle chassis component corresponds to a first intelligent agent, and every two first intelligent agents are associated and connected by a second intelligent agent.
[0148] The second acquisition module 320 is used to obtain the local collaborative maintenance evaluation value corresponding to each first intelligent agent when the special vehicle is in the current operating state of the intelligent agent composition system. The local collaborative maintenance evaluation value corresponding to the first intelligent agent is the system reward value obtained after each first intelligent agent executes the preset maintenance strategy for the corresponding vehicle chassis component.
[0149] The first determination module 330 is used to determine the global maintenance evaluation value of the agent composition system corresponding to the current operating state based on the local collaborative maintenance evaluation value corresponding to each first agent.
[0150] The second determination module 340 is used to determine the target maintenance strategy corresponding to each vehicle chassis component based on the global maintenance evaluation value of the agent composition system corresponding to the current operating state, so that each first agent performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
[0151] In this embodiment, optionally, each vehicle chassis component corresponds to a different preset maintenance strategy, and the preset maintenance strategy is determined by the second intelligent agent associated with the first intelligent agent corresponding to the vehicle chassis component based on the component status of the vehicle chassis component and the component status of adjacent chassis components; the preset maintenance strategy includes: at least one candidate maintenance strategy.
[0152] The second acquisition module 320 includes: a control unit, an acquisition unit and a determination unit.
[0153] The control unit is used for controlling each first intelligent agent to sequentially execute each candidate maintenance strategy for the corresponding vehicle chassis component when the special vehicle is in the current operating state of the intelligent agent composition system.
[0154] The acquisition unit is used to acquire the system operation data obtained after each first intelligent agent executes all corresponding candidate maintenance strategies.
[0155] The determination unit is used to determine the local collaborative maintenance evaluation value corresponding to each first intelligent agent based on the system operation data obtained after each first intelligent agent executes all corresponding candidate maintenance strategies.
[0156] In this embodiment, optionally, the determining unit is specifically configured to:
[0157] Based on the system operation data obtained after each first intelligent agent executes all corresponding candidate maintenance strategies, the system operation risk value of the overall operation system of the special vehicle is determined; based on the system operation risk value of the overall operation system of the special vehicle, the local collaborative maintenance assessment value corresponding to each first intelligent agent is determined.
[0158] In this embodiment, optionally, the determining unit is specifically configured to:
[0159] Based on the system operation risk value of the overall operation system of the special vehicle, the local reward value corresponding to each first intelligent agent is determined; based on the system operation risk value of the overall operation system of the special vehicle and the collaborative influence value between each first intelligent agent and the adjacent intelligent agent, the global reward value corresponding to each first intelligent agent is determined; based on the local reward value and the global reward value corresponding to each first intelligent agent, the local collaborative maintenance evaluation value corresponding to each first intelligent agent is determined.
[0160] In this embodiment, optionally, the second determining module 340 is specifically configured to:
[0161] The candidate maintenance strategy executed by each first agent is obtained when the global maintenance evaluation value of the agent composition system corresponding to the current operating state; and the candidate maintenance strategy executed by the first agent is determined to be the target maintenance strategy corresponding to the corresponding vehicle chassis component.
[0162] In this embodiment, optionally, the first determining module 330 is specifically configured to:
[0163] Obtain the system setting weight corresponding to each first agent; based on the system setting weight corresponding to each first agent and the local collaborative maintenance evaluation value corresponding to each first agent, determine the global maintenance evaluation value of the agent composition system corresponding to the current operating state.
[0164] In this embodiment, optionally, the vehicle chassis components included in the intelligent body composition system correspond to upstream production units and downstream production units respectively, and the upstream production unit corresponds to a buffer inventory, which is used to transfer the stored production products of the upstream production unit to the downstream production unit.
[0165] Also includes: Lifting module.
[0166] The upgrading module is used to upgrade the maintenance level of the vehicle chassis components corresponding to the upstream production unit if the buffer stock does not meet the preset inventory level; if the buffer stock meets the preset inventory level, the maintenance level of the vehicle chassis components corresponding to the downstream production unit is upgraded.
[0167] The multi-agent collaborative vehicle chassis maintenance device provided in the present disclosure can execute the above method embodiments. Its specific implementation principles and technical effects can be found in the above method embodiments, and the present disclosure will not repeat them here.
[0168] The present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0169] The computer device includes a memory 410 and a processor 420 that are connected to each other through a system bus. It should be noted that the figure only shows a computer device with a memory 410 and a processor 420, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.
[0170] Computer devices can be computing devices such as desktop computers, notebooks, PDAs, and cloud servers. Computer devices can interact with users through keyboards, mice, remote controls, touch pads, or voice control devices.
[0171] The memory 410 includes at least one type of readable storage medium, and the readable storage medium includes a non-volatile memory or a volatile memory, for example, a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc., and the RAM may include a static RAM or a dynamic RAM. In some embodiments, the memory 410 may be an internal storage unit of a computer device, for example, a hard disk or a memory of the computer device. In other embodiments, the memory 410 may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card or a flash card (FlashCard) equipped on the computer device. Of course, the memory 410 may also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the memory 410 is generally used to store the operating system and various application software installed on the computer device, such as the program code of the above method. In addition, the memory 410 may also be used to temporarily store various data that have been output or are to be output.
[0172] The processor 420 is generally used to perform the overall operation of the computer device. In this embodiment, the memory 410 is used to store program codes or instructions, the program code includes computer operation instructions, and the processor 420 is used to execute the program codes or instructions stored in the memory 410 or process data, such as running the program code of the above method.
[0173] In this article, the bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus system can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0174] Another embodiment of the present application also provides a computer-readable medium, which may be a computer-readable signal medium or a computer-readable medium. A processor in a computer reads a computer-readable program code stored in the computer-readable medium, so that the processor can execute the functional actions specified in each step or a combination of steps in the above method; and generate a device for implementing the functional actions specified in each block or a combination of blocks in the block diagram.
[0175] Computer-readable media include but are not limited to electronic, magnetic, optical, electromagnetic, infrared memory or semiconductor systems, devices or apparatuses, or any appropriate combination of the foregoing, the memory is used to store program codes or instructions, the program codes include computer operating instructions, and the processor is used to execute the program codes or instructions of the above methods stored in the memory.
[0176] For the definitions of memory and processor, please refer to the description of the aforementioned computer device embodiment and will not be repeated here.
[0177] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0178] Each functional unit or module in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0179] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to perform all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, ROM), random access memory (Random Access Memory, RAM), disk or optical disk and other media that can store program code.
[0180] In the claims, any reference symbols placed between brackets shall not be construed as limiting the claims. The word "comprising" described in the present application does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application can be implemented with the aid of hardware comprising several different elements and with the aid of a suitably programmed computer. In a unit claim that lists a number of devices, several units of these devices may be embodied by the same hardware item. The use of first, second, and third, etc. does not indicate any order, and these words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be understood as limitations on the order of execution.
[0181] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A multi-agent collaborative vehicle chassis maintenance method, characterized in that: include: Acquire an agent composition system corresponding to a special vehicle chassis system, wherein the agent composition system is used to describe an intelligent operation system composed of a plurality of vehicle chassis components, each of the vehicle chassis components corresponds to a first agent, and every two of the first agents are associated and connected by a second agent; When the special vehicle is in the current running state of the agent system, a local collaborative maintenance evaluation value corresponding to each of the first agents is obtained, where the local collaborative maintenance evaluation value corresponding to the first agent is a system reward value obtained after each of the first agents respectively executes a preset maintenance strategy for the corresponding vehicle chassis component; Determine, based on the local collaborative maintenance evaluation value corresponding to each of the first agents, a global maintenance evaluation value of the agent composition system corresponding to the current operating state; Based on the global maintenance evaluation value of the agent composition system corresponding to the current operating state, the target maintenance strategy corresponding to each of the vehicle chassis components is determined, so that each of the first agents performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
2. The method according to claim 1, characterized in that Each of the vehicle chassis components corresponds to a different preset maintenance strategy, and the preset maintenance strategy is determined by a second agent associated with a first agent corresponding to the vehicle chassis component based on a component state of the vehicle chassis component and a component state of an adjacent chassis component; The preset maintenance strategy includes: at least one candidate maintenance strategy; The obtaining, when the special vehicle is in the current operating state of the agent composition system, a local collaborative maintenance evaluation value corresponding to each of the first agents comprises: The special vehicle controls each of the first agents to sequentially execute each candidate maintenance strategy for the corresponding vehicle chassis component under the current operation state of the agent composition system; Obtaining system operation data obtained after each of the first intelligent agents executes all corresponding candidate maintenance strategies; Based on the system operation data obtained after each of the first intelligent agents executes all corresponding candidate maintenance strategies, a local collaborative maintenance evaluation value corresponding to each of the first intelligent agents is determined.
3. The method according to claim 2, characterized in that The determining of the local collaborative maintenance evaluation value corresponding to each of the first agents based on the system operation data obtained after each of the first agents executes all corresponding candidate maintenance strategies includes: Determining a system operation risk value of the overall operation system of the special vehicle based on system operation data obtained after each of the first intelligent agents executes all corresponding candidate maintenance strategies; Based on the system operation risk value of the overall operation system of the special vehicle, a local collaborative maintenance assessment value corresponding to each of the first intelligent agents is determined.
4. The method according to claim 3, characterized in that The determining of the local collaborative maintenance assessment value corresponding to each of the first intelligent agents based on the system operation risk value of the overall operation system of the special vehicle includes: Determining a local reward value corresponding to each of the first intelligent agents based on a system operation risk value of the overall operation system of the special vehicle; Determine a global reward value corresponding to each of the first intelligent agents based on the system operation risk value of the overall operation system of the special vehicle and the synergistic influence value between each of the first intelligent agents and the adjacent intelligent agents; Based on the local reward value and the global reward value corresponding to each of the first agents, a local collaborative maintenance evaluation value corresponding to each of the first agents is determined.
5. The method according to claim 1, characterized in that The determining of a target maintenance strategy corresponding to each of the vehicle chassis components based on the global maintenance evaluation value of the intelligent agent composition system corresponding to the current operating state includes: acquiring a candidate maintenance strategy executed by each of the first agents when the global maintenance evaluation value of the agent-composed system corresponding to the current operating state; The candidate maintenance strategy executed by the first agent is determined to be a target maintenance strategy corresponding to the vehicle chassis component.
6. The method according to claim 1, characterized in that The determining, based on the local collaborative maintenance evaluation value corresponding to each of the first agents, a global maintenance evaluation value of the agent composition system corresponding to the current operating state comprises: Obtaining a system-set weight corresponding to each of the first agents; Based on the system setting weight corresponding to each of the first agents and the local collaborative maintenance evaluation value corresponding to each of the first agents, a global maintenance evaluation value of the agent composition system corresponding to the current operating state is determined.
7. The method according to claim 1, characterized in that The vehicle chassis components included in the intelligent agent composition system correspond to upstream production units and downstream production units respectively, and the upstream production unit corresponds to a buffer stock, and the buffer stock is used to transfer the stored production products of the upstream production unit to the downstream production unit; The method further comprises: If the buffer stock does not meet the preset stock volume, the maintenance level of the vehicle chassis component corresponding to the upstream production unit is increased; If the buffer stock meets the preset stock volume, the maintenance level of the vehicle chassis component corresponding to the downstream production unit is increased.
8. A multi-agent collaborative vehicle chassis maintenance device, characterized in that: include: A first acquisition module is used to acquire an agent composition system corresponding to a special vehicle chassis system, wherein the agent composition system is used to describe an intelligent operation system composed of a plurality of vehicle chassis components, each of the vehicle chassis components corresponds to a first agent, and every two of the first agents are associated and connected by a second agent; A second acquisition module is used to acquire a local collaborative maintenance evaluation value corresponding to each of the first agents when the special vehicle is in the current operating state of the agent composition system, wherein the local collaborative maintenance evaluation value corresponding to the first agent is a system reward value obtained after each of the first agents respectively executes a preset maintenance strategy for the corresponding vehicle chassis component; A first determination module, configured to determine a global maintenance evaluation value of the agent composition system corresponding to a current operating state based on a local collaborative maintenance evaluation value corresponding to each of the first agents; The second determination module is used to determine the target maintenance strategy corresponding to each of the vehicle chassis components based on the global maintenance evaluation value of the intelligent agent composition system corresponding to the current operating state, so that each of the first intelligent agents performs component maintenance on the corresponding vehicle chassis component based on the corresponding target maintenance strategy.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, a vehicle chassis maintenance method of multi-agent collaboration as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-agent collaborative vehicle chassis maintenance method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
A multi-agent based bi-directional joint scheduling strategy decision method for maintenance resources
CN109190995A
Task scheduling method based on multi-agent auxiliary edge cloud server
CN116974751A
Task reliability modeling and simulation method and system considering guarantee resources
CN118095903A
Intelligent maintenance decision-making method and system for underwater production control system
CN118096121A
Equipment optimal maintenance strategy searching method and system based on reinforcement learning
CN118941275A
Cited By
Vehicle formation communication resource allocation method and device and electronic equipment
CN120857283A
Vehicle platoon communication resource allocation method and device, and electronic equipment
CN120857283B