An rl control method for driving mode switching of an amphibious vehicle based on fuzzy logic and mfac
By combining fuzzy logic and MFAC with reinforcement learning control methods, a smooth, intelligent, adaptive, and dynamically stable switching of the driving mode of amphibious vehicles was achieved. This solved the stability and energy efficiency problems of the switching process in existing technologies, and improved the driving comfort and equipment reliability of the vehicles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-10
AI Technical Summary
Existing amphibious vehicles suffer from problems such as abrupt switching of drive modes, poor stability, failure to consider dynamic operating conditions, and low energy efficiency.
A RL control method based on fuzzy logic and MFAC is adopted. By monitoring the vehicle status in real time, the mode switching coefficient is calculated using IF-THEN rules and membership functions. Combined with model-free adaptive control and deep reinforcement learning, dynamic continuous adjustment and adaptive optimization are achieved.
It achieves smooth, intelligent, adaptive, and dynamic stable switching of amphibious vehicle drive modes, improving the stability of switching and energy utilization efficiency, and reducing mechanical shock and system losses.
Smart Images

Figure CN122354532A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle control technology, and in particular to a drive mode switching control method for amphibious vehicles. Specifically, it is a hybrid model-free intelligent control method based on fuzzy logic and model-free adaptive control (MFAC) combined with reinforcement learning (RL) to achieve smooth and dynamically stable switching between land drive and water propulsion modes. Background Technology
[0002] Amphibious vehicles, due to their unique mobility, have significant application value in fields such as emergency rescue, military reconnaissance, and ecological monitoring. These vehicles need to operate on two completely different media—land and water—and therefore are typically equipped with two drive systems: a wheeled or tracked drive system for land travel, and a waterjet propulsion or propeller propulsion system for water travel. Currently, the drive mode switching of amphibious vehicles mostly relies on simple threshold judgment or manual switching. For example, when the vehicle's sensors detect that the water depth exceeds a predetermined threshold, the controller disengages the land drive clutch and simultaneously engages the water propulsion system. This switching method has significant drawbacks: 1. Abrupt switching process and poor stability: Hard switching mode will cause sudden changes in vehicle driving force and braking force, causing vehicle impact, pitching or yaw, which seriously affects ride comfort and handling stability, and may even pose a risk of rollover in complex water entry and exit conditions.
[0003] 2. Failure to consider dynamic operating conditions: Simple threshold judgment cannot adapt to dynamic environmental factors such as vehicle speed, attitude, and water surface waves. For example, when entering water on a bumpy road, the wheels may be briefly submerged and then exposed, causing the drive system to switch repeatedly at high frequency, which aggravates system wear and instability.
[0004] 3. Low energy efficiency: Abrupt switching cannot optimize and coordinate the power system, which may lead to power interruption or a sharp deterioration in engine operating conditions, reducing energy utilization efficiency. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, this invention provides an RL control method for switching the drive mode of amphibious vehicles based on fuzzy logic and MFAC, aiming to achieve smooth, intelligent, adaptive, and dynamically stable switching of the drive mode of amphibious vehicles, and improve the stability and energy utilization efficiency of the switching, thereby solving the stability and smoothness problems in the drive mode switching process of amphibious vehicles.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: The present invention provides an RL control method for switching drive modes of amphibious vehicles based on fuzzy logic and MFAC, characterized by the following steps: Step 1: Sensors on the amphibious vehicle monitor the current situation in real time. Draft at any time Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value ; Step 2: Based on the data collected in Step 1, calculate the current status of the amphibious vehicle. Membership in real-time water mode Membership degree in land mode Membership degree in transition mode : Step 3: Based on the IF-THEN rule, apply the current rules to amphibious vehicles. Membership in real-time water mode Membership degree in land mode Membership degree in transition mode Perform constraint corrections to obtain the current... Membership degree in the water mode after time constraint correction Membership degree in land mode Membership degree in transition mode ; Step 4: Based on the current Membership degree in the water mode after time constraint correction Membership degree in land mode Membership degree in transition mode Calculate the current Mode switching coefficient at any time ,like , representing the current If the amphibious vehicle is always in pure land mode, , representing the current The amphibious vehicle is always in pure water mode. , representing the current The amphibious vehicle is always in transition mode. Step 5: Combining the dynamic and continuous adjustment requirements for the switching of the amphibious vehicle's drive mode and the requirements for adaptive operation, the model-free adaptive control model MFAC and deep deterministic policy gradient reinforcement learning RL are used to perform adaptive optimization of the amphibious vehicle's state vector to obtain the amphibious vehicle's control vector.
[0007] The RL control method for switching drive modes of amphibious vehicles based on fuzzy logic and MFAC described in this invention is characterized in that step 2 includes the following steps: Step 2.1: Calculate the current... Draft in real-time water mode Membership function Road surface adhesion coefficient Membership function The ratio of engine torque-speed ratio to its rated value Membership function ; Step 2.2: Calculate the current... Draft in land mode Membership function Road surface adhesion coefficient Membership function The ratio of engine torque-speed ratio to its rated value Membership function ; Step 2.3: Calculate the current... Draft in time transition mode Membership function Road surface adhesion coefficient Membership function The ratio of engine torque-speed ratio to its rated value Membership function ; Step 2.4: Use equation (1) to obtain the current status of amphibious vehicles. The membership degree of the vehicle to the i-th mode at any given time : + , (1) In equation (1), Represents draft Membership degree of the i-th mode, Represents the road surface adhesion coefficient Membership degree of the i-th mode, Represents the engine torque-speed ratio and its rated value. The membership degree of the ratio to the i-th mode; , , Represents the normalization coefficient, satisfying .
[0008] Furthermore, step 4 includes the following steps: Step 4.1: Membership degree of the i-th pattern After normalization, the normalized membership degree of the amphibious vehicle to the i-th mode is obtained. ; Step 4.2: Calculate the current value using equation (2). Mode switching coefficient at any time : (2) In equation (2), This represents a transitional bias in weighting; This represents the membership degree of amphibious vehicles to the land-based model after normalization. This represents the degree of membership of a vehicle to the transition mode after normalization.
[0009] Furthermore, step 5 includes the following steps: Step 5.1: Construct a deep neural network consisting of an Actor network and a Critic network, and make the actions of the Actor network adjustable parameters; Build the current Vehicle state vector at time 1 ,in, Indicates the current longitudinal velocity at time, Indicates the current Lateral velocity at any moment Indicates the current The yaw rate at any given moment; Indicates transpose. Build the current Control vector at time ,in, Representing the present Engine torque at any given moment Representing the present The thruster power at any given moment, Representing the present The wheel steering angle at any given moment, Representing the present The thruster deflection angle at any given moment; Initialize current The pseudo-partial derivative matrix at time t ; Step 5.2: Set the current Mode switching coefficient at any time State tracking error The input is processed in the Actor network, and the current value is output. Adjustable parameters of action at any moment Including: synovial gain 2 weighting coefficients and Step size factor and forgetting factor ,in, Indicates preset The reference status of the amphibious vehicle at any given time; Step 5.3: Use equation (3) to obtain the current... Control vector at time Increment : (3) In equation (3), Represents a constant. represent The reference status of the amphibious vehicle at any given time; Step 5.4: Calculate the current Multidimensional weighted reward function at time step ; Step 5.5: The Critic network is based on... To obtain the current The state at any given moment—action value Used to guide the Actor network on adjustable parameters Perform the update and obtain the updated adjustable parameters. ; Step 5.6: Based on the updated adjustable parameters Using MFAC to obtain Control vector at time Used for the current Vehicles are controlled at specific times; Step 5.7: Obtain using equation (4) Vehicle state vector at time 1 : (4) In equation (4), Representing the present The perturbation term at any given time; Step 5.8: Use equation (5) to... Update and get The pseudo-partial derivative matrix at time t : (5) In equation (5), Represents the Euclidean norm; Step 5.9: Obtain the results following steps 1-4. Mode switching coefficient at any time ,Will Assign to ,Will +1 is assigned to Then return to step 5.2 and execute sequentially until the vehicle stops moving, thereby realizing the rolling real-time adaptive control of the amphibious vehicle.
[0010] Furthermore, step 5.4 includes the following steps: Step 5.4.1: Construct the current equation using equation (4) Real-time status tracking rewards : (4) In equation (4), This represents the first adjustment coefficient; Step 5.4.2: Construct the current equation using equation (5) Time-of-use transition smoothness reward : (5) In equation (5), This represents the second adjustment coefficient; Step 5.4.3: Construct the current equation using equation (6) Energy reward at any moment : (6) In equation (6), Representing the present Engine speed at any given moment. This represents the third adjustment coefficient; Step 5.4.4: Construct the current equation using equation (7) Moment-time stability reward : (7) In equation (7), Represents the maximum yaw rate. Represents the minimum lateral velocity. Represents the maximum lateral velocity; Step 5.4.5: Construct the current equation using equation (8) Multidimensional weighted reward function at time step : (8) In equation (8), This represents four weighting coefficients.
[0011] The present invention provides an electronic device, including a memory and a processor, characterized in that the memory is used to store a program supporting the processor in performing the method described therein, and the processor is configured to execute the program stored in the memory.
[0012] The present invention discloses a computer-readable storage medium storing a computer program, characterized in that the computer program is executed by a processor to perform the steps of the method described thereon.
[0013] Compared with existing technologies, the present invention has the following beneficial effects: 1. This invention abandons the traditional "land-water" 0 / 1 hard switch control mode. Based on the fuzzy logic algorithm, it outputs continuously changing mode switching coefficients, transforming the power system switching process into continuous weight adjustment. This eliminates the mechanical shock and power interruption caused by hard switching, significantly reduces the wear of vehicle transmission components, and improves driving comfort and equipment reliability.
[0014] 2. The hybrid model-free intelligent control method of the present invention, which combines model-free adaptive control (MFAC) with reinforcement learning (RL), can solve the problem of difficult parameter tuning in traditional MPC control methods and significantly improve the dynamic stability and safety of the switching process.
[0015] 3. This invention employs a fuzzy logic algorithm, thus possessing the ability to handle uncertainties and nonlinear information, and can effectively adapt to complex water-land boundary conditions such as waves, water skidding, and rugged riverbanks. By fusing membership functions with multiple parameters (draft, adhesion coefficient, and the ratio of engine torque-speed ratio to rated value) and using IF-THEN rule reasoning, it avoids the frequent oscillations and switching of traditional threshold control under critical conditions, ensuring the system's control robustness under extreme conditions.
[0016] 4. This invention constructs a multi-objective optimization function to optimize the power distribution ratio of track drive and propeller propulsion in real time, reducing unnecessary power output and redundant energy consumption during mode switching, and achieving a synergistic improvement in stability and energy efficiency. Attached Figure Description
[0017] Figure 1 A schematic diagram of a tracked amphibious vehicle. Figure 2 This is a block diagram of the overall architecture of the control system described in this invention; Figure 3 This is a flowchart illustrating the internal structure and working principle of the MFAC+DDPG controller. Detailed Implementation
[0018] The following description, with reference to the accompanying drawings in the embodiments of the present invention, uses a tracked amphibious vehicle (total weight > 5t) as an example, whose structure is as follows: Figure 1 As shown, the technical solution will be further described in detail.
[0019] In this embodiment, an RL control method for switching the drive mode of an amphibious vehicle based on fuzzy logic and MFAC is described in the following process: Figure 2 The diagram illustrates a method for achieving dynamic and stable switching between land-based drive and water propulsion, and includes the following steps: Step 1: The system starts working. Sensors monitor the draft in real time. Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value .
[0020] Step 2: Based on the data collected in Step 1, calculate the current status of the amphibious vehicle. Membership in real-time water mode Membership degree in land mode Membership degree in transition mode The segmentation points of each piecewise function are determined according to the fuzzy rule set.
[0021] Step 2.1: Draft in surface mode Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value The membership functions are calculated as shown in equations (1) to (3): (1) In equation (1), Represents draft. This represents the membership degree of the draft at time k to the surface pattern; (2) In equation (2), Represents the road surface adhesion coefficient. The degree of membership of the road surface adhesion coefficient at time k to the waterborne mode; (3) In equation (3), This represents the ratio of the torque-speed ratio to the rated value. This represents the membership degree of the engine torque-speed ratio at time k to the rated value in the waterborne mode.
[0022] Step 2.2: Draft in Land Mode Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value The membership functions are calculated according to equations (4) to (6): (4) In equation (4), This represents the membership degree of the current draft at time k to the land pattern; (5) In equation (5), This represents the membership degree of the road surface adhesion coefficient at time k to the land model; (6) In equation (6), This represents the membership degree of the engine torque-speed ratio to the rated value at time k to the land mode.
[0023] Step 2.3: Draft in transition mode Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value The membership functions are calculated as shown in equations (7) to (9): (7) In equation (7), The degree of membership of the draft at time k to the transition mode; (8) In equation (8), This represents the membership degree of the road surface adhesion coefficient at time k to the transition mode; (9) In equation (9), This represents the membership degree of the transition mode to the ratio of the engine torque-speed ratio to the rated value at the current time k.
[0024] Step 2.4: The final membership degree of the vehicle to each mode is determined by the weighted average method, i.e., equation (10): + ( (10) In equation (10), This represents the membership degree of the vehicle at time k to the i-th mode. This represents the membership degree of the draft at time k to the i-th mode. This represents the membership degree of the road surface adhesion coefficient at time k to the i-th mode. The degree of membership of the engine torque-speed ratio to the rated value at time k to the i-th mode; , , Represents the normalization coefficient, satisfying Recommended , Take 0.3, Take 0.5.
[0025] Step 3: Based on the IF-THEN rule, apply the current rules to amphibious vehicles. Membership in real-time water mode Membership degree in land mode Membership degree in transition mode Perform constraint corrections to obtain the current... Membership degree in the water mode after time constraint correction Membership degree in land mode Membership degree in transition mode This allows the output value of the membership function to preferentially match the triggering conditions of high-priority rules, while weakening unreasonable membership values without rule support.
[0026] Step 3.1: Define the fuzzy subsets of each input and output variable. The fuzzy subsets of the input variables are shown in Table 1: Table 1
[0027] The fuzzy subset of the output variable (pattern) is “land”, “water”, and “transition”.
[0028] Step 3.2: Establish the core IF-THEN rule set. The rules are ordered by priority (core rules > modification rules > boundary rules). The core rules are R1-R3, which are dominated by draft; the modification rules are R4-R12, which are modified in combination with auxiliary variables; and the boundary rules are R13-R15, which are used to handle critical fuzzy conditions. The specific contents are shown in Table 2. Table 2
[0029] If multiple rules are triggered simultaneously and the results conflict, the multi-condition rule will be executed first rather than the single-condition rule, because multi-condition rules are usually closer to the actual working conditions.
[0030] Step 3.3: After constraint correction, the draft of the surface mode Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value The membership functions are calculated as shown in equation (11): (11) In equation (11), Draft in surface mode The rule constraint factor is 1.4 when R4 is triggered, 0.6 when R12 is triggered, and 1.0 in other cases. To determine the membership degree of the draft of the surface mode at time k after constraint correction.
[0031] (12) In equation (12), The road surface adhesion coefficient in the waterborne mode The rule constraint factor is 1.5 when R8 is triggered, 0.5 when R11 is triggered, and 1.0 in other cases. The membership degree of the surface adhesion coefficient in the waterborne mode at time k after constraint correction.
[0032] (13) In equation (13), The ratio of engine torque-speed ratio to rated value in water mode The rule constraint factor is 1.5 when R8 is triggered, 0.5 when R14 is triggered, and 1.0 in other cases. The membership degree of the ratio of the engine torque-speed ratio to the rated value at time k in the water mode after constraint correction.
[0033] Draft in land mode Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value The membership functions are calculated as shown in equation (14): (14) In equation (14), Draft for land mode The rule constraint factor is set to 1.2 when R5 is triggered, 0.7 when R11 is triggered, and 1.0 in all other cases. To constrain the membership degree of the land pattern draft at time k after the correction.
[0034] (15) In equation (15), For the road surface adhesion coefficient in land mode The rule constraint factor is 1.2 when R5 is triggered, 1.5 when R9 is triggered, and 1.0 in other cases. This is the membership degree of the road surface adhesion coefficient in the land mode at time k after constraint correction.
[0035] (16) In equation (16), The ratio of engine torque-speed ratio to rated value in land mode. The rule constraint factor is 1.2 when R7 is triggered, 1.5 when R9 is triggered, and 1.0 in other cases. To constrain the correction, the membership degree of the ratio of the torque-speed ratio of the land-mode engine to its rated value at time k.
[0036] Draft of transition mode Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value The membership functions are calculated as shown in equation (17): (17) In equation (17), Draft for transition mode The rule constraint factor is 0.5 when R4, R5, R6, and R7 are triggered, 1.3 when R11, R12, and R13 are triggered, and 1.0 in other cases. To constrain the correction, the membership degree of the draft of the transition mode at time k.
[0037] (18) In equation (18), For the transition mode, the road surface adhesion coefficient The rule constraint factor is 0.6 when R8 and R9 are triggered, 1.2 when R13 is triggered, and 1.0 in other cases. This is the membership degree of the transition mode road surface adhesion coefficient at time k after constraint correction.
[0038] (19) In equation (19), The ratio of engine torque-speed ratio to rated value in transition mode The rule constraint factor is 0.6 when R8 and R9 are triggered, 1.3 when R13 and R14 are triggered, and 1.0 in other cases. The membership degree of the ratio of the transition mode engine torque-speed ratio to the rated value at the current time k after constraint correction.
[0039] Step 3.4: The final membership degree of the vehicle to the i-th mode It is calculated using the same weighted average method as Equation (10).
[0040] Step 4: Based on the current Membership degree in the water mode after time constraint correction Membership degree in land mode Membership degree in transition mode Calculate the current Mode switching coefficient at any time ,like , representing the current If the amphibious vehicle is always in pure land mode, , representing the current The amphibious vehicle is always in pure water mode. , representing the current The amphibious vehicle is in transition mode at all times.
[0041] Step 4.1: Normalize the membership degrees under the three modes using equation (20) to ensure the logical completeness of fuzzy reasoning: (20) In equation (20), This represents the normalized membership degree of the vehicle to the i-th mode at time k.
[0042] Step 4.2: Mode Switching Coefficient Calculate according to formula (21): (twenty one) In equation (21), This represents a biased transition weight. In actual working conditions, the transition mode has a bias, such as in shallow water or swamp scenarios. Values higher than 0.5 are suitable for water-filled surfaces or riverbanks. Values below 0.5 are generally used for water-land interface scenarios. ; This represents the normalized membership degree of the amphibious vehicle to the land mode at time k. This represents the membership degree of the vehicle at time k to the normalized transition mode.
[0043] Step 5: Combining the dynamic and continuous adjustment requirements of amphibious vehicle drive mode switching and the requirements of adaptive operation, the state vector of the amphibious vehicle is adaptively optimized using a model-free adaptive control algorithm and deep deterministic policy gradient reinforcement learning to obtain an executable amphibious control vector. The process is as follows: Figure 3 As shown.
[0044] Step 5.1: Construct a deep neural network consisting of an Actor network and a Critic network, making the actions of the Actor network adjustable parameters. Including: synovial gain 2 weighting coefficients and Step size factor and forgetting factor ,in Represents synovial gain used for , , Represents two weighting coefficients used for adaptation. Bias under different working conditions (e.g., shallow water) Smaller than average Increase; water-crossing road surface Too big (increase) This represents the step size factor, whose value is between 0 and 1, used to control the estimation speed. This represents the forgetting factor, which has a value greater than 0. It is used to suppress the influence of historical data to improve real-time performance. and The value should take into account both the estimated speed and the adaptability to the operating conditions.
[0045] Build the current Vehicle state vector at time 1 ,in, Indicates the current longitudinal velocity at time, Indicates the current Lateral velocity at any moment Indicates the current The yaw rate at any given moment; This indicates transpose.
[0046] Build the current Control vector at time ,in, Representing the present Engine torque at any given moment Representing the present The thruster power at any given moment, Representing the present The wheel steering angle at any given moment, Representing the present The thruster deflection angle at any given moment; Initialize current The pseudo-partial derivative matrix at time t .
[0047] Step 5.2: Set the current Mode switching coefficient at any time State tracking error The input is processed in the Actor network, and the current value is output. Adjustable parameters of action at any moment ,in, Indicates preset The reference status of the amphibious vehicle at any given time.
[0048] Step 5.3: To ensure the robustness and tracking performance of the control, a sliding mode control law based on MFAC is designed to make the actual vehicle state track the reference state, while introducing a mode switching coefficient. Coupling land and water dynamic characteristics, MFAC is based on a preset pseudo-partial derivative matrix. And use equation (22) to obtain the current Control vector at time Increment : (twenty two) In equation (22), represent Vehicle reference status at +1 moment This represents a constant, used to avoid the denominator being zero.
[0049] Step 5.4: Build the current Multidimensional weighted reward function at time step : Step 5.4.1: Construct the current equation using equation (23) Real-time status tracking rewards : (twenty three) In equation (23), This represents the first adjustment coefficient.
[0050] Step 5.4.2: Construct the current equation using equation (24) Time-of-use transition smoothness reward : (twenty four) In equation (24), This represents the second adjustment coefficient.
[0051] Step 5.4.3: Construct the current equation using equation (25) Energy reward at any moment : (25) In equation (25), Representing the present Engine speed at any given moment. This represents the third adjustment coefficient.
[0052] Step 5.4.4: Construct the current equation using equation (26) Moment-time stability reward : (26) In equation (26), Represents the maximum yaw rate. Represents the minimum lateral velocity. This represents the maximum lateral velocity.
[0053] Step 5.4.5: Construct the current equation using equation (27) Multidimensional weighted reward function at time step : (27) In equation (27), This represents four weighting coefficients.
[0054] Step 5.5: The Critic network is based on... To obtain the current The state at any given moment—action value Used to guide the Actor network on adjustable parameters Perform the update and obtain the updated adjustable parameters. .
[0055] Step 5.6: Based on the updated adjustable parameters Using MFAC to obtain Control vector at time Used for the current Vehicles are controlled at specific times.
[0056] Step 5.7: For the drive system of the amphibious vehicle, define its pseudo-linear relationship between input and output, and obtain it using equation (28). Vehicle state vector at time 1 : (28) In equation (28), Representing the present Small disturbances caused by environmental factors such as road surface and waves at all times.
[0057] Step 5.8: Employ a recursive estimation algorithm with a forgetting factor to update in real time. Based on the dynamic changes of the system during the adaptation mode switching process, equation (5) is used to analyze the changes. Update to obtain The pseudo-partial derivative matrix at time t : (29) In equation (29), Represents step size factor Represents the forgetting factor. It represents the Euclidean norm.
[0058] Step 5.9: Obtain the results following steps 1-4. Mode switching coefficient at any time At the same time Assign to ,Will +1 is assigned to Then return to step 5.2 and execute sequentially until the vehicle stops moving, thereby realizing the rolling real-time adaptive control of the amphibious vehicle.
[0059] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.
[0060] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
Claims
1. A method for RL control of amphibious vehicle drive mode switching based on fuzzy logic and MFAC, characterized in that, Includes the following steps: Step 1: Sensors on the amphibious vehicle monitor the current situation in real time. Draft at any time Road surface adhesion coefficient The ratio of engine torque-speed ratio to its rated value ; Step 2: Based on the data collected in Step 1, calculate the current status of the amphibious vehicle. Membership in real-time water mode Membership degree in land mode Membership degree in transition mode : Step 3: Based on the IF-THEN rule, apply the current rules to amphibious vehicles. Membership in real-time water mode Membership degree in land mode Membership degree in transition mode Perform constraint corrections to obtain the current... Membership degree in the water mode after time constraint correction Membership degree in land mode Membership degree in transition mode ; Step 4: Based on the current Membership degree in the water mode after time constraint correction Membership degree in land mode Membership degree in transition mode Calculate the current Mode switching coefficient at any time ,like , representing the current If the amphibious vehicle is always in pure land mode, , representing the current The amphibious vehicle is always in pure water mode. , representing the current The amphibious vehicle is always in transition mode. Step 5: Combining the dynamic and continuous adjustment requirements for the switching of the amphibious vehicle's drive mode and the requirements for adaptive operation, the model-free adaptive control model MFAC and deep deterministic policy gradient reinforcement learning RL are used to perform adaptive optimization of the amphibious vehicle's state vector to obtain the amphibious vehicle's control vector.
2. The RL control method for switching drive modes of amphibious vehicles based on fuzzy logic and MFAC according to claim 1, characterized in that, Step 2 includes the following steps: Step 2.1: Calculate the current... Draft in real-time water mode Membership function Road surface adhesion coefficient Membership function The ratio of engine torque-speed ratio to its rated value Membership function ; Step 2.2: Calculate the current... Draft in land mode Membership function Road surface adhesion coefficient Membership function The ratio of engine torque-speed ratio to its rated value Membership function ; Step 2.3: Calculate the current... Draft in time transition mode Membership function Road surface adhesion coefficient Membership function The ratio of engine torque-speed ratio to its rated value Membership function ; Step 2.4: Use equation (1) to obtain the current status of amphibious vehicles. The membership degree of the vehicle to the i-th mode at any given time : + , ; (1) In equation (1), Represents draft Membership degree of the i-th mode, Represents the road surface adhesion coefficient Membership degree of the i-th mode, Represents the engine torque-speed ratio and its rated value. The membership degree of the ratio to the i-th mode; , , Represents the normalization coefficient, satisfying .
3. The RL control method for switching drive modes of amphibious vehicles based on fuzzy logic and MFAC according to claim 2, characterized in that, Step 4 includes the following steps: Step 4.1: Membership degree of the i-th pattern After normalization, the normalized membership degree of the amphibious vehicle to the i-th mode is obtained. ; Step 4.2: Calculate the current value using equation (2). Mode switching coefficient at any time : (2) In equation (2), This represents a transitional bias in weighting; This represents the membership degree of amphibious vehicles to the land-based model after normalization. This represents the degree of membership of a vehicle to the transition mode after normalization.
4. The RL control method for switching drive modes of amphibious vehicles based on fuzzy logic and MFAC according to claim 3, characterized in that, Step 5 includes the following steps: Step 5.1: Construct a deep neural network consisting of an Actor network and a Critic network, and make the actions of the Actor network adjustable parameters; Build the current Vehicle state vector at time 1 ,in, Indicates the current longitudinal velocity at time, Indicates the current Lateral velocity at any moment Indicates the current The yaw rate at any given moment; Indicates transpose. Build the current Control vector at time ,in, Representing the present Engine torque at any given moment Representing the present The thruster power at any given moment, Representing the present The wheel steering angle at any given moment, Representing the present The thruster deflection angle at any given moment; Initialize current The pseudo-partial derivative matrix at time t ; Step 5.2: Set the current Mode switching coefficient at any time State tracking error The input is processed in the Actor network, and the current value is output. Adjustable parameters of action at any moment Including: synovial gain 2 weighting coefficients and Step size factor and forgetting factor ,in, Indicates preset The reference status of the amphibious vehicle at any given time; Step 5.3: Use equation (3) to obtain the current... Control vector at time Increment : (3) In equation (3), Represents a constant. represent The reference status of the amphibious vehicle at any given time; Step 5.4: Calculate the current Multidimensional weighted reward function at time step ; Step 5.5: The Critic network is based on... To obtain the current The state at any given moment—action value Used to guide the Actor network on adjustable parameters Perform the update and obtain the updated adjustable parameters. ; Step 5.6: Based on the updated adjustable parameters Using MFAC to obtain Control vector at time Used for the current Vehicles are controlled at specific times; Step 5.7: Obtain using equation (4) Vehicle state vector at time 1 : (4) In equation (4), Representing the present The perturbation term at any given time; Step 5.8: Use equation (5) to... Update and get The pseudo-partial derivative matrix at time t : (5) In equation (5), Represents the Euclidean norm; Step 5.9: Obtain the results following steps 1-4. Mode switching coefficient at any time ,Will Assign to ,Will +1 is assigned to Then return to step 5.2 and execute sequentially until the vehicle stops moving, thereby realizing the rolling real-time adaptive control of the amphibious vehicle.
5. The RL control method for switching drive modes of amphibious vehicles based on fuzzy logic and MFAC according to claim 4, characterized in that, Step 5.4 includes the following steps: Step 5.4.1: Construct the current equation using equation (4) Real-time status tracking rewards : (4) In equation (4), This represents the first adjustment coefficient; Step 5.4.2: Construct the current equation using equation (5) Time-of-use transition smoothness reward : (5) In equation (5), This represents the second adjustment coefficient; Step 5.4.3: Construct the current equation using equation (6) Energy reward at any moment : (6) In equation (6), Representing the present Engine speed at any given moment. This represents the third adjustment coefficient; Step 5.4.4: Construct the current equation using equation (7) Moment-time stability reward : (7) In equation (7), Represents the maximum yaw rate. Represents the minimum lateral velocity. Represents the maximum lateral velocity; Step 5.4.5: Construct the current equation using equation (8) Multidimensional weighted reward function at time step : (8) In equation (8), This represents four weighting coefficients.
6. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports a processor in executing the method of any one of claims 1-5, the processor being configured to execute the program stored in the memory.
7. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method according to any one of claims 1-5.