Parameter-adjustable threshing system and intelligent regulation and control method thereof
By applying an intelligent threshing model based on reinforcement learning in the harvester, the parameters of the threshing system are adjusted in real time, and the problem of agricultural machinery operators need to get off the vehicle to check and adjust frequently, improving the threshing efficiency and quality.
Patent Information
- Application Number
- CN202510357299.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
Smart Images

Figure CN120202833A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of agricultural machinery and mechanical intelligent control technology, and particularly relates to a parameter-adjustable threshing system and an intelligent control method thereof. Background Art
[0002] During the operation of the harvester, the operator needs to get off the vehicle regularly to check the operation performance of the harvester, such as the loss rate and the breakage rate. If one or more of these parameters do not meet the operation requirements, it is necessary to gradually check and adjust parameters such as the roller speed and the concave clearance of the harvester according to past adjustment experience, so that the operation performance of the harvester can meet the actual operation requirements.
[0003] However, this process is not only time-consuming and laborious, but also depends on the personal experience and skill level of the operator, and the adjustment effect often has uncertainty. Summary of the Invention
[0004] The purpose of the present application is to provide a parameter-adjustable threshing system and an intelligent control method thereof. Through the intelligent control method, the operation parameters of the threshing system can be adjusted in real time, thereby improving the threshing efficiency and operation quality during the operation of the harvester.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] In a first aspect, the present application provides an intelligent control method for a parameter-adjustable threshing system. The parameter-adjustable threshing system includes a roller, a concave screen, a concave screen clearance adjustment mechanism, an upper cover of the roller, and a roller speed adjustment mechanism. The intelligent control method includes:
[0007] Obtain the current operation data of the parameter-adjustable threshing system and the previous operation reward. The operation data includes the roller speed, the concave clearance, and the operation speed. The operation reward includes the breakage rate and the threshing loss rate of crop grains.
[0008] Input the operation data and the operation reward into an intelligent threshing model based on reinforcement learning to obtain the current control action of the parameter-adjustable threshing system. The control action includes increasing the roller speed, decreasing the roller speed, increasing the concave clearance, decreasing the concave clearance, and maintaining the state. The intelligent threshing model based on reinforcement learning is used to be trained according to historical operation data and a preset reward function, and aims to maximize the long-term operation reward and output the optimal control action.
[0009] Optionally, after obtaining the current operation data of the parameter-adjustable threshing system and the previous operation reward, it further includes:
[0010] Perform state division on the roller speed, the concave clearance, and the operation speed in the operation data respectively, specifically including:
[0011] According to the rotational speed range of the drum, it is divided into a high rotational speed state, a medium-high rotational speed state, a medium rotational speed state, a medium-low rotational speed state, and a low rotational speed state;
[0012] According to the size range of the concave clearance, it is divided into a large clearance state, a medium-large clearance state, a medium clearance state, a medium-small clearance state, and a small clearance state;
[0013] According to the speed range of the operating speed, it is divided into a high speed state, a medium-high speed state, a medium speed state, a medium-low speed state, and a low speed state.
[0014] Optionally, after obtaining the current operating data and the previous operating reward of the parameter-adjustable threshing system, it further includes:
[0015] Performing state division on the operating reward, specifically including:
[0016] According to the output range of the breakage rate, it is divided into a high breakage rate state, a medium breakage rate state, and a low breakage rate state;
[0017] According to the output range of the threshing loss rate, it is divided into a high loss rate state, a medium loss rate state, and a low loss rate state.
[0018] Optionally, inputting the operating data and the operating reward into an intelligent threshing model based on reinforcement learning to obtain the current regulation action of the parameter-adjustable threshing system, specifically including:
[0019] Inputting the operating data and the operating reward into an intelligent threshing model based on reinforcement learning to obtain the current operating reward under different regulation actions;
[0020] For each regulation action, based on the look-up table method, determine the reward value of the parameter-adjustable threshing system according to the current operating reward under the regulation action and the input previous operating reward;
[0021] Determine the regulation action corresponding to the maximum value of the reward value of the parameter-adjustable threshing system under each regulation action as the current regulation action.
[0022] Optionally, the calculation formula for the reward value of the parameter-adjustable threshing system is:
[0023] R = R1 + R2;
[0024] Wherein, R is the threshing system reward value, R1 is the breakage state change reward value, and R2 is the loss state change reward value.
[0025] Optionally, the calculation method for the breakage state change reward value is:
[0026] If the breakage rate in the current job reward is in the high breakage rate state and the breakage rate in the previous job reward was also in the high breakage rate state, the breakage state change reward value is -2;
[0027] If the breakage rate in the current job reward is in the high breakage rate state and the breakage rate in the previous job reward was in the medium breakage rate state, the breakage state change reward value is -2;
[0028] If the breakage rate in the current job reward is in the high breakage rate state and the breakage rate in the previous job reward was in the low breakage rate state, the breakage state change reward value is -4;
[0029] If the breakage rate in the current job reward is in the medium breakage rate state and the breakage rate in the previous job reward was in the high breakage rate state, the breakage state change reward value is 1;
[0030] If the breakage rate in the current job reward is in the medium breakage rate state and the breakage rate in the previous job reward was also in the medium breakage rate state, the breakage state change reward value is -1;
[0031] If the breakage rate in the current job reward is in the medium breakage rate state and the breakage rate in the previous job reward was in the low breakage rate state, the breakage state change reward value is -2;
[0032] If the breakage rate in the current job reward is in the low breakage rate state and the breakage rate in the previous job reward was in the high breakage rate state, the breakage state change reward value is 2;
[0033] If the breakage rate in the current job reward is in the low breakage rate state and the breakage rate in the previous job reward was in the medium breakage rate state, the breakage state change reward value is 1;
[0034] If the breakage rate in the current job reward is in the low breakage rate state and the breakage rate in the previous job reward was also in the low breakage rate state, the breakage state change reward value is 0.
[0035] Optionally, the calculation method of the loss state change reward value is as follows:
[0036] If the threshing loss rate in the current job reward is in the high loss rate state and the threshing loss rate in the previous job reward was also in the high loss rate state, the loss state change reward value is -2;
[0037] If the threshing loss rate in the current job reward is in the high loss rate state and the threshing loss rate in the previous job reward was in the medium loss rate state, the loss state change reward value is -2;
[0038] If the threshing loss rate in the current job reward is in the high loss rate state and the threshing loss rate in the previous job reward was in the low loss rate state, the loss state change reward value is -4;
[0039] If the threshing loss rate in the current job reward is in the medium loss rate state and the threshing loss rate in the previous job reward is in the high loss rate state, the loss state change reward value is 1;
[0040] If the threshing loss rate in the current job reward is in the medium loss rate state and the threshing loss rate in the previous job reward is in the medium loss rate state, the loss state change reward value is -1;
[0041] If the threshing loss rate in the current job reward is in the medium loss rate state and the threshing loss rate in the previous job reward is in the low loss rate state, the loss state change reward value is -2;
[0042] If the threshing loss rate in the current job reward is in the low loss rate state and the threshing loss rate in the previous job reward is in the high loss rate state, the loss state change reward value is 2;
[0043] If the threshing loss rate in the current job reward is in the low loss rate state and the threshing loss rate in the previous job reward is in the medium loss rate state, the loss state change reward value is 1;
[0044] If the threshing loss rate in the current job reward is in the low loss rate state and the threshing loss rate in the previous job reward is in the low loss rate state, the loss state change reward value is 0.
[0045] In a second aspect, the present application provides a threshing system with adjustable parameters, including: a cylinder, a concave screen, a concave screen gap adjusting mechanism, a cylinder upper cover, a cylinder speed adjusting mechanism, and an electronic control system;
[0046] One end of the cylinder is connected to the frame through a bearing seat, and the other end is connected to the cylinder speed regulating mechanism through a gearbox; the cylinder speed regulating mechanism is connected to the engine; a cylinder speed sensor is provided on the cylinder speed regulating mechanism; the cylinder speed sensor is used to feedback the measured cylinder speed to the electronic control system;
[0047] The cylinder upper cover is fixedly installed on the frame and is used to enclose the upper part of the cylinder;
[0048] The concave screen is arranged on one side of the frame, and the concave screen is connected to the concave screen gap adjusting mechanism through a connecting rod;
[0049] A position feedback sensor is installed on the concave screen gap adjusting mechanism, and the position feedback sensor is used to feedback the position of the concave screen gap adjusting mechanism to the electronic control system;
[0050] The electronic control system is used for the intelligent control method of the threshing system with adjustable parameters described above.
[0051] Optionally, the drum speed regulation mechanism consists of a belt pulley, an input shaft, a hydraulic drive mechanism, a driving disc, a belt, a driven disc, an output shaft, a gearbox, and a drum speed sensor; the belt pulley is connected to the engine through a belt to provide power for the parameter-adjustable threshing system; the input shaft is installed on the frame through a bearing seat, one end of the input shaft is connected to the input end of the hydraulic continuously variable transmission, and the other end is connected to the belt pulley; the hydraulic drive mechanism, the driving disc, the belt, and the driven disc are combined to form a hydraulic continuously variable transmission; the hydraulic continuously variable transmission realizes speed regulation by changing the oil pressure of the hydraulic driver; the output shaft is fixed on the bracket, one end of the output shaft is connected to the driven disc in the hydraulic continuously variable transmission, and the other end is connected to the gearbox; the gearbox is fixedly connected to the output shaft and fixedly connected to the drum shaft; the drum speed sensor is connected to the output shaft of the drum through a coupling.
[0052] Optionally, the concave clearance of the concave sieve is the distance difference between the inner diameter of the concave and the outer diameter of the drum; the concave sieve consists of a hinge point, a concave, a connecting rod, a left crank, a regulating mechanism fixing frame, a motor, and a gear; one end of the concave sieve is connected to the frame through a hinge point, and the other end is connected to the concave sieve clearance regulating mechanism through a connecting rod.
[0053] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application:
[0054] The present application provides a parameter-adjustable threshing system and its intelligent control method. First, the current operation data of the parameter-adjustable threshing system is obtained, including the drum speed, concave clearance, and operation speed, as well as the previous operation rewards, such as the crushing rate and threshing loss rate of crop grains. These data provide a comprehensive description of the current operation state for the intelligent threshing model. Then, the operation data and operation rewards are input into the intelligent threshing model based on reinforcement learning. The model is trained using historical operation data and a preset reward function with the goal of maximizing the long-term operation reward. Through reinforcement learning, the model can learn how to adjust parameters such as the drum speed, concave clearance, and operation speed in different operation states to achieve the optimal threshing effect. Then, the intelligent threshing model outputs the current optimal control actions, including drum speed increase, drum speed decrease, concave clearance increase, concave clearance decrease, and maintaining the state. These control actions are calculated based on the current operation data and the goal of maximizing the long-term operation reward, aiming to optimize the threshing process and reduce the crushing rate and threshing loss rate of crop grains. Finally, by executing the optimal control actions output by the intelligent threshing model, the parameters of the threshing system can be adjusted in real time, thereby improving the threshing efficiency and operation quality. The intelligent threshing model can continuously learn and optimize the control strategy to adapt to different operation environments and crop types, ensuring the high efficiency and stability of the threshing process. Description of the Drawings
[0055] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0056] Figure 1 Schematic flow chart of an intelligent control method for a parameter-adjustable threshing system provided by an embodiment of the present application;
[0057] Figure 2 Schematic diagram of concave plate gap fitting provided by an embodiment of the present application;
[0058] Figure 3 Schematic structural diagram of a parameter-adjustable threshing system provided by an embodiment of the present application;
[0059] Figure 4 Schematic structural diagram of a drum speed regulation mechanism provided by an embodiment of the present application;
[0060] Figure 5 Schematic structural diagram of a concave plate screen provided by an embodiment of the present application;
[0061] Figure 6 Schematic structural diagram of a concave plate gap adjustment mechanism provided by an embodiment of the present application;
[0062] Figure 7 Schematic flow chart of drum speed regulation provided by an embodiment of the present application;
[0063] Figure 8 Schematic flow chart of concave plate gap adjustment provided by an embodiment of the present application;
[0064] Figure 9 Schematic structural diagram of an intelligent threshing system based on reinforcement learning provided by an embodiment of the present application;
[0065] Figure 10 Schematic diagram of setting output state adjustment actions provided by an embodiment of the present application;
[0066] Figure 11 Schematic diagram of steps for intelligent threshing control based on reinforcement learning provided by an embodiment of the present application. Detailed implementation manners
[0067] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0068] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0069] Embodiment 1
[0070] As Figure 1 shown, this embodiment provides an intelligent control method for a parameter-adjustable threshing system. The parameter-adjustable threshing system includes a drum, a concave screen, a concave screen gap adjustment mechanism, a drum upper cover, and a drum speed adjustment mechanism. The intelligent control method includes:
[0071] Step 101: Obtain the current operation data of the parameter-adjustable threshing system and the previous operation reward. The operation data includes the drum speed, the concave gap, and the operation speed. The operation reward includes the crushing rate and the threshing loss rate of crop particles.
[0072] Step 102: Input the operation data and the operation reward into an intelligent threshing model based on reinforcement learning to obtain the current control action of the parameter-adjustable threshing system. The control action includes drum speed increase, drum speed decrease, concave gap increase, concave gap decrease, and hold state. The intelligent threshing model based on reinforcement learning is used to train according to historical operation data and a preset reward function, aiming to maximize the long-term operation reward and output the optimal control action.
[0073] Among them, in some embodiments, after obtaining the current operation data of the parameter-adjustable threshing system and the previous operation reward, it further includes:
[0074] As Figure 10 shown, the drum speed, the concave gap, and the operation speed in the operation data are respectively divided into five states A, B, C, D, and E, specifically including:
[0075] According to the speed range of the drum speed, it is divided into a high-speed state A, a medium-high-speed state B, a medium-speed state C, a medium-low-speed state D, and a low-speed state E.
[0076] According to the size range of the concave gap, it is divided into a large-gap state A, a medium-large-gap state B, an intermediate-gap state C, a medium-small-gap state D, and a small-gap state E.
[0077] According to the speed range of the operation speed, it is divided into a high-speed state A, a medium-high speed state B, a medium speed state C, a medium-low speed state D, and a low speed state E.
[0078] For example, if the drum rotation speed range is 300 - 800 revolutions per minute, then 300 - 400 corresponds to A, 400 - 500 corresponds to B, 500 - 600 corresponds to C, 600 - 700 corresponds to D, and 700 - 800 corresponds to E.
[0079] It also includes dividing the operation rewards into states, specifically including:
[0080] According to the output range of the crushing rate, it is divided into a high crushing rate state, a medium crushing rate state, and a low crushing rate state; which can be simply referred to as the high, medium, and low levels of the crushing rate.
[0081] According to the output range of the threshing loss rate, it is divided into a high loss rate state, a medium loss rate state, and a low loss rate state, which can be simply referred to as the high, medium, and low levels of the threshing loss rate.
[0082] For example, if the crushing rate output range is 0 - 1, then 0 - 0.5 is low, 0.5 - 0.75 is medium, and 0.75 - 1 is high.
[0083] In some embodiments, the operation data and the operation rewards are input into an intelligent threshing model based on reinforcement learning to obtain the current adjustment actions of the parameter-adjustable threshing system, specifically including:
[0084] Input the operation data and the operation rewards into an intelligent threshing model based on reinforcement learning to obtain the current operation rewards under different adjustment actions;
[0085] For each adjustment action, based on the look-up table method, determine the reward value of the parameter-adjustable threshing system according to the current operation rewards under the adjustment action and the input previous operation rewards;
[0086] Determine the adjustment action corresponding to the maximum value of the reward values of the parameter-adjustable threshing system under each adjustment action as the current adjustment action.
[0087] Among them, the calculation formula for the reward value of the parameter-adjustable threshing system is:
[0088] R = R1 + R2;
[0089] Among them, R is the threshing system reward value, R1 is the crushing state change reward value, and R2 is the loss state change reward value.
[0090] During the regulation of the intelligent threshing system, the corresponding rewards are obtained by judging the state transition of the quality of the previous operation (breakage rate and loss rate). If the breakage rate changes from high to low, it indicates that the system regulation is effective and the operation effect is improved, and the system rewards it. If the breakage rate changes from low to high, it indicates that the system regulation is ineffective and the operation effect is deteriorated, and the system punishes this regulation. Since the state transition process is relatively complex, the look-up table method is used to implement it. Specifically as follows:
[0091] The calculation method of the reward value for the change in breakage state is as follows:
[0092] If the breakage rate in the current operation reward is in the high breakage rate state and the breakage rate in the previous operation reward is in the high breakage rate state, the reward value for the change in breakage state is -2;
[0093] If the breakage rate in the current operation reward is in the high breakage rate state and the breakage rate in the previous operation reward is in the medium breakage rate state, the reward value for the change in breakage state is -2;
[0094] If the breakage rate in the current operation reward is in the high breakage rate state and the breakage rate in the previous operation reward is in the low breakage rate state, the reward value for the change in breakage state is -4;
[0095] If the breakage rate in the current operation reward is in the medium breakage rate state and the breakage rate in the previous operation reward is in the high breakage rate state, the reward value for the change in breakage state is 1;
[0096] If the breakage rate in the current operation reward is in the medium breakage rate state and the breakage rate in the previous operation reward is in the medium breakage rate state, the reward value for the change in breakage state is -1;
[0097] If the breakage rate in the current operation reward is in the medium breakage rate state and the breakage rate in the previous operation reward is in the low breakage rate state, the reward value for the change in breakage state is -2;
[0098] If the breakage rate in the current operation reward is in the low breakage rate state and the breakage rate in the previous operation reward is in the high breakage rate state, the reward value for the change in breakage state is 2;
[0099] If the breakage rate in the current operation reward is in the low breakage rate state and the breakage rate in the previous operation reward is in the medium breakage rate state, the reward value for the change in breakage state is 1;
[0100] If the breakage rate in the current operation reward is in the low breakage rate state and the breakage rate in the previous operation reward is in the low breakage rate state, the reward value for the change in breakage state is 0.
[0101] Specifically, the reward value for the change in breakage state is shown in Table 1:
[0102] Table 1 Reference Table for Reward Value of Change in Breakage State
[0103]
[0104] Among them, the calculation method of the loss state change reward value is as follows:
[0105] If the threshing loss rate in the current job reward is in the high loss rate state and the threshing loss rate in the previous job reward is in the high loss rate state, the loss state change reward value is -2;
[0106] If the threshing loss rate in the current job reward is in the high loss rate state and the threshing loss rate in the previous job reward is in the medium loss rate state, the loss state change reward value is -2;
[0107] If the threshing loss rate in the current job reward is in the high loss rate state and the threshing loss rate in the previous job reward is in the low loss rate state, the loss state change reward value is -4;
[0108] If the threshing loss rate in the current job reward is in the medium loss rate state and the threshing loss rate in the previous job reward is in the high loss rate state, the loss state change reward value is 1;
[0109] If the threshing loss rate in the current job reward is in the medium loss rate state and the threshing loss rate in the previous job reward is in the medium loss rate state, the loss state change reward value is -1;
[0110] If the threshing loss rate in the current job reward is in the medium loss rate state and the threshing loss rate in the previous job reward is in the low loss rate state, the loss state change reward value is -2;
[0111] If the threshing loss rate in the current job reward is in the low loss rate state and the threshing loss rate in the previous job reward is in the high loss rate state, the loss state change reward value is 2;
[0112] If the threshing loss rate in the current job reward is in the low loss rate state and the threshing loss rate in the previous job reward is in the medium loss rate state, the loss state change reward value is 1;
[0113] If the threshing loss rate in the current job reward is in the low loss rate state and the threshing loss rate in the previous job reward is in the low loss rate state, the loss state change reward value is 0.
[0114] Specifically, the loss state change reward value is shown in Table 2:
[0115] Table 2 Loss State Change Reward Value Reference Table
[0116]
[0117] For example, if the breakage rate changes from low to high, the reward value for the change in breakage rate is R1 = -4. If the breakage rate changes from high to low, the reward value for the change in breakage rate is R1 = 2.
[0118] For example, if the previous state of the breakage rate is high and the loss rate is medium, and after performing an action A, the breakage rate becomes medium and the loss rate becomes low, then R1 = 1, R2 = 1, and R = R1 + R2 = 2 can be calculated.
[0119] For example, if the previous state of the breakage rate is medium and the loss rate is high, and after performing an action A, the breakage rate becomes low and the loss rate remains high, then R1 = 1, R2 = -2, and R = R1 + R2 = -1 can be calculated.
[0120] In some embodiments, the control action corresponding to the maximum reward value of the parameter - adjustable threshing system under each control action is determined as the current control action, which specifically includes:
[0121] As shown in Table 3, it is used to maintain the value in the corresponding state. The number of rows in the table is the number of states, a total of 5 * 5 * 5 = 125 rows. The state set is {S1S2S3}, where S1 is the drum speed ∈ (A, B, C, D, E), S2 is the concave plate clearance ∈ (A, B, C, D, E), and S3 is the operating speed ∈ (A, B, C, D, E).
[0122] The number of columns in the table is the number of actions, a total of 5 columns A1, A2, A3, A4, A5, representing the action set {drum increase, drum decrease, concave plate increase, concave plate decrease, maintain state}.
[0123] Table 3 Intelligent Threshing System State Table
[0124]
[0125] In some embodiments, the intelligent threshing control steps based on reinforcement learning are as Figure 11 shown, and specifically can be as follows:
[0126] Step 1: Perform action selection. Specifically: By random sampling, generate a random number between 0 and 1. If the random number W is less than α, then select the action with the maximum reward from the Q - table as the next action to be executed. Otherwise, randomly select a set of actions from the available actions for execution. During the initial learning process of the system, α is generally selected as a relatively small value, such as 0.5, to enhance the exploration of the system. After the system parameter update is completed, α is generally selected as a relatively large value, such as 0.9, to improve the convergence time of the system. If during the operation process, it is found that the adjustment effect becomes worse, the value of α can be reduced to enable the system to quickly adapt to the operation environment.
[0127] Step 2: Execute an action and update the status table. Specifically: Perform a harvesting operation for 20 seconds, obtain the status NS, reward R, action A after the execution of the action, calculate the value V of the current state through the Q-table, V = Q[S, A], and the target value T = R + γ * max Continue the operation.
[0128] Among them, parameter description:
[0129] The α value is to balance experience and exploration. The larger the α value, the more conservative the system is, and it tends to adopt past data. The smaller the α value, the greater the exploratory nature of the system, and it tends to find new optimal solutions. Generally, it is 0.5 - 0.9.
[0130] The γ value is to balance the current reward and future rewards. The larger the γ value, the greater the weight of future rewards. The smaller the γ value, the greater the weight of the current reward. Generally, it is 0.6 - 0.95. Here, the reward of the action with the maximum value in the next state is selected as the future reward, that is, max(Q[NS]).
[0131] β is the learning rate. The smaller the learning rate, the slower the Q-table is updated, the smaller the oscillation, and the longer the learning time. The larger the learning rate, the shorter the learning time, the faster the Q-table is updated, and the larger the oscillation. Generally, it is 0.01 - 0.1.
[0132] Adjustment process case:
[0133] For example, the current state is AAB. If the W sampling value is less than α, then from the row where AAB is located, select action A1 to execute. If W is greater than α, then from the row where AAB is located, randomly select an action from (A1, A2, A3, A4, A5) to execute. Here, assume that action A1 is selected. Action A = A1, state N = AAB, and the current input reward is low fragmentation and high loss. After 20 seconds of operation, state NS = BAB, and the input reward is low fragmentation and low loss. It can be obtained that:
[0134] Reward R = 0 + 2 = 2, and the current action value V = 0.5.
[0135] Set γ = 0.9, and the target state value T = R + γ * max(Q[BAB]), that is, T = 2 + 0.9 * 0.5 = 2.5.
[0136] Set β = 0.1, and calculate U = (T - V) * β = 2 * 0.1 = 0.2.
[0137] Update the Q-table Q[AAB, A1] = Q[AAB, A1] + U = 0.5 + 0.2 = 0.8.
[0138] By comparison, it can be seen that before adjustment, Q[AAB,A1] = 0.5, and after adjustment, Q[AAB,A1] = 0.8. This value has been strengthened, indicating that in this state, taking this action is effective, that is, when the system reaches this state again, this action will still be preferentially selected for execution. If the reward value is negative, the value of this action will be weakened, and other actions with higher value will be preferentially selected for execution next time.
[0139] Among them, in some embodiments, when performing the electric control adjustment step of the drum speed, it can specifically be as follows:
[0140] Step 1: Determine whether the condition step is satisfied. Specifically: Obtain the actual speed of the drum 100 through the drum speed sensor 509, and determine whether the speed is greater than 50 revolutions per minute. If so, execute Step 2; otherwise, close the hydraulic output to end the adjustment. Due to the characteristics of the hydraulic CVT speed adjustment mechanism, the speed adjustment can only be carried out under the condition that the mechanism is in motion, otherwise it will cause damage to the mechanism;
[0141] Step 2: Determine whether the actual speed is higher than the set value. Specifically: Determine whether the difference between the actual speed and the target speed is greater than 1. If so, execute Step 3; otherwise, execute Step 4.
[0142] Step 3: Execute the speed reduction step. Specifically: Open the pressure reducing valve so that the pressure of the hydraulic drive mechanism 503 in the drum speed adjustment mechanism 500 decreases. Under the action of the spring, the diameter of the driving wheel of the CVT transmission system decreases, and the diameter of the driven wheel increases, thereby achieving the reduction of the drum speed.
[0143] Step 4: Determine whether the speed is lower than the set value and whether it is necessary to execute the speed increase action. Specifically: Determine whether the difference between the actual speed and the target speed is less than 1. If so, open the pressure increasing valve so that the pressure of the hydraulic drive mechanism 503 in the drum speed adjustment mechanism 500 increases. The diameter of the driving wheel of the CVT transmission system increases, and the diameter of the driven wheel decreases, thereby achieving the increase of the drum speed. Otherwise, close the hydraulic output to end the adjustment.
[0144] Among them, in some embodiments, when performing the electric control adjustment step of the concave plate gap, it can specifically be as follows:
[0145] Step 1: Concave plate gap acquisition step. Specifically, obtain the position data of the concave plate gap encoder, and convert it into the measured value of the concave plate gap through a linear transformation method, that is: Concave plate gap = a * position of the encoder + b, and the parameters a and b are obtained through a calibration method, such as Figure 2 shown.
[0146] Example of calibration process: Assume that in one implementation, within the movement range of the concave clearance adjustment mechanism, the numerical range of the encoder is 100 - 200. Then, through manual adjustment, within the range of 100 - 200, collect some encoder data and concave clearance data. The concave clearance data is obtained using a vernier caliper. Obtained: encoder position = {101, 120, 143, 160, 181, 200}, concave clearance = {5, 13, 17, 21, 26, 30}, and for the least squares fit, a = 0.2416, b = 17.77474.
[0147] Step 2: Determine whether concave clearance adjustment is required. Specifically: Obtain the set value of the concave clearance, and then determine whether the set value is less than the threshold value compared to the measured value. If so, end the adjustment; if not, execute Step 3.
[0148] Step 3: Execute the concave clearance adjustment action. Specifically: Calculate the control quantity of the concave clearance adjustment motor through PID and send the control quantity to the motor driver. The driver drives the motor to perform forward and reverse movements, thereby achieving concave clearance adjustment.
[0149] Embodiment 2
[0150] As Figure 3 shown, this embodiment provides a parameter - adjustable threshing system, including: a cylinder 100, a concave screen 200, a concave screen clearance adjustment mechanism 300, a cylinder upper cover 400, a cylinder speed - regulating mechanism 500, and an electronic control system;
[0151] One end of the cylinder 100 is connected to the frame through a bearing seat, and the other end is connected to the cylinder speed - regulating mechanism 500 through a gearbox; the cylinder speed - regulating mechanism 500 is connected to the engine; a cylinder speed sensor is provided on the cylinder speed - regulating mechanism 500; the cylinder speed sensor is used to feedback the measured cylinder speed to the electronic control system;
[0152] The cylinder upper cover 400 is fixedly installed on the frame and is used to enclose the upper part of the cylinder 100;
[0153] The concave screen 200 is arranged on one side of the frame, and the concave screen 200 is connected to the concave screen clearance adjustment mechanism 300 through a connecting rod;
[0154] A position feedback sensor is installed on the concave screen clearance adjustment mechanism 300, and the position feedback sensor is used to feedback the position of the concave screen clearance adjustment mechanism 300 to the electronic control system;
[0155] The electronic control system is used to execute the intelligent control method of the parameter - adjustable threshing system described above.
[0156] Among them, as Figure 4As shown, the drum speed regulating mechanism 500 mainly consists of a belt pulley 501, an input shaft 502, a hydraulic driving mechanism 503, a driving disc 504, a belt 505, a driven disc 506, an output shaft 507, a gearbox 508, and a drum speed sensor 509. The belt pulley 501 is connected to the engine through a belt to provide power for the system. The input shaft 502 is installed on the frame through a bearing block, with one end connected to the input end of the hydraulic continuously variable transmission and the other end connected to the belt pulley 501. The hydraulic driving mechanism 503, the driving disc 504, the belt 505, and the driven disc 506 are combined to form a hydraulic continuously variable transmission (CVT), and speed regulation is achieved by changing the oil pressure of the hydraulic driver 503. The output shaft 507 is fixed to the bracket, with one end connected to the driven disc 506 of the continuously variable transmission and the other end connected to the gearbox 508. The gearbox 508 is fixedly connected to the output shaft 507 and fixedly connected to the drum shaft, and its main function is to reduce speed and change the power transmission direction. The drum speed sensor 509 is connected to the drum output shaft through a coupling, and its main function is to convert the drum speed into an electrical signal.
[0157] Among them, as Figure 5 shown, the concave screen gap d of the concave screen 200 refers to the distance difference between the inner diameter of the concave plate 202 and the outer diameter of the drum 100. The adjustable concave screen mainly consists of a hinge point 201, a concave plate 202, a connecting rod 203, a left crank 308, a regulating mechanism fixing frame 307, a motor 303, and a gear 305. One end of the concave screen is connected to the frame through the hinge point 201, and the other end is connected to the concave screen gap regulating mechanism through the connecting rod 203. The crank 308 of the concave screen gap regulating mechanism is used to drive the connecting rod 203 to move up and down, driving the concave plate 202 to rotate around the hinge point 201.
[0158] Among them, as Figure 6As shown in the figure, a concave plate gap adjustment mechanism mainly consists of a right crank 301, a shaft 302, a driving motor 303, a gear 304, a gear 305, a worm 306, an adjustment mechanism fixing bracket 307, a left crank 308, a worm gear 309, an encoder 310, a synchronous belt 311, a synchronous belt pulley 312, and a synchronous belt pulley 313. The driving motor is fixed on the frame and drives the gear 305 to rotate through the gear 304. The gear 305 is fixedly connected to the worm 306. When the motor rotates, the worm 306 is driven to rotate synchronously. The rotation of the worm 306 drives the worm gear 309 to rotate. The worm gear 309 is fixedly connected to the shaft 302, and the right crank 301 is fixedly connected to the left crank 308. When the worm gear 309 rotates, the cranks 301 and 308 are driven to rotate synchronously. After the motor stops rotating, due to the self-locking effect of the worm and worm gear mechanism, it is ensured that the crank will not be displaced under the action of external forces. The encoder 310 is fixedly connected to the synchronous belt pulley 313 through a coupling. The synchronous belt pulley 312 is fixedly connected to the shaft 302. When the shaft 302 rotates, it drives the synchronous belt pulley 312 to rotate. The synchronous belt pulley 312 drives the synchronous belt pulley 313 to rotate through the synchronous belt 311, and the synchronous belt pulley 313 drives the encoder 310 to rotate, thereby indirectly measuring the rotation angles of the cranks 301 and 308. The crankshafts 301 and 308 are connected to the concave plate screen through a connection. Through on-site calibration, the rotation angle of the encoder 310 can be mapped to the gap of the concave plate screen by a linear variation method.
[0159] In summary, the present application has the following technical effects:
[0160] The advantage of the method provided by the present application is that there is no need to formulate many complex control strategies (knowledge bases) in advance. Only need to consider how to find the optimal control strategy by optimizing the control method. When the external environment changes (crop variety, harvesting area), the user can enable the method of reinforcement learning to let the harvester re-learn and update the system parameters, so that the harvester can adapt to the current operating environment.
[0161] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0162] Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An intelligent control method for a parameter-adjustable threshing system, characterized in that: The parameter-adjustable threshing system comprises a drum, a concave plate screen, a concave plate screen gap adjustment mechanism, a drum upper cover and a drum speed adjustment mechanism; The intelligent control method comprises: Acquire the current operation data of the parameter-adjustable threshing system and the last operation reward; the operation data includes the drum rotation speed, the concave plate gap and the operation speed; the operation reward includes the crop particle breakage rate and the threshing loss rate; The operation data and the operation reward are input into the intelligent threshing model based on reinforcement learning to obtain the current control action of the parameter-adjustable threshing system; the control action includes drum speed increase, drum speed decrease, concave plate gap increase, concave plate gap decrease and maintaining state; the intelligent threshing model based on reinforcement learning is used to train according to historical operation data and a preset reward function, with the goal of maximizing long-term operation rewards, and outputting the optimal control action.
2. The intelligent control method of a parameter-adjustable threshing system according to claim 1, characterized in that: After obtaining the current operation data of the parameter adjustable threshing system and the last operation reward, it also includes: The drum rotation speed, concave plate gap and operation speed in the operation data are respectively divided into states, specifically including: According to the speed range of the drum speed, it is divided into a high speed state, a medium-high speed state, a medium speed state, a medium-low speed state and a low speed state; According to the size range of the concave plate gap, it is divided into a large gap state, a medium-large gap state, a medium gap state, a medium-small gap state and a small gap state; According to the speed range of the operating speed, it is divided into a high speed state, a medium-high speed state, a medium speed state, a medium-low speed state and a low speed state.
3. The intelligent control method of a parameter-adjustable threshing system according to claim 1, characterized in that: After obtaining the current operation data of the parameter adjustable threshing system and the last operation reward, it also includes: The job rewards are divided into status categories, including: According to the output range of the crushing rate, it is divided into a high crushing rate state, a medium crushing rate state and a low crushing rate state; According to the output range of the threshing loss rate, it is divided into a high loss rate state, a medium loss rate state and a low loss rate state.
4. The intelligent control method of a parameter-adjustable threshing system according to claim 1, characterized in that: The operation data and the operation reward are input into the intelligent threshing model based on reinforcement learning to obtain the current control action of the parameter-adjustable threshing system, which specifically includes: Inputting the operation data and the operation reward into an intelligent threshing model based on reinforcement learning to obtain current operation rewards under different control actions; For each control action, according to the current operation reward under the control action and the input last operation reward, based on the table lookup method, determine the reward value of the parameter adjustable threshing system; The control action corresponding to the maximum value of the reward value of the parameter-adjustable threshing system under each control action is determined as the current control action.
5. The intelligent control method of a parameter-adjustable threshing system according to claim 4, characterized in that: The calculation formula of the reward value of the parameter adjustable threshing system is: R = R1 + R2; Among them, R is the threshing system reward value, R1 is the broken state change reward value, and R2 is the loss state change reward value.
6. The intelligent control method of a parameter-adjustable threshing system according to claim 5, characterized in that: The calculation method of the broken state change reward value is: If the current job reward has a high breakage rate, and the previous job reward has a high breakage rate, the breakage state change reward value is -2; If the current job reward has a high shattering rate and the previous job reward has a medium shattering rate, the shattering rate change reward value is -2; If the current job reward has a high shattering rate, and the previous job reward has a low shattering rate, the shattering rate change reward value is -4; If the current job reward has a medium breakage rate, and the previous job reward has a high breakage rate, the breakage state change reward value is 1; If the current job reward has a medium breakage rate, and the previous job reward has a medium breakage rate, the breakage status change reward value is -1; If the current job reward has a medium shattering rate, and the previous job reward has a low shattering rate, the shattering rate change reward value is -2; If the current job reward has a low breakage rate and the previous job reward has a high breakage rate, the breakage state change reward value is 2; If the current job reward has a low breakage rate and the previous job reward has a medium breakage rate, the breakage state change reward value is 1; If the fragmentation rate in the current job reward is in a low fragmentation rate state, and the fragmentation rate in the previous job reward was in a low fragmentation rate state, the fragmentation state change reward value is 0.
7. The intelligent control method of a parameter-adjustable threshing system according to claim 5, characterized in that: The calculation method of the loss state change reward value is: If the threshing loss rate in the current job reward is in the high loss rate state, and the threshing loss rate in the previous job reward was in the high loss rate state, the loss state change reward value is -2; If the threshing loss rate in the current job reward is in the high loss rate state, and the threshing loss rate in the previous job reward is in the medium loss rate state, the loss state change reward value is -2; If the threshing loss rate in the current job reward is in the high loss rate state, and the threshing loss rate in the previous job reward is in the low loss rate state, the loss state change reward value is -4; If the threshing loss rate in the current job reward is in the medium loss rate state, and the threshing loss rate in the previous job reward is in the high loss rate state, the loss state change reward value is 1; If the threshing loss rate in the current job reward is in the medium loss rate state, and the threshing loss rate in the previous job reward is in the medium loss rate state, the loss state change reward value is -1; If the threshing loss rate in the current job reward is in the medium loss rate state, and the threshing loss rate in the previous job reward is in the low loss rate state, the loss state change reward value is -2; If the threshing loss rate in the current job reward is in the low loss rate state, and the threshing loss rate in the previous job reward is in the high loss rate state, the loss state change reward value is 2; If the threshing loss rate in the current job reward is in the low loss rate state, and the threshing loss rate in the previous job reward is in the medium loss rate state, the loss state change reward value is 1; If the threshing loss rate in the current job reward is in a low loss rate state, and the threshing loss rate in the previous job reward was in a low loss rate state, the loss state change reward value is 0.
8. A parameter-adjustable threshing system, characterized in that: include: Drum, concave plate screen, concave plate screen gap adjustment mechanism, drum cover, drum speed adjustment mechanism and electronic control system; One end of the drum is connected to the frame through a bearing seat, and the other end is connected to the drum speed regulating mechanism through a gearbox; the drum speed regulating mechanism is connected to the engine; a drum speed sensor is provided on the drum speed regulating mechanism; the drum speed sensor is used to feed back the measured drum speed to the electronic control system; The drum upper cover is fixedly mounted on the frame and is used to close the upper part of the drum; The concave plate screen is arranged on one side of the frame, and the concave plate screen is connected to the concave plate screen gap adjustment mechanism through a connecting rod; A position feedback sensor is installed on the concave plate screen gap adjustment mechanism, and the position feedback sensor is used to feed back the position of the concave plate screen gap adjustment mechanism to the electronic control system; The electronic control system is used to execute the intelligent control method of a parameter-adjustable threshing system as described in any one of claims 1-7.
9. A parameter-adjustable threshing system according to claim 8, characterized in that: The drum speed regulating mechanism is composed of a belt pulley, an input shaft, a hydraulic drive mechanism, an active disk, a belt, a driven disk, an output shaft, a gear box and a drum speed sensor; the belt pulley is connected to the engine through a belt to provide power for the parameter-adjustable threshing system; the input shaft is installed on the frame through a bearing seat, one end of the input shaft is connected to the input end of the hydraulic continuously variable transmission, and the other end is connected to the belt pulley; the hydraulic drive mechanism, the active disk, the belt and the driven disk are combined to form a hydraulic continuously variable transmission; the hydraulic continuously variable transmission realizes speed adjustment by changing the oil pressure of the hydraulic drive; the output shaft is fixed on the bracket, one end of the output shaft is connected to the driven disk in the hydraulic continuously variable transmission, and the other end is connected to the gear box; the gear box is fixedly connected to the output shaft and the drum shaft; the drum speed sensor is connected to the output shaft of the drum through a coupling.
10. The parameter-adjustable threshing system according to claim 8, characterized in that: The concave plate gap of the concave plate screen is the distance difference between the inner diameter of the concave plate and the outer diameter of the drum; the concave plate screen is composed of a hinge point, a concave plate, a connecting rod, a left crank, an adjustment mechanism fixing frame, a motor and a gear; one end of the concave plate screen is connected to the frame through a hinge point, and the other end is connected to the concave plate screen gap adjustment mechanism through a connecting rod.
Citation Information
Patent Citations
Adjusting device and adjusting method for cereal threshing cylinder concave clearance
CN106664990A
Combine harvester including machine feedback control
CN110740635A
Supervisory and improvement system for machine control
CN112445130A
Wheat low-loss threshing adaptive control system and method based on multi-objective particle swarm optimization algorithm
CN118131614A
Multi-parameter regulation and control system for entrainment loss of threshing device and combine harvester
CN213127154U