Reinforcement learning based multi-degree of freedom shaker control method and system

By using a multi-degree-of-freedom vibrating screen control method based on reinforcement learning, the distribution of material on the screen surface and grain loss can be monitored and adjusted in real time. This solves the problem of uneven distribution of the screen surface in combine harvesters, improves operational applicability and stability, and reduces grain loss rate.

CN116416509BActive Publication Date: 2026-05-12JIANGSU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU UNIV
Filing Date
2023-03-22
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

The existing vibrating screening devices of combine harvesters are mostly single-degree-of-freedom, which cannot effectively solve the problem of uneven distribution of threshed mixture on the screen surface, resulting in high cleaning loss rate and poor applicability and stability in complex environments.

Method used

A multi-degree-of-freedom vibrating screening device control method based on reinforcement learning is adopted. By monitoring the material distribution on the screen surface and the grain loss, a deep reinforcement learning model is established to adjust the screening state in real time to achieve the optimal screening effect, including the dynamic adjustment of the screen surface inclination angle, horizontal attitude angle and vibration frequency.

Benefits of technology

It improves the applicability and stability of combine harvesters in different environments, reduces grain loss rate, and enhances screening efficiency and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116416509B_ABST
    Figure CN116416509B_ABST
Patent Text Reader

Abstract

The application provides a multi-degree-of-freedom vibrating screen control method and system based on reinforcement learning, a screen material distribution monitoring device collects the image of the screen material distribution state of the multi-degree-of-freedom mixed vibrating screen and transmits it to the controller, the controller monitors the uniformity and dispersion degree of the screen material according to the screen material distribution monitoring model; a grain loss monitoring device monitors the grain loss of the multi-degree-of-freedom mixed vibrating screen and transmits it to the controller; the controller obtains the current screening state of the degree-of-freedom vibrating screen from the screen material distribution monitoring device, the grain loss monitoring device and the actuator, and controls the actuator to adjust the screening state to be optimal according to the reinforcement learning model, thereby improving the convergence speed and stability of the model and realizing the adaptive control of the multi-degree-of-freedom vibrating screen. The control system does not depend on the traditional control rule, but has self-learning ability, thereby improving the operation applicability and stability of the harvester in different environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of screening mechanism control technology, specifically relating to a control method and system for a multi-degree-of-freedom vibrating screening device based on reinforcement learning. Background Technology

[0002] Cleaning is a crucial step in grain combine harvesting, directly impacting the overall machine performance. The cleaning unit of a combine harvester mainly consists of a blower and a vibrating screen. During combine harvesting, after threshing, the crop forms a threshed mixture of grains and short straw. This mixture falls onto the vibrating screen surface under the action of the shaking and return plates. Driven by the combined action of the blower and the vibrating screen, it moves backward through the screen. Grains pass through the screen openings into the grain collecting device, while straw and other impurities are discharged from the rear of the screen. A small amount of grains that fail to pass through the screen will also be discharged from the rear of the vibrating screen, resulting in cleaning losses.

[0003] Due to the working principle of the threshing components in combine harvesters, the threshed mixture exhibits a non-uniform distribution below the threshing concave plate, directly leading to uneven distribution of the threshed mixture when it enters the vibrating screen. Furthermore, in hilly and mountainous areas, and in deep, muddy, and wet fields, combine harvesters exhibit poor horizontal stability and significant horizontal posture fluctuations, which exacerbate the uneven distribution of the threshed mixture entering the screen. Currently, most combine harvesters use single-degree-of-freedom reciprocating vibrating screens, whose motion trajectory is singular and cannot effectively solve the problem of uneven distribution of the threshed mixture on the screen surface. Multi-degree-of-freedom motion of the screen surface can effectively promote rapid and uniform dispersion of materials, and is an effective way to efficiently screen non-uniformly input materials. Currently, there is a lack of multi-degree-of-freedom vibrating screen control systems suitable for grain combine harvesting operations. my country has a wide variety of grains, a broad geographical distribution, and complex operating environments. Even for the same crop, the screening and cleaning performance of combine harvesters varies significantly depending on the harvesting season. Currently, a universally applicable multi-degree-of-freedom vibrating screen control method and system has not yet been established to improve the applicability and stability of harvesters in different environments. Summary of the Invention

[0004] To address the aforementioned technical problems, one objective of this invention is to provide a control method and system for a multi-degree-of-freedom vibrating screen device based on reinforcement learning. This method controls the actuator to achieve the optimal screening state based on the material distribution on the screen surface of the multi-degree-of-freedom vibrating screen, thereby improving the applicability and stability of the harvester in different environments.

[0005] Note that the description of these objectives does not preclude the existence of other objectives. One aspect of the invention does not require achieving all of the above objectives. Objectives other than those described above can be extracted from the description, drawings, and claims.

[0006] The present invention achieves the above-mentioned technical objectives through the following technical means.

[0007] A control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning includes the following steps:

[0008] The screen surface material distribution monitoring device collects images of the screen surface material distribution state of a multi-degree-of-freedom hybrid vibrating screen and transmits them to the controller. The controller monitors the screen surface material uniformity K and dispersion degree P according to the screen surface material distribution monitoring model.

[0009] The grain loss monitoring device monitors the grain loss of the multi-degree-of-freedom hybrid vibrating screen and transmits the data to the controller;

[0010] The controller obtains the current screening state s of the multi-degree-of-freedom vibrating screen from the screen material distribution monitoring device, the grain loss monitoring device, and the actuator of the multi-degree-of-freedom hybrid vibrating screen. s = [α, β, f, μ, K, P], including grain loss rate μ, uniformity K, dispersion coefficient P, screen inclination angle α, screen horizontal attitude angle β, and vibration frequency f. Based on the control model of the multi-degree-of-freedom vibrating screen device, the controller controls the actuator to adjust the screening state so that the loss rate μ, the absolute value of uniformity |K|, and the dispersion coefficient P are optimal, achieving the optimal screening state.

[0011] In the above scheme, the screen surface material distribution monitoring model is established through the following steps:

[0012] Images of the material distribution on the screen surface of a multi-degree-of-freedom vibrating screen are acquired. The acquired images are binarized to extract particle distribution information. The binarized images are then subjected to dilation and erosion processing to establish a material distribution model on the screen surface. A rectangular plane C with length and width h and l is selected on the image, located at a distance d1 from the front of the screen and d2 from the inner side of the screen wall. A planar rectangular coordinate system is established, and samples are collected sequentially and uniformly on plane C. A linear binary classification model is established using a two-input, one-output neural network. The network contains only one input layer and one output layer. The gradient ascent method based on maximum likelihood estimation is used to train the model until convergence, obtaining the optimal linear equation.

[0013] Furthermore, the optimal straight line equation is:

[0014] W1·x+W2·y+B=0

[0015] Where W1 is the connection weight of the optimal input layer neuron, W2 is the connection weight of the optimal output layer neuron, and B is the bias term.

[0016] Further, definition Screen surface uniformity refers to the degree of uniformity of material distribution along the Y-axis of the screen surface. definition Let P be the dispersion coefficient of the sieve surface, which is P∈(0,h).

[0017] Furthermore, the optimal sieving state is as follows: the loss rate μ approaches its minimum value, the absolute value of uniformity |K| approaches π / 2, and the degree of dispersion P approaches h.

[0018] In the above scheme, the control model of the multi-degree-of-freedom vibrating screen is established using deep reinforcement learning. The control model includes an agent, an environment, the state of the agent, the actions of the agent, and a reward function R. The environment is defined as the working environment of the multi-degree-of-freedom vibrating screen control system, i.e., the interior of the cleaning chamber. The agent is defined as the entire multi-degree-of-freedom vibrating screen control system. The agent's state is the current working state of the multi-degree-of-freedom vibrating screen. The agent's actions are defined as the changes in the working parameters of the multi-degree-of-freedom vibrating screen. The multi-degree-of-freedom vibrating screen control system obtains the current working state of the multi-degree-of-freedom vibrating screen through the screen surface material distribution monitoring device, the grain loss monitoring device, and the actuator. Based on feedback from the working environment, it calculates a quantifiable reward signal through the reward function R and outputs the changes in the working parameters of the multi-degree-of-freedom vibrating screen. The working environment of the multi-degree-of-freedom vibrating screen is affected by the actions, the screening state changes, and a new reward is generated. The action strategy is then updated using the feedback screening state and reward. This iterative process continues until the multi-degree-of-freedom vibrating screen reaches its optimal screening state.

[0019] In the above scheme, in the sieving state s=[θ,β,f,μ,K,P], α∈(-10°,10°), β∈(-10°,10°), f∈

[0020] (10Hz, 15Hz), μ∈(0, 0.1),

[0021] Furthermore, the changes in the screen surface inclination angle α, the screen surface horizontal attitude angle β, and the vibration frequency f of the actuator of the multi-degree-of-freedom vibrating screen are Δα, Δβ, and Δf, respectively. The motion space of the multi-degree-of-freedom vibrating screen is defined as A = [Δα, Δβ, Δf], and the motion space is discretized, where:

[0022] Δα=[0, ±0.5°, ±1°, ±1.5°, ±2°],

[0023] Δβ=[0,±0.5°,±1°,±1.5°,±2°],

[0024] Δf=[0, ±0.5Hz, ±1Hz, ±1.5Hz, ±2Hz].

[0025] In the above scheme, the reward function R is a function relating the current screening states μ, K, and P, as well as the screening state change rates Δμ, ΔK, and ΔP. The reward function R includes the state reward / penalty function Ri.s and action reward / punishment function R a :

[0026] R = R s +R a ;

[0027] State reward and punishment function R s The value of evaluating the current screening state s is expressed as:

[0028] R s =ρ1·F1(μ)+ρ2·F2(|K|)+ρ3·F3(P);

[0029] Where ρ1, ρ2, and ρ3 are positive constants, F1(μ) is a decreasing function of the loss rate μ, F2(|K|) is an increasing function of the absolute value of uniformity |K|, and F3(P) is an increasing function of the dispersion degree P.

[0030] Action reward and punishment function R a It evaluates the value brought by action 'a', expressed as:

[0031] R a =σ1·G1(Δμ)+σ2·G2(Δ|K|)+σ3·G3(ΔP);

[0032] Where σ1, σ2, and σ3 are positive constants, G1(Δμ) is an odd function of the rate of change of loss Δμ, and G1(Δμ) is continuous and monotonically decreasing; G2(Δ|K|) is an odd function of the rate of change of the absolute value of uniformity Δ|K|, and G3(ΔP) is an odd function of the rate of change of dispersion ΔP, and G2(Δ|K|) and G3(ΔP) are continuous and monotonically increasing.

[0033] In the above scheme, the control model of the multi-degree-of-freedom vibrating screening device is a reinforcement learning model.

[0034] Furthermore, the reinforcement learning model is a DQN model; the controller is a DQN controller;

[0035] The DQN model is built using the following steps:

[0036] The DQN controller reads the current screening state s of the multi-degree-of-freedom vibrating screen from the screen material distribution monitoring device and the actuator, and predicts the value of each action in the current state through the eval_net network. It selects the next action a to be executed using an ε-greedy strategy and sends it to the vibrating screen actuator for execution. After time t, the screening state is sampled again to obtain the next screening state s'. The reward r is calculated according to the reward function R. The acquired experience s, a, s', and r are stored in the experience database. A portion of experience is randomly extracted from the experience database to train the deep Q network. According to the loss function LW, the eval_net network is trained using SGD stochastic gradient descent. Every N training iterations, W' = W is set to update the eval_net network parameters W. s = s' is set to iteratively train the reinforcement learning model until it converges. When the reinforcement learning model converges, the DQN controller controls the actuator to adjust the screening state according to the current screening state s, so that the loss rate μ, the absolute value of uniformity |K|, and the dispersion degree P are optimal, thus achieving the optimal screening state.

[0037] A system for implementing the reinforcement learning-based control method for a multi-degree-of-freedom vibrating screen includes a multi-degree-of-freedom vibrating screen, a screen surface material distribution monitoring device, a grain loss monitoring device, and a controller.

[0038] The multi-degree-of-freedom vibrating screen includes a vibrating screen, a parallel drive mechanism, a series drive mechanism, and a constraint link, which can realize two translations and two rotations. The parallel drive mechanism realizes three degrees of freedom motion of the screen surface, namely rotation around the X-axis and Y-axis and translation around the Z-axis. The series mechanism realizes one degree of freedom reciprocating motion of the screen surface. One end of the constraint link is connected to the frame, and the other end is connected to the side of the vibrating screen.

[0039] The screen surface material distribution monitoring device collects images of the screen surface material distribution state of the multi-degree-of-freedom vibrating screen and transmits them to the controller. The controller monitors the screen surface material uniformity K and dispersion degree P according to the screen surface material distribution monitoring model.

[0040] The grain loss monitoring device monitors the grain loss of the multi-degree-of-freedom vibrating screen and transmits the data to the controller;

[0041] The controller obtains the current screening state s of the multi-degree-of-freedom vibrating screen from the screen material distribution monitoring device and the actuator of the multi-degree-of-freedom vibrating screen, s=[α,β,f,μ,K,P], including the grain loss rate μ, uniformity K, dispersion coefficient P monitored by the vibrating screen monitoring system, as well as the screen surface inclination angle α, screen surface horizontal attitude angle β and vibration frequency f of the actuator, and controls the actuator to adjust the screening state according to the reinforcement learning model, so that the loss rate μ, the absolute value of uniformity |K|, and the dispersion P are optimal, thus achieving the optimal screening state.

[0042] In the above scheme, the parallel drive mechanism includes four sets of parallel drive components, namely a first parallel drive component, a second parallel drive component, a third parallel drive component, and a fourth parallel drive component;

[0043] Each set of drive components includes a stepper motor, a lead screw, a slider, a slide base, and a laser displacement sensor; the slide base of each set of drive components is mounted on the frame, the slider is connected to one end of the boom via a sixth fisheye bearing, and the other end of the boom is connected to the vibrating screen via a fifth fisheye bearing; the laser displacement sensor emitter is mounted vertically downward on the slider, and the displacement measuring plate is mounted on the lower end face of the slide base.

[0044] In the above scheme, the series drive mechanism includes a drive link, an eccentric rotating disk, and a DC drive motor;

[0045] One end of the drive link is connected to the vibrating screen via a third fisheye bearing, and the other end of the drive link is connected to the eccentric rotating disk via a fourth fisheye bearing; the eccentric rotating disk is mounted on the output shaft of the DC drive motor, and the DC drive motor is mounted on the frame.

[0046] In the above scheme, the number of constraint links is two;

[0047] One end of each of the two identical constraint rods is connected to the frame via a second fisheye bearing, and the other end of each constraint rod is connected to the side of the vibrating screen via a first fisheye bearing.

[0048] In the above scheme, the controller calculates the screen surface inclination angle α and the screen surface horizontal attitude angle β according to the following formula;

[0049]

[0050]

[0051] Wherein, H1, H2, H3, and H4 are the distances from their emitting ends to the displacement measuring plate detected by the four laser displacement sensors, respectively; L X and L Y These are the center distances of the parallel drive components along the X and Y axes, respectively.

[0052] In the above scheme, the controller calculates the vibration frequency of the vibrating screen according to the following formula:

[0053]

[0054] ω is the rotational speed of the DC drive motor.

[0055] In the above scheme, the screen surface material distribution monitoring device includes a screen camera, multiple screen surface photosensors, and screen surface RBG supplementary lights. The screen surface RBG supplementary lights are installed above the screen surface to provide real-time supplementary lighting to the screen surface. The screen surface photosensors are used to detect the light intensity of the screen surface and transmit it to the controller. The controller adjusts the brightness of the RBG supplementary lights according to the light intensity. The camera is used to capture images of the material distribution state on the screen surface and transmit them to the controller.

[0056] Compared with the prior art, the beneficial effects of the present invention are:

[0057] 1. This invention controls the actuator to achieve the optimal screening state based on the material distribution state on the screen surface of a multi-degree-of-freedom vibrating screen, thereby improving the applicability and stability of the harvester in different environments.

[0058] 2. This invention establishes a screen surface material distribution monitoring model, uses image processing to acquire the screen surface material distribution state, and proposes uniformity and dispersion descriptive indicators. A grain loss monitoring device monitors the loss of particles at the screen tail in real time. A reinforcement learning reward function is constructed by integrating the screen surface material distribution state and grain loss information, establishing a deep reinforcement learning model. This improves the model's convergence speed and stability, achieving adaptive control of a multi-degree-of-freedom vibrating screen. The control system of this invention does not rely on traditional control rules but possesses self-learning capabilities. It can gradually optimize the control scheme through practical experience accumulation in different operating environments, reducing screening losses and improving operational efficiency and adaptability.

[0059] 3. The multi-degree-of-freedom vibrating screen of the present invention is a multi-degree-of-freedom hybrid vibrating screen that adopts a series-parallel hybrid mechanism, which achieves two translations and two rotations of the screen surface while ensuring the strength of the mechanism.

[0060] Note that the description of these effects does not preclude the existence of other effects. One aspect of the invention does not necessarily have all the aforementioned effects. Effects other than those described above can be readily observed and extracted from the description, drawings, claims, etc. Attached Figure Description

[0061] Figure 1 This is a schematic diagram of the structure of a multi-degree-of-freedom vibrating screening device according to an embodiment of the present invention.

[0062] Figure 2 This is a schematic diagram of the parallel drive component structure according to an embodiment of the present invention.

[0063] Figure 3 This is a schematic diagram of the HSV image of the screen material according to an embodiment of the present invention.

[0064] Figure 4 This is a schematic diagram of a material distribution model on a screen surface according to an embodiment of the present invention.

[0065] Figure 5 This is a binary classification neural network model according to one embodiment of the present invention.

[0066] Figure 6 This is a schematic diagram illustrating the working principle of a loss sensor according to an embodiment of the present invention.

[0067] Figure 7 This is a flowchart of the DQN model training process according to one embodiment of the present invention.

[0068] Figure 8 This is a schematic diagram of the control principle of a multi-degree-of-freedom vibrating screen based on deep reinforcement learning according to one embodiment of the present invention.

[0069] In the diagram: 1. Vibrating screen; 2. First parallel drive component; 3. First fisheye bearing; 4. Constraint link; 5. Second fisheye bearing; 6. Third fisheye bearing; 7. Drive link; 8. Fourth fisheye bearing; 9. DC drive motor; 10. Eccentric rotating disk; 11. Grain loss sensor; 12. Second parallel drive component; 13. Third parallel drive component; 14. Screen surface photosensitive sensor; 15. Screen camera; 16. Screen surface RBG supplementary light; 17. Camera field of view; 18. Fourth parallel drive component; 19. Fifth fisheye bearing; 20. Suspension rod; 21. Sixth fisheye bearing; 22. Stepper drive motor; 23. Lead screw; 24. Slide table base; 25. Slider; 26. Laser displacement sensor; 27. Displacement measuring plate; 28. DQN controller. Detailed Implementation

[0070] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0071] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "front," "rear," "left," "right," "upper," "lower," "axial," "radial," "vertical," "horizontal," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0072] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0073] Combination Figure 1 and Figure 2As shown, in one embodiment of the present invention, a system for controlling a multi-degree-of-freedom vibrating screen device based on reinforcement learning includes a multi-degree-of-freedom vibrating screen, a screen surface material distribution monitoring device, a grain loss monitoring device, and a controller. The multi-degree-of-freedom vibrating screen includes a vibrating screen 1, a parallel drive mechanism, a series drive mechanism, and a constraint link 4, capable of two translations and two rotations. The parallel drive mechanism realizes three-degree-of-freedom motion of the screen surface, namely rotation around the X and Y axes and translation around the Z axis, while the series mechanism realizes one-degree-of-freedom reciprocating motion of the screen surface. One end of the constraint link 4 is connected to the frame, and the other end is connected to the side of the vibrating screen 1. The screen surface material distribution monitoring device acquires and transmits images of the screen surface material distribution state of the multi-degree-of-freedom vibrating screen. The controller monitors the uniformity K and dispersion P of the material on the screen surface according to the screen surface material distribution monitoring model. The grain loss monitoring device monitors the grain loss of the multi-degree-of-freedom vibrating screen and transmits it to the controller. The controller obtains the current screening state s of the multi-degree-of-freedom vibrating screen from the screen surface material distribution monitoring device and the actuator of the multi-degree-of-freedom vibrating screen, s=[α,β,f,μ,K,P], including the grain loss rate μ, uniformity K, dispersion coefficient P monitored by the vibrating screen monitoring system, as well as the screen surface inclination angle α, screen surface horizontal attitude angle β and vibration frequency f of the actuator. Based on the reinforcement learning model, the controller controls the actuator to adjust the screening state so that the loss rate μ, the absolute value of uniformity |K|, and the dispersion P are optimal, thus achieving the optimal screening state.

[0074] In one embodiment of the present invention, the parallel drive mechanism includes four sets of parallel drive components, namely a first parallel drive component 2, a second parallel drive component 12, a third parallel drive component 13, and a fourth parallel drive component 18; each set of drive components includes a stepper drive motor 22, a lead screw 23, a slider 25, a slide base 24, and a laser displacement sensor 26; the slide base 24 of each set of drive components is mounted on the frame, the slider 25 is connected to one end of the suspension rod 20 through a sixth fisheye bearing 21, and the other end of the suspension rod 20 is connected to the vibrating screen 1 through a fifth fisheye bearing 19; the emitting end of the laser displacement sensor 26 is vertically downward mounted on the slider 25, and the displacement measuring plate 27 is mounted on the lower end face of the slide base 24.

[0075] In one embodiment of the present invention, after installation, the four suspension rods 20 should be kept parallel to each other; the screen surface should be kept horizontal; the emitting ends of the four laser displacement sensors 26 should be vertically downward; and the four displacement measuring plates 27 should be on the same horizontal plane.

[0076] In one embodiment of the present invention, the series drive mechanism includes a drive link 7, an eccentric rotating disk 10, and a DC drive motor 9; one end of the drive link 7 is connected to the vibrating screen 1 via a third fisheye bearing 6, and the other end of the drive link 7 is connected to the eccentric rotating disk 10 via a fourth fisheye bearing 8; the eccentric rotating disk 10 is mounted on the output shaft of the DC drive motor 9, and the DC drive motor 9 is mounted on the frame.

[0077] In one embodiment of the present invention, there are two constraint links 4; one end of the two identical constraint links 4 is connected to the frame through a second fisheye bearing 5, and the other end of the two constraint links 4 is connected to the side of the vibrating screen 1 through a first fisheye bearing 3.

[0078] In one embodiment of the present invention, after installation, it should be ensured that: the two constraint links 4 remain parallel to each other, forming a parallelogram mechanism; the length L of the constraint link 4 should be much greater than the eccentricity of the eccentric rotating disk 10.

[0079] Screen movement process:

[0080] The screen surface inclination angle α is defined as the angle of rotation of the vibrating screen around its central X-axis, with clockwise being positive and counterclockwise being negative; the screen surface horizontal attitude angle β is the angle of rotation of the vibrating screen around its central Z-axis, with clockwise being positive and counterclockwise being negative. v1, v2, v3, and v4 are the rotational speeds of the stepper drive motors 22 of the first, second, third, and fourth parallel drive components, respectively; when v2 = v3 = v0, v1 = v4 = -v0, the screen surface rotates around the Y-axis, which changes the screen surface inclination angle α of the vibrating screen; when v1 = v2 = v0, v3 = v4 = -v0, the screen surface rotates around the X-axis, which is the horizontal attitude angle β of the vibrating screen.

[0081] Using four sets of displacement sensors 26, the current screen surface inclination angle α and screen surface horizontal attitude angle β can be calculated. The calculation formula is as follows:

[0082]

[0083]

[0084] Wherein, H1, H2, H3, and H4 are the distances from their emitting ends to the displacement measuring plate 27 detected by the laser displacement sensors 26 of the first, second, third, and fourth parallel driving components, respectively; L X and L Y These are the center distances of the parallel drive components along the X and Y axes, respectively.

[0085] The series drive mechanism drives the vibrating screen 1 to reciprocate. The speed ω of the DC drive motor 9 can be directly controlled by controlling the input voltage of the DC drive motor 9. The DC drive motor 9 integrates a speed encoder, which can directly output the current speed ω and form a closed-loop speed control to realize the vibration frequency of the vibrating screen 1. Real-time monitoring and adjustment.

[0086] In one embodiment of the present invention, the screen surface material distribution monitoring device includes a screen-mounted camera 15, multiple screen surface photosensitive sensors 14, and a screen surface RBG supplementary light 16. The screen surface RBG supplementary light 16 is installed directly above the screen surface to provide real-time supplementary lighting to the screen surface. The screen surface photosensitive sensors 14 are used to detect the light intensity of the screen surface and transmit it to the controller. The controller adjusts the brightness of the RBG supplementary light 16 according to the light intensity. The camera 15 is used to capture images of the material distribution state on the screen surface and transmit them to the controller.

[0087] Material enters from the front of the screen and moves along the X-axis under the vibration of the screen surface. As the material moves towards the rear of the screen, more particles pass through the screen surface, and fewer particles remain on the screen surface. Under uneven feeding conditions, the uneven distribution of particles on the screen surface becomes more pronounced behind the screen surface. Therefore, a camera 15 is installed above and behind the screen, with its field of view vertically downward, ensuring that the camera's field of view 17 covers the middle and rear part of the screen. Within the camera's field of view 17, four screen surface photosensors 14 are arranged on both sides of the screen surface and fixed to the frame to sense the illumination conditions within the current camera's field of view 17.

[0088] A control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning includes the following steps:

[0089] The screen surface material distribution monitoring device collects images of the screen surface material distribution state of the multi-degree-of-freedom vibrating screen and transmits them to the controller. The controller monitors the screen surface material uniformity K and dispersion degree P according to the screen surface material distribution monitoring model.

[0090] The grain loss monitoring device monitors the grain loss of the multi-degree-of-freedom vibrating screen and transmits the data to the controller;

[0091] The controller obtains the current screening state s of the multi-degree-of-freedom vibrating screen from the screen material distribution monitoring device, the grain loss monitoring device, and the actuator of the multi-degree-of-freedom vibrating screen, where s = [α, β, f, μ, K, P]. This includes the grain loss rate μ, uniformity K, and dispersion coefficient P monitored by the vibrating screen monitoring system, as well as the screen surface inclination angle α, screen surface horizontal attitude angle β, and vibration frequency f of the actuator. Based on the control model of the multi-degree-of-freedom vibrating screen device, the actuator adjusts the screening state to optimize the loss rate μ, the absolute value of uniformity |K|, and the dispersion coefficient P, thus achieving the optimal screening state.

[0092] The screen surface material distribution monitoring model is established through the following steps:

[0093] Images of the material distribution on the screen surface of a multi-degree-of-freedom vibrating screen are acquired. The acquired images are binarized to extract particle distribution information. The binarized images are then subjected to dilation and erosion processing to establish a material distribution model on the screen surface. A rectangular plane C with length and width h and l is selected on the image, located at a distance d1 from the front of the screen and d2 from the inner side of the screen wall. A planar rectangular coordinate system is established, and samples are collected sequentially and uniformly on plane C. A linear binary classification model is established using a two-input, one-output neural network. The network contains only one input layer and one output layer. The gradient ascent method based on maximum likelihood estimation is used to train the model until convergence, obtaining the optimal linear equation.

[0094] In one embodiment of the present invention, a method for acquiring an image of the material distribution state on the screen surface is described below. Figure 3 , Figure 4 and Figure 5 The images of the material distribution on the screen surface are binarized to extract particle distribution information. The processed images are then segmented, a linear binary classification problem involving the distribution of black and white pixels, to establish a monitoring model for the material distribution on the screen surface. The specific steps include:

[0095] Adjust the light source:

[0096] Illumination conditions have a significant impact on the acquisition of the screen surface distribution and also affect the adjustment of image processing parameters. To ensure the stability of the screen surface illumination conditions, an adaptive adjustment system for screen surface illumination is established using a screen surface photosensor 14 and a screen surface array RBG supplementary light 16.

[0097] The PID control method is adopted. Based on the difference between the signal collected by the screen surface photosensitive sensor 14 and the set target signal, the voltage value of each lamp of the screen surface array RBG supplementary light 16 is controlled, and the PID model parameters are adjusted to improve the accuracy and sensitivity of the adaptive adjustment of the illumination conditions in each area, so as to ensure the clarity of the monitoring image of the material distribution on the screen surface.

[0098] Image preprocessing:

[0099] The images captured by camera 15 on the screen are RGB images. To improve the robustness of detection, they are converted into HSV images. The conversion formula is as follows:

[0100] R′=R / 255

[0101] G′=G / 255

[0102] B′=B / 255

[0103] Cmax = max(R′, G′, B′)

[0104] Cmin=min(R′,G′,B′)

[0105] Δ=Cmax-Cmin

[0106] H calculation:

[0107]

[0108] S calculation:

[0109]

[0110] V calculation:

[0111] V = Cmax

[0112] The image is binarized based on the material color using an HSV threshold to extract particle distribution information. White pixels represent particles present on the screen surface, while black pixels represent the absence of particles. Since particles are primarily aggregated at the front of the screen with gaps between them, a small number of black pixels exist. As particles move backward, more particles pass through the screen, fewer particles remain on the screen surface, and particle diffusion becomes more complete, resulting in a gradual decrease and eventual disappearance of white pixels. The distribution of white pixels reflects the current material distribution on the screen surface. The binarized image is then subjected to dilation and erosion processing to reduce the influence of particle aggregation gaps on the white pixel distribution. The dilation formula is:

[0113]

[0114] This formula represents the dilation process of image T by F, where T is the image matrix, representing the binarized image of the material distribution on the screen surface, and F is a convolution kernel, preferably with a size of 5×5; the erosion formula is:

[0115]

[0116] This formula represents using F to erode the image T, where T is the image matrix, representing the image of the material distribution on the screen surface after expansion, and F is a convolution kernel, preferably with a size of 5×5.

[0117] Establish a material distribution model on the screen surface:

[0118] To quantify the distribution of materials on the screen surface, a screen surface material distribution model is established. The processed image is then segmented, which is a linear binary classification problem concerning the distribution of black and white pixels. A planar coordinate system is established in the field of view of the processed image, where each pixel represents a coordinate point. There are two types of coordinate points: white pixels and black pixels. The goal is to find an optimal straight line to separate the two types of samples. The optimal straight line is the screen surface material distribution model.

[0119] Since there are walls on both sides of the sieve surface, material tends to accumulate on both sides, affecting the calculation results. Therefore, a rectangular plane C is selected on the image, with a distance of d1 from the front of the sieve and d2 from the inner side of the sieve wall, and its length and width are h and l. A planar rectangular coordinate system (origin and coordinate axes defined) is established, and samples are uniformly collected sequentially on the rectangular plane C to obtain a certain number of samples. The sample elements can be represented as:

[0120] (x i y i , z i )

[0121] Where, x i y is the x-coordinate of the pixel. i Z represents the ordinate of the pixel; when the sample pixel is white, z... i =1, when the sample pixel is black, z i =0.

[0122] A linear binary classification model is established using a two-input, one-output neural network. The network contains only one input layer and one output layer. The input layer contains two input neurons, representing the x-coordinate of the pixel. i and the vertical coordinate y i The output layer has only one output neuron. w1 and w2 are the connection weights between the input layer neuron and the output layer neuron, respectively, and b is the bias term. The Logistic function set is used as the activation function, which can also be called the binary classification function. Therefore, the output of the neural network can be expressed as:

[0123] A i =Sigmoid(w1·x) i +w2·y i +b)

[0124] The binary classification cross-entropy loss function can be expressed as:

[0125]

[0126] Where j is the number of samples; z is the number of samples when the sample pixel is white. i =1, when the sample pixel is black. i =0. The model is trained using the gradient ascent method with maximum likelihood estimation until convergence, thus obtaining the optimal linear equation:

[0127] W1·x+W2·y+B=0

[0128] Among them, the definition Screen surface uniformity refers to the degree of uniformity of material distribution along the Y-axis of the screen surface. The smaller |K| is, the more uneven the material distribution on the screen surface; the larger |K| is, the more uniform the material distribution on the screen surface. definition P is the screen surface dispersion coefficient, which represents the degree of material dispersion along the X-axis of the screen surface. The smaller P is, the lower the degree of dispersion, and the larger P is, the higher the degree of dispersion. P∈(0,h).

[0129] In one embodiment of the present invention, the grain loss monitoring device is described herein. Figure 6 :

[0130] A grain loss monitoring sensor 11 is installed at the tail of the sieve to calculate the number of grains impacted per unit time, N, in grains / s. The grain loss rate μ can be expressed as:

[0131]

[0132] Where, m grain The value is the weight per thousand grains in grams, and M is the feed rate per unit time in kg / s.

[0133] In one embodiment of the present invention, the grain loss monitoring sensor 11 is essentially a grain counting sensor, mainly comprising a sensing element, a metal sensitive plate, and a signal processing circuit. The sensing element uses piezoelectric materials such as piezoelectric ceramics and piezoelectric films. The sensing element is fixed at the center of one side of the metal sensitive plate, while the other side is the grain impact surface. When a grain impacts the sensitive plate, the sensing element converts the periodic oscillation signal generated by the sensitive plate into a charge signal q. The signal processing circuit converts the received charge signal into dual signals, a digital square wave v1 and an analog voltage v2, for output. When the grain impact frequency is low, the digital square wave signal provides higher accuracy for counting grain impacts; when the grain impact frequency is high, the analog voltage signal provides even higher accuracy. The ANFIS algorithm is used to fuse these two signals, which can improve the counting accuracy at different grain impact frequencies. During operation, the monitoring unit is the number of digital square waves v1 and the average value of the analog voltage v2 within a time period. These are input into the ANFIS algorithm model to obtain the number of grain impacts N per unit time, and the current loss rate is calculated according to the formula.

[0134] In one embodiment of the present invention, the control method for a multi-degree-of-freedom vibrating screen device is a deep reinforcement learning-based control method for a multi-degree-of-freedom vibrating screen device, see reference. Figure 7 and Figure 8 :

[0135] This invention employs deep reinforcement learning to establish a control model for a multi-degree-of-freedom vibrating screen device. The control model includes an agent, an environment, the agent's state, the agent's actions, and a reward function R. The environment is defined as the working environment of the multi-degree-of-freedom vibrating screen device control system, i.e., the interior of the cleaning chamber. The agent is defined as the entire multi-degree-of-freedom vibrating screen device control system. The agent's state is the current working state of the multi-degree-of-freedom vibrating screen. The agent's actions are defined as the changes in the multi-degree-of-freedom vibrating screen's working parameters. The multi-degree-of-freedom vibrating screen device control system obtains the current working state of the multi-degree-of-freedom vibrating screen through a screen surface material distribution monitoring device, a grain loss monitoring device, and an actuator. Based on feedback from the working environment, it calculates a quantifiable reward signal using the reward function R and outputs the changes in the multi-degree-of-freedom vibrating screen's working parameters. The multi-degree-of-freedom vibrating screen's working environment is affected by the actions, changing the screening state and generating a new reward. The action strategy is then updated using the feedback screening state and reward, and this iterative process continues until the multi-degree-of-freedom vibrating screen reaches its optimal screening state.

[0136] In one embodiment of the invention, a control model for a multi-degree-of-freedom vibrating screening device is established using deep reinforcement learning (DRL). DRL is an end-to-end perception and control system that mainly includes an agent and an environment. The agent obtains the current state by perceiving the environment and generates an action based on the quantifiable reward signal fed back by the environment. The environment is affected by the action, the state changes, and a new reward is generated. The action strategy is then updated with the feedback state and reward. This process is repeated iteratively until the agent finally learns the optimal action strategy required to complete the task.

[0137] The environment refers to the working environment of the control system of the multi-degree-of-freedom vibrating screen device, which is the working environment inside the cleaning room.

[0138] In this invention, the intelligent agent is the control system of the entire multi-degree-of-freedom vibrating screening device, mainly including a vibrating screen monitoring system, a vibrating screen actuator, and a DQN controller 28. The vibrating screen monitoring system includes a screen surface material distribution monitoring device and a grain loss monitoring device, which can acquire the current status information of the vibrating screen: grain loss rate μ, uniformity K, and dispersion coefficient P. The vibrating screen actuator can monitor and adjust the screen surface inclination angle α, screen surface horizontal attitude angle β, and vibration frequency f, enabling interaction with the environment.

[0139] The state of the agent is the current operating parameters of the vibrating screen.

[0140] The state space of the agent can be represented as S = [α, β, f, μ, K, P], with a spatial dimension of 5; the state space is a continuous space, where:

[0141] α∈(-10°,10°),β∈(-10°,10°),f∈(10Hz,15Hz),μ∈(0,0.1), P∈

[0142] (0,h).

[0143] The agent's action is the change in the vibrating screen's operating parameters (Δα, Δβ, and Δf). The action space of the multi-free vibrating screen is defined as A = [Δα, Δβ, Δf], with a spatial dimension of 3. The action space is then discretized, where:

[0144] Δα=[0,±0.5°,±1°,±1.5°,±2°], Δβ=[0,±0.5°,±1°,±1.5°,±2°], Δf=[0,±0.5Hz,±1Hz,±1.5Hz,±2Hz].

[0145] The reward signal is a correlation function between the current screening state (μ, K, and P) and the rate of change of the screening state (Δμ, ΔK, and ΔP). For the entire screening system, the lower the loss rate μ, the better; the larger the absolute value of uniformity |K|, the better; and the higher the dispersion degree P, the better. Therefore, the lower the loss rate μ, the larger the absolute value of uniformity |K|, and the higher the dispersion degree P, the greater the reward should be, and vice versa. Thus, two types of reward functions are defined: the state reward function R. s and action reward / punishment function R a .

[0146] State reward and punishment function R s The value of evaluating the current state s is only related to the loss rate μ, the absolute value of uniformity |K|, and the dispersion degree P, and is expressed as:

[0147] R s =ρ1·F1(μ)+ρ2·F2(|K|)+ρ3·F3(P);

[0148] Where ρ1, ρ2, and ρ3 are positive constants, F1(μ) is a decreasing function of the loss rate μ, and F2(|K|) and F3(P) are increasing functions of the absolute value of uniformity |K| and the degree of dispersion P, respectively.

[0149] Action reward and punishment function R a The value of action a is evaluated, which is related to the rate of change of loss Δμ, the rate of change of uniformity ΔK, and the rate of change of dispersion ΔP, and is expressed as:

[0150] Ra =σ1·G1(Δμ)+σ2·G2(Δ|K|)+σ3·G3(ΔP);

[0151] Where σ1, σ2, and σ3 are positive constants, G1(Δμ) is an odd function of the rate of change of loss Δμ, and G1(Δμ) is continuous and monotonically decreasing; G2(Δ|K|) and G3(ΔP) are odd functions of the rate of change of the absolute value of uniformity Δ|K| and the rate of change of dispersion ΔP, respectively, and G2(Δ|K|) and G3(ΔP) are continuous and monotonically increasing. The total prize function R can be expressed as: R = R s +R a .

[0152] The total prize function R can be expressed as:

[0153] R = R s +R a

[0154] In one embodiment of the invention, the control model of the multi-degree-of-freedom vibrating screening device is a reinforcement learning model.

[0155] In one embodiment of the invention, the reinforcement learning model is a DQN model, but is not limited thereto; the controller is a DQN controller 28.

[0156] The reinforcement learning model is built through the following steps:

[0157] The DQN controller 28 reads the current screening state s of the multi-degree-of-freedom vibrating screen from the screen material distribution monitoring device and the actuator, and predicts the value of each action in the current state through the eval_net network. It adopts the ε-greedy strategy to select the next action a to be executed and sends it to the vibrating screen actuator for execution. After time t, the screening state is sampled again to obtain the next screening state s'. The reward r is calculated according to the reward function R. The acquired experience s, a, s', and r are stored in the experience database. A portion of experience is randomly extracted from the experience database to train the deep Q network. According to the loss function L(W), the eval_net network is trained using SGD stochastic gradient descent. Every N training iterations, W' = W is set to update the eval_net network parameters W. S = s' is set to iteratively train the reinforcement learning model until the reinforcement learning model converges. When the reinforcement learning model converges, the DQN controller 28 controls the actuator to adjust the screening state according to the current screening state s, so that the loss rate μ, the absolute value of uniformity |K|, and the dispersion degree P are optimal, thus achieving the optimal screening state.

[0158] The reinforcement learning model used in this invention is the DQN model. Compared with traditional reinforcement learning models, the DQN algorithm combines q-learning algorithm with deep learning. It uses a neural network to approximate the Q-value function, thus solving the limitation of Q-table dimension. At the same time, DQN adopts an experience replay mechanism during the learning process and uses another neural network to generate the target Q-value, which greatly improves the stability of the algorithm.

[0159] The DQN model mainly consists of an experience base, a deep Q-network, and a selection policy. The experience base is used to store past experiences, and the experience storage format is: [s, a, s', r], where s is the current screening state, s∈S; a is the action selected by the agent, a∈A; s' is the next screening state after executing the action, s'∈S; and r is the reward value obtained after executing the action, which is the output value of the reward function R.

[0160] In one embodiment of the present invention, the deep Q-network includes two neural networks, eval_net and target_net, with identical structures but different parameters. Let the parameters of eval_net be W, and the parameters of target_net be W'. The number of nodes in their input and output layers is determined by the agent's state space S and action space A. The spatial dimension of state space S is 5, so the number of nodes in the neural network's input layer is 5. The spatial dimension of state space A is 3, and each action space is discretized into 5 sub-actions, therefore there are a total of 5 sub-actions. 3 = 125 actions, so the output layer of the neural network has 125 nodes. The output value of each node corresponds to the Q-value of one action. The number of hidden layer nodes should be greater than the number of output layer nodes. Define the network loss function as:

[0161] L(W)=E[(r+γmax a’ target_net(s',a',W'))-eval_net(s,a,W)] 2

[0162] Where r is the calculated reward value, γ is the learning step size, and max is the maximum reward value. a’ `target_net(s', a', W')` represents the maximum Q-value of the `target_net` network when the parameters are W' and the input is s'. `a'` is the action corresponding to the maximum Q-value, and `r+γmax` is the maximum Q-value. a’ target_net(s', a', W') is also called the target Q value; eval_net(s, a, W) is the Q value of the eval_net network for action a when the parameters are W and the input is s. eval_net(s, a, W) is also called the estimated Q value.

[0163] An ε-greedy strategy is used to select actions. The 1-ε probability selects the action a with the largest Q value according to the Q value calculated by eval_net, and the ε probability randomly selects an action a, where ε < 1 and decreases with the increase of training times. This action selection strategy can improve the exploration ability of the agent, avoid getting trapped in local optima, and ensure the stability of the system operation.

[0164] In one embodiment of the present invention, the DQN model training process is as follows:

[0165] During training, the agent randomly extracts some experience from the experience base to train the deep Q-network. The more experiences the experience base traversed, the higher the prediction accuracy of the deep Q-network, and the stronger the stability and adaptability of the system. The specific steps are as follows:

[0166] ① Initialize the Q network, that is, generate two neural networks with the same structure but different parameters: eval_net (parameter W) and target_net (parameter W').

[0167] ② The intelligent agent (i.e., the control system of the multi-degree-of-freedom vibrating screen) collects the current screening status s.

[0168] ③ Use the ε-greedy strategy to select action a.

[0169] ④ After time t, the intelligent agent (i.e., the control system of the multi-degree-of-freedom vibrating screening device) collects the next screening state s', calculates the reward value r according to the reward function R, and stores the obtained experience in the form of [s,a,s',r] into the experience database.

[0170] ⑤ If the amount of experience data is less than n, then let s = s′ and repeat steps ③, ④, and ⑤ until the requirements are met. The experience database is a finite database. When the storage limit is reached, new experiences will replace old experiences. This is why the DQN model can continuously learn.

[0171] ⑥ Randomly select m memories from the experience base, calculate the loss function L(W), update the eval_net network parameters W using SGD stochastic gradient descent, let s = s', repeat steps ③④⑤⑥ to iteratively train the model, and let W' = W every N training iterations until the model converges.

[0172] In one embodiment of the present invention, the operation process of the control system of the multi-degree-of-freedom vibrating screening device is as follows:

[0173] The established DQN model is embedded in the DQN controller 28. During operation, the DQN controller 28 reads the current screening state s of the multi-degree-of-freedom vibrating screen from the screen material distribution monitoring device and the actuator, and predicts the value of each action in the current state through the eval_net network. It then selects the next action a to be executed using an ε-greedy strategy and sends it to the vibrating screen actuator for execution. After time t, the screening state s' is sampled again, the reward r is calculated, and the acquired experience is stored in the experience base. The agent then retrieves data from the experience base... The machine extracts some experience to train the deep Q-Net network. According to the defined loss function, SGD stochastic gradient descent is used to train the eval_net network. Every N training iterations, W' = W is set to update the eval_net network parameters. Then, s = s' is set to iterate the training of the model until the model converges. When the model converges, the DQN controller can automatically adjust the screen surface inclination angle α and the screen surface horizontal attitude angle β according to the current screening state s, so that the loss rate tends to the minimum value, the absolute value of uniformity |K| tends to π / 2, and the dispersion degree P tends to h, thus achieving the optimal screening state.

[0174] In one embodiment of the present invention, a camera 15 is used to acquire the motion of the threshing mixture on the vibrating screen surface in real time. An image processing technology is used to establish a monitoring model for the material distribution on the screen surface, and an evaluation index for the material distribution on the screen surface is defined. Reinforcement learning is introduced into the control of the multi-degree-of-freedom vibrating screening device, and information fusion is performed on the material distribution on the screen surface and the screening loss rate, realizing the self-learning of the control system. This has important theoretical research significance and practical value for improving the applicability and stability of the harvester in different environments. Currently, no publicly available research reports have been found.

[0175] In one embodiment of the present invention, a control system for a multi-degree-of-freedom vibrating screen of a grain combine harvester is established using deep reinforcement learning. This system includes a multi-degree-of-freedom vibrating screen, a screen surface material distribution monitoring device, a grain loss monitoring device, and a deep reinforcement learning-based intelligent controller. The multi-degree-of-freedom vibrating screen employs a series-parallel hybrid mechanism, achieving two translations and two rotations of the screen surface while ensuring structural strength. The screen surface monitoring system includes a screen-mounted camera 15, a screen surface photosensor 14, and a screen surface supplementary light 16. A screen surface material distribution monitoring model is established, enabling real-time monitoring of the screen surface distribution. The grain loss monitoring system mainly includes a grain counting sensor and a signal conditioning circuit, enabling real-time monitoring of grain loss at the tail end of the screen. The controller utilizes the Jetson Xavier NX deep learning platform, which can collect real-time data on the material movement and grain loss on the screen surface, and obtain the screen surface material distribution status through image processing. The controller integrates a deep reinforcement learning model, constructing a reward function based on the screen surface distribution status and grain loss. It possesses self-learning capabilities, continuously optimizing the control model during operation to improve the harvester's applicability and stability in different operating environments.

[0176] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This way of describing the specification is only for clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

[0177] The detailed descriptions listed above are merely specific illustrations of feasible embodiments of the present invention and are not intended to limit the scope of protection of the present invention. All equivalent embodiments or modifications made without departing from the spirit of the present invention should be included within the scope of protection of the present invention.

Claims

1. A control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning, characterized in that, Includes the following steps: The screen surface material distribution monitoring device collects images of the screen surface material distribution state of a multi-degree-of-freedom hybrid vibrating screen and transmits them to the controller. The controller monitors the screen surface material uniformity K and dispersion degree P according to the screen surface material distribution monitoring model. The grain loss monitoring device monitors the grain loss of the multi-degree-of-freedom hybrid vibrating screen and transmits the data to the controller; The controller obtains the current screening state s of the multi-degree-of-freedom vibrating screen from the screen material distribution monitoring device, the grain loss monitoring device, and the actuator of the multi-degree-of-freedom hybrid vibrating screen, where s = [α, β, f, μ, K, P], including grain loss rate μ, uniformity K, dispersion coefficient P, screen surface inclination angle α, screen surface horizontal attitude angle β, and vibration frequency f. Based on the control model of the multi-degree-of-freedom vibrating screen device based on reinforcement learning, the controller controls the actuator to adjust the screening state so that the loss rate μ, the absolute value of uniformity |K|, and the dispersion P are optimal, thus achieving the optimal screening state. The screen surface material distribution monitoring model is established through the following steps: Images of the material distribution on the screen surface of a multi-degree-of-freedom vibrating screen are acquired. The acquired images are binarized to extract particle distribution information. The binarized images are then subjected to dilation and erosion processing to establish a material distribution model on the screen surface. A rectangular plane C with length and width h and l is selected on the image, located at a distance d1 from the front of the screen and d2 from the inner side of the screen wall. A planar rectangular coordinate system is established, and samples are collected sequentially and uniformly on plane C. A linear binary classification model is established using a two-input, one-output neural network. The network contains only one input layer and one output layer. The gradient ascent method based on maximum likelihood estimation is used to train the model until convergence, obtaining the optimal linear equation. The control model of the multi-degree-of-freedom vibrating screen is established using deep reinforcement learning. The control model includes an agent, an environment, the agent's state, the agent's actions, and a reward function R. The environment is defined as the working environment of the multi-degree-of-freedom vibrating screen control system, i.e., the interior of the cleaning chamber. The agent is defined as the entire multi-degree-of-freedom vibrating screen control system. The agent's state is the current working state of the multi-degree-of-freedom vibrating screen. The agent's actions are defined as the changes in the working parameters of the multi-degree-of-freedom vibrating screen. The multi-degree-of-freedom vibrating screen control system obtains the current working state of the multi-degree-of-freedom vibrating screen through the screen surface material distribution monitoring device, the grain loss monitoring device, and the actuator. Based on feedback from the working environment, it calculates a quantifiable reward signal through the reward function R and outputs the changes in the working parameters of the multi-degree-of-freedom vibrating screen. The working environment of the multi-degree-of-freedom vibrating screen is affected by the actions, the screening state changes, and a new reward is generated. The action strategy is then updated using the feedback screening state and reward, and this process is iterated until the multi-degree-of-freedom vibrating screen reaches its optimal screening state.

2. The control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning according to claim 1, characterized in that, The optimal straight line equation is: W1·x+W2·y+B=0 Where W1 is the connection weight of the optimal input layer neuron, W2 is the connection weight of the optimal output layer neuron, B is the bias term, and x and y are the horizontal and vertical coordinates of the pixel in the image coordinate system, respectively.

3. The control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning according to claim 2, characterized in that, definition Screen surface uniformity refers to the degree of uniformity of material distribution along the Y-axis of the screen surface. definition Let P be the dispersion coefficient of the sieve surface, which is P∈(0,h).

4. The control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning according to claim 3, characterized in that, The optimal sieving state is as follows: the loss rate μ approaches its minimum value, the absolute value of uniformity |K| approaches π / 2, and the degree of dispersion P approaches h.

5. The control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning according to claim 1, characterized in that, In the sieving state s=[α,β,f,μ,K,P], α∈(-10°,10°), β∈(-10°,10°), f∈(10Hz,15Hz), μ∈(0,0.1), P∈(0,h).

6. The control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning according to claim 5, characterized in that, The changes in the screen surface inclination angle α, screen surface horizontal attitude angle β, and vibration frequency f of the actuator of the multi-degree-of-freedom vibrating screen are Δα, Δβ, and Δf, respectively. The motion space of the multi-degree-of-freedom vibrating screen is defined as A = [Δα, Δβ, Δf], and the motion space is discretized, where: Δα=[0,±0.5°,±1°,±1.5°,±2°], Δβ=[0,±0.5°,±1°,±1.5°,±2°], Δf=[0,±0.5Hz,±1Hz,±1.5Hz,±2Hz].

7. The control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning according to claim 4, characterized in that, The reward function R is a correlation function of the current grain loss rate μ, uniformity K, dispersion coefficient P, and the sieving state change rates Δμ, ΔK, and ΔP. The reward function R includes the state reward / penalty function R. s and action reward / punishment function R a : R=R s +R a ; State reward and punishment function R s The value of evaluating the current screening state s is expressed as: R s =ρ1·F1(μ)+ρ2·F2(|K|)+ρ3·F3(P); Where ρ1, ρ2, ρ3 are positive constants, F1(μ) is a decreasing function of the loss rate μ, F2(|K|) is an increasing function of the absolute value of uniformity |K|, and F3(P) is an increasing function of the dispersion degree P. Action reward and punishment function R a It evaluates the value brought by action 'a', expressed as: R a =σ1·G1(Δμ)+σ2·G2(Δ|K|)+σ3·G3(ΔP); Where σ1, σ2, and σ3 are positive constants, G1(Δμ) is an odd function of the rate of change of loss Δμ, and G1(Δμ) is continuous and monotonically decreasing; G2(Δ|K|) is an odd function of the rate of change of the absolute value of uniformity Δ|K|, and G3(ΔP) is an odd function of the rate of change of dispersion ΔP, and G2(Δ|K|) and G3(ΔP) are continuous and monotonically increasing.

8. The control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning according to claim 1, characterized in that, The control model of the multi-degree-of-freedom vibrating screening device is a reinforcement learning model.

9. The control method for a multi-degree-of-freedom vibrating screening device based on reinforcement learning according to claim 8, characterized in that, The reinforcement learning model is a DQN model, and the controller is a DQN controller (28); The DQN model is built using the following steps: The DQN controller (28) reads the current screening state s of the multi-degree-of-freedom vibrating screen from the screen surface material distribution monitoring device and the actuator, and predicts the value of each action in the current state through the eval_net network. It adopts the ε-greedy strategy to select the next action a to be executed and sends it to the vibrating screen actuator for execution. After time t, the screening state is sampled again to obtain the next screening state s'. The reward r is calculated according to the reward function R. The obtained experience s, a, s', and r are stored in the experience database. The DQN controller (28) randomly extracts the next action a from the experience database. Some experience is used to train the deep Q network. According to the loss function L(W), the eval_net network is trained by SGD stochastic gradient descent. Every N training iterations, W' = W is set to update the eval_net network parameters W. Then, s = s' is set to iteratively train the reinforcement learning model until the reinforcement learning model converges. After the reinforcement learning model converges, the DQN controller (28) controls the actuator to adjust the screening state according to the current screening state s, so that the loss rate μ, the absolute value of uniformity |K|, and the dispersion degree P are optimal, thus achieving the optimal screening state.

10. A system for implementing the control method of a multi-degree-of-freedom vibrating screening device based on reinforcement learning as described in any one of claims 1-9, characterized in that, This includes a multi-degree-of-freedom vibrating screen, a screen surface material distribution monitoring device, a grain loss monitoring device, and a controller; The multi-degree-of-freedom vibrating screen includes a vibrating screen (1), a parallel drive mechanism, a series drive mechanism, and a constraint link (4), which can realize two translations and two rotations. The parallel drive mechanism realizes three-degree-of-freedom motion of the screen surface around the X-axis, Y-axis and Z-axis rotation, and the series mechanism realizes one-degree-of-freedom reciprocating motion of the screen surface. One end of the constraint link (4) is connected to the frame and the other end is connected to the side of the vibrating screen (1). The screen surface material distribution monitoring device collects images of the screen surface material distribution state of the multi-degree-of-freedom vibrating screen and transmits them to the controller. The controller monitors the screen surface material uniformity K and dispersion degree P according to the screen surface material distribution monitoring model. The grain loss monitoring device monitors the grain loss of the multi-degree-of-freedom vibrating screen and transmits the data to the controller; The controller obtains the current screening state s of the multi-degree-of-freedom vibrating screen from the screen material distribution monitoring device and the actuator of the multi-degree-of-freedom vibrating screen, s=[α,β,f,μ,K,P], including the grain loss rate μ, uniformity K, dispersion coefficient P monitored by the vibrating screen monitoring system, as well as the screen surface inclination angle α, screen surface horizontal attitude angle β and vibration frequency f of the actuator, and controls the actuator to adjust the screening state according to the reinforcement learning model, so that the loss rate μ, the absolute value of uniformity |K|, and the dispersion P are optimal, thus achieving the optimal screening state.

11. The system of the multi-degree-of-freedom vibrating screening device control method based on reinforcement learning according to claim 10, characterized in that, The parallel drive mechanism includes four sets of parallel drive components, namely the first parallel drive component (2), the second parallel drive component (12), the third parallel drive component (13) and the fourth parallel drive component (18); Each set of driving components includes a stepper motor (22), a lead screw (23), a slider (25), a slide base (24), and a laser displacement sensor (26); the slide base (24) of each set of driving components is mounted on the frame, the slider (25) is connected to one end of the boom (20) through the sixth fisheye bearing (21), and the other end of the boom (20) is connected to the vibrating screen (1) through the fifth fisheye bearing (19); the laser displacement sensor (26) is mounted vertically downward on the slider (25), and the displacement measuring plate (27) is mounted on the lower end face of the slide base (24).

12. The system of the multi-degree-of-freedom vibrating screening device control method based on reinforcement learning according to claim 10, characterized in that, The series drive mechanism includes a drive link (7), an eccentric rotating disk (10), and a DC drive motor (9); One end of the drive link (7) is connected to the vibrating screen (1) through the third fisheye bearing (6), and the other end of the drive link (7) is connected to the eccentric rotating disk (10) through the fourth fisheye bearing (8); the eccentric rotating disk (10) is mounted on the output shaft of the DC drive motor (9), and the DC drive motor (9) is mounted on the frame.

13. The system of the multi-degree-of-freedom vibrating screening device control method based on reinforcement learning according to claim 10, characterized in that, The number of constraint links (4) is two; One end of each of the two identical constraint rods (4) is connected to the frame via a second fisheye bearing (5), and the other end of each of the two constraint rods (4) is connected to the side of the vibrating screen (1) via a first fisheye bearing (3).

14. The system of the multi-degree-of-freedom vibrating screening device control method based on reinforcement learning according to claim 10, characterized in that, The controller calculates the screen surface tilt angle α and the screen surface horizontal attitude angle β according to the following formulas; Among them, H1, H2, H3, and H4 are the distances from their emitting ends to the displacement measuring plate (27) detected by the four laser displacement sensors (26); L X and L Y These are the center distances of the parallel drive components along the X and Y axes, respectively.

15. The system of the multi-degree-of-freedom vibrating screening device control method based on reinforcement learning according to claim 10, characterized in that, The controller calculates the vibration frequency of the vibrating screen (1) according to the following formula: ω is the rotational speed of the DC drive motor (9).

16. The system of the multi-degree-of-freedom vibrating screening device control method based on reinforcement learning according to claim 10, characterized in that, The screen surface material distribution monitoring device includes a screen camera (15), multiple screen surface photosensitive sensors (14), and a screen surface RBG supplementary light (16). The screen surface RBG supplementary light (16) is installed above the screen surface to provide real-time supplementary lighting to the screen surface. The screen surface photosensitive sensors (14) are used to detect the light intensity of the screen surface and transmit it to the controller. The controller adjusts the brightness of the RBG supplementary light (16) according to the light intensity. The camera (15) is used to capture images of the screen surface material distribution status and transmit them to the controller.