Control method of field automatic following transportation platform based on human body posture interaction
Through the combination of depth cameras and RGB cameras, real-time identification of dynamic postures of workers and precise control of the platform are achieved, and the problems of insufficient positioning accuracy and lack of dynamic interaction capabilities in the existing technology are solved, and the efficiency and safety of agricultural operations are improved.
Patent Information
- Application Number
- CN202510435363.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-08-12
AI Technical Summary
The existing agricultural automatic follower platform lacks positioning and navigation accuracy in complex agricultural environments, lacks real-time identification of the dynamic postures of workers, resulting in insufficient environmental adaptability and human-machine collaboration flexibility, and high cross-scene migration costs.
The depth camera and RGB camera work together, combined with deep learning human posture recognition technology and adaptive control algorithm, through quantitative analysis of the geometric relationship between human key points, natural movements are transformed into control instructions, real-time analysis of the dynamic behavior of workers and precise control of platform movement.
It improves the reliability of human-computer interaction in complex farmland scenarios, improves agricultural operation efficiency, reduces the labor intensity of workers, and maintains stable working performance under variable light and terrain conditions.
Smart Images

Figure CN120469398A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of agricultural production technology, relates to the field of automatic following and human posture judgment, and specifically relates to a control method for a large-scale automatic following transportation platform based on human posture interaction. Background Art
[0002] With the rapid development of agricultural automation and intelligence, agricultural automatic follow-up transportation platforms have shown great potential in improving operational efficiency and reducing labor intensity. However, existing technologies generally have the following core problems:
[0003] 1. Inadequate environmental adaptability: In complex agricultural environments such as crop occlusion, undulating terrain, and variable lighting, the system exhibits poor robustness, resulting in a significant decrease in positioning and navigation accuracy.
[0004] 2. Lack of dynamic interaction capabilities: Existing agricultural automatic following platforms lack the ability to accurately perceive the real-time posture of operators and are unable to dynamically adjust control strategies based on the operator's movements, resulting in insufficient flexibility in human-machine collaboration.
[0005] 3. Limited scenario generalization: Existing methods are mostly designed for specific crops or single scenarios. Cross-scenario migration requires readjustment of hardware or algorithms, which is costly and inefficient.
[0006] For example, CN105404299A discloses a labor-saving automatic following operation platform for greenhouses based on somatosensory sensors. It obtains the position of the operator through a combination of skeletal tracking and depth images, and avoids obstacles based on the height and length of obstacles. However, this solution relies on preset slope judgment for obstacle recognition, which is difficult to adapt to complex agricultural environments, such as different terrains, changes in crop growth status, etc. For example, when crops are blocked or the terrain is undulating, the accuracy of obstacle recognition will be greatly reduced, affecting the stability and operation efficiency of the platform. At the same time, the solution lacks real-time posture recognition function and cannot accurately perceive the dynamic behavior of the operator, which in turn affects the smoothness and safety of the operation.
[0007] In addition, CN115568332A discloses an automatic following transportation platform for field environments and its control method, which combines UWB sensors and machine vision technology, locates workers through UWB tags, and uses image processing to extract crop row navigation baselines, thereby realizing automatic following transportation. However, the positioning accuracy of UWB sensors is easily affected by crop occlusion, terrain undulations and multipath interference, resulting in fluctuations in positioning accuracy, making it difficult to ensure stable following of the platform. In terms of image processing, traditional methods have poor adaptability to environmental conditions such as complex lighting, crop density, and terrain changes, making it difficult to stably identify crop rows, affecting the navigation accuracy of the platform and the automatic driving of crop rows. At the same time, this technology also lacks the real-time recognition function of the dynamic posture of the workers, and cannot respond quickly and accurately according to the actual actions of the workers. Summary of the Invention
[0008] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a control method for a large-scale automatic following transport platform based on human posture interaction, aiming to solve the problems of the existing automatic following platform in complex agricultural operation environments with insufficient accurate tracking, environmental adaptability and intelligent control. The present invention realizes real-time analysis of the dynamic behavior of operators and precise control of platform movement through the collaborative work of depth cameras and RGB cameras, multi-perspective perception fusion, combined with deep learning human posture recognition technology and adaptive control algorithms. Through quantitative analysis of the geometric relationship of key points of the human body, natural movements are converted into control instructions, significantly improving the reliability of human-computer interaction in complex farmland scenes.
[0009] The purpose of the present invention also includes: to improve the efficiency of agricultural operations and reduce the labor intensity of operators through an intelligent human-machine collaborative mechanism. Based on the safety distance threshold control and real-time posture response strategy, an adaptive safety buffer zone is constructed between the operator and the transportation platform to effectively avoid operational risks caused by sudden movement or environmental interference. Through the modular visual perception architecture and scalable control logic design, the system has the ability to migrate across scenarios. It is not only suitable for the harvesting and transportation of field crops, but also can adapt to diversified agricultural operation needs such as plant management in greenhouse environments and fruit handling in orchard scenes, and maintain stable working performance under complex terrain, variable lighting and crop growth differences.
[0010] The hardware used in this invention primarily consists of front and rear RGB cameras, a depth camera, a two-axis gimbal, a host computer, a slave computer (STM32 chip), and an electric crawler chassis. The front and rear RGB cameras and the host computer form the crop row detection module, while the depth camera and the host computer form the human object detection module and the human posture detection module. The depth camera, two-axis gimbal, and the host computer form the human tracking module and the human position detection module.
[0011] The technical solution includes the following steps:
[0012] Step S1: Crop row detection:
[0013] The front RGB camera and rear RGB camera are used to collect images in front of and behind the crop rows at preset downward tilt angles. The front camera is used for visual alignment when the platform moves forward, and the rear camera is used for visual alignment when the platform moves backward. The camera field of view is adjusted to cover three ridges of crop rows according to the height of the vehicle body, and the crop rows occupy 60%-80% of the lower edge of the camera field of view. The image is segmented in real time based on the Fast-SCNN image segmentation algorithm, and the crop rows and background are binarized, with the crop rows in white and the background in black. The ROI area is set at 70%-90% below the image to focus on the turning trend of the nearest crop row. The largest white connected domain is calculated, and its center of mass point C is extracted as the crop row wheel D based on the contour centroid algorithm. x Calculate the pixel horizontal coordinate C of the crop row contour center point C x and the center point R of the RGB camera field of view 0' The pixel horizontal coordinate R 0'x Deviation value D x , the deviation value D x Directly reflects the degree of deviation between the platform and the center line of the crop row. The deviation value D x The formula is as follows:
[0014] D x =R 0'x -C x
[0015] The deviation value D x The PID controller is input to generate a control signal for the platform's direction of travel through dynamic adjustment of the proportional (P), integral (I), and differential (D) parameters, enabling precise navigation in both forward and reverse directions.
[0016] Step S2: Human target detection:
[0017] A two-axis gimbal is installed at the center of the platform surface at a height of h from the platform. The depth camera is installed on the two-axis gimbal, and the image data of the operator is collected in real time through but not limited to the Intel D435 depth camera to ensure stable depth and RGB information in complex farmland environments. The image data is processed using but not limited to the YOLO11-Pose deep learning model to achieve real-time detection and key point recognition of human targets; the YOLO11-Pose model is trained based on the COCO dataset and can accurately locate 17 key points of the human body (including shoulders, elbows, wrists, hips and knees, etc.). Based on the human key points defined in the COCO dataset, the midpoint of the hip joint is extracted as the human center point H is calculated as follows:
[0018]
[0019] Among them, the coordinates of H are (Hx, Hy), K 11_x is the horizontal coordinate of the left hip, K 11_y is the vertical coordinate of the left hip, K 12_x is the horizontal coordinate of the right hip, K 12_y is the vertical coordinate of the right hip.
[0020] Calculate the pixel coordinates (Hx, Hy) of the center point H of the human body and the pixel coordinates (R 0x ,R 0y )'s lateral offset value D HR_x With the longitudinal offset value D HR_y The lateral offset value D is transmitted to the human body tracking module. HR_x With the longitudinal offset value D HR_y Calculated by the following formula:
[0021] D HR_x =H x -R 0x
[0022] D HR_y =H y -R 0y ;
[0023] Step S3: Human body tracking and pan / tilt control:
[0024] Based on the lateral displacement D between the center point H of the human body and the center point R0 of the depth camera field of view in step S2 HR_x With the longitudinal offset value D HR_y The PID control algorithm dynamically adjusts the pitch angle θ and yaw angle φ of the two-axis gimbal. Specifically, it includes the following: the proportional term (P): generates an immediate adjustment signal based on the current offset value; the integral term (I): accumulates historical offset errors to eliminate static deviations; and the differential term (D): predicts the offset change trend and suppresses overshoot oscillations. The two-axis gimbal rotates to follow the human target, and the human target is always located in the center of the depth camera's field of view. By coordinating the pitch angle θ and yaw angle φ of the two-axis gimbal, the human target is kept in the center of the depth camera's field of view in real time, ensuring the continuity and stability of tracking. The adjusted pitch angle θ and yaw angle φ are output to the human position detection module in real time, providing dynamic parameters for subsequent three-dimensional coordinate conversion.
[0025] S4. Calculation of the distance between the human body and the platform:
[0026] The distance D between the human body and the platform is defined as the difference between the projection distance ox from the center point H of the operator's body in the forward direction of the platform and the installation distance q between the depth camera and the front end of the platform. The distance D between the human body and the platform is calculated as follows:
[0027] D=ox-q
[0028] The projection distance ox between the operator and the depth camera in the platform's forward direction is calculated using a three-dimensional coordinate conversion formula. Combining the depth value d from the center of the human body to the depth camera obtained by the depth camera and the real-time rotation angle of the gimbal (pitch angle θ and yaw angle φ), the three-dimensional coordinate conversion formula is:
[0029] The calculation formula for the three-dimensional coordinate transformation is:
[0030]
[0031] Among them, based on the right-hand coordinate system definition, the origin o is the position of the platform, the x-axis points to the front of the platform, the y-axis points to the left of the platform, and the z-axis points to the top of the platform; d is the depth value of the operator's body center point, θ is the gimbal pitch angle, φ is the gimbal yaw angle, and ox is the projection distance of the operator in the vehicle's forward direction.
[0032] S5. Human posture detection and command output:
[0033] Based on the key point coordinates, the operator's posture is judged and instructions are issued according to the following conditional rules:
[0034] (1) When a bending posture is detected, the platform brakes immediately. This state corresponds to the operator performing a bending cutting action. The stationary platform can provide a safe and stable working space for the operator. When the lowest wrist is lower than the highest knee or the lowest elbow is lower than the highest hip, the system determines it as the bending posture. The judgment formula is:
[0035] (max(K 9_y ,K 10_y )>min(K 13_y ,K 14_y )) or (max(K 7_y ,K 8_y )>min(K 11_y ,K 12_y ))
[0036] (2) When one arm is raised above the shoulder, this state corresponds to the operator needing to move the platform while it is stationary, including moving away from the platform or approaching the platform. In addition to braking, the gimbal continues to track the operator to ensure that the camera's field of view is not lost. This design is aimed at the pause operation scenario to ensure that the platform continues to obtain the latest action instructions of the mobile operator during braking. When the left wrist is higher than the left shoulder or the right wrist is higher than the right shoulder, the system determines it as the single-handed posture, and the judgment formula is:
[0037] (K 9_y <K 5_y ) or (K 10_y <K 6_y )
[0038] (3) When the key points of both wrists are higher than the shoulders, this state corresponds to a full load and needs to be withdrawn to the edge of the field for unloading, and the platform switches to the reverse mode. This state is suitable for harvesting a full load and returning to unload. The operator raises both hands to command the retreat. If he needs to stop, he can quickly switch to raising one hand for braking, which is both flexible and safe. When the key points of both wrists are higher than the shoulders, the system determines that it is the above-mentioned raised hands posture, and the judgment formula is:
[0039] (K 9_y <K 5_y ) and (K 10_y <K 6_y )
[0040] (4) When the operator is in a standing position, the system determines that the operator is searching for mature crops or fruits while moving. At this time, the platform follows the operator and moves forward. The platform is always behind the operator and the distance between the platform and the operator is always within the safety threshold D to avoid interfering with the operator's search field of view. When a human target is detected and none of the above conditions are met, the system determines that it is in the standing position.
[0041] S6. Platform motion control:
[0042] A safety distance threshold D' is preset based on the requirements of the operation scenario. The threshold is determined through experimental data and ergonomic analysis to ensure that the safe distance between the platform and the operator meets both safety requirements and the convenience requirements of selective harvesting. Setting the safety distance threshold D' receives the real-time distance ox output in step S4 and the control instruction generated in step S5:
[0043] (1) Row-to-row advance condition: When D>D', and the crop row end has not been reached, and a standing posture is detected, the row-to-row advance command is triggered;
[0044] (2) Braking conditions: When any one or more of the following conditions are met, the platform braking is triggered immediately: when D ≤ D', or no human target is detected, or a hand-raising gesture is detected, or the crop row detection is lost, that is, no crop row is detected in the ROI area;
[0045] (3) Row-by-row retreat condition: When the rear camera detects a crop row and detects that the operator is in the posture of raising both hands, the row-by-row retreat command is triggered.
[0046] Furthermore, the image segmentation algorithm described in step S1 is a Fast-SCNN. Its training and optimization methods are as follows: A training dataset with high generalization capabilities is constructed by collecting crop row images in a variety of weather conditions (sunny, cloudy, rainy), different lighting conditions (strong light, weak light, backlight), and diverse terrain scenarios (flat farmland, sloping fields, and undulating ridges). Data augmentation techniques (such as random rotation, brightness adjustment, and noise addition) are used during training to improve model robustness and ensure crop row recognition accuracy in complex environments.
[0047] Furthermore, the human body detection model described in step S2 is YOLO11-Pose, and the human body detection weights used are trained by the COCO dataset, which contains standardized human body key point annotation information (17 key points).
[0048] Furthermore, the human body key points defined in the COCO dataset in steps S2 and S5 include:
[0049] Shoulder key points: key points 5 and 6;
[0050] Elbow key points: key points 7 and 8;
[0051] Wrist key points: key points 9 and 10;
[0052] Hip key points: key points 11 and 12;
[0053] Knee key points: key points 13 and 14.
[0054] Furthermore, in steps S2 and S5, the coordinates of the key points of the human body and the coordinate system are defined as follows:
[0055] Left shoulder coordinates: (K 5_x ,K 5_y ), right shoulder coordinates: (K 6_x ,K 6_y );
[0056] Left elbow coordinates: (K 7_x ,K 7_y ), right elbow coordinates: (K 8_x ,K 8_y );
[0057] Left wrist coordinates: (K 9_x ,K 9_y ), right wrist coordinates: (K 10_x ,K 10_y );
[0058] Left hip coordinates: (K 11_x ,K 11_y ), right hip coordinates: (K 12_x ,K 12_y );
[0059] Left knee coordinates: (K 13_x ,K 13_y ), right knee coordinates: (K 14_x ,K 14_y );
[0060] The coordinate system is a pixel coordinate system, with its origin at the upper left corner of the image and the positive y-axis pointing downward.
[0061] Furthermore, in step S6, the safety distance threshold D' is set by the following experiment:
[0062] The minimum safe operating space W was measured for 20 workers of different body types (1.60-1.85m tall) while they performed agricultural operations (standing, bending, swinging a knife, harvesting, carrying, etc.). Combined with the emergency braking distance S when the platform was fully loaded, the safety threshold D' was ultimately determined using the following formula:
[0063] D'=W+S.
[0064] The beneficial effect of the present invention is that the present invention significantly improves the efficiency of agricultural operations and reduces the labor intensity of operators by combining an automatic alignment system, a human tracking system, a position detection system and a human posture interaction system. In agricultural operations, the transport platform can accurately and intelligently follow the operators, reduce manual intervention, and improve the progress of operations. At the same time, the platform can dynamically adjust its motion state to ensure a safe distance between the operators and the platform, thereby improving the safety of the operation. The combination of a depth camera and an RGB camera enables the platform to have good adaptability in complex agricultural environments, and can effectively respond to changes in lighting, terrain, crop row spacing, etc., to ensure the efficiency and stability of the operation process. Through real-time posture recognition and intelligent response, the present invention enables the platform to have higher flexibility and can make accurate responses according to the behavior of the operators. In addition, the present invention has broad application prospects and is suitable for various scenarios such as field crop harvesting, greenhouse management, orchard transportation, etc., and promotes agricultural operations to develop towards higher automation and intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1Schematic diagram of the working steps of the control method of the large-scale automatic following transportation platform based on human posture interaction.
[0066] Figure 2 Schematic diagram of the hardware layout and principle of the large-scale automatic following transportation platform based on human posture interaction.
[0067] Figure 3 Schematic diagram of the control flow of the control method for the field automatic following transport platform based on human posture interaction.
[0068] Figure 4 Schematic diagram of the interface for front and rear camera raw images, image segmentation, and crop row midpoint deviation calculation.
[0069] Figure 5 Schematic diagram of the depth camera's human body key point recognition, posture recognition, distance calculation, and depth map interface. DETAILED DESCRIPTION
[0070] The specific embodiments of the present invention will be further described below in conjunction with the accompanying drawings. The following examples are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention. The following examples take the automatic follow-up selective harvesting and transportation of field broccoli as an example for detailed description.
[0071] Figure 1 A schematic diagram of the working steps of a control method for a field automatic following transport platform based on human posture interaction provided by an embodiment of the present invention is provided. The method includes the following steps:
[0072] Step S1. Crop row detection:
[0073] In this embodiment, the front RGB camera and the rear RGB camera are used to collect images of the front and rear of the broccoli row at a preset downward tilt angle. When the platform moves forward to align the rows, the front camera is enabled for visual alignment, and when the platform moves backward to align the rows, the rear camera is enabled for visual alignment. The camera field of view is adjusted to cover three ridges of broccoli rows according to the height of the vehicle body, and the broccoli rows occupied 60%-80% of the lower edge of the camera field of view. The image is segmented in real time based on the Fast-SCNN image segmentation algorithm, and the crop rows and the background are binarized, with the crop rows being white and the background being black. The ROI area is set at 70%-90% below the image to calculate the white connected domain with the largest area, and its center of mass point C is extracted as the crop row wheel D based on the contour centroid algorithm. x Calculate the pixel horizontal coordinate C of the crop row contour center point C x and the center point R of the RGB camera field of view 0' The pixel horizontal coordinate R 0'x Deviation value D x , the deviation value D x The formula is as follows:
[0074] D x =R 0'x -C x .
[0075] The installation position and working principle diagram of the front and rear RGB cameras are as follows: Figure 2 、 Figure 3 As shown, the images captured by the front and rear RGB cameras and the post-processing effects are as follows Figure 4 The pixel deviation value is input into the PID controller, and the forward and backward navigation of the platform is controlled based on a dynamic feedback adjustment mechanism.
[0076] Step S2. Human target detection:
[0077] In this embodiment, a two-axis gimbal is installed at the center of the upper surface of the platform and at a height h from the platform, and a depth camera is installed on the two-axis gimbal. The depth camera installation position and working principle diagram are shown in the figure below. Figure 3 As shown. The image data of broccoli harvesters is collected in real time using the Intel D435 depth camera. The YOLO11-Pose deep learning model is used to infer human targets and key points in the camera's field of view in real time. The YOLO11-Pose model is trained based on the COCO dataset and can accurately locate 17 key points of the human body (including shoulders, elbows, wrists, hips, and knees, etc.). Based on the key points of the human body defined in the COCO dataset, the midpoint of the hip joint is extracted as the center point of the human body H, which is calculated as follows:
[0078]
[0079] Among them, the coordinates of H are (Hx, Hy), K 11_x is the horizontal coordinate of the left hip, K 11_y is the vertical coordinate of the left hip, K 12_x is the horizontal coordinate of the right hip, K 12_y is the vertical coordinate of the right hip.
[0080] Calculate the pixel coordinates (Hx, Hy) of the center point H of the human body and the pixel coordinates (R 0x ,R 0y )'s lateral offset value D HR_x With the longitudinal offset value D HR_y The lateral offset value D is transmitted to the human body tracking module. HR_x With the longitudinal offset value D HR_y Calculated by the following formula:
[0081] D HR_x =H x -R 0x
[0082] D HR_y =H y -R 0y .
[0083] The human targets and key points in the camera's field of view are inferred in real time based on the YOLO11-Pose deep learning model. Figure 5 shown.
[0084] Step S3. Human body tracking and pan / tilt control:
[0085] Based on the lateral displacement D between the center point H of the human body and the center point R0 of the depth camera field of view in step S2 HR_x With the longitudinal offset value D HR_y The PID control algorithm dynamically adjusts the two-axis gimbal's pitch angle θ and yaw angle φ. This allows the gimbal to follow the human target, keeping it at the center of the depth camera's field of view. By coordinating the two-axis gimbal's pitch angle θ and yaw angle φ, the human target remains in the center of the depth camera's field of view in real time, ensuring tracking continuity and stability. The adjusted pitch angle θ and yaw angle φ are output to the human position detection module in real time, providing dynamic parameters for subsequent three-dimensional coordinate conversion.
[0086] Step S4. Calculation of the distance between the human body and the platform:
[0087] The distance D between the broccoli harvester and the platform is defined as the difference between the projection distance ox from the center point H of the human body to the depth camera in the platform's forward direction and the installation distance S between the depth camera and the front end of the platform. The distance q between the human body and the platform is calculated as follows:
[0088] D=ox-q.
[0089] The projection distance ox between the operator and the depth camera in the platform's forward direction is calculated using a three-dimensional coordinate conversion formula. Combining the depth value d from the center of the human body to the depth camera obtained by the depth camera and the real-time rotation angle of the gimbal (pitch angle θ and yaw angle φ), the three-dimensional coordinate conversion formula is:
[0090] The calculation formula for the three-dimensional coordinate transformation is:
[0091]
[0092] Among them, based on the right-hand coordinate system definition, the origin o is the location of the platform, the x-axis points to the front of the platform, the y-axis points to the left of the platform, and the z-axis points to the top of the platform; d is the depth value of the operator's body center point, θ is the pitch angle of the gimbal, φ is the yaw angle of the gimbal, and ox is the projection distance of the operator in the direction of vehicle movement. The depth value d from the center point of the body to the depth camera is as follows: Figure 5As shown in the red word Distance in the upper left corner, the projection distance ox from the center point H of the human body to the depth camera in the direction of the platform is as follows: Figure 5 The red word cal_distance in the upper left corner is in meters.
[0093] Step S5. Human body posture detection and command output:
[0094] Based on the key point coordinates, the broccoli harvester's posture is determined and instructions are issued according to the following conditional rules:
[0095] (1) When a bending posture is detected, the platform brakes immediately. This state corresponds to the broccoli harvester performing a bending and cutting action. The stationary platform can provide a safe and stable working space for the operator. When the lowest wrist is lower than the highest knee or the lowest elbow is lower than the highest hip, the system determines that it is the bending posture. The judgment formula is:
[0096] (max(K 9_y ,K 10_y )>min(K 13_y ,K 14_y )) or (max(K 7_y ,K 8_y )>min(K 11_y ,K 12_y ));
[0097] (2) When one arm is raised above the shoulder, this state corresponds to the broccoli harvester needing to move while the platform is stationary, including moving away from the platform or approaching the platform. In addition to braking, the gimbal continues to track the harvester to ensure that the camera's field of view is not lost. This design is aimed at the pause operation scenario to ensure that the platform continues to obtain the latest action instructions of the mobile broccoli harvester during braking. When the left wrist is higher than the left shoulder or the right wrist is higher than the right shoulder, the system determines it as the single-handed posture, and the judgment formula is:
[0098] (K 9_y <K 5_y ) or (K 10_y <K 6_y );
[0099] (3) When the key points of both wrists are higher than the shoulders, this state corresponds to a full load and needs to be returned to the field for unloading, and the platform switches to the reverse mode. This state is suitable for the scene of harvesting a full load and returning to unload. The broccoli harvester can command to retreat by raising both hands. If he needs to stop, he can quickly switch to raising one hand for braking, which is both flexible and safe. When the key points of both wrists are higher than the shoulders, the system determines that it is the above-mentioned raised hands posture, and the judgment formula is:
[0100] (K 9_y <K 5_y ) and (K10_y <K 6_y );
[0101] (4) When the broccoli harvester is in a standing position, the system determines that the operator is searching for mature broccoli while moving. At this time, the platform follows the human body and moves forward. The platform is always behind the operator and the distance between the platform and the human body is always within the safety threshold D to avoid interfering with the operator's search field of view. When a human target is detected and none of the above conditions are met, the system determines that it is in the standing position.
[0102] According to the requirements of the broccoli selective harvesting scenario, a safety distance threshold D' is preset. The threshold is determined by experimental data analysis to ensure that the safety distance between the platform and the broccoli harvester meets both safety requirements and the convenience requirements of selective harvesting. Setting the safety distance threshold D' receives the real-time distance ox output by step S4 and the control instruction generated by step S5. The detailed control flow diagram is shown as follows: Figure 3 As shown:
[0103] (1) Row-to-row advance condition: When D>D', and the crop row end has not been reached, and a standing posture is detected, the row-to-row advance command is triggered;
[0104] (2) Braking conditions: When any one or more of the following conditions are met, the platform braking is triggered immediately: when D ≤ D', or no human target is detected, or a hand-raising gesture is detected, or the crop row detection is lost, that is, no crop row is detected in the ROI area;
[0105] (3) Row-by-row retreat condition: When the rear camera detects a crop row and detects that the operator is in the posture of raising both hands, the row-by-row retreat command is triggered.
[0106] In the crop row detection (step S1) of a specific embodiment of the present invention, KC-WQ cameras are used as the front and rear cameras, with a resolution of 640×480 and a frame rate of 30 fps. The semantic segmentation module uses the Fast-SCNN image segmentation algorithm, manually annotated and trained on 3,000 real-world images of broccoli crop rows in a field covering various weather conditions, lighting conditions, and terrain scenarios. After extracting the ROI, the center of mass deviation and the pixel deviation of the field of view center are calculated, and steering instructions are generated using a PID controller.
[0107] In the human object detection and key point detection (steps S2 and S5) of the specific embodiment of the present invention, the operator key point detection uses the Pose model with an input size of 640×480. Inference is performed on the COCO dataset pre-trained model. There are 17 human key points in total. Following the definition method of the COCO dataset, the present invention selects 10 of them as follows:
[0108] Shoulder key points: Key point No. 5 is the left shoulder, and key point No. 6 is the right shoulder;
[0109] Elbow key points: Key point No. 7 is the left elbow, and key point No. 8 is the right elbow;
[0110] Wrist key points: Key point No. 9 is the left wrist, and key point No. 10 is the right wrist;
[0111] Hip key points: Key point 11 is the left hip, and key point 12 is the right hip;
[0112] Knee key points: Key point No. 13 is the left knee, and key point No. 14 is the right knee.
[0113] In the human target detection (steps S2 and S5) of the specific embodiment of the present invention, the coordinates of the key points of the human body and the coordinate system are defined as follows:
[0114] Left shoulder coordinates: (K 5_x ,K 5_y ), right shoulder coordinates: (K 6_x ,K 6_y );
[0115] Left elbow coordinates: (K 7_x ,K 7_y ), right elbow coordinates: (K 8_x ,K 8_y );
[0116] Left wrist coordinates: (K 9_x ,K 9_y ), right wrist coordinates: (K 10_x ,K 10_y );
[0117] Left hip coordinates: (K 11_x ,K 11_y ), right hip coordinates: (K 12_x ,K 12_y );
[0118] Left knee coordinates: (K 13_x ,K 13_y ), right knee coordinates: (K 14_x ,K 14_y );
[0119] The coordinate system is a pixel coordinate system, with its origin at the upper left corner of the image and the positive y-axis pointing downward.
[0120] In the human body tracking and human body position detection (steps S3 and S4) of the specific embodiment of the present invention, the depth camera uses the Intel D435 depth camera, which has high resolution and strong depth perception capabilities, can work stably under complex outdoor lighting conditions, and provides support for human target detection, position detection and posture recognition.
[0121] In the human body tracking and human body position detection (step S6) of the specific embodiment of the present invention, the setting experiment of the safety distance threshold D' is as follows:
[0122] The minimum safe operating space (W) was measured for 20 broccoli harvesters of different body shapes (1.60-1.85m tall) while they performed agricultural operations (standing, bending, swinging a knife, harvesting, carrying, etc.). Combined with the emergency braking distance (S) when the platform was fully loaded, the safety threshold (D') was determined using the following formula:
[0123] D'=W+S.
[0124] The various technical features of the above-described embodiments can be further combined. To make the description concise, not all possible combinations of the various technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0125] The above-described embodiments merely represent several implementations of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that a person skilled in the art would be able to make numerous modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of the present invention. The scope of the present invention is defined by the appended claims and any equivalents thereof.
Claims
1. A control method for a field automatic following transport platform based on human posture interaction, characterized in that: The method comprises the following steps: S1. Crop row detection: The front and rear RGB cameras are used to capture images of the front and rear of the crop rows at preset downward angles. The front camera is used for visual alignment when the platform is moving forward, and the rear camera is used for visual alignment when the platform is moving backward. The machine's field of view is adjusted according to the vehicle body height to cover three ridges of crop rows, and the crop rows occupy 60%-80% of the bottom edge of the camera's field of view. The Fast-SCNN image segmentation algorithm segments the image in real time and binarizes the crop rows and background, with the crop rows in white and the background in black. The ROI area is set at 70%-90% below the image, and the largest white connected domain is calculated. Its center of mass point C is extracted as the crop row wheel D based on the contour centroid algorithm. x Contour center point; calculate the pixel horizontal coordinate C of the crop line contour center point C x and the center point R of the RGB camera field of view 0' The pixel horizontal coordinate R 0'x Deviation value D x , the deviation value D x The formula is as follows: D x =R 0'x -C x The pixel deviation value is input into a PID controller to control the forward and backward navigation of the platform based on a dynamic feedback adjustment mechanism; S2. Human target detection: A two-axis gimbal is installed at the center of the platform surface at a height h from the platform. A depth camera is installed on the two-axis gimbal and collects image data of the operator in real time. The YOLO11-Pose deep learning model is used for human target detection and key point recognition. Based on the human key points defined in the COCO dataset, the midpoint of the hip joint is extracted as the human center point H, which is calculated as follows: Among them, the coordinates of H are (Hx, Hy), K 11_x is the horizontal coordinate of the left hip, K 11_y is the vertical coordinate of the left hip, K 12_x is the horizontal coordinate of the right hip, K 12_y is the vertical coordinate of the right hip; Calculate the pixel coordinates (Hx, Hy) of the center point H of the human body and the pixel coordinates (R 0x ,R 0y )'s lateral offset value D HR_x With the longitudinal offset value D HR_y and transmit the offset value to the human body tracking module; the lateral offset value D HR_x With the longitudinal offset value D HR_y Calculated by the following formula: D HR_x =H x -R 0x D HR_y =H y -R 0y ; S3. Human tracking and pan / tilt control: Based on the lateral displacement D between the center point H of the human body and the center point R0 of the depth camera field of view in step S2 HR_x With the longitudinal offset value D HR_y ,The PID control algorithm is used to dynamically adjust the pitch angle θ and yaw angle φ of the two-axis gimbal, so that the two-axis gimbal rotates along with the human target, and the human target is always located in the center of the depth camera's field of view; At the same time, the gimbal angle information is output in real time, that is, the real-time pitch angle θ and yaw angle φ of the two-axis gimbal; S4. Calculation of the distance between the human body and the platform: The distance D between the human body and the platform is defined as the difference between the projection distance ox from the center point H of the operator's body in the forward direction of the platform and the installation distance q between the depth camera and the front end of the platform. The distance D between the human body and the platform is calculated as follows: D=ox-q The projection distance ox between the operator and the depth camera in the forward direction of the platform is calculated by the three-dimensional coordinate conversion formula; combined with the depth value d from the center point of the human body to the depth camera obtained by the depth camera and the real-time rotation angle of the gimbal, including the pitch angle θ and the yaw angle φ, the three-dimensional coordinate conversion formula is: The calculation formula for the three-dimensional coordinate transformation is: Wherein, the origin o is the location of the platform, the x-axis points to the front of the platform, the y-axis points to the left of the platform, and the z-axis points to the top of the platform; d is the depth value of the operator's body center point, θ is the pitch angle of the gimbal, φ is the yaw angle of the gimbal, and ox is the projected distance of the operator in the vehicle's forward direction; S5. Human posture detection and command output: Based on the key point coordinates, the operator's posture is judged and instructions are issued according to the following conditional rules: (S5.1) When a bending posture is detected, the platform brakes immediately; the state corresponds to the operator performing a bending cutting action, and the platform is stationary to provide a safe and stable working space for the operator; if the lowest wrist is lower than the highest knee or the lowest elbow is lower than the highest hip, the system determines that it is a bending posture, and the judgment formula is: (max(K 9_y ,K 10_y )>min(K 13_y ,K 14_y )) or (max(K 7_y ,K 8_y )>min(K 11_y ,K 12_y )); (S5.2) When one arm is raised above the shoulder, this state corresponds to the operator requiring the platform to move while the platform is stationary, including moving away from the platform or approaching the platform. In addition to triggering the brake, the gimbal continues to track the operator to ensure that the camera's field of view is not lost. This design is for paused operation scenarios, ensuring that the platform continues to obtain the latest movement instructions from the mobile operator during braking. When the left wrist is higher than the left shoulder or the right wrist is higher than the right shoulder, the system determines that it is the single-handed posture described above, and the judgment formula is: (K 9_y <K 5_y ) or (K 10_y <K 6_y ); (S5.3) When the wrist key points are higher than the shoulders, the platform switches to reverse mode when fully loaded and needs to be unloaded at the edge of the field. This state is suitable for harvesting a full load and returning to unload. The operator raises both hands to command the retreat. If stopping is required, the operator can quickly switch to raising one hand to brake. When the wrist key points are higher than the shoulders, the system determines that the posture is raised with both hands. The judgment formula is: (K 9_y <K 5_y ) and (K 10_y <K 6_y ); (S5.4) When the operator is in a standing position, the system determines that the operator is searching for mature crops or fruits while moving. At this time, the platform follows the operator's movement, and the platform is always behind the operator and the distance between the platform and the operator is always within the safety threshold D to avoid interfering with the operator's search field of view. When a human target is detected and none of the above conditions are met, the system determines that the operator is in the standing position described above; S6. Platform motion control: Set the safety distance threshold D' to receive the real-time distance ox output in step S4 and the control instruction generated in step S5: (S6.1) Row-to-row advance condition: When D>D', and the crop row end has not been reached, and a standing posture is detected, the row-to-row advance command is triggered; (S6.2) Braking conditions: When any one or more of the following conditions are met, the platform braking is immediately triggered: when D ≤ D', or no human target is detected, or a hand-raising gesture is detected, or crop row detection is lost, that is, no crop row is detected in the ROI area; (S6.3) Row retreat condition: When the rear camera detects a crop row and detects that the operator is in a hands-raised posture, the row retreat command is triggered.
2. The control method according to claim 1, characterized in that: In step S1, the weights trained based on the Fast-SCNN image segmentation algorithm are labeled and trained using a field-collected dataset covering a variety of weather conditions, lighting conditions, and terrain scenes.
3. The control method according to claim 1, characterized in that: In step S2, the image segmentation model adopts the YOLO11-Pose model.
4. The control method according to claim 1, wherein: In steps S2 and S5, the COCO dataset defines 17 key points for the human key point detection task, and the specific numbering of the 10 key points used is as follows: Shoulder key points: Key point No. 5 is the left shoulder, and key point No. 6 is the right shoulder; Elbow key points: Key point No. 7 is the left elbow, and key point No. 8 is the right elbow; Wrist key points: Key point No. 9 is the left wrist, and key point No. 10 is the right wrist; Hip key points: Key point 11 is the left hip, and key point 12 is the right hip; Knee key points: Key point No. 13 is the left knee, and key point No. 14 is the right knee.
5. The control method according to claim 1, characterized in that: In steps S2 and S5, the coordinates of the key points of the human body and the coordinate system are defined as follows: Left shoulder coordinates: (K 5_x ,K 5_y ), right shoulder coordinates: (K 6_x ,K 6_y ); Left elbow coordinates: (K 7_x ,K 7_y ), right elbow coordinates: (K 8_x ,K 8_y ); Left wrist coordinates: (K 9_x ,K 9_y ), right wrist coordinates: (K 10_x ,K 10_y ); Left hip coordinates: (K 11_x ,K 11_y ), right hip coordinates: (K 12_x ,K 12_y ); Left knee coordinates: (K 13_x ,K 13_y ), right knee coordinates: (K 14_x ,K 14_y ); The coordinate system is a pixel coordinate system, with the origin at the upper left corner of the image and the positive y-axis pointing downward.
6. The control method according to claim 1, characterized in that: In step S6, the safety distance threshold D' is set by the following experiment: The minimum safe operating space W was measured for 20 workers of different body types, ranging in height from 1.60 to 1.85 meters, while performing agricultural operations. Combined with the emergency braking distance S when the platform was fully loaded, the safety threshold D' was ultimately determined using the following formula: D'=W+S.
Citation Information
Patent Citations
Greenhouse labour-saving automatic following work platform based on somatosensory inductor
CN105404299A