Strawberry garden inter-ridge visual navigation based on improved Unet network and MPC control method based on PFT preview theory
By improving the Unet network and adaptive micro ROI navigation path extraction algorithm, combined with the MPC controller of PFT pre-targeting theory, the complexity of the ridge road environment in strawberry garden was solved, high-precision and robust navigation was achieved, and the "bumping of the ridge" phenomenon was avoided, and it was suitable for a variety of crops and ridge road scenarios.
Patent Information
- Application Number
- CN202510362148.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-18
AI Technical Summary
The ridge road environment in the strawberry garden is complex, with problems such as uneven light, crop occlusion, and blurred ridge road boundaries. Traditional navigation methods are difficult to meet the needs of high precision and high robustness.
The MPC controller that uses an improved Unet network for ridge-track feature segmentation, combines the adaptive micro ROI navigation path extraction algorithm and PFT pre-targeting theory to achieve high-precision navigation through visual perception and path planning.
It realizes high-precision and robust navigation in complex environments, avoids the phenomenon of "bumping ridges". It is suitable for a variety of crops and ridge scenes, and provides reliable technical support for automated operations between farmland ridges.
Smart Images

Figure CN120335349A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of agricultural intelligence, and particularly relates to a 2D visual navigation between ridges in a strawberry field based on an improved Unet network and adaptive micro ROI for feature point extraction, and a control algorithm based on the PFT preview theory. Background Art
[0002] With the rapid development of agricultural intelligence, the application of automated agricultural machinery in inter-ridge operations in farmland is becoming increasingly widespread. Especially in fine agricultural scenarios such as strawberry cultivation, the inter-ridge navigation technology has become the key to achieving efficient and precise operations. However, the ridge environment in a strawberry field is complex, with problems such as uneven lighting, crop occlusion, and blurred ridge boundaries. Traditional navigation methods are difficult to meet the requirements of high precision and high robustness. Summary of the Invention
[0003] To solve the above problems, the present invention proposes a 2D visual navigation between ridges in a strawberry field based on an improved Unet network and adaptive micro ROI for feature point extraction, and a control algorithm based on the PFT preview theory. By using the improved Unet network to segment the ridge features and extracting navigation feature points based on the adaptive micro ROI, and then obtaining the planned path through polynomial fitting based on discrete points after coordinate transformation, and using the MPC control algorithm based on the PFT preview theory to track the given path, the navigation function between ridges is realized.
[0004] The present invention adopts the following technical solutions: a visual navigation between ridges in a strawberry field based on an improved Unet network and an MPC control method based on the PFT preview theory; including the following steps:
[0005] Step 1, obtain a ridge field image through the visual perception module of the RealSense D435 depth camera, input the image into the improved Unet network for image semantic segmentation and extract the feasible area of the inter-ridge road. The improved Unet network enables the network to more accurately segment the feasible area of the inter-ridge road by using a coordinated attention encoding module, a multi-axis Hadamard product attention module, a group aggregation bridging module, and an atrous spatial pyramid pooling layer module, and performs excellently in complex scenarios;
[0006] Step 2, design an adaptive micro ROI navigation path extraction algorithm based on 2D vision. By dynamically adjusting the ROI area boundary to adapt to different crop structures and terrain conditions, the extracted navigation planned path is transformed to the robot base coordinate system through a 2D to 3D conversion from the pixel coordinate system to the camera coordinate system and a rigid body transformation from the camera coordinate system to the robot base coordinate system, and the position information of the coordinate points of the navigation line is obtained;
[0007] Step 3: Build a kinematic model of the strawberry spraying robot, obtain an error prediction state model through linearization, adopt an incremental control strategy combined with an MPC controller to optimize the objective function, and solve the input control variables at the next moment in real time to achieve path tracking control of the given input quantity; propose an MPC controller based on the PFT preview control algorithm, judge and call the preview control mode through the error threshold, and impose constraints on the input control quantity in combination with the special environmental conditions of the strawberry field ridge to avoid the "ridge collision" phenomenon. The strawberry field inter-ridge image segmentation method using the improved Unet network mainly includes:
[0008] Obtain the RGB image of the farmland scene and perform preprocessing through light compensation and distortion correction; input the preprocessed image into the improved Unet network and output the ridge semantic segmentation result; among them, the improved Unet network includes:
[0009] Coordinated attention encoding module: The coordinated attention module generates channel-spatial attention weights through adaptive pooling in the horizontal and vertical directions. This module first inputs an image with a size of C×H×W, where C is the number of channels, and H and W are the height and width respectively. The image passes through two parallel average pooling layers, aggregates features along the H and W spatial directions, and obtains two direction-aware feature maps with dimensions compressed to C×H×1 and C×1×W. Then, through the dimension transformation of the image, dimension splicing in the third dimension and a two-dimensional convolution transformation are fused into an image, and then through normalization operation and non-linear connection, and finally through two two-dimensional convolutions and activation processes, two attention weight maps with the original dimensions (C×H×1 and C×1×W) after segmentation are obtained and fused with the original image on the input side to finally output a feature map with an attention mechanism.
[0010] Multi-axis Hadamard product attention module: The multi-axis Hadamard product attention module groups the feature maps by channel, applies parameterized convolutional kernels in the XY plane, ZX axis, and ZY axis respectively (O-XYZ is the spatial Cartesian coordinate system), and fuses multi-dimensional features through the Hadamard product. This module first decomposes the input image with a size of C×H×W along the channel dimension into 4 equal sub-features, and the dimension of the decomposed features is B×(C / 4)×H×W, where B represents the data batch size. Among them, 3 equal sub-features respectively pass through the Hadamard attention mechanism under the XY, ZX, and ZY axes, and the 4th sub-feature only applies depthwise separable convolution. Finally, the dimensions are spliced and layer-normalized, and then output after depthwise separable convolution.
[0011] Group aggregation bridging module: The group aggregation bridging module fuses high-level semantic features and low-level detail features through a multi-dilation rate dilated convolution group. This module first uses depthwise separable convolution and bilinear interpolation to input the size C of the image with high-level features h ×Hh ×W h Adjusted to be consistent with the low-level feature C l ×H l ×W l to be consistent, where C h , H h , W h respectively represent the number of channels, height, and width of the high-level features. The adjusted low-level features and high-level features are respectively divided into four groups along the channel dimension, and then each group of low-level features is concatenated with each group of high-level features to obtain four groups of fused features. Then, the masked image is fused with the four groups of features, and dilated convolutions with dilation rates of 1, 3, 5, and 7 are respectively used for feature extraction. Finally, the four groups of features are concatenated along the channel dimension, and 1×1 convolution is used for feature integration and output.
[0012] Atrous Spatial Pyramid Pooling layer: The atrous spatial pyramid pooling fuses multi-scale information to output the segmentation result. For a given input, parallel sampling is performed using four atrous convolutions with dilation rates of 6, 12, 18, and 24 and a 1×1 convolution, and a global average pooling branch is added. Finally, the above 6 branches are concatenated along the channel dimension and the number of channels is reduced to the expected value through 1×1 convolution and then output.
[0013] The process of the adaptive micro ROI navigation path extraction algorithm based on 2D vision includes:
[0014] Step 2.1: First, input the binarized segmentation image with a relatively smooth edge obtained after image preprocessing operations. Mark the label center of the sub-strip according to the projection peak of the white pixels of the sub-strip, denoted as r n . Sub-ROI regions are framed on the left and right sides of r n . Appropriate values can be selected according to the crop type to determine the widths of these two sub-ROIs centered on the initial center points c l and c r to obtain the best detection area.
[0015] Step 2.2: Along the direction of the sub-strip, the sub-ROI regions on the left and right sides of the label center r n are respectively translated in opposite directions with a constant pixel value increment, and the moving pixel value increment is defined as i.
[0016] Step 2.3: The sub-ROI regions will perform a sequential search along the strip until it finds the edge of the label or until it reaches the boundary of the nth sub-strip. If the current sub-ROI region overlaps with the already fitted sub-ROI regions from neighboring regions, the search for this loop will also stop.
[0017] Step 2.4: According to formula (1), the sub-strip stripn And the mask image mask composed of sub - ROI regions n Perform a bit - wise AND operation, and the resulting res n Is the overlapping part of the strip and the sub - ROI region (make the overlapping part white pixels), and calculate the percentage of white pixels in the strip. If it is less than the threshold T, stop the inspection, which means the threshold condition is met. If it meets this threshold condition, it indicates that the sub - ROI region reaches the edge of the label. The formula is:
[0018] res n = strip n ∧ mask n (1)
[0019] Step 2.5: If the percentage of white pixels in the strip does not meet the requirements of the threshold condition, it is regarded as not reaching the edge, and it is considered that the set pixel step size is too small. Go back to step 2 and set it to 2i again for searching. The search based on the sub - ROI will continue to use the step size of ki (k = 1, 2, 3, …, 16) for searching until it meets any one of the stop conditions in 3.
[0020] Step 2.6: When the search stops due to meeting the stop condition and meets the threshold condition, according to the label center r n And the center points c l And c r Of the left and right sub - ROI regions, update the boundaries m l And m r Of the micro - ROI region as the new micro - ROI region. Calculate the average value of white pixels in this region according to the formula, update the label center r n Of the nth sub - strip and repeat the whole process starting from step 1 for the next sub - strip until the last sub - strip is searched.
[0021] The steps of converting the extracted navigation planning path from the pixel coordinate system to the camera coordinate system through 2D - to - 3D transformation and the rigid - body transformation from the camera coordinate system to the robot base coordinate system include:
[0022] The navigation path planned through the depth - camera vision perception module. Since the volume of the vehicle body is large, it cannot be regarded as a particle. After the depth camera captures the navigation path, the planned path needs to be first transformed to the front - wheel drive - wheel coordinate system through coordinate transformation to obtain the coordinate points and slope data of the navigation line in the front - wheel drive - wheel coordinate system.
[0023] The specific steps of the coordinate transformation process are:
[0024] Step 2.7: Calculate the internal parameter matrix of the camera through camera calibration. Based on the coordinate points on the navigation line in the pixel coordinate system extracted by the visual perception algorithm, denoted as (U, V), perform an inverse coordinate transformation of the internal parameter matrix to obtain the coordinate points in the camera coordinate system, denoted as (X c , Y c , Z c ). The above coordinate transformation is essentially a transformation from a 2D to a 3D coordinate system, and the expression is as follows:
[0025]
[0026]
[0027] Among them, f represents the focal length, in millimeters; dx represents the width of the pixel in the x direction; dy represents the width of the pixel in the y direction; f x represents the length of the focal length in the x-axis direction described in pixels, that is, the normalized focal length in the x-axis direction; f y represents the length of the focal length in the y-axis direction described in pixels, that is, the normalized focal length in the y-axis direction; (u0, v0) represents the optical center, that is, the intersection point of the camera optical axis and the image plane, usually located at the center of the image, so its value is often taken as half of the resolution. z c represents.
[0028] Step 2.8: In order to obtain the navigation path points in the world coordinate system, let the center axis of the front driving wheel of the vehicle be the origin of the world coordinate system, also known as the base coordinate system of this robot. In the transformation algorithm between the camera coordinate system O c -X c Y c Z c and the robot base coordinate system O v -X v Y v Z v , this is actually an essential rigid body transformation. By constructing the camera coordinate system O c -X c Y c Z c , where O w is the optical center position of the camera, X c is perpendicular to Y v Z v inward. Due to a certain tilt angle during the installation of the camera, the pitch angle of the O c Z c axis in the horizontal direction is defined as α. Ideally, the baseline of the camera coincides with the OX axis. From the base coordinate system O v -X v Y v Z v, where O v is located at the midpoint of the rear wheel connection line, X v The positive direction is perpendicular to Y v Z v The plane faces outward, the installation height of the camera is H, and the longitudinal distance from the camera to the driving wheel is L.
[0029] Step 2.9: Rotate by α + 90° around the X c axis to obtain the coordinate system O1-X1Y1Z1:
[0030]
[0031] Step 2.10: According to the coordinate system O1-X1Y1Z1, translate it according to the following steps: Move H in the negative Z1 direction to obtain the translated coordinate system O2-X2Y2Z2:
[0032]
[0033] Step 2.11: Finally, rotate 180° around the Z1 axis to obtain the coordinate system O v -X v Y v Z v :
[0034]
[0035] It is obtained that the transformation relationship from the O c -X c Y c Z c system to the O v -X v Y v Z v system can be written as:
[0036]
[0037] Finally, the conversion relationship between the navigation point in the robot base coordinate system and the camera coordinate system is obtained as:
[0038]
[0039] Construct a kinematic model of the strawberry spraying robot, obtain an error prediction state model through linearization, adopt an incremental control strategy combined with an MPC controller to optimize the objective function, solve the input control variable at the next moment in real time, and achieve path tracking control of the given input quantity. A MPC controller based on the PFT preview control algorithm is proposed. The preview control mode is called through the error threshold judgment, and constraints are imposed on the input control quantity in combination with the special environmental conditions of the strawberry field ridge to avoid the "ridge collision" phenomenon; specifically including:
[0040] Step 3.1: Obtain two input control variables and three system state variables of the autonomous navigation strawberry spraying robot at the current moment. The two input control variables are the centroid motion speed of the robot and the angular velocity of the centroid rotating around the instantaneous center. The three system state variables are the x-axis coordinate, y-axis coordinate of the robot in the world coordinate system, and the angle of rotation around the z-axis coordinate of the centroid.
[0041] Before the said Step 3.1, it further includes:
[0042] Obtain the hub motor parameters, wheel and robot parameters of the autonomous navigation strawberry spraying robot; the motor parameters mainly include the motor reduction ratio, encoder count value, and number of lines of the encoder; the wheel parameters mainly include the wheel radius and wheel rotation speed; the robot parameters mainly include the robot yaw angular velocity and the position of the robot in the world coordinate system, including the x-axis coordinate and y-axis coordinate.
[0043] According to the hub motor parameters, the rotational speed of a single motor can be obtained as:
[0044]
[0045] Where M is the count value of the sampling period encoder, P is the number of lines of the encoder, that is, the total number of pulses triggered by one revolution of the code disk, N is the motor reduction ratio, R is the radius of the wheel, and Δt is the sampling period of the encoder.
[0046] Step 3.2: Construct the kinematic model of the autonomous navigation strawberry spraying robot according to the wheel parameters and robot parameters:
[0047]
[0048] Where the instantaneous linear velocity of the robot centroid's motion is v, x is the lateral displacement; y is the longitudinal displacement; is the vehicle yaw angle; respectively represent the lateral and longitudinal velocities; ω represents the angular velocity of the robot centroid rotating around the instantaneous center; the state variables can be represented by and the control variables can be represented by C = [v, ω] T for representation.
[0049] Since this model is a non-linear model, after linearization at the target state and proposing a set of new error state variables, a linearized error model is constructed:
[0050]
[0051] Where A and B are Jacobian matrices, and the state variables defining the linear error model are:
[0052]
[0053] where x r is the reference lateral displacement; y r is the reference longitudinal displacement; is the reference vehicle yaw angle; all are obtained from the reference values calculated by this control algorithm.
[0054] The input control quantity of the linear error model is:
[0055]
[0056] Assume that the control quantity input of the target point is 0.
[0057]
[0058] Using the forward Euler method for discretization, the discretized error model is:
[0059]
[0060] where: where T is the sampling time and k is the sampling moment.
[0061] By introducing an incremental model, that is, changing the system input to incremental control, this can make the system input smoother. By defining new state variables:
[0062] Substituting the new state variables into formula (11) to obtain the state equation of the discretized incremental prediction model:
[0063]
[0064] where
[0065]
[0066] where C = [I O], I is the identity matrix, and O is the all-zero matrix.
[0067] Through the recurrence relation, iterating the values at times k+1, k+2,..., k+N p to obtain the final incremental prediction model:
[0068] Y = Ψξ(k) + ΘΔU(15)
[0069] where,
[0070] N c is the control time domain, and N p is the prediction time domain.
[0071] Step 3.3: Since the control input is the centroid velocity vc The increment of and the increment of yaw rate ω, according to the formula
[0072]
[0073] where, v1 and v2 are the output rotational speeds of the left and right drive wheels respectively, and l icr is the radius of curvature of the robot's center of mass rotating around the instantaneous center.
[0074] Further decoupling gives the speeds v1 and v2 acting on each drive wheel as respectively:[[]]
[0075]
[0076] According to formula (8), by changing the duty ratio of the hub motor drive signal, and further changing the system input to the speeds of the two drive wheels, the robot realizes two-wheel differential operation during operation.
[0077] Step 3.4: In order to ensure that the robot can more quickly track the target navigation line under large errors, the present invention is based on the PFT preview control algorithm to judge that the cumulative sum of the reference point errors within the preview distance range reaches the error threshold δ, so as to determine whether the condition for calling the PFT preview control algorithm is satisfied. This threshold δ needs to be given in advance manually.
[0078] When the error is greater than the error threshold δ, it can be obtained from the kinematic model of the two-wheel differential mechanism followed by the autonomous navigation strawberry spraying robot in the present invention example that:[[]]
[0079]
[0080] where, a y is the lateral acceleration of the tire, V c is the speed of the center of mass movement, and l icr is defined the same as in formula (15).
[0081] A fixed origin of the world coordinate system is specified. Under this world coordinate system, the speed V of the robot's center of mass c is decomposed into
[0082] and The specific formula is:[[]]
[0083]
[0084] where the variable definitions are the same as those defined in formula (9).
[0085] In Figure 17 , P(X) is the target navigation line planned after the ridge channel semantic segmentation, t is the current moment, and d is the preview distance given when the robot moves.
[0086] Defined by the formula:
[0087]
[0088] Define t p as the preview time.
[0089] At time t, the lateral position y and lateral velocity v of the robot y are respectively:
[0090]
[0091] After the preview time interval, under the action of the lateral velocity and lateral acceleration of the robot, it reaches the preview point P(X(t + t p ))), and its lateral displacement Y(t + t p ) can be obtained by numerical integration using the second-order Runge-Kutta method:
[0092]
[0093] Under the action of the desired lateral velocity and lateral acceleration of the robot, the lateral displacement is Y(t + t p ) and the preview point P(X(t + t p )) coincide, that is, Y(t + t p ) = P(X(t + t p )), and the lateral acceleration a y is:
[0094]
[0095] where T is the sampling period. From the lateral acceleration formula (22) and the preview time formula (19), the radius of curvature l icr of the robot's center of mass rotating around the instantaneous center can be obtained as:
[0096]
[0097] Since the yaw angular velocity during the robot's driving is determined by the center-of-mass velocity ω and the radius of curvature l icr :
[0098]
[0099] During the sampling period T of the MPC controller, the yaw angle can be obtained by numerical integration of the yaw angular velocity ω, and this yaw angle with preview information is used as the desired reference yaw angle The specific calculation formula is:
[0100]
[0101] Among them, t0 represents the sampling moment of the controller.
[0102] Substitute this reference yaw angle into the error prediction state model of the discrete autonomous navigation strawberry spraying robot. By judging whether the Euclidean distance from the current position to the target navigation line trajectory is less than a given threshold, in the area where the distance is greater than the given threshold δ, the MPC controller in the preview following mode is used to approach the target navigation line trajectory more quickly. If the distance is less than the given threshold δ, only the MPC controller is used for path tracking.
[0103] Step 3.5: According to the error prediction state model of the discrete autonomous navigation strawberry spraying robot, establish the objective function of the linear model predictive controller MPC.
[0104] Based on the kinematic characteristics of the robot and considering the minimization of tracking error and control input increment, the following objective function of the MPC path tracking controller is constructed:
[0105]
[0106] where t = 1, 2, …, N c , …, N p , y t is the value of the state variable at time t, y t is the value of the target reference state variable at time t, u t is the control input at time t, u t - is the control input at time t - 1, ε is the slack variable introduced to ensure a feasible solution, Q, P, F, ρ are the weight matrices of the system state tracking error, incremental control input, system terminal control input, and slack factor respectively, N p is the MPC prediction time domain length, N c is the MPC control time domain length.
[0107] Step 3.6: Based on the target navigation path, solve the objective function to determine the input control variable of the MPC at the next moment; the input control variable is used to minimize the error between the robot and the target navigation path with the minimum control energy to control the motion trajectory of the robot in real time.
[0108] Based on the above objective function, the MPC control problem can be summarized as follows:
[0109]
[0110] where u * represents the optimal input control quantity obtained by solving the optimization objective function; u and u are the upper and lower limits of the control input respectively; Δu min and Δu max are the maximum and minimum values of the calculated control increment respectively; The above constraint critical parameters can be obtained by constructing the extreme conditions when the strawberry spray robot actually operates on the ridges of the strawberry field.
[0111] The following is the extreme condition 1 when the strawberry spray robot actually operates on the ridges of the strawberry field:
[0112] In this case, since the rotation center of the robot is at the suspension center of the front-wheel drive wheels and the lateral displacement of the robot does not deviate significantly from the trajectory center line during driving, the rear-wheel driven wheels are more likely to touch the strawberry plant ridges during rotation. Therefore, in this extreme condition, the maximum yaw angle of the vehicle, denoted as
[0113] The following is the extreme condition 2 when the strawberry spray robot actually operates on the ridges of the strawberry field:
[0114] In this case, since the lateral displacement of the robot deviates significantly from the trajectory center line during driving, the front-wheel drive wheels are more likely to touch the strawberry plant ridges during rotation. Therefore, in this extreme condition, the maximum yaw angle of the vehicle is denoted as
[0115] Based on the error prediction state model of the discretized autonomous navigation strawberry spray robot, the present invention relies on an independently proposed path planning feature point extraction algorithm for adaptive micro ROI navigation between ridges in a strawberry field based on 2D vision. Based on this, it is judged whether the robot meets the activation conditions of the algorithm, a linear MPC controller is constructed, the constrained optimization problem is solved, and according to the optimal control quantity obtained by optimizing the objective function, through further decoupling calculation, the control drive signal acting on the underlying hub motor is obtained, realizing the control goal of integrating path planning control and tracking.
[0116] The present invention has the following beneficial effects:
[0117] A ridge image segmentation technology based on an improved Unet neural network model for complex environments is proposed. The coordinated attention coding module, multi-axis Hadamard product attention module, group aggregation bridging module, and atrous spatial pyramid pooling layer module are introduced, significantly improving the extraction accuracy and robustness of ridge features, and being applicable to various crop and ridge scenarios. An adaptive micro ROI navigation path extraction algorithm is proposed. For the extracted ridge feature framework, a feasible trajectory route can be planned, and it is applicable to various crops and ridge scenarios. This algorithm can dynamically adjust the center point of the ROI area label, thereby adaptively planning a feasible trajectory route and showing good robustness and stability in different scenarios.
[0118] To achieve precise tracking of the planned path, an MPC controller based on the combination of preview control theory and incremental control strategy is proposed as the path tracking controller of this system. At the same time, combining the mechanical characteristics of the chassis structure of the physical machine and the feasibility analysis of driving between ridges, a specific steering yaw angle constraint applicable to navigation on the inter-ridge road is constructed through geometric constraints, and a special controller is built to make it applicable to the navigation process on the inter-ridge road. Through the close combination of deep learning, path planning, and control strategy, this system realizes high-precision, high-robustness, and wide applicability of inter-ridge road navigation, providing reliable technical support for automated operations between farmland ridges. Description of the Drawings
[0119] Figure 1 Schematic diagram of the model architecture of the improved Unet neural network in the embodiment of the present invention
[0120] Figure 2 Schematic diagram of the structure of the coordinated attention mechanism CA module in the embodiment of the present invention
[0121] Figure 3 Schematic diagram of the structure of the multi-axis Hadamard product attention module in the embodiment of the present invention
[0122] Figure 4 Schematic diagram of the structure of the group aggregation bridge GAB module in the embodiment of the present invention
[0123] Figure 5 Schematic diagram of the structure of the atrous spatial pyramid pooling layer ASPP module in the embodiment of the present invention
[0124] Figure 6 Actual effect diagram of ridge feature extraction in the embodiment of the present invention
[0125] Figure 7 Actual effect diagram of ridge navigation route planning in the embodiment of the present invention
[0126] Figure 8 Schematic diagram of the transformation between the pixel coordinate system and the robot base coordinate system in the embodiment of the present invention
[0127] Figure 9 Schematic diagram of ultrasonic ranging calculation in the embodiment of the present invention
[0128] Figure 10 Schematic diagram of path tracking control based on the MPC controller in the embodiment of the present invention
[0129] Figure 11 Flowchart of ridge semantic segmentation and path planning of the cross-ridge strawberry variable spray width spraying robot in the embodiment of the present invention
[0130] Figure 12Flowchart of path tracking control for the cross-ridge variable-rate and variable-spray-width strawberry spraying robot in the embodiments of the present invention
[0131] Figure 13 Figure 1 of the extreme case of collision during the operation of the strawberry variable-rate and variable-spray-width spraying robot in the embodiments of the present invention
[0132] Figure 14 Simplified mechanical structure model diagram under extreme case 1 of collision during operation in the embodiments of the present invention
[0133] Figure 15 Figure 2 of the extreme case of collision during the operation of the strawberry variable-rate and variable-spray-width spraying robot in the embodiments of the present invention
[0134] Figure 16 Simplified mechanical structure model diagram under extreme case 2 of collision during operation in the embodiments of the present invention
[0135] Figure 17 Schematic diagram of the PFT preview theory algorithm based on the carrier model in the embodiments of the present invention Detailed implementation manners
[0136] In this embodiment, the images of the strawberry garden ridges are collected by the RealSense D435 depth camera vision perception module installed on the robot. The collected images may contain interference factors such as uneven illumination and noise, so preprocessing is required. The preprocessing steps include:
[0137] Enhancing the image contrast using histogram equalization;
[0138] Removing noise using image preprocessing methods such as Gaussian filtering;
[0139] Normalizing the image and scaling the pixel values to the range of [0, 1].
[0140] The preprocessed image is input into the improved Unet neural network model to extract the semantic segmentation results of the navigation lines between the ridges. In the present invention, a coordinated attention module (CA module) is introduced into the shallow encoder of the improved Unet neural network model, a multi-axis Hadamard product attention module is added to the deep encoder and deep decoder, and multi-scale features are fused through a group aggregation bridging module (GAB module) between the encoder and decoder, which improves the network's ability to process complex perception environments, enhances the accuracy of image segmentation, reduces the number of model parameters, and improves the real-time performance. The specific steps are as follows:
[0141] As Figure 1The structural diagram of the improved Unet network is shown. First, the RGB image of the strawberry field ridge path is input into the network. It first undergoes downsampling through three encoder modules with residual connections and coordinated attention mechanisms to gradually obtain ridge path image features at different scales. The encoder module of the coordinated attention mechanism is as shown in Figure 2 shown. The coordinated attention module generates channel-spatial attention weights through adaptive pooling in the horizontal and vertical directions. This module first inputs an image with a size of C×H×W, where C is the number of channels, and H and W are the height and width respectively. The image passes through two parallel average pooling layers to aggregate features along the two spatial directions of H and W, obtaining two direction-aware feature maps with dimensions compressed to C×H×1 and C×1×W. Then, through the dimensional transformation of the image, they are fused into one image through dimension concatenation in the third dimension and a two-dimensional convolution transformation. After that, through normalization operations and non-linear connections, and finally, two two-dimensional convolutions and activation processes are performed to obtain two attention weight maps after segmentation that maintain the original dimensions (C×H×1 and C×1×W), which are then fused with the original image on the input side to finally output a feature map with an attention mechanism.
[0142] The improved Unet network passes through 3 multi-axis Hadamard product attention modules in the deep encoder to gradually obtain image features. During the upsampling process, the deep decoder first passes through 3 groups of multi-axis Hadamard product attention modules, and then through three decoder modules to restore the feature maps of each layer to the original resolution for subsequent operations. The structure of the multi-axis Hadamard product attention module is as shown in Figure 3 shown. The multi-axis Hadamard product attention module groups the feature maps by channel and applies parameterized convolutional kernels on the XY plane, ZX axis, and ZY axis respectively (O-XYZ is the spatial Cartesian coordinate system), and fuses multi-dimensional features through the Hadamard product.
[0143] This module first decomposes the input image with a size of C×H×W into 4 equal sub-features along the channel dimension. The dimension of the decomposed features is B×(C / 4)×H×W, where B represents the data batch size. Among them, 3 equal sub-features respectively pass through the Hadamard attention mechanism under the XY, ZX, and ZY axes, and the 4th sub-feature only applies depthwise separable convolution. Finally, the dimensions are concatenated and layer-normalized, and then output after depthwise separable convolution.
[0144] The network optimizes the skip connection module of the original Unet network. The features of each layer after 5 times of downsampling are respectively upsampled after being skip-connected to the corresponding layer, and pass through a group aggregation bridge (GAB) module in the middle. The structure of the GAB module is as shown in Figure 4As shown. The group aggregation bridging module fuses high-level semantic features and low-level detail features through a multi-dilation rate atrous convolution group. This module first adjusts the size C of the input image with high-level features to be consistent with that of the low-level features through the use of depthwise separable convolution and bilinear interpolation. The input image with high-level features has dimensions C h ×H h ×W h and is adjusted to be the same as that of the low-level features C l ×H l ×W l where C h 、H h 、W h represent the number of high-level feature channels, height, and width respectively; the adjusted low-level features and high-level features are divided into four groups along the channel dimension respectively, and then each group of low-level features is concatenated with each group of high-level features to obtain four groups of fused features; then the mask image is fused with the four groups of features, and atrous convolutions with dilation rates of 1, 3, 5, and 7 are used for feature extraction respectively. Finally, the four groups of features are concatenated along the channel dimension, and 1×1 convolution is used for feature integration and output;
[0145] After the upsampling process of the network, atrous spatial pyramid pooling is used to fuse multi-scale information. After passing through a 3×3 convolution and batch normalization, the image segmentation result is output through a 1×1 convolution. The structure of the atrous spatial pyramid pooling layer is as Figure 5 shown. The atrous spatial pyramid pooling fuses multi-scale information and outputs the segmentation result. For a given input, parallel sampling is performed using four atrous convolutions with dilation rates of 6, 12, 18, and 24 and a 1×1 convolution, and a global average pooling branch is added. Finally, the above 6 branches are concatenated along the channel dimension and the number of channels is reduced to the expected value through a 1×1 convolution and then output.
[0146] In the encoder part, spatial features are extracted through the coordinated attention module to generate attention weights and enhance the feature representation ability; in the decoder part, multi-dimensional features are fused through the multi-axis Hadamard product attention module to improve the robustness of feature extraction; a group aggregation bridging module is introduced. After the skip connection layer, high-level features and low-level features are fused through a multi-dilation rate atrous convolution group. The formula for the output x fuse of this module is:
[0147]
[0148] where the TailConv function represents the tail convolution. In this example of the present invention, depthwise separable convolution is used as the tail convolution; the Concat function represents concatenating the feature images along a specific dimension; {g i |i = 1, 2, 3, 4} are atrous convolution groups with dilation rates of 1, 3, 5, and 7 respectively; Represents high-level features, that is, deep features of the network; Represents low-level features, that is, shallow features of the network.
[0149] In the example of the present invention, the improved Unet network uses a residual convolution module with coordinated attention for feature extraction in the shallow encoder (1x, 2x, 4x downsampling layers): In the shallow encoder part, the preprocessed strawberry ridge image is first convolved once, and after using group normalization GroupNorm and GELU activation, a coordinated attention module is used for the activated part to generate attention weights. In order to reduce the number of parameters in the deep network and achieve the lightweight of the network. A multi-axis Hadamard product attention module is used in the deep encoder and deep decoder to fuse multi-dimensional features.
[0150] The formula is:
[0151] x out = Concat[H xy (x1), H zx (x2), H zy (x3), H dw (x4)](30)
[0152] Among them, the Concat function means to splice the feature images along a specific dimension, and H xy (x1), H zx (x2), H zy (x3) represents the equal quantum features obtained by decomposing the image along the channel dimension, and are the results obtained through the Hadamard attention mechanism under the XY, ZX, and ZY axes and after depthwise separable convolution respectively. H dw (x4) is the result obtained by only using depthwise separable convolution for the equal quantum features.
[0153] In the design of the network loss function, in order to accelerate the model training and constrain the model convergence, a multi-level hybrid loss function is used for deep supervision training. The cross-entropy loss function L CE is used for constraint for the intermediate layer of the network, and the Lovasz loss function L Lovasz is used for boundary accuracy constraint for the final output of the network. In addition, the DiceLoss loss function L Dice is added to enhance the similarity measurement between inter-class pixels. The optimized loss function L total has the following calculation formula:
[0154]
[0155] Among them, λ ∈ [0, 1] is a dynamic balance coefficient, which is used to adjust the weight distribution between boundary optimization and region segmentation. For Among them, ω kIn deep supervision training, it is the weight distribution of the output results of each layer of the network in the total deep supervision loss. L CE Calculation formula of cross-entropy loss function:
[0156]
[0157] Among them, M represents the number of categories; s ic represents the sign function, which takes 1 if the category c of the predicted sample i is the same as the true category, and 0 otherwise; p ic represents the predicted probability that the predicted sample i belongs to the category c.
[0158] L Dice Loss function formula:
[0159]
[0160] Among them, p i represents the probability value predicted by the model for the i-th pixel; g i represents the true label value of the i-th pixel.
[0161] L Lovasz Loss function formula:
[0162]
[0163] Among them, represents the true category; c represents the predicted category of the image; f i (c) represents the probability that the pixel point is classified as the c-th category; k + 1 represents the total number of categories.
[0164] The following specifically describes in combination with the schematic diagram an algorithm for strawberry field ridge 2D vision navigation based on an improved Unet network and adaptive micro ROI feature point extraction and path tracking control based on the PFT preview theory.
[0165] In the work of shallow and deep feature fusion, the present invention improves the high-low layer feature fusion ability through the Group Aggregation Bridging Residual Module (GAB module). Compared with the traditional Unet network, the ridge boundary segmentation accuracy is improved. In the work of channel space feature enhancement, the present invention uses the Coordinate Attention Mechanism (CA module) to enhance the effective features of the network for the number of picture channels and the picture space. The recognition accuracy is improved, and high segmentation stability is maintained even in the case of foliage occlusion. Multi-scale feature fusion segmentation: The Atrous Spatial Pyramid Pooling layer structure fuses multi-scale features to improve the semantic segmentation accuracy of strawberry ridges. Due to the different sizes and shapes of plants and the hilly terrain of the outdoor strawberry garden planting area, the present invention designs an adaptive micro ROI navigation path extraction method based on 2D vision, which explores the ridge boundary by using two sub-ROIs, thereby dynamically updating the edge of the micro ROI, and finally can better adapt to the influence brought by the crop structure change in each sub-strip.
[0166] As Figure 6 shown in the ridge feature extraction result, on the Jetson nano B01 development board equipped with an integrated deep learning environment, by calling the trained strawberry garden ridge semantic segmentation model based on the improved Unet network, using the RealSenseD435 camera to capture ridge image data in real time, and sending it into the model to complete the ridge feature extraction task.
[0167] As Figure 7 shown in the ridge navigation route planning result, after using image preprocessing operations such as dilation, erosion, and image binarization on the image with the ridge features extracted, use the strawberry garden inter-ridge navigation route planning algorithm based on the adaptive micro ROI.
[0168] First, preprocess the input image to obtain a binary segmentation image with smooth edges; determine the label center r n according to the peak value of the white pixel projection in the sub-strip, and set the initial ROI regions on the left and right of the label center respectively, with the center points c l and c r as the benchmarks to determine the ROI width to optimize the detection area; then, translate the sub-ROI regions to the left and right sides along the strip direction respectively with a constant pixel step i for sequential search until reaching the label edge or the boundary of the sub-strip; during the search process, the strip n and the mask image mask n composed of the sub-ROI regions are subjected to a bitwise AND operation to obtain the image res n .
[0169] Calculate the percentage P n of the white pixels in the bitwise AND operation result res i in the total micro ROI area.
[0170] When P i <T, where T is the threshold for white pixel setting, stop the search; if the threshold condition is not met, double the pixel step size and repeat the search until the stop condition is met; finally, according to the distance between the label center r n and the centers c l and c r of the left and right ROI regions, update the boundaries m l and m r , and calculate the average value of white pixels in the new region to update the position of the label center r n . The formula is as follows:
[0171]
[0172]
[0173] where S is the total number of white pixels in the micro ROI region; each point in the region is (x i , y i ); and the width of each sub-strip is defined as d s .
[0174] According to the target point (x avg , y avg +d s ), update the starting iteration coordinate point of the next strip. Then repeat the process of extracting the micro ROI until the search and path extraction processes for all sub-strips are completed. Finally, use polynomial regression to fit the set of all label centers r n into the planned path at the current moment.
[0175] As Figure 8 shown in the schematic diagram of the transformation between the pixel coordinate system and the robot base coordinate system, the planned path is transformed to the front-wheel drive wheel coordinate system through coordinate transformation to obtain the coordinate points and slope data of the navigation line in the front-wheel drive wheel coordinate system. The robot base coordinate system is O v -X v Y v Z v , the camera coordinate system is O c -X c Y c Z c , the coordinate points (U, V) of the navigation line in the pixel coordinate system, the pitch angle of the camera in the horizontal direction is defined as α, the horizontal installation height of the camera is H, and the longitudinal distance from the camera to the drive wheel is L. Camera hardware parameter: f x represents the length of the focal length in the x-axis direction described in pixels, that is, the normalized focal length in the x-axis direction; f yIt represents the length of the focal length in the y-axis direction described in pixels, that is, the normalized focal length in the y-axis direction. (u0, v0) represents the optical center, that is, the intersection point of the camera optical axis and the image plane. All camera hardware parameters can be measured through camera calibration. The conversion relationship between the navigation points in the robot base coordinate system and the camera coordinate system is as follows:
[0176]
[0177] Such as Figure 9 The schematic diagram of ultrasonic ranging calculation shown. If d b1 and d f1 are the ultrasonic ranging results fixed on both sides of the left front and left rear drive wheels respectively, and l is the wheelbase of the vehicle body, that is, the vertical distance from the center of the front wheel to the center of the rear wheel; it is assumed that the front and rear spacing of the side ultrasonic sensors is equal to the wheelbase l of the vehicle body, and it is assumed that the ultrasonic sensors measure distance at a frequency of 20HZ. Calculate the distance d l between the vehicle and the left ridge according to the ultrasonic ranging results, and obtain the heading angle of the vehicle body attitude, that is, the angle between the vehicle and the ridge path There is the following relationship:
[0178]
[0179] Calculate in the same way to obtain the ultrasonic ranging results fixed on both sides of the right front and right rear drive wheels, and similarly obtain the distance d r between the vehicle and the right ridge calculated from the ultrasonic ranging results, and obtain the distance d c from the center line of the suspension center of the front drive wheel. There is the following relationship:
[0180] d c = |d l - d r | (39)
[0181] The pose of the trolley is calculated through ultrasonic ranging, and the angle between the vehicle and the ridge path and the distance d c from the center line of the ridge path are obtained.
[0182] The cross-ridge mobile chassis uses an MPC controller to control the motion trajectory, and realizes steering and speed control by controlling the rotation speeds of the two drive wheel groups of the cross-ridge mobile chassis, and further controls the motion trajectory of the cross-ridge mobile chassis. The cross-ridge mobile chassis has strong passability and can achieve in-situ steering, and can operate in an environment with relatively limited space.
[0183] The target navigation line planned through the semantic segmentation of the ridge provides the planned path for the robot. The preview time is determined by the preview distance given when the robot is moving. At the current moment, the lateral position and lateral velocity of the robot are known. After the preview time, the robot reaches the preview point under the action of the lateral velocity and lateral acceleration, and its lateral displacement can be calculated by the numerical integration method. In order to make the robot reach the preview point, the required lateral acceleration is calculated according to the desired lateral velocity and lateral acceleration, and further the curvature radius of the robot's centroid rotating around the instantaneous center is obtained. The yaw angular velocity is determined by the centroid velocity and the curvature radius, and the reference yaw angle is obtained by integrating within the sampling period of the controller.
[0184] Substitute this reference yaw angle into the error prediction state model of the discrete autonomous navigation strawberry spraying robot. By judging whether the Euclidean distance from the current position to the target navigation line trajectory is less than the given threshold, in the area where the distance is greater than the given threshold δ, the MPC controller in the preview following mode is used to more quickly approach the target navigation line trajectory. If it is less than the given threshold δ, only the MPC controller is used for path tracking.
[0185] As Figure 10 shown in the schematic diagram of path tracking control based on the MPC controller, (g x ,g y ) is a path point on the path within the forward viewing distance range. The planned path within this path point is the effective path. The coordinate point positions and the heading angle data of the effective path are used as the expected values of the MPC controller to generate the error prediction state model of the discrete autonomous navigation strawberry spraying robot, establish the objective function of the linear model predictive controller MPC, and solve the optimization problem with inequality constraints:
[0186] Based on the above objective function, the MPC control problem can be summarized as follows:
[0187]
[0188] where Δu min and Δu max are the minimum and maximum values of the control input increment, determined according to the two extreme cases solved; t = 1, 2,..., N c ,…,N p , N p is the MPC prediction time domain length, N c is the MPC control time domain length; y t is the state variable value at time t; is the target reference state variable value at time t; u t is the control input at time t; Δu tis the incremental control input at time t; ε is the slack variable introduced to ensure a feasible solution, and Q, P, F, and ρ are the weight matrices of the system state tracking error, incremental control input, system terminal control input, and slack factor, respectively.
[0189] A set of optimal control quantities u is obtained by solving * =[v * , ω * T , and then the control input quantity is decoupled to obtain the actual front-wheel drive wheel speed v * =[v1 * , v2 * T
[0190] The optimal control signal is connected to the signal conversion module through serial communication. Using its TTL-RS485 signal conversion function, the control signal is input into the hub motor driver to drive the front-wheel motor to rotate.
[0191] As Figure 11 shown, it is the flow chart of ridge channel semantic segmentation and path planning for the cross-ridge strawberry variable spray width spraying robot. First, the ridge channel image information is captured in real time by a RealSense D435 camera, and the improved Unet network is used for semantic segmentation to extract the ridge channel features. Secondly, after image preprocessing operations such as dilation, erosion, and image binarization are performed according to the results of semantic segmentation, the feature points of the ridge channel navigation line are extracted using an adaptive micro ROI strawberry garden inter-ridge navigation line path planning algorithm. After that. According to the coordinate mapping relationship from the pixel coordinate system to the robot's front-wheel base coordinate system, the feature points of the navigation line based on the robot's base coordinate system are obtained, and then this set of navigation line feature points is inversely synthesized into the target navigation trajectory using polynomial interpolation for path planning between ridge channels.
[0192] As Figure 12 shown, it is the flow chart of path tracking control for the cross-ridge strawberry variable spray width spraying robot. First, it is judged whether the current robot reaches the end point or is on the ridge through GNSS positioning, and the distance d l between the vehicle and the left ridge and the distance d r between the vehicle and the right ridge are updated using ultrasonic ranging, and the heading angle of the vehicle body attitude, that is, the angle between the vehicle and the ridge channel, is obtained Further, an MPC controller based on the PFT preview control algorithm is used according to the planned path. Whether to use the PFT preview control algorithm is determined by judging whether the cumulative sum of the reference point errors within the forward view distance range meets the set threshold: if the cumulative sum of the errors is greater than the set error threshold, the PFT preview control algorithm is used for fast tracking; if the cumulative sum of the errors is less than the set error threshold, only the MPC controller is used for tracking. After obtaining a set of optimal control quantities based on the above control process, the actual control quantity input is obtained through control quantity decoupling, and the level signal is converted, and finally the motor is driven to complete the path tracking control.
[0193] Specific extreme case 1, such as Figure 13 the maximum yaw angle shown The calculation process is as follows:
[0194] First, simplify the mechanical structure model into a simplified schematic diagram with only the chassis suspension, as Figure 14 shown. Since the rear wheel size is small, the radius of the rear wheel driven wheel is ignored. Only when the hub center touches the boundary of the ridge channel, extreme case 1 is triggered;
[0195] Assume that the wheel travels along the center line of the ridge channel, and draw PF perpendicular to the extension line of ED at F. From the geometric relationship, it can be obtained that BO is equal to half of the width of the robot chassis, OD = BP is equal to the length of the robot chassis. According to ΔOAB~ΔOCD and the mathematical relationship, AE is equal to the length from the center point of the robot to the outermost boundary of the ridge channel, that is, AE = l + d / 2, where l is the distance from the center line to the boundary of the ridge channel and d is the width of the ridge channel. Denote BP = L0 as the length of the robot chassis and BO = D0 as half of the width of the robot chassis.
[0196] To obtain The specific formula is:
[0197]
[0198] Furthermore, solve the variable at the analytical solution.
[0199]
[0200] The maximum yaw angle of specific extreme case 2 The calculation process is as follows:
[0201] Since the front wheel size is large and cannot be ignored, considering the radius r of the front wheel drive wheel, simplify the mechanical structure model into a simplified schematic diagram with only the chassis suspension. When the front wheel drive wheel touches the boundary of the ridge channel, extreme case 2 is triggered, as Figure 15 shown.
[0202] According to asFigure 16 From the geometric relationship shown, the maximum yaw angle is the angle between the vehicle body movement direction and the axis parallel to the ridge channel The angle between the contact point of the front drive wheel of the chassis with the inner wall of the ridge channel and the front axle of the chassis of the front drive wheel is ∠QPN = β, the angle between the contact point of the front drive wheel of the chassis with the inner wall of the ridge channel and the direction perpendicular to the ridge channel is ∠MPN = γ, and the distance between the suspension center of the front drive wheel and the center line is d c The distance from point P to the boundary of the ridge channel is PM = l + d c
[0203] Thus, the relationship formula can be obtained:
[0204]
[0205] Furthermore, the maximum yaw angle is obtained :
[0206]
[0207] Denote the maximum value of the robot yaw angle as The minimum value of the robot yaw angle is
[0208] The range of the robot yaw angle is finally determined as:
[0209]
[0210] Where and are solved according to formula (42) and formula (44) respectively
[0211] By solving its differential within a sampling period T, the range of the angular velocity is:
[0212]
[0213] Where the maximum value of the angular velocity of the robot's center of mass is denoted as The minimum value of the angular velocity of the robot's center of mass is denoted as
[0214] The range of the input speed is given artificially. Given a constant tracking linear velocity of the center of mass, a set of actual control quantities u t is obtained
[0215]
[0216] Where is the minimum value of the increment of the linear velocity of the robot's center of mass; is the maximum value of the increment of the linear velocity of the robot's center of mass
[0217] Finally, through the transformation of the incremental control quantity, a set of control quantity increments Δu are obtained t constraint:
[0218]
[0219] where is the minimum value of the increment of the linear velocity of the robot's center of mass; the maximum value of the increment of the linear velocity of the robot's center of mass; is the minimum value of the increment of the angular velocity of the robot's center of mass; is the minimum value of the increment of the angular velocity of the robot's center of mass.
[0220] By constraining the input quantity according to the special environmental conditions of the strawberry field ridge, the phenomenon of "hitting the ridge" caused by large adjustments is avoided. Solving the above MPC problem can obtain the system input control variable at the next moment, so as to minimize the tracking distance error; the input is as small as possible; the change rate of the input is as small as possible, achieving the purpose of reducing power consumption.
Claims
1. A strawberry field ridge visual navigation method based on an improved Unet network and an MPC control method based on the PFT preview theory, mainly including the following steps: Step 1, obtain the ridge field image through the RealSense D435 depth camera vision perception module, input the image into the improved Unet network for image semantic segmentation and extract the feasible area of the ridge road. The improved Unet network enables the network to more accurately segment the feasible area of the ridge road by using a coordinated attention encoding module, a multi-axis Hadamard product attention module, a group aggregation bridging module, and an atrous spatial pyramid pooling layer module, and performs excellently in complex scenarios; Step 2, design an adaptive micro ROI navigation path extraction algorithm based on 2D vision. By dynamically adjusting the ROI area boundary, it adapts to different crop structures and terrain conditions. The extracted navigation planning path is transformed to the robot base coordinate system through the 2D to 3D conversion from the pixel coordinate system to the camera coordinate system and the rigid body transformation from the camera coordinate system to the robot base coordinate system, and the position information of the coordinate points of the navigation line is obtained; Step 3, construct the kinematic model of the strawberry spraying robot, and obtain the error prediction state model through linearization. Adopt an incremental control strategy combined with the MPC controller to optimize the objective function, and solve the input control variable at the next moment in real time to achieve path tracking control of the given input quantity; propose an MPC controller based on the PFT preview control algorithm, judge and call the preview control mode through the error threshold, and impose constraints on the input control quantity in combination with the special environmental conditions of the strawberry field ridge road to avoid the "ridge collision" phenomenon.
2. The control method according to claim 1, characterized in that, In Step 1, the improved Unet network includes: Coordinated attention encoding module: The coordinated attention module generates attention weights through adaptive pooling in the horizontal and vertical directions. First, average pool the input image with size C×H×W along the height and width directions to obtain direction-aware feature maps with sizes C×H×1 and C×1×W. Then, splice the two and fuse them into a single feature map through a two-dimensional convolution. After normalization and non-linear activation, generate an attention weight map with the same size as the input image through two two-dimensional convolutions. Finally, fuse it with the original image to output a feature map with an attention mechanism; Multi-axis Hadamard product attention module: The multi-axis Hadamard product attention module groups the feature map by channels, applies parameterized convolutional kernels on the XY plane, ZX axis, and ZY axis respectively, and fuses multi-dimensional features through the Hadamard product. This module first decomposes the input image with size C×H×W along the channel dimension into 4 equal quantum feature maps, and the decomposed feature dimension is B×(C / 4)×H×W, where B represents the data batch size; among them, 3 equal quantum feature maps respectively pass through the Hadamard attention mechanism under the XY, ZX, and ZY axes, and the 4th sub-feature map only applies depthwise separable convolution. Finally, splice the dimensions, perform layer normalization, and then output after depthwise separable convolution; Group Aggregation Bridging Module: The group aggregation bridging module fuses high-level semantic features and low-level detail features through multi-dilation rate atrous convolution. First, the module adjusts the size C h ×H h ×W h of the input image with high-level features to be consistent with that of the low-level features C l ×H l ×W l by using depthwise separable convolution and bilinear interpolation, where C h 、H h 、W h represent the number of high-level feature channels, height, and width respectively. Then, the adjusted low-level features and high-level features are divided into four groups along the channel dimension respectively, and then each group of low-level features is concatenated with each group of high-level features to obtain four groups of fused features. Then, the mask image is fused with the four groups of features, and atrous convolutions with dilation rates of 1, 3, 5, and 7 are used for feature extraction respectively. Finally, the four groups of features are concatenated along the channel dimension, and 1×1 convolution is used for feature integration and output; Atrous Spatial Pyramid Pooling layer: The atrous spatial pyramid pooling fuses multi-scale information to output the segmentation result; for a given input, parallel sampling is performed using four atrous convolutions with dilation rates of 6, 12, 18, 24 and a 1×1 convolution, and in addition, a global average pooling branch is added. Finally, the above 6 branches are concatenated along the channel dimension and the number of channels is reduced to the expected value through a 1×1 convolution and then output.
3. The control method according to claim 1, characterized in that In step 2, the process of the adaptive micro ROI navigation path extraction algorithm based on 2D vision includes: Step 2.1: First, input the binarized segmentation image with smoother edges obtained after image preprocessing operations. Mark the label center of the sub-stripes based on the projection peaks of the white pixels of the sub-stripes, denoted as r n ; Frame the sub-ROI regions on the left and right sides of r n . Appropriate values can be selected according to the crop type to determine the widths of these two sub-ROIs centered on the initial center points c l and c r respectively, with the aim of obtaining the best detection area; Step 2.2: Along the direction of the sub-strip, at the label center r n The sub-ROI regions on the left and right sides are respectively translated in opposite directions with a constant pixel value increment along the sub-strip direction, and the moving pixel value increment is defined as i; Step 2.3: The sub-ROI area will perform a sequential search along the strip until it finds the edge of the label or until it reaches the boundary of the nth sub-strip. If the current sub-ROI area overlaps with the already fitted sub-ROI area from the neighboring area, the search for this loop will also stop; Step 2.4: Take the sub-strip n and perform a bitwise AND operation with the mask image mask composed of sub-ROI regions n The result obtained is denoted as res n , and the result is the overlapping part of the character strip and the sub-ROI area (let the overlapping part be white pixels), and calculate the percentage of white pixels in the strip. If it is less than the threshold T, stop the inspection, which means that the threshold condition is met. If it meets this threshold condition, it indicates that the sub-ROI area reaches the edge of the label; Step 2.5: If the percentage of white pixels in the strip does not meet the requirements of the threshold condition, it is considered that the edge has not been reached and the set pixel step size is too small. Go back to step 2 and set it to 2i and search again. The search based on the sub-ROI will continue to search with a step size of ki (k = 1, 2, 3,..., 16) until it meets any one of the stop conditions in 2.3; Step 2.6: When the stop condition is met to stop the search and the threshold condition is satisfied, according to the distance between the label center r n and the center points c l and c r of the left and right sub-ROI regions, update the boundaries m l and m r of the micro-ROI region as the new micro-ROI region, calculate the average value of white pixels within this region according to the formula, and update the label center r n of the nth sub-strip, and repeat the entire process starting from Step 1 for the next sub-strip until the last sub-strip is searched.
4. The control method according to claim 1, wherein The specific process of step 3 is as follows: Step 3.1: Obtain 2 input control variables and 3 system state variables of the autonomous navigation strawberry spraying robot at the current moment. The 2 input control variables are the centroid movement speed of the robot and the angular velocity of the centroid rotating around the instantaneous center. The 3 system state variables are the x-axis coordinate and y-axis coordinate of the robot in the world coordinate system, and the angle of rotation around the z-axis coordinate of the centroid; Step 3.2: Construct a kinematic model of the autonomous navigation strawberry spraying robot according to the wheel parameters and robot parameters; Step 3.3: Control the increment of the input as the centroid velocity v c and the increment of the yaw angular velocity ω, and further decouple according to the formula to obtain the speeds v1 and v2 acting on each driving wheel; by changing the duty ratio of the driving signal of the in-wheel motor, and further changing the system input as the speeds of the two driving wheels, so that the robot realizes two-wheel differential operation when running; Step 3.4: In order to ensure that the robot can more quickly track the target navigation line under large errors, the present invention is based on the PFT preview control algorithm to judge that the cumulative sum of the reference point errors within the preview distance range reaches the error threshold δ, so as to determine whether the condition for calling the PFT preview control algorithm is satisfied. This threshold δ needs to be given in advance manually; Step 3.5: According to the discretized error prediction state model of the autonomous navigation strawberry spraying robot, establish the objective function of the linear model predictive controller MPC; Step 3.6: Based on the target navigation path, solve the objective function to determine the input control variable of the MPC at the next moment; the input control variable is used to minimize the error between the robot and the target navigation path with the minimum control energy to real-time control the movement trajectory of the robot.