Multi-modal feature fusion apple recognition method based on YOLOv5s
Through multimodal feature fusion and model optimization, the problems of insufficient multimodal feature fusion accuracy and inaccurate point cloud matching in Apple recognition are solved, high-precision Apple recognition and three-dimensional visualization are realized, and agricultural intelligence development is promoted.
Patent Information
- Application Number
- CN202510356855.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has insufficient multimodal feature fusion accuracy in Apple recognition, the point clouds match with color images inaccurately, and the YOLOv5s model input layer structure is fixed and it is difficult to process multimodal features, resulting in limited recognition accuracy and generalization capabilities.
By integrating color images, key fruit feature maps, depth images and point cloud coordinate information, the YOLOv5s model input layer is optimized, and a variety of data enhancement methods and mixed precision training techniques are used to achieve accurate matching and efficient training of multimodal features.
It significantly improves the accuracy and robustness of apple recognition, realizes three-dimensional point cloud visualization, supports agricultural picking and sorting, and improves picking efficiency and fruit integrity.
Smart Images

Figure CN120340020A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a multi-modal feature fusion apple recognition method based on YOLOv5s. Background Art
[0002] As one of the most widely grown fruits in the world, apple picking, sorting and quality inspection have a significant impact on agricultural production and economic benefits. In recent years, with the development of deep learning technology, object detection algorithms based on convolutional neural networks (CNNs) have been widely used in the field of apple recognition, especially the YOLO series of algorithms, which have attracted attention due to their high efficiency and real-time performance.
[0003] However, existing technologies still face many challenges in complex scenarios. On the one hand, the accuracy of multimodal feature fusion is insufficient, and existing fusion methods are difficult to effectively integrate RGB images, depth images, and point cloud data, resulting in limited recognition accuracy. On the other hand, in the process of fusion of point cloud and color image, there is a technical bottleneck in the precise matching of point cloud coordinates and color image pixel coordinates, which affects the accuracy and reliability of fused data.
[0004] In addition, the input layer structure of the YOLOv5s model is fixed, which makes it difficult to directly process multimodal feature inputs, resulting in the lack of efficient training techniques for apple recognition scenarios in existing training methods, further limiting the generalization ability and training efficiency of the model under complex lighting conditions and background interference. Therefore, how to improve the accuracy of multimodal feature fusion, optimize the correction method of point clouds and color images, improve the YOLOv5s model to adapt to multimodal inputs, and improve model performance through multimodal data have become key issues that need to be solved in current apple recognition technology. Summary of the invention
[0005] In view of the above technical problems, this technical solution provides a multi-modal feature fusion apple recognition method based on YOLOv5s, which significantly improves the accuracy and robustness of apple recognition by fusing multiple modal features such as color images, key fruit feature maps, depth images and point cloud coordinate information, and realizes the three-dimensional point cloud visualization of apples; it effectively solves the above problems. The present invention is implemented through the following technical solutions:
[0006] A multimodal feature fusion apple recognition method based on YOLOv5s, comprising the following steps:
[0007] Step 1. Image acquisition: obtain color images, depth images and point cloud images through the depth camera;
[0008] Step 2. Decomposition and reconstruction of color image: decompose the acquired color image into three color components of R, G, and B; reconstruct the three channels using RG color difference operator to obtain RG color difference map;
[0009] Step 3. Edge feature imaging: Grayscale the color image; Use the Laplacian operator for edge detection to obtain an edge map;
[0010] Step 4. Obtaining the key fruit feature map: Fuse the color difference map and the edge map according to a weight ratio of 6:4 to obtain the key fruit feature map;
[0011] Step 5. Obtaining point cloud coordinate information: Obtain the point cloud image in Step 1 and find the corresponding color image; Project the point cloud image onto the color image: First, convert the point cloud coordinates to camera coordinates and project the points from the camera coordinate system to the image plane; Convert the processed camera coordinates to pixel coordinates using the camera intrinsic matrix and normalize them; Flip the image horizontally and correct the projection offset to obtain the color image pixel coordinates corresponding to the point cloud coordinate information XYZ, and compare with the original pixel coordinates at the same position to determine whether the point cloud coordinate information XYZ matches;
[0012] Step 6. Multimodal feature fusion: Synthesize an 8-channel image from R, G, B, the key fruit feature map, the depth image, and the point cloud coordinate information XYZ;
[0013] Step 7. Model optimization: Modify the number of channels of the input layer of YOLOv5s from 3 to 8 to adapt to the fused multimodal features;
[0014] Step 8. Model training: Use the fused multimodal features as input for deep learning training. During the training process, adopt data augmentation methods, including random cropping, color jittering, and Gaussian noise addition, to improve the generalization ability of the model; Use the mixed-precision training technique to accelerate the training process and reduce memory occupancy;
[0015] Step 9. Model testing: Use a new dataset to verify and evaluate the trained model, and visualize the recognition results in 3D point cloud.
[0016] Further, the calculation formula of the R-G color difference operator described in Step 2 is:
[0017] R-G(x,y) = R(x,y) - G(x,y)
[0018] where R-G(x,y) represents the R-G color difference map, and R(x,y) and G(x,y) represent the red component and the green component of the image at the pixel position (x,y), respectively.
[0019] Further, the calculation formula of the Laplacian operator described in Step 3 is:
[0020]
[0021] Among them, Δf(x,y) represents the value of the Laplacian operator at the image position (x,y), that is, the sum of the second-order derivatives of the image gray value at this position. represents the second-order partial derivative of the image gray value f(x,y) with respect to the x-axis, reflecting the curvature change of the image in the x-axis direction. represents the second-order partial derivative of the image gray value f(x,y) with respect to the y-axis, reflecting the curvature change of the image in the y-axis direction. f(x,y) represents the gray value of the grayscale image at the pixel position (x,y).
[0022] Furthermore, the chromatic aberration map and the edge map are fused according to a weight ratio of 6:4 in step 4. The specific fusion formula is:
[0023] F(x,y) = 0.6×(R - G(x,y)) + 0.4×L(x,y)
[0024] Among them, F(x,y) represents the value of the key fruit feature map at the pixel position (x,y), and L(x,y) represents the edge map.
[0025] Furthermore, the point cloud image is projected onto the color image in step 5. The specific operation method is as follows:
[0026] First, convert the point cloud coordinates to camera coordinates:
[0027]
[0028] Among them, E is the external parameter matrix, which is initialized as the identity matrix here; P cam is the 3D point in the camera coordinate system, and [x y z 1] is the homogeneous coordinate of the 3D point in the point cloud.
[0029] Next, project the point from the camera coordinate system to the image plane. Here, the last component of the homogeneous coordinate needs to be removed, and only the first three components of P cam are used for projection, and it is represented by P cam,3 as follows:
[0030]
[0031] Among them, x cam , y cam , z cam are the coordinates of the point in the camera coordinate system.
[0032] Furthermore, the processed camera coordinates are converted to pixel coordinates using the camera internal parameter matrix. The specific operation method is as follows:
[0033] P pixel= K·P cam,3
[0034]
[0035] where K is the intrinsic parameter matrix, and f x and f y are the focal lengths of the camera on the x-axis and y-axis respectively; c x and c y are the coordinates of the optical center of the camera; P cam,3 are the first three components of the 3D point in the camera coordinate system; P pixel are the pixel coordinates of the point on the image plane.
[0036] Then, normalize P pixel to obtain the actual pixel coordinates:
[0037]
[0038] where u and v are the pixel coordinates on the image plane after transformation by the intrinsic parameter matrix and normalization, corresponding to the abscissa and ordinate of the image respectively.
[0039] Furthermore, the mirror flipping of the image described in step 5 is specifically performed as follows:
[0040] u′ = m w - 1 - u
[0041] where u' is the abscissa of the image after mirror flipping; m w is the width of the image (number of pixels);
[0042] Finally, correct the projection offset:
[0043]
[0044] where x offset and y offset are the correction parameters for correcting the projection offset, representing the offset amounts in the x-axis and y-axis directions respectively; u o and v o are the pixel coordinates on the corrected image plane, corresponding to the accurate abscissa and accurate ordinate of the image respectively.
[0045] Further, the specific operation method of step 7 is as follows: The dimension of the convolution kernel corresponding to the input layer is represented as W×H×m×n, where W×H represents the length and width of the convolution kernel; for the YOLOv5s network, W×H is 6×6; m represents the number of convolution kernel layers, which should be consistent with the number of channels of the input image, and m is adjusted to 8 here; n represents the number of convolution kernels, that is, n different convolution kernels are used to extract features from the image. For the YOLOv5s network, n is 32 and remains unchanged here.
[0046] Further, the data augmentation method in step 8 also includes random flipping and rotation to further improve the robustness of the model.
[0047] Further, the use of the mixed-precision training technique in step 8 to accelerate the training process and reduce memory occupancy is achieved by dynamically adjusting the numerical precision in model training, realizing the acceleration of the training process and a significant reduction in memory occupancy; the specific operation method is as follows: Automatically switch to use single-precision floating-point numbers (FP32) and half-precision floating-point numbers (FP16) during the training phase, and while maintaining the convergence performance of the model, make full use of the computing power of modern GPUs; at the same time, introduce a loss scaling mechanism to avoid numerical underflow problems that may occur in half-precision calculations.
[0048] Further, during the model testing process in step 9, the reverse mapping technique is used to fuse the recognition results with the three-dimensional point cloud data to generate an intuitive three-dimensional visualization image. This visualization method can display the appearance, position, and shape of apples in real time, providing accurate navigation and decision-making support for agricultural picking robots. In addition, this method supports multi-view interactive viewing, enabling operators to evaluate the distribution of apples from different angles, thereby optimizing the picking path and improving the picking efficiency and fruit integrity.
[0049] Beneficial effects
[0050] A multi-modal feature fusion apple recognition method based on YOLOv5s proposed by the present invention has the following beneficial effects compared with the prior art:
[0051] (1) By fusing multiple modal features such as R, G, B, key fruit feature maps, depth images, and point cloud coordinate information, the present invention constructs an 8-channel multi-modal input, significantly improving the accuracy and robustness of apple recognition. This fusion method not only uses color differences and edge detection to highlight the contrast between the fruit and the background, but also combines depth and three-dimensional space information, effectively solving the problem of insufficient recognition ability of the prior art under complex backgrounds and different lighting conditions.
[0052] (2) The key fruit feature extraction method proposed by the present invention effectively reduces the mutual redundancy and interference of traditional feature information by weighted fusion of the color difference map and the edge map. An improved point cloud and color image calibration projection method is also proposed, which realizes the accurate matching of point cloud coordinates and color image pixel coordinates through precise coordinate transformation and projection offset correction, further improving the accuracy and consistency of multi-modal data fusion.
[0053] (3) The present invention optimizes the input layer of the YOLOv5s model, adjusts the input channels of the convolutional layer to adapt to multi-modal feature input, and enhances the adaptability of the model to complex scenes. During the training process, various data augmentation methods such as random cropping, color jittering, Gaussian noise addition, random flipping and rotation are adopted, and combined with the mixed precision training technology, which not only improves the generalization ability of the model, but also accelerates the training process and reduces the memory occupation.
[0054] (4) The present invention improves the apple recognition accuracy, realizes the three-dimensional point cloud visualization of the apple recognition result, provides more intuitive and efficient technical support for apple picking, sorting and quality detection in precision agriculture, and promotes the development of agricultural intelligence. Brief Description of the Drawings
[0055] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0056] Figure 2 It is a color image of an apple in the present invention.
[0057] Figure 3 It is a depth image of an apple in the present invention.
[0058] Figure 4 It is a point cloud image of an apple in the present invention.
[0059] Figure 5 It is an R-G color difference map of an apple in the present invention.
[0060] Figure 6 It is an edge map of an apple in the present invention.
[0061] Figure 7 It is a key fruit feature map of an apple in the present invention.
[0062] Figure 8 It is a multi-channel image synthesis map of the present invention.
[0063] Figure 9 It is a schematic diagram of the architecture optimization of the input layer of the YOLOv5s object detection model in the present invention.
[0064] Figure 10 It is a view set of result annotation and three-dimensional point cloud visualization of model testing in the present invention. Detailed implementation manners
[0065] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Without departing from the design concept of the present invention, various variations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention should all fall within the protection scope of the present invention.
[0066] Embodiment 1:
[0067] As Figure 1 shown, a multi-modal feature fusion apple recognition method based on YOLOv5s includes the following steps:
[0068] Step 1. Image acquisition: Obtain the apple color image as Figure 2 shown, the apple depth image as Figure 3 shown, and the apple point cloud image as Figure 4 shown through a depth camera.
[0069] Step 2. Decomposition and reconstruction of the color image: Decompose the obtained color image into three color components of R, G, and B; reconstruct the three channels using the R-G color difference operator to obtain an R-G color difference map; as Figure 5 shown.
[0070] The calculation formula of the R-G color difference operator is:
[0071] R-G(x,y) = R(x,y) - G(x,y)
[0072] where R-G(x,y) represents the R-G color difference map, and R(x,y) and G(x,y) respectively represent the red component and the green component of the image at the pixel position (x,y).
[0073] Step 3. Edge feature imaging: Grayscale the color image; perform edge detection using the Laplacian operator to obtain an edge map; as Figure 6 shown.
[0074] The calculation formula of the Laplacian operator is:
[0075]
[0076] where Δf(x,y) represents the value of the Laplacian operator at the image position (x,y), that is, the sum of the second-order derivatives of the image grayscale value at this position, represents the second-order partial derivative of the image grayscale value f(x,y) with respect to the x-axis, reflecting the curvature change of the image in the x-axis direction, Represents the second-order partial derivative of the image grayscale value f(x, y) with respect to the y-axis, reflecting the curvature change of the image in the y-axis direction. f(x, y) represents the grayscale value of the grayscale image at the pixel position (x, y).
[0077] Step 4. Obtaining the key fruit feature map: Fuse the color difference map and the edge map according to a weight ratio of 6:4 to obtain the key fruit feature map; as Figure 7 shown.
[0078] The specific fusion formula is:[[]]
[0079] F(x,y) = 0.6×(R - G(x,y)) + 0.4×L(x,y)
[0080] where F(x, y) represents the value of the key fruit feature map at the pixel position (x, y), and L(x, y) represents the edge map.
[0081] Step 5. Obtaining the point cloud coordinate information: Obtain the point cloud image in Step 1 and find the corresponding color image; Project the point cloud image onto the color image: First, convert the point cloud coordinates to camera coordinates:
[0082]
[0083] where E is the extrinsic parameter matrix, which is initialized as the identity matrix here; P cam is the 3D point in the camera coordinate system, and [x y z1] is the homogeneous coordinate of the 3D point in the point cloud.
[0084] Next, project the point from the camera coordinate system to the image plane. Here, the last component of the homogeneous coordinate needs to be removed, and only the first three components of P cam are used for projection, represented by P cam,3 as follows:
[0085]
[0086] where x cam and y cam and z cam are the coordinates of the point in the camera coordinate system.
[0087] Project the point from the camera coordinate system to the image plane; Convert the processed camera coordinates to pixel coordinates using the camera intrinsic parameter matrix. The specific operation method is:
[0088] P pixel = K·P cam,3
[0089]
[0090] Among them, K is the intrinsic parameter matrix, and f x and f y are the focal lengths of the camera on the x-axis and y-axis respectively; c x and c y are the coordinates of the optical center of the camera; P cam,3 are the first three components of the 3D point in the camera coordinate system; P pixel are the pixel coordinates of the point on the image plane.
[0091] Then, P pixel is normalized to obtain the actual pixel coordinates:
[0092]
[0093] Among them, u and v are the pixel coordinates on the image plane after transformation by the intrinsic parameter matrix and normalization, corresponding to the abscissa and ordinate of the image respectively.
[0094] The image is flipped horizontally. The specific operation method is as follows:
[0095] u′ = m w -1 - u
[0096] Among them, u' is the abscissa of the image after horizontal flipping; m w is the width (number of pixels) of the image;
[0097] Finally, the correction projection offset is performed:
[0098]
[0099] Among them, x offset and y offset are the correction parameters for correcting the projection offset, representing the offset amounts in the x-axis and y-axis directions respectively; u o and v o are the pixel coordinates on the corrected image plane, corresponding to the accurate abscissa and accurate ordinate of the image respectively.
[0100] The color image pixel coordinates corresponding to the point cloud coordinate information XYZ are obtained through the above steps, and whether the point cloud coordinate information XYZ matches is determined by comparing with the original pixel coordinates at the same position.
[0101] Step 6. Multi-modal feature fusion: Synthesize the R, G, B, key fruit feature map, depth image, and point cloud coordinate information XYZ into an 8-channel image; as Figure 8 shown; Modify the number of channels of the input layer of YOLOv5s, as Figure 9As shown in the figure, the number of channels in the input layer is modified from 3 to 8 to adapt to the fused multi-modal features. Specifically, the modification is to adjust the number of input channels of the convolutional layer. The dimension of the convolutional kernel corresponding to the input layer can be expressed as W×H×m×n, where W×H represents the length and width of the convolutional kernel. For the YOLOv5s network, W×H is 6×6; m represents the number of convolutional kernel layers, which should be consistent with the number of channels of the input image. Here, m needs to be adjusted to 8; n represents the number of convolutional kernels, that is, n different convolutional kernels are used to extract features from the image. For the YOLOv5s network, n is 32 and remains unchanged here.
[0102] Step 7. Model optimization: Modify the number of channels in the input layer of YOLOv5s, changing the number of channels in the input layer from 3 to 8 to adapt to the fused multi-modal features;
[0103] Step 8. Model training: Use the fused multi-modal features as the input for deep learning training. During the training process, data augmentation methods are adopted, including random cropping, color jittering, and Gaussian noise addition, to improve the generalization ability of the model; The data augmentation method also includes random flipping and rotation to further improve the robustness of the model.
[0104] Use the mixed-precision training technique to accelerate the training process and reduce memory occupancy; During the model training process, the mixed-precision training technique is adopted. By dynamically adjusting the numerical precision in model training, the acceleration of the training process and a significant reduction in memory occupancy are achieved. Specifically, this technique automatically switches between single-precision floating-point numbers (FP32) and half-precision floating-point numbers (FP16) during the training phase. While maintaining the convergence performance of the model, it fully utilizes the computing power of modern GPUs to significantly improve the training efficiency. In addition, by introducing a loss scaling mechanism, the problem of numerical underflow that may occur in half-precision calculations is avoided, ensuring the stability and accuracy of the model.
[0105] Step 9. Model testing: Use a new dataset to verify and evaluate the trained model, and visualize the recognition results in 3D point clouds to generate an intuitive 3D visualization image; As Figure 10 shown. This visualization method can display the appearance, position, and shape of apples in real time, providing precise navigation and decision-making support for agricultural picking robots. In addition, this method supports multi-view interactive viewing, enabling operators to evaluate the distribution of apples from different angles, thereby optimizing the picking path and improving the picking efficiency and fruit integrity.
Claims
1. A multi-modal feature fusion apple recognition method based on YOLOv5s, characterized in that: It includes the following steps: Step 1. Image acquisition: Obtain color images, depth images, and point cloud images through a depth camera; Step 2. Decomposition and reconstruction of color images: Decompose the obtained color image into three color components of R, G, and B; Reconstruct the three channels using the R - G color difference operator to obtain an R - G color difference map; Step 3. Edge feature imaging: Grayscale the color image; Perform edge detection using the Laplacian operator to obtain an edge map; Step 4. Obtaining the key fruit feature map: Fuse the color difference map and the edge map according to a weight ratio of 6:4 to obtain the key fruit feature map; Step 5. Obtaining point cloud coordinate information: Obtain the point cloud image in Step 1 and find the corresponding color image; Project the point cloud image onto the color image: First, convert the point cloud coordinates to camera coordinates and project the points from the camera coordinate system to the image plane; Convert the processed camera coordinates to pixel coordinates using the camera internal parameter matrix and normalize them; Mirror - flip the image, correct the projection offset to obtain the color image pixel coordinates corresponding to the point cloud coordinate information XYZ, and compare with the original pixel coordinates at the same position to determine whether the point cloud coordinate information XYZ matches; Step 6. Multimodal feature fusion: Synthesize an 8 - channel image from R, G, B, the key fruit feature map, the depth image, and the point cloud coordinate information XYZ; Step 7. Model optimization: Modify the number of channels of the input layer of YOLOv5s, change the number of channels of the input layer from 3 to 8 to adapt to the fused multimodal features; Step 8. Model training: Use the fused multimodal features as input for deep learning training. During the training process, adopt data augmentation methods, including random cropping, color jittering, and Gaussian noise addition, to improve the generalization ability of the model; Use the mixed - precision training technique to accelerate the training process and reduce memory occupancy; Step 9. Model testing: Use a new dataset to verify and evaluate the trained model, and visualize the recognition results in 3D point cloud; 2. The multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, wherein: The calculation formula of the R - G color difference operator described in Step 2 is: R - G(x,y) = R(x,y) - G(x,y) where R - G(x,y) represents the R - G color difference map, and R(x,y) and G(x,y) respectively represent the red component and the green component of the image at the pixel position (x,y).
3. A multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, characterized in that: The calculation formula of the Laplacian operator described in Step 3 is: Among them, Δf(x,y) represents the value of the Laplacian operator at the image position (x,y), that is, the sum of the second-order derivatives of the image gray value at this position. represents the second-order partial derivative of the image gray value f(x,y) with respect to the x-axis, reflecting the curvature change of the image in the x-axis direction. represents the second-order partial derivative of the image gray value f(x,y) with respect to the y-axis, reflecting the curvature change of the image in the y-axis direction. f(x,y) represents the gray value of the grayscale image at the pixel position (x,y).
4. A multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, characterized in that: The specific fusion formula for fusing the color difference map and the edge map according to a weight ratio of 6:4 described in Step 4 is: F(x,y) = 0.6×(R - G(x,y)) + 0.4×L(x,y) where F(x,y) represents the value of the key fruit feature map at the pixel position (x,y), and L(x,y) represents the edge map.
5. The multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, characterized in that: The specific operation method for projecting the point cloud image onto the color image described in Step 5 is: First, convert the point cloud coordinates to camera coordinates: where E is the external parameter matrix, which is initialized as the identity matrix here; P cam is a 3D point in the camera coordinate system, and [x y z 1] is the homogeneous coordinate of the 3D point in the point cloud; Next, project the point from the camera coordinate system to the image plane, where the last component of the homogeneous coordinates needs to be removed, and only the first three components of P cam are used for projection, and it is represented by P cam,3 as follows: where x cam , y cam , z cam are the coordinates of the point in the camera coordinate system.
6. A multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, characterized in that: The specific operation method for converting the processed camera coordinates to pixel coordinates using the camera internal parameter matrix described in Step 5 is: P pixel = K·P cam,3 where K is the intrinsic matrix, f x and f y are the focal lengths of the camera along the x-axis and y-axis respectively; c x and c y are the coordinates of the optical center of the camera; P cam,3 are the first three components of the 3D point in the camera coordinate system; P pixel are the pixel coordinates of the point on the image plane. Then normalize P pixel to obtain the actual pixel coordinates: Among them, u and v are pixel coordinates on the image plane after being transformed by the intrinsic matrix and normalized, corresponding to the abscissa and ordinate of the image respectively.
7. A multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, characterized in that: The mirror flipping of the image described in step 5 is specifically operated as follows: u′ = m w -1 - u where u' is the abscissa of the image after mirror flipping; m w is the width of the image (number of pixels); Finally, correct the projection offset: where x offset and y offset are correction parameters for correcting projection offsets, representing the offsets in the x-axis and y-axis directions respectively; u o and v o are pixel coordinates on the corrected image plane, corresponding to the accurate abscissa and accurate ordinate of the image respectively.
8. A multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, characterized in that: The specific operation method of step 7 is as follows: The convolutional kernel dimension corresponding to the input layer is represented as W×H×m×n, where W×H represents the length and width of the convolutional kernel; for the YOLOv5s network, W×H is 6×6; m represents the number of convolutional kernel layers, which should be consistent with the number of channels of the input image, and m is adjusted to 8 here; n represents the number of convolutional kernels, that is, n different convolutional kernels are used to extract features from the image. For the YOLOv5s network, n is 32 and remains unchanged here.
9. A multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, characterized in that: The data augmentation method described in step 8 also includes random flipping and rotation to further improve the robustness of the model; The use of the mixed-precision training technique to accelerate the training process and reduce memory occupancy is achieved by dynamically adjusting the numerical precision in model training, realizing the acceleration of the training process and a significant reduction in memory occupancy; the specific operation method is: automatically switch to use single-precision floating-point numbers (FP32) and half-precision floating-point numbers (FP16) during the training phase, while maintaining the convergence performance of the model, making full use of the computing power of modern GPUs; at the same time, introduce a loss scaling mechanism to avoid numerical underflow problems that may occur in half-precision calculations.
10. A multi-modal feature fusion apple recognition method based on YOLOv5s according to claim 1, characterized in that: During the model testing process described in step 9, the reverse mapping technology is used to fuse the recognition results with the three-dimensional point cloud data to generate an intuitive three-dimensional visualization image. This visualization method can display the appearance, position and shape of apples in real time, providing accurate navigation and decision-making support for agricultural picking robots. In addition, this method supports multi-view interactive viewing, enabling operators to evaluate the distribution of apples from different angles, thereby optimizing the picking path and improving the picking efficiency and fruit integrity.