Cotton bale detection method based on multi-modal data timestamp synchronization and GPU / CPU collaborative acceleration

Through the cotton bag detection method of multimodal data timestamp synchronization and GPU/CPU collaborative acceleration, the problem of data synchronization and real-time processing in a multi-sensor fusion environment is solved, and efficient and real-time cotton bag detection is achieved.

CN120047911AActive Publication Date: 2025-05-27TIANJIN UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510103087.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The cotton bale detection system is difficult to achieve efficient data synchronization and real-time processing in a multi-sensor fusion environment, resulting in limited detection accuracy and real-time performance.

Method used

The multimodal data timestamp synchronization technology is adopted, and the precise synchronization of different sensor data flows is achieved through the message_filters module of ROS, combined with GPU/CPU collaborative acceleration, and the cotton bag detection model is optimized using TensorRT, and the Nodelet framework and C++ multithreading technology are used for parallel processing.

Benefits of technology

It realizes efficient synchronization of multi-sensor data, significantly reduces detection time, improves the system's response speed and detection accuracy, and ensures efficient and real-time performance of cotton bag detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047911A_ABST
    Figure CN120047911A_ABST
Patent Text Reader

Abstract

The invention discloses a cotton bale detection method based on multi-modal data timestamp synchronization and GPU / CPU cooperative acceleration, and the method comprises the steps: collecting multi-modal data, marking a timestamp, and carrying out the GPU acceleration distortion correction of a cotton bale image; timestamp synchronization and integration are carried out on the multi-modal data; the main program adopts Nodelet in combination with a C + + multi-thread technology to carry out parallel processing on receiving of multi-modal data and issuing of a multi-modal sensor coordinate system, and the main program runs to optimize the multi-modal data transmission efficiency; and constructing a cotton bale detection model, inputting the optimized multi-modal data for training, optimizing the trained cotton bale detection model by using TensorRT, inputting a cotton bale image for reasoning, obtaining cotton bale target identification data, and transmitting the cotton bale target identification data to a CPU for post-processing operation. Point cloud traversal in an OpenMP acceleration main program is used to carry out cotton bale target identification data and laser radar point cloud data matching; and performing point cloud feature and plane fitting calculation on the matched multi-modal data to obtain a rectangular frame, pose information and a score of the cotton bale target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking algorithms, and particularly to a bale detection method based on multi-modal data timestamp synchronization and GPU / CPU collaborative acceleration. Background Art

[0002] In recent years, the application of modern full-process mechanization and intelligent technologies in cotton planting and harvesting has been increasing year by year. Automatic operations have been realized from sowing, field management to the harvesting stage. Such efficient management not only reduces the dependence on labor, but also effectively improves the cotton production efficiency and the stability of quality. In addition, with the technological progress of domestic agricultural machinery equipment, the localization of agricultural machinery equipment has been gradually realized, effectively reducing the equipment cost and maintenance difficulty, and enhancing the sustainability of cotton production. Under this vast cotton market, with the increasing demand for agricultural automation, the application of intelligent perception technology in agricultural scenarios has been continuously deepened. In the process of automatic cotton picking and transportation, the high-precision detection and multi-target tracking of cotton bales face various challenges. These problems not only affect the accuracy of cotton bale detection technology, but also pose severe requirements on the real-time performance and stability of the entire system.

[0003] In recent years, the fusion of vision and radar sensors has become a hot research direction in autonomous driving technology. By combining visual information with radar information, the deficiencies of each sensor can be compensated. Radar provides relatively accurate distance and speed information, while vision provides richer object category and texture information. Such fusion solutions can improve the accuracy of target recognition and tracking in complex environments with poor lighting.

[0004] For example, DeepFusion [Li Y, Yu A W, Meng T, et al. Deepfusion: Lidar-camera deepfusion for multi-modal 3d object detection [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022: 17182-17191.] and CLOCs [Dhall A, Chelani K, Radhakrishnan V, et al. LiDAR-camera calibration using 3D-3D point correspondences [J]. arXiv preprint arXiv:1705.09785, 2017] proposed methods for visual-lidar deep fusion. By jointly processing visual images and lidar point cloud data, the robustness of object detection can be enhanced. In the field of autonomous driving, deep learning models that combine visual and lidar data can provide more comprehensive object detection and tracking capabilities. By using the joint feature maps of vision and lidar, objects can be located more precisely, improving the accuracy of multi-object detection and tracking.

[0005] Although the fusion of vision and radar can significantly improve detection accuracy, it still faces many challenges. For example, how to efficiently synchronize the timestamp data of different sensors, how to eliminate redundancy and noise in multi-sensor data fusion, and how to achieve efficient information sharing and processing between different data sources. In recent years, sensor fusion methods combined with deep learning, such as BEVfusion [Liu Z, Tang H, Amini A, et al. Bevfusion: Multi-task multi-sensor fusion with unified bird's-eye view representation [C] / / 2023 IEEE international conference on robotics and automation (ICRA). IEEE, 2023: 2774-2781] and Fast-BEV [Li Y, Huang B, Chen Z, et al. Fast-BEV: A Fast and Strong Bird's-Eye View Perception Baseline [J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024.], use multi-sensor fusion and bird's-eye view technology to perform more comprehensive perception and prediction of targets, further improving the accuracy and real-time performance of the fusion system.

[0006] With the continuous development of autonomous driving technology, vision-radar fusion solutions play an increasingly important role in object detection and multi-object tracking. Vision-radar fusion solutions have shown great potential in improving multi-object tracking, reducing occlusion and separation problems. However, there are also accompanying problems: Optimization of multi-modal data fusion: Further optimize the fusion method of vision and radar data to improve the accuracy and real-time performance of multi-sensor systems. Robustness improvement: In complex environments, how to improve the robustness of the system to factors such as occlusion, reflection, and low light is an important research topic.

[0007] Therefore, in the detection task of high-precision identification of cotton bales in a cotton gin, the following difficult problems are accompanied: The cotton bale detection system not only requires high precision but also sufficient real-time performance. Cotton bale detection involves multiple computational steps such as image acquisition, point cloud processing, feature extraction, object detection and tracking, and each link has different computational delays, which makes it difficult to ensure the real-time performance of the system. Especially in the case of multi-sensor fusion, the data frequencies from vision sensors and point cloud sensors are different. How to achieve efficient synchronization and fusion of data to ensure that no key information is lost in real-time decision-making is a huge challenge.

[0008] In addition, due to the complexity of the computational tasks at each stage and the huge amount of computation, how to optimize the algorithm, reduce the consumption of computing resources, and ensure that the system can run in real time under limited hardware conditions has become a key issue that must be solved. How to achieve efficient real-time operation of the system through hardware acceleration and algorithm optimization, and further improve the real-time performance of detection and tracking, is a core issue that needs to be solved urgently. Summary of the invention

[0009] The purpose of the present invention is to provide a cotton bale detection method based on multimodal data timestamp synchronization and GPU / CPU collaborative acceleration in order to address the technical defects existing in the prior art.

[0010] The technical solution adopted to achieve the purpose of the present invention is:

[0011] A cotton bale detection method based on multimodal data timestamp synchronization and GPU / CPU collaborative acceleration includes the following steps:

[0012] Step 1, using a multimodal sensor to collect multimodal data, the multimodal data including cotton bale images in the cotton ginning mill, depth information of the cotton bales and their surroundings, bale clamping vehicle positioning and posture data, and fork arm pitch data, and timestamping the collected multimodal data;

[0013] Step 2: Based on the main thread, a sub-thread is opened to call the message_filters module of ROS to synchronize the timestamps of the multimodal data marked by the timestamps in step 1, and integrate the multimodal data after the timestamp synchronization;

[0014] Step 3: The main program uses the Nodelet framework of ROS and combines C++ multi-threading technology to parallelly process the integrated multi-modal data reception and multi-modal sensor coordinate system release. The main program runs to optimize the integrated multi-modal data transmission efficiency.

[0015] Step 4: Build a cotton bale detection model, input the multimodal data optimized in step 3 into the cotton bale detection model for training, use TensorRT to optimize the trained cotton bale detection model, input the cotton bale image into the optimized cotton bale detection model for reasoning, obtain cotton bale target recognition data, and store it in the output buffer of the GPU;

[0016] Step 5, transferring the cotton bale target recognition data stored in the GPU output buffer to the CPU for post-processing, and using the C++ OpenMP framework to accelerate the point cloud traversal in the main program to match the cotton bale target recognition data with the lidar point cloud data;

[0017] Step 6: Calculate the point cloud features and plane fitting of the multimodal data after matching in Step 5 to obtain the rectangular box, pose information, and score of the bale target.

[0018] In the above technical solution, the multimodal sensor includes multiple cameras, multiple lidars, an integrated inertial navigation system, and a fork arm pitch height sensor.

[0019] In the above technical solution, the camera is a fish-eye camera. Calibrate the fish-eye camera to determine the internal parameters (camera matrix) and distortion coefficients of the camera. The internal parameters are the bale images collected, and the distortion coefficients are used to correct the collected bale images to eliminate distortion.

[0020] In the above technical solution, Step 2 includes: Create a child thread in the main thread to call the message_filters module to achieve precise synchronization of the bale images, depth information of the bale and its surrounding environment, the positioning and attitude data of the forklift truck, and the pitch data of the fork arm collected by the camera, lidar, integrated inertial navigation system, and fork arm pitch height sensor, so that the error between the bale images, depth information of the bale and its surrounding environment, the positioning and attitude data of the forklift truck, and the pitch data of the fork arm is within 0.01s - 0.04s.

[0021] In the above technical solution, using TensorRT to optimize the bale detection model after training includes: Optimize the inference process of the bale target recognition data through technologies such as mixed-precision calculation, layer fusion, and memory reuse.

[0022] In the above technical solution, the post-processing operation includes screening, sorting, and NMS of the bale target recognition boxes for the bale target recognition data.

[0023] In the above technical solution, Step 6 includes the following steps:

[0024] S6.1: Calculate the centroid of the point cloud of the multimodal data after matching in Step 5 to obtain the position information of the bale in the global coordinate system;

[0025] S6.2: Calculate the width and height of the point cloud of the multimodal data to obtain the width and height of the bale;

[0026] S6.3: Randomly select three non-coplanar points from the point cloud of the multimodal data for plane fitting to obtain the corresponding plane normal vector, normalize the plane normal vector, and calculate the dot product and cross product of the normalized plane normal vector and the reference axis;

[0027] S6.4: Construct a quaternion and convert the quaternion to Euler angles to extract the pose information corresponding to the bale.

[0028] In the above technical solution, the plane fitting formula is as follows:

[0029]

[0030]

[0031] In the formula, P 1 (x 1 , y 1 , z 1 ), P 2 (x 2 , y 2 , z 2 ), P 3 (x 3 , y 3 , z 3 ); P 1 , P 2 , P 3 respectively represent three non-coplanar points randomly selected from the point cloud of the multimodal data, respectively represent the corresponding direction vectors, represents the corresponding plane normal vector.

[0032] In the above technical solution, the calculation formula for the dot product of the normalized plane normal vector and the reference axis is as follows:

[0033]

[0034] In the formula, represents the unit direction vector of the reference axis, dot product represents the dot product of the normalized plane normal vector and the reference axis, represents the corresponding plane normal vector;

[0035] The calculation formula for the cross product of the normalized plane normal vector and the reference axis is as follows:

[0036]

[0037] In the formula, represents the cross product of the normalized plane normal vector and the reference axis, n x represents the value of the normal vector on the x-axis, n y represents the value of the normal vector on the y-axis, n z represents the value of the normal vector on the z-axis, represents the unit direction vector of the reference axis, represents the corresponding plane normal vector.

[0038] In the above technical solution, the formula for extracting the attitude information corresponding to the bale of cotton is as follows:

[0039]

[0040] q = (cos(θ / 2), r x sin(θ / 2), r y sin(θ / 2), r z sin(θ / 2))

[0041] Roll = atan2(2(wx + yz), 1 - 2(x 2 + y 2 ))

[0042] Pitch = asin(2(wy - zx))

[0043] Yaw = atan2(2(wz + xy), 1 - 2(y 2 + z 2 ))

[0044] In the formula, represents the modulus of the cross product, θ represents the angle of the plane normal vector with respect to the x-axis of the unit direction vector. Among them, the function can use both sine and cosine values to find the correct quadrant, represents the unit vector obtained after normalizing the cross product vector, which is used to determine the rotation axis of the quaternion. q represents the quaternion, where q = (w, x, y, z), r x represents the value of the x-axis of the unit direction vector, r y represents the value of the y-axis of the unit direction vector, r z represents the value of the z-axis of the unit direction vector. w represents the rotation angle or scale factor, x represents the x-axis direction and the rotation angle, y represents the y-axis direction and the rotation angle, z represents the z-axis direction and the rotation angle. Roll, Pitch, and Yaw represent the Euler angles calculated according to the quaternion, which are the roll angle, pitch angle, and yaw angle respectively.

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0046] 1. The present invention realizes the precise synchronization of different sensor data streams through the message_filters library, controls the error between 0.01 s and 0.04 s, and effectively avoids the problem of data inconsistency caused by time delay.

[0047] 2. The present invention uses the Nodelet framework to run the camera node and the bale detection algorithm node within the same process. Data transmission does not require TCP communication and is directly carried out within the same thread through C++ shared pointers, significantly reducing ROS communication latency and improving data processing efficiency. The number of threads is customized through the OpenMP technology of C++ to parallelize the data processing tasks, further reducing the calculation time and improving the system response speed.

[0048] 3. The present invention uses TensorRT to optimize the bale detection model, which can increase the detection speed of a single bale image by 4 to 5 times, significantly reducing the detection time and improving the system response speed. That is, in practical applications, the bale detection model can process more images in a shorter time, improving the system throughput and real-time performance. Moreover, the accelerated bale detection model has almost no loss in detection accuracy compared with the original detection model, ensuring high accuracy and high reliability of bale detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 The following shows the workflow of multi-modal data timestamp synchronization and GPU / CPU collaborative acceleration according to the present invention.

[0050] Figure 2 The following shows the flowchart of multi-modal data processing according to the present invention.

[0051] Figure 3 The following shows the comparison diagram between the Nodelet node and the Node node according to the present invention.

[0052] Figure 4 The following shows the block diagram of C++ multi-thread technology according to the present invention.

[0053] Figure 5 The following shows the flowchart of the bale detection model according to the present invention.

[0054] Figure 6 The following shows the flowchart of TensorRT model inference acceleration according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following further elaborates on the present invention in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0056] A bale detection method based on multi-modal data timestamp synchronization and GPU / CPU collaborative acceleration, see Figure 1 , including the following steps:

[0057] Step 1: Use a multi-modal sensor to collect multi-modal data, where the multi-modal data includes cotton bale images in a ginning mill, depth information of the cotton bale and its surrounding environment, positioning and attitude data of a clamp truck, and fork arm pitch data. Timestamp the collected multi-modal data.

[0058] The multi-modal sensor in this embodiment includes multiple cameras, multiple lidars (LiDARs), an integrated inertial navigation system (INS), and a fork arm pitch height sensor. The cameras are fisheye cameras. Calibrate the fisheye cameras to determine the internal parameters (camera matrix) and distortion coefficients of the cameras. The internal parameters are the cotton bale images collected, and the distortion coefficients are used to correct the collected cotton bale images to eliminate distortion. Based on the fisheye cameras in this embodiment, GPU-accelerated distortion correction can be performed on the cotton bale images in the multi-modal data in real time. Among them, use the cameras to capture cotton bale images with an update frequency of 15 Hz. Use the lidars to collect depth information of the cotton bale and its surrounding environment at a frequency of 15 Hz for constructing a three-dimensional environment model. Based on the integrated inertial navigation system to provide accurate vehicle positioning and its attitude data with a maximum update frequency of 10 Hz for vehicle navigation. Use the fork arm pitch height sensor to monitor the position and height data of the fork arm and provide accurate pitch data at a high frequency of 30 Hz. The cameras and lidars are fixedly installed on the fork arm to adjust the dynamic coordinates of the cameras and lidars.

[0059] Step 2: On the basis of the main thread (CPU), create a child thread to call the message_filters module of ROS (Robot Operating System) to synchronize the timestamps of the multi-modal data timestamped in Step 1, for ensuring the time consistency between each frame of cotton bale image and lidar data.

[0060] See Figure 2 , the synchronization of the timestamps of the multi-modal data timestamped in Step 1 includes: creating a child thread in the main thread to call the message_filters module to achieve precise synchronization of the cotton bale images, depth information of the cotton bale and its surrounding environment, positioning and attitude data of the clamp truck, and pitch data of the fork arm collected by the cameras, lidars (LiDARs), integrated inertial navigation system (INS), and fork arm pitch height sensor, so that the error between the cotton bale images, depth information of the cotton bale and its surrounding environment, positioning and attitude data of the clamp truck, and pitch data of the fork arm is within 0.01 s - 0.04 s, to avoid data inconsistency problems caused by time delay received by the program.

[0061] Step 3: The main program uses ROS's Nodelet framework combined with C++ multi-threading technology to parallelly process the reception of multi-modal data after timestamp synchronization and the publication of multi-modal sensor coordinate systems. The multi-modal data after timestamp synchronization is directly obtained through shared pointers. Among them, directly obtaining the multi-modal data after timestamp synchronization through shared pointers can optimize the transmission efficiency of multi-modal data.

[0062] See Figure 3 、 Figure 4 The Nodelet framework of this embodiment is used to run the camera node and the bale detection model node within the same process, which can achieve data transmission without TCP communication. And direct data copying is performed within the same thread using C++ shared pointers (i.e., multiple pointers perform data copying within the same thread), which can reduce ROS communication latency and improve data processing efficiency.

[0063] Step 4: Build a bale detection model (Yolov8s neural network model), input the multi-modal data optimized in Step 3 into the bale detection model for training (the trained bale detection model can achieve efficient and accurate bale detection in various environments, and identify the position information and category of bales in the bale image). Use TensorRT (TensorRT is a deep learning inference acceleration library launched by NVIDIA) to optimize the trained bale detection model, which can improve the actual inference operation speed of the bale detection model on the GPU. See Figure 5 Input the bale image into the optimized bale detection model for inference to obtain bale target recognition data (the bale target recognition data includes bale target recognition boxes (rectangular boxes), confidence values, and category information), and store the bale target recognition data in the output buffer of the GPU.

[0064] See Figure 6 The optimization of the trained bale detection model using TensorRT includes: optimizing the inference process of bale target recognition data through techniques such as mixed-precision calculation, layer fusion, and memory reuse, reducing unnecessary calculations and memory accesses, thereby greatly accelerating the inference process.

[0065] Step 5: Transmit the bale target recognition data stored in the GPU output buffer to the CPU (central processing unit) for post-processing operations, and use C++'s OpenMP (Open Multi-Processing) framework to accelerate the point cloud traversal in the main program to match the bale target recognition data with the lidar point cloud data, which can improve the processing speed of multi-modal data.

[0066] The post - processing operation includes screening, sorting, and NMS (Non - Maximum Suppression) of the cotton bale target recognition data, which can ensure the accuracy and integrity of the detection results.

[0067] The OpenMP framework of C++ custom - allocates the required number of threads for the code block in the main program. By distributing the post - processing tasks of the cotton bale target recognition data to multiple sub - threads in this code block, it can improve the parallelism of data transmission and processing, thereby reducing the overall calculation time, that is, improving the data processing speed.

[0068] Step 6: Perform point - cloud feature calculation and plane fitting calculation on the multi - modal data after matching in Step 5 to obtain the rectangular frame, pose information, and score of the cotton bale target.

[0069] The said Step 6 includes the following steps:

[0070] S6.1: Calculate the centroid of the point cloud of the multi - modal data after matching in Step 5 to obtain the position information (x, y, z) of the cotton bale in the global coordinate system.

[0071] S6.2: Calculate the width and height of the point cloud of the multi - modal data to obtain the width and height of the cotton bale.

[0072] S6.3: Randomly select three non - coplanar points from the point cloud of the multi - modal data for plane fitting to obtain the corresponding plane normal vector, normalize the plane normal vector, and calculate the dot product and cross product of the normalized plane normal vector and the reference axis (the reference axis is the x - axis).

[0073] The plane fitting formula is as follows:

[0074]

[0075] In the formula, P 1 (x 1 , y 1 , z 1 ), P 2 (x 2 , y 2 , z 2 ), P 3 (x 3 , y 3 , z 3 ); P 1 , P 2 , P 3 respectively represent three non - coplanar points randomly selected from the point cloud of the multi - modal data, respectively represent the corresponding direction vectors, represents the corresponding plane normal vector.

[0076] The dot product calculation formula of the normalized plane normal vector and the reference axis is as follows:

[0077]

[0078] In the formula, represents the unit direction vector of the reference axis (x-axis), and dot product represents the dot product of the normalized plane normal vector and the reference axis, represents the corresponding plane normal vector.

[0079] The cross product calculation formula of the normalized plane normal vector and the reference axis is as follows:

[0080]

[0081] In the formula, represents the cross product of the normalized plane normal vector and the reference axis, n x represents the value of the normal vector on the x-axis, n y represents the value of the normal vector on the y-axis, n z represents the value of the normal vector on the z-axis, represents the unit direction vector of the reference axis (x-axis), represents the corresponding plane normal vector.

[0082] S6.4: Construct a quaternion, convert the quaternion into Euler angles, and extract the corresponding attitude information (pitch, roll, yaw) of the bale.

[0083] The formula for extracting the corresponding attitude information of the bale is as follows:

[0084]

[0085] q = (cos(θ / 2), r x sin(θ / 2), r y sin(θ / 2), r z sin(θ / 2))

[0086] Roll = atan2(2(wx + yz), 1 - 2(x 2 + y 2 ))

[0087] Pitch = asin(2(wy - zx))

[0088] Yaw = atan2(2(wz + xy), 1 - 2(y 2 + z 2 ))

[0089] In the formula, represents the modulus of the cross product, and θ represents the plane normal vector The angle with respect to the x-axis of the unit direction vector, where the function can use both the sine value and the cosine value to find the correct quadrant. Represents the unit vector obtained by normalizing the cross product vector, which is used to determine the rotation axis of the quaternion. q represents the quaternion, where q = (w, x, y, z), r x Represents the value of the unit direction vector on the x-axis, r y Represents the value of the unit direction vector on the y-axis, r z Represents the value of the unit direction vector on the z-axis. w represents the rotation angle or scale factor, similar to the scaling part in the rotation matrix. x represents the x-axis direction and the rotation angle, y represents the y-axis direction and the rotation angle, z represents the z-axis direction and the rotation angle. Roll, Pitch, and Yaw represent the Euler angles calculated from the quaternion, which are the roll angle, pitch angle, and yaw angle respectively.

[0090] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A cotton bale detection method based on multimodal data timestamp synchronization and GPU / CPU collaborative acceleration, characterized in that: The following steps are involved: Step 1, using a multimodal sensor to collect multimodal data, the multimodal data including cotton bale images in the cotton ginning mill, depth information of the cotton bales and their surroundings, bale clamping vehicle positioning and posture data, and fork arm pitch data, and timestamping the collected multimodal data; Step 2: Based on the main thread, a sub-thread is opened to call the message_filters module of ROS to synchronize the timestamps of the multimodal data marked by the timestamps in step 1, and integrate the multimodal data after the timestamp synchronization; Step 3: The main program uses the Nodelet framework of ROS and combines C++ multi-threading technology to parallelly process the integrated multi-modal data reception and multi-modal sensor coordinate system release. The main program runs to optimize the integrated multi-modal data transmission efficiency. Step 4: Build a cotton bale detection model, input the multimodal data optimized in step 3 into the cotton bale detection model for training, use TensorRT to optimize the trained cotton bale detection model, input the cotton bale image into the optimized cotton bale detection model for reasoning, obtain cotton bale target recognition data, and store it in the output buffer of the GPU; Step 5, transferring the cotton bale target recognition data stored in the GPU output buffer to the CPU for post-processing, and using the C++ OpenMP framework to accelerate the point cloud traversal in the main program to match the cotton bale target recognition data with the lidar point cloud data; Step 6: Perform point cloud feature calculation and plane fitting calculation on the multimodal data matched in step 5 to obtain the rectangular frame, position information and score of the cotton bale target.

2. The cotton bale detection method according to claim 1, characterized in that: The multimodal sensor includes multiple cameras, multiple laser radars, a combined inertial navigation system and a fork arm pitch height sensor.

3. The cotton bale detection method according to claim 2, characterized in that: The camera is a fisheye camera, which is calibrated using the fisheye camera to determine the camera's internal parameters (camera matrix) and distortion coefficient (distortion), wherein the internal parameters are the collected cotton bale images, and the distortion coefficient is used to correct the collected cotton bale images to eliminate distortion.

4. The cotton bale detection method according to claim 1, characterized in that: The step 2 includes: opening a subthread in the main thread to call the message_filters module to accurately synchronize the cotton bale image, the depth information of the cotton bale and its surrounding environment, the positioning and attitude data of the bale clamping vehicle, and the pitch data of the fork arm collected by the camera, lidar, combined inertial navigation system and fork arm pitch height sensor, so that the error between the cotton bale image, the depth information of the cotton bale and its surrounding environment, the positioning and attitude data of the bale clamping vehicle, and the pitch data of the fork arm is 0.01s-0.04s.

5. The cotton bale detection method according to claim 1, characterized in that: The use of TensorRT to optimize the trained cotton bale detection model includes: optimizing the reasoning process of cotton bale target recognition data through technologies such as mixed precision computing, layer fusion, and memory reuse.

6. The cotton bale detection method according to claim 1, characterized in that: The post-processing operation includes screening, sorting and NMS of cotton bale target recognition frames of cotton bale target recognition data.

7. The cotton bale detection method according to claim 1, characterized in that: The step 6 comprises the following steps: S6.1: Calculate the point cloud centroid of the multimodal data after matching in step 5 to obtain the position information of the cotton bale in the global coordinate system; S6.2: Calculate the width and height of the multimodal data point cloud to obtain the width and height of the cotton bale; S6.3: randomly selecting three non-coplanar points in the point cloud of the multimodal data to perform plane fitting, obtain corresponding plane normal vectors, normalize the plane normal vectors, and calculate the dot product and cross product of the normalized plane normal vectors with the reference axis; S6.4: Construct a quaternion, and convert the quaternion into Euler angles to extract the posture information corresponding to the cotton bale.

8. The cotton bale detection method according to claim 1, characterized in that: The plane fitting formula is as follows: Wherein, P1(x1, y1, z1), P2(x2, y2, z2), P3(x3, y3, z3); ​​P1, P2, P3 represent three non-coplanar points randomly selected from the point cloud of the multimodal data, respectively. Represent the corresponding direction vectors, Represents the corresponding plane normal vector.

9. The cotton bale detection method according to claim 1, characterized in that: The dot product calculation formula of the normalized plane normal vector and the reference axis is as follows: In the formula, Represents the unit direction vector of the reference axis, dot product represents the dot product of the normalized plane normal vector and the reference axis, Represents the corresponding plane normal vector; The cross product calculation formula of the normalized plane normal vector and the reference axis is as follows: In the formula, represents the cross product of the normalized plane normal vector and the reference axis, n x Represents the value of the normal vector x-axis, n y Represents the value of the normal vector on the y axis, n z Represents the value of the normal vector z-axis, represents the unit direction vector of the reference axis, Represents the corresponding plane normal vector.

10. The cotton bale detection method according to claim 1, characterized in that: The posture information extraction formula corresponding to the cotton bale is as follows: q=(cos(θ / 2),r x sin(θ / 2),r y sin(θ / 2),r z sin(θ / 2)) Roll=atan2(2(wx+yz),1-2(x 2 +y 2 )) Pitch = asin(2(wy-zx)) Yaw=atan2(2(wz+xy),1-2(y 2 +z 2 )) In the formula, represents the modulus of the cross product, and θ represents the plane normal vector The angle about the unit direction vector x-axis, where the function can use both sine and cosine values ​​to find the correct quadrant, Represents the unit vector obtained after the cross product vector is normalized, which is used to determine the rotation axis of the quaternion. q represents the quaternion, where q = (w, x, y, z), r x Represents the value of the unit direction vector x-axis, r y Represents the value of the unit direction vector y-axis, r z Represents the value of the unit direction vector on the z-axis, w represents the angle of rotation or scale factor, x represents the x-axis direction and rotation angle, y represents the y-axis direction and rotation angle, z represents the z-axis direction and rotation angle, Roll, Pitch, and Yaw represent the Euler angles calculated from the quaternion, which are roll, pitch, and yaw, respectively.

Citation Information

Patent Citations

  • A multi-sensor timestamp acquisition synchronizing device

    CN109729277A

  • Millimeter wave radar and vision fused three-dimensional target detection method based on attention mechanism

    CN114708585A

  • Multi-modal sensor hardware time synchronization method

    CN117118554A

  • Multi-modal 3D target detection method and device based on cloud edge collaboration

    CN117496322A

  • Visual SLAM optimization method in dynamic scene based on deep learning and GPU acceleration

    CN118097265A