Lightweight unmanned forklift AI visual anti-collision method based on domestic embedded platform
By combining lightweight deep learning and adaptive feature fusion technology on domestic embedded platforms, the real-time and accuracy problems of pedestrian detection and anti-collision functions on the embedded platform are solved, efficient multi-pegeot detection and distance estimation are achieved, and the safety and adaptability of smart forklifts are improved.
Patent Information
- Application Number
- CN202510301154.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is difficult to achieve efficient pedestrian detection and anti-collision functions on low-cost embedded platforms, especially in complex environments, detection accuracy and real-time performance are difficult to take into account, and the existing lightweight models are insufficient in multi-pegeot detection and distance estimation.
The lightweight deep learning algorithm and adaptive feature fusion method are used to deeply integrate pedestrian detection and distance estimation features through a dynamic adaptive sparse transformation network, and combined with model pruning and compression technology, it is deployed on a domestic embedded platform, and the software functions are optimized using the Hongmeng operating system.
Implementing high-precision multi-peer detection and distance estimation in complex scenarios improves the real-time and adaptability of the system, reduces power consumption, and significantly improves the safety and operation efficiency of smart forklifts.
Smart Images

Figure CN120236264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and AI vision anti-collision technology for unmanned forklifts, and particularly to a lightweight AI vision anti-collision method for unmanned forklifts based on a domestic embedded platform. Background Art
[0002] With the development of industrial automation technology, intelligent forklifts are increasingly widely used in fields such as logistics and warehousing. In these intelligent forklift systems, ensuring the safety of pedestrians and equipment is a key issue. Therefore, vision anti-collision technology based on artificial intelligence (AI) has been proposed to real-time monitor pedestrians or obstacles that may exist in the surrounding environment of the forklift to achieve early prevention of collision risks. Currently, anti-collision systems mainly adopt detection methods based on lidar, ultrasonic sensors, and computer vision. Among them, lidar and ultrasonic sensors identify obstacles by measuring distances, but their costs are relatively high, and they are affected by various factors in complex environments, such as reflectivity and material properties. In addition, these methods have certain limitations in pedestrian recognition and feature extraction.
[0003] With the progress of deep learning technology, object detection methods based on computer vision have been significantly developed. The basic process of object detection includes the generation of candidate regions, feature extraction, and object classification and bounding box regression. Currently, it is mainly divided into object detection methods based on traditional machine learning and deep learning. Traditional object detection methods, including selective search, sliding window, HOG (Histogram of Oriented Gradients), SIFT (Scale-Invariant Feature Transform), etc., rely on manually extracted features, and their robustness and adaptability are poor. Especially when facing complex working scenarios and external interferences such as light changes, the detection effect is difficult to guarantee. With the development of deep learning technology, object detection methods based on convolutional neural network (CNN) have been widely applied to pedestrian detection. Common deep learning models include Faster R-CNN, YOLO, and SSD, etc. These models extract features from the input image through a multi-layer convolutional network, have strong representation capabilities, and can also achieve good detection effects in complex environments. However, the computational complexity of these models is relatively high, and the required hardware resources are large. Especially on mobile and embedded devices, it is difficult to meet the real-time requirements, which become the main limitations of the existing technology.
[0004] To achieve efficient pedestrian detection and anti-collision functions on low-cost embedded platforms, researchers have started to explore lightweight deep learning models and algorithm optimization techniques. Lightweight networks such as MobileNet and ShuffleNet reduce the complexity of the model by means of reducing the number of parameters and pruning, etc., so as to deploy these models on embedded devices. However, while these methods reduce the computational complexity, they also lead to a decline in the detection accuracy of the model. Therefore, there is still a large room for improvement in the existing technology in terms of balancing detection accuracy and computational resources. In addition, existing lightweight models are difficult to ensure accuracy and stability when facing multi-pedestrian detection and distance estimation. Especially when dealing with complex working scenarios, the detection model is easily affected by factors such as occlusion, lighting changes, and cluttered backgrounds.
[0005] To achieve lightweight unmanned forklift AI vision anti-collision based on an embedded platform, the existing technology usually uses preprocessed RGB images as input to a deep learning model for detection and estimation. However, in practical applications, how to balance accuracy and real-time performance on domestic embedded platforms remains a difficult problem. Especially based on existing object detection and distance estimation models, there are the following main defects: First, the complexity of pedestrian detection is relatively high. Existing deep convolutional networks require a large amount of computational resources and memory, which is a huge challenge for embedded systems. Second, the existing technology lacks an efficient fusion and adaptive optimization mechanism for detection features, which results in poor adaptability of the model when facing occlusion and pose changes, and it is difficult for the detection accuracy to meet the requirements of industrial applications.
[0006] In existing embedded anti-collision systems, the vision technologies used generally can only complete single pedestrian detection or distance estimation, and it is difficult to simultaneously achieve comprehensive analysis of pedestrian postures, distances, and dynamic behaviors. In addition, most existing embedded solutions rely on externally input hardware modules rather than achieving autonomous operation on domestic platforms, resulting in high costs and strong dependencies and making it difficult to break through technical blockades. At the same time, due to the limited computing power of embedded platforms, there is an irreconcilable contradiction between the real-time performance and accuracy of existing detection algorithms. Especially in complex environments, the accuracy of target detection is affected by various external factors such as light conditions and occlusion phenomena, resulting in unsatisfactory application effects of existing systems in complex environments.
[0007] Therefore, how to provide a lightweight unmanned forklift AI vision anti-collision method based on a domestic embedded platform is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0008] An object of the present invention is to propose a lightweight anti-collision method for unmanned forklifts based on a domestic embedded platform. The present invention adopts a lightweight deep learning algorithm and an adaptive feature fusion method, and successfully realizes the deployment of a multi-pedestrian detection and anti-collision system on a domestic embedded platform. Through the dynamic adaptive sparse transformation network, the pedestrian detection and distance estimation features are deeply fused, and combined with efficient model pruning and compression technologies, to ensure that the system has high precision and high real-time performance in complex scenarios. At the same time, by using the HarmonyOS to trim software function modules, the system has low power consumption, high reliability and good adaptability, significantly improving the safety and operation efficiency of intelligent forklifts.
[0009] The lightweight anti-collision method for unmanned forklifts based on a domestic embedded platform according to an embodiment of the present invention includes the following steps:
[0010] S1. Preprocess the RGB image through a lightweight pedestrian detection module, based on an image recognition algorithm of a lightweight convolutional neural network, automatically prune and compress the deep convolutional network, and extract lightweight pedestrian detection features;
[0011] S2. Use a lightweight pedestrian distance estimation module, through camera calibration, image correction, stereo matching and distance acquisition, and adopt model compression to extract lightweight pedestrian distance estimation features;
[0012] S3. Fuse the lightweight pedestrian detection features and the lightweight pedestrian distance estimation features, perform feature analysis in combination with the dynamic adaptive sparse transformation network, and complete multi-pedestrian detection and pedestrian distance estimation by introducing a sparse feature adaptive reconstruction mechanism;
[0013] S4. Deploy on a domestic embedded platform, implement anti-collision detection and rapid response through hardware, and use the HarmonyOS to trim software function modules.
[0014] Optionally, the image recognition algorithm of the lightweight convolutional neural network in S1 specifically includes: optimizing the binary convolutional neural network based on symbol distribution adjustment, performing binary processing on the input RGB image; automatically pruning the deep convolutional network by using similarity clustering and swarm intelligence optimization; fusing the binary processing and the pruning result; a deep convolutional network compression method based on the fusion of binary and pruning.
[0015] Optionally, the camera calibration in S2 includes the external parameters and internal parameters of the camera. The Zhang Zhengyou calibration method is used to calibrate the external parameters and internal parameters of the camera. The Zhang Zhengyou calibration method is a camera calibration based on a 2D planar target, and a calibration board composed of two-dimensional grids is used for calibration.
[0016] Optionally, the image correction in S2 processes the images collected by the left and right cameras in the binocular vision system according to the rules of epipolar geometry constraints. Assuming that when the target point corresponding to the point to be matched in the left camera image exists in the right camera image, the target point must be located on the epipolar line corresponding to the point to be matched;
[0017] The Bouget algorithm in the OpenCV vision library is used to correct the collected data to eliminate the error caused by the position offset during the three-dimensional imaging process.
[0018] Optionally, the stereo matching in S2 performs pedestrian ranging based on the deep learning pyramid stereo matching network, specifically including:
[0019] Use the convolutional layer for feature extraction;
[0020] Use spatial pyramid pooling to unify the tensors of the pictures to form a matching cost volume;
[0021] Adjust the matching cost volume through 3D convolution. The 3D convolution uses a stacked hourglass structure, and an output of a disparity map is formed behind each hourglass;
[0022] Calculate the loss functions respectively, and perform weighted summation on the loss functions to obtain the total loss function.
[0023] Optionally, the compression method in S2 for model compression is 3D convolution channel compression. The feature extraction module of the original model is retained, and the cost aggregation module using 3D convolution is compressed.
[0024] Optionally, S3 specifically includes:
[0025] S31. Fuse the lightweight pedestrian detection feature matrix and the lightweight pedestrian distance estimation feature matrix to generate a comprehensive feature matrix F;
[0026] S32. Construct an adaptive sparse transformation model, use the dynamic adaptive sparse transformation network to perform feature analysis on the comprehensive feature matrix F, perform multi-level adaptive reconstruction, and generate a sparse feature matrix F ′ ;
[0027] S33. Combine the multi-layer optimization of matrix tensor transformation and high-order feature coupling, and define an adaptive sparse reconstruction objective function L:
[0028]
[0029] Among them, W represents the feature weight matrix, b represents the bias vector, α and β represent the adaptive weight coefficients, tanh(·) represents the hyperbolic tangent function, represents the Kronecker product, denotes the Frobenius norm, λ and η denote the sparse regularization and low-rank regularization coefficients respectively, and ∑ i,j |W ij | p denotes the sparse regularization term, p ∈ [0, 1] denotes the sparsity parameter, and SVD(W) denotes the singular value decomposition of the weight matrix W;
[0030] S34. Adaptively reconstruct the sparse feature matrix F ′ to generate the reconstructed feature matrix F rec ;
[0031] S35. Use the reconstructed feature matrix F rec to perform multi-person detection, and use a deep convolutional neural network to classify and label the features of different pedestrians;
[0032] S36. Based on the detected pedestrian features, estimate the relative distance of each pedestrian to generate a distance estimate value for the pedestrian, and combine with anti-collision rules to determine whether to generate an alarm signal.
[0033] Optionally, the hardware in S4 specifically includes a housing, a liquid crystal touch screen, a video camera, a voice intelligent switching button, an external speaker, and a data interface;
[0034] The development board is ATK-DLRV1126, which is equipped with a domestic Rockchip RV1126 processor; the RV1126 processor integrates a CPU and an NPU. The CPU is a quad-core ARM Cortex-A7 processor with a main frequency of up to 1.5 GHz, and an NPU with 2 TOPS is built-in for accelerating deep learning inference tasks; the RV1126 processor provides communication interfaces including USB, I2C, and SPI; the development board ATK-DLRV1126 communicates with other functional modules through UART.
[0035] Optionally, the software in S4 is divided into four layers, including a kernel layer, a system service layer, an application framework layer, and an application layer. The application layer includes a system basic capability subsystem set, a basic software service subsystem set, an enhanced software service subsystem set, and a hardware service subsystem set.
[0036] The beneficial effects of the present invention are:
[0037] First, through the pruning and compression techniques of lightweight convolutional neural networks, the present invention successfully reduces the computational complexity of the model, enabling deep learning algorithms to run in real time on embedded devices, significantly enhancing the real-time performance and resource utilization efficiency of the system. By combining lightweight pedestrian detection and distance estimation modules, the present invention can achieve high-precision detection and distance estimation of multiple pedestrian targets in complex working scenarios, solving the problems of excessive computational burden and the irreconcilable contradiction between detection accuracy and real-time performance in traditional methods.
[0038] In addition, the present invention introduces a Dynamic Adaptive Sparse Transform Network (DAST-Net) to deeply fuse detection features and distance features. Through an adaptive sparse feature reconstruction mechanism, it enhances the representation ability of pedestrian features in complex scenarios, enabling the system to still have good detection effects under complex conditions such as pedestrian occlusion, pose changes, and cluttered backgrounds. Sparse feature reconstruction effectively reduces redundant features in the model while improving the expressive ability of features, making the results of detection and distance estimation more accurate and reliable. Especially in the case of multi-pedestrian detection, through deep coupling and optimization of the fused comprehensive feature matrix, the system can effectively handle feature interactions in multi-target situations, ensuring accurate monitoring of pedestrians in a busy working environment.
[0039] In terms of hardware implementation, the present invention is deployed on a domestic embedded platform. Utilizing the powerful computing power of the RV1126 processor and through deep learning inference acceleration, the entire system not only significantly improves detection accuracy but also features low power consumption and high efficiency. Based on the trimming and optimization of software functional modules on the HarmonyOS, the system can better adapt to the actual application requirements of unmanned forklifts, ensuring the efficient operation of the entire system on the embedded platform. At the same time, through the integration of multiple modules such as a liquid crystal touch screen, a camera, and voice intelligent switching buttons in the hardware design, more intuitive user interaction and diverse anti-collision strategies are realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0041] Figure 1 is a flowchart of a lightweight unmanned forklift AI vision anti-collision method based on a domestic embedded platform proposed by the present invention;
[0042] Figure 2 is a flowchart of a lightweight pedestrian detection algorithm based on binaryzation and clustering pruning for the lightweight unmanned forklift AI vision anti-collision method proposed by the present invention;
[0043] Figure 3 Schematic diagram of a stereo matching algorithm based on channel compression of a stereo matching model for the lightweight unmanned forklift AI vision anti-collision method proposed by the present invention based on a domestic embedded platform;
[0044] Figure 4 Schematic diagram of the system hardware structure of the lightweight unmanned forklift AI vision anti-collision method proposed by the present invention based on a domestic embedded platform. Detailed implementation manners
[0045] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.
[0046] Refer to Figures 1-4 , the lightweight unmanned forklift AI vision anti-collision method based on a domestic embedded platform includes the following steps:
[0047] S1. Preprocess the RGB image through a lightweight pedestrian detection module, perform automatic pruning and compression on the deep convolutional network based on the image recognition algorithm of the lightweight convolutional neural network, and extract lightweight pedestrian detection features;
[0048] S2. Use a lightweight pedestrian distance estimation module, through camera calibration, image correction, stereo matching and distance acquisition, and adopt model compression to extract lightweight pedestrian distance estimation features;
[0049] S3. Fuse the lightweight pedestrian detection features and the lightweight pedestrian distance estimation features, perform feature analysis in combination with a dynamic adaptive sparse transformation network, and complete multi-pedestrian detection and pedestrian distance estimation by introducing a sparse feature adaptive reconstruction mechanism;
[0050] S4. Deploy on a domestic embedded platform, implement anti-collision detection and rapid response through hardware, and use the HarmonyOS to perform software function module trimming.
[0051] In this embodiment, the image recognition algorithm of the lightweight convolutional neural network in S1 specifically includes: optimizing the binary convolutional neural network based on symbol distribution adjustment, performing binarization processing on the input RGB image; automatically pruning the deep convolutional network by using similarity clustering and swarm intelligence optimization; fusing the binarization processing and the pruning result; and a deep convolutional network compression method based on the fusion of binarization and pruning.
[0052] In this embodiment, the camera calibration in S2 includes the external parameters and internal parameters of the camera. The Zhang Zhengyou calibration method is used to calibrate the external parameters and internal parameters of the camera. The Zhang Zhengyou calibration method is a camera calibration based on a 2D planar target, and a calibration board composed of two-dimensional grids is used for calibration.
[0053] In this embodiment, the image correction in S2 processes the images collected by the left and right cameras in the binocular vision system according to the rules of epipolar geometry constraints. Assuming that when the target point corresponding to the point to be matched in the left camera image exists in the right camera image, the target point must be located on the epipolar line corresponding to the point to be matched;
[0054] The Bouget algorithm in the OpenCV vision library is used to correct the collected data to eliminate the error caused by the position offset during the three-dimensional imaging process.
[0055] In this embodiment, the stereo matching in S2 is based on the deep learning-based Pyramid Stereo Matching Network for pedestrian ranging, which specifically includes:
[0056] Use the convolutional layer for feature extraction;
[0057] Use spatial pyramid pooling to unify the tensors of the pictures to form a matching cost volume;
[0058] Adjust the matching cost volume through 3D convolution. The 3D convolution uses a stacked hourglass structure, and an output of a disparity map is formed behind each hourglass;
[0059] Calculate the loss functions respectively, and perform weighted summation on the loss functions to obtain the total loss function.
[0060] In this embodiment, the compression method of model compression in S2 is 3D convolution channel compression. The feature extraction module of the original model is retained, and the cost aggregation module using 3D convolution is compressed.
[0061] In this embodiment, S3 specifically includes:
[0062] S31. Fuse the lightweight pedestrian detection feature matrix and the lightweight pedestrian distance estimation feature matrix to generate a comprehensive feature matrix F;
[0063] S32. Construct an adaptive sparse transformation model, use the dynamic adaptive sparse transformation network to perform feature analysis on the comprehensive feature matrix F, perform multi-level adaptive reconstruction, and generate a sparse feature matrix F ′ ;
[0064] S33. Combine the multi-layer optimization of matrix tensor transformation and high-order feature coupling, and define an adaptive sparse reconstruction objective function L:
[0065]
[0066] Among them, W represents the feature weight matrix, b represents the bias vector, α and β represent the adaptive weight coefficients, and tanh(·) represents the hyperbolic tangent function. denotes the Kronecker product, denotes the Frobenius norm, λ and η respectively denote the sparse regularization and low-rank regularization coefficients, ∑ i,j |W ij | p denotes the sparse regularization term, p ∈ [0, 1] denotes the sparsity parameter, and SVD(W) denotes the singular value decomposition of the weight matrix W;
[0067] S34. Adaptive reconstruction is performed on the sparse feature matrix F ′ to generate the reconstructed feature matrix F rec ;
[0068] S35. Using the reconstructed feature matrix F rec to perform multi-person detection, and a deep convolutional neural network is used to classify and label the features of different pedestrians;
[0069] S36. Based on the detected pedestrian features, estimate the relative distance of each pedestrian to generate a pedestrian distance estimate value, and combine the anti-collision rules to determine whether to generate an alarm signal.
[0070] In this embodiment, the hardware in S4 specifically includes a housing, a liquid crystal touch screen, a video camera, a voice intelligent switching button, an external speaker, and a data interface;
[0071] The development board is ATK-DLRV1126, which is equipped with a domestic Rockchip RV1126 processor; the RV1126 processor integrates a CPU and an NPU. The CPU is a quad-core ARM Cortex-A7 processor with a main frequency of up to 1.5 GHz, and an NPU with 2 TOPS is built-in for accelerating deep learning inference tasks; the RV1126 processor provides communication interfaces including USB, I2C, and SPI; the development board ATK-DLRV1126 communicates with other functional modules through UART.
[0072] In this embodiment, the software in S4 is divided into four layers, including a kernel layer, a system service layer, an application framework layer, and an application layer. The application layer includes a system basic capability subsystem set, a basic software service subsystem set, an enhanced software service subsystem set, and a hardware service subsystem set.
[0073] Example 1:
[0074] To verify the feasibility of the present invention in implementation, the present invention is applied to the actual working scenario of a large-scale warehousing and logistics park. The entire experiment lasted for 8 weeks, and multiple verifications were carried out on the operation of the driverless forklift to evaluate the actual effect of the present invention.
[0075] In this logistics park, the warehouses are vast with numerous passageways, and the transportation of goods is frequent. The interaction between pedestrians and forklifts is relatively complex and frequent. The operating environment includes changes in light during day and night, crowded areas, and open working areas. In this scenario, the anti-collision systems of ordinary forklifts often rely on lidar or ultrasonic sensors, which are costly and vulnerable to environmental factors, resulting in unsatisfactory detection effects. Additionally, existing computer vision solutions are usually difficult to meet the high-efficiency real-time requirements in embedded systems due to their high computational complexity. To address these issues, we adopted a lightweight AI vision anti-collision method based on a domestic embedded platform. A camera and a domestic Rockchip RV1126 processor were installed on the forklift, and combined with lightweight deep learning algorithms to achieve real-time detection and distance estimation of pedestrian targets.
[0076] In practical applications, the camera captures the RGB image in front of the forklift. The image data is preprocessed by the RV1126 processor, and lightweight convolutional neural networks are used for pedestrian detection feature extraction. The entire system optimizes the neural network structure through pruning and compression, reducing the consumption of computing resources to a level that can be tolerated by the embedded platform. The lightweight pedestrian detection module can detect the image of the current frame within 0.15 seconds, significantly improving the processing speed. After detecting a pedestrian, the system will use the lightweight pedestrian distance estimation module to perform real-time estimation of the distance between the pedestrian and the forklift. Binocular vision technology is adopted and combined with image correction and stereo matching algorithms to maintain high-precision distance estimation under different light conditions and complex backgrounds. The system finally fuses the pedestrian detection features and distance estimation features, and uses a dynamic adaptive sparse transform network for multi-level feature analysis to finally determine whether the pedestrians around the forklift are within the dangerous range.
[0077] In the experiment, the forklift traveled for a total of 1,200 hours, passing through different areas of the logistics park, including the goods placement area, loading and unloading areas, and main roads. To evaluate the actual performance of the system, we recorded the multi-pedestrian detection accuracy, distance estimation accuracy, and the overall response time of the system. In the area where goods are densely placed, the environmental light changes frequently, and the forklift often needs to interact with pedestrians, other forklifts, and goods. In this case, the pedestrian detection accuracy of this system reached 93.7%, showing a significant improvement compared with 82.4% of the ordinary lidar detection system. Especially in the case of partial occlusion, the dynamic adaptive sparse transform network can effectively identify and distinguish the features of different pedestrians, ensuring the stability of detection.
[0078] In the case of poor lighting at night, through image correction and stereo matching, the system can still maintain a high-precision estimation of the distance to pedestrians, with the error rate remaining within ±7.8%, showing a significant improvement compared to the ±15% error rate of traditional methods. By real-time estimating the relative distance between pedestrians and forklifts, when the distance to pedestrians is less than 1.5 meters, the system can respond within 0.3 seconds, emit audible and visual alarms, and automatically decelerate, effectively reducing the risk of collisions between forklifts and pedestrians. In the tests conducted in the main road area, the system operates stably at a speed of 10 frames per second, ensuring that the forklift can detect pedestrian targets in all directions in real time and optimize the path to avoid obstacles.
[0079] The lightweight AI vision anti-collision system in the present invention not only improves the detection accuracy but also significantly reduces the cost. Deployed on a domestic embedded platform based on the RV1126 processor, compared with imported lidar systems, the hardware cost is reduced by approximately 65%. During the entire experimental period, the average response time of the traditional lidar anti-collision system is 0.5 seconds, while the average response time of this system is 0.3 seconds, significantly shortening the reaction time of the forklift after detecting potential dangers. Especially when operating in crowded areas, this system can effectively identify multiple targets, preventing missed detections caused by dense crowds. No safety accidents occurred during the experimental process due to interactions between forklifts and pedestrians, demonstrating the effectiveness and reliability of the system.
[0080] By deploying this system on a domestic embedded platform, the intelligent anti-collision function of forklift operation is realized, enabling it to adapt to the working requirements in different environments, including daytime, night, crowded places, and wide open spaces, etc. Through module trimming and optimization of the HarmonyOS, the system maintains low power consumption while operating efficiently, with the average power consumption reduced by 30%, thereby extending the working hours of the forklift and reducing power consumption. In addition, the stability of the system has been fully verified. During the 1200-hour test, no system crashes or abnormal detections occurred.
[0081] This embodiment shows that through a series of innovative technologies such as lightweight convolutional neural network pruning and compression, binocular vision stereo matching, and dynamic adaptive sparse transformation network, the present invention not only solves the problems of high computational complexity, high cost, and insufficient accuracy of existing forklift anti-collision systems but also ensures the autonomy and customization of the system through successful deployment on a domestic embedded platform, providing a practical path for future large-scale applications in the logistics and warehousing industries.
[0082] Table 1 Comparison table of lightweight pedestrian detection and distance estimation performance based on ETH and INRIA datasets
[0083]
[0084] Table 1 shows the experimental results on two benchmark datasets, ETH and INRIA, which compare the performance of YOLO, MobileNet, and the proposed solution of the present invention in multiple aspects. First, in terms of pedestrian detection accuracy, the proposed solution of the present invention significantly outperforms the other two models with an accuracy of 93.7%. The accuracy of YOLO is 85.3%, while that of MobileNet is 88.1%. Although MobileNet has an improvement compared to YOLO, there is still a large gap compared to the proposed solution of the present invention. This indicates that the proposed solution of the present invention has higher reliability and accuracy in detecting pedestrians and can more effectively identify target objects.
[0085] In terms of the pedestrian distance estimation error rate, the proposed solution of the present invention also performs better than the other two models. The error rate of YOLO is 12.5%, the error rate of MobileNet is 10.8%, while the error rate of the proposed solution of the present invention is only 7.8%. This shows that the proposed solution of the present invention has higher accuracy in pedestrian distance estimation, which helps to reduce the distance error in the actual application of the system and improve the safety performance.
[0086] In terms of the model size, the model of the proposed solution of the present invention is only 75MB, significantly smaller than 220MB of YOLO and 90MB of MobileNet. This advantage makes the proposed solution of the present invention more applicable to resource-constrained embedded devices, especially in application scenarios that require storage and processing efficiency, where it can play a greater advantage.
[0087] In terms of the detection speed (frames per second), the proposed solution of the present invention also performs excellently, reaching 40 frames per second, significantly higher than 25 frames per second of YOLO and 35 frames per second of MobileNet. This means that the proposed solution of the present invention has excellent real-time performance, can process image data faster, improve the response speed of the system, and thus detect and identify pedestrians faster in dynamic scenarios.
[0088] Finally, in terms of power consumption, the power consumption of the proposed solution of the present invention is 2.5W, much lower than 4.8W of YOLO and 3.2W of MobileNet. The lower power consumption is of significant importance for extending the battery life of the device and reducing the energy consumption overhead. Especially in battery-powered application scenarios, the low power consumption advantage of the proposed solution of the present invention is very obvious.
[0089] In summary, the proposed solution of the present invention is superior to YOLO and MobileNet in various indicators such as pedestrian detection accuracy, pedestrian distance estimation error rate, model size, detection speed, and power consumption. These advantages make the proposed solution of the present invention more suitable for application scenarios with high requirements for accuracy, speed, and energy consumption in practical applications, reflecting significant technological progress and innovation.
[0090] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A lightweight unmanned forklift AI visual collision avoidance method based on a domestic embedded platform, characterized in that: The steps include: S1. Preprocess the RGB image through the lightweight pedestrian detection module, automatically prune and compress the deep convolutional network based on the image recognition algorithm of the lightweight convolutional neural network, and extract lightweight pedestrian detection features; S2, using the lightweight pedestrian distance estimation module, through camera calibration, image correction, stereo matching and distance acquisition, using model compression to extract lightweight pedestrian distance estimation features; S3, the lightweight pedestrian detection features and lightweight pedestrian distance estimation features are integrated, and feature analysis is performed in combination with a dynamic adaptive sparse transform network. The sparse feature adaptive reconstruction mechanism is introduced to complete multi-pedestrian detection and pedestrian distance estimation. S4. Deploy on domestic embedded platforms, implement anti-collision detection and rapid response through hardware, and use Hongmeng operating system to tailor software functional modules.
2. The lightweight unmanned forklift AI visual collision avoidance method based on a domestic embedded platform according to claim 1 is characterized in that: The image recognition algorithm of the lightweight convolutional neural network of S1 specifically includes: binary convolutional neural network optimization based on symbol distribution adjustment to binarize the input RGB image; automatic pruning of the deep convolutional network using similarity clustering and swarm intelligence optimization; fusion of the binarization processing and pruning results; and a deep convolutional network compression method based on the fusion of binarization and pruning.
3. The lightweight unmanned forklift AI visual collision avoidance method based on a domestic embedded platform according to claim 1 is characterized in that: The camera calibration in S2 includes external parameters and internal parameters of the camera. The external parameters and internal parameters of the camera are calibrated using the Zhang Zhengyou calibration method. The Zhang Zhengyou calibration method is a camera calibration based on a 2D plane target, and is calibrated using a calibration plate composed of two-dimensional grids.
4. The lightweight unmanned forklift AI visual collision avoidance method based on a domestic embedded platform according to claim 1 is characterized in that: The image correction in S2 is performed in accordance with the rules of epipolar geometric constraints, and the images captured by the left and right cameras are processed in the binocular vision system. Assuming that the target point corresponding to the to-be-matched point in the left camera image exists in the right camera image, the target point must be located on the epipolar line corresponding to the to-be-matched point; The Bouget algorithm in the OpenCV vision library is used to correct the collected data to eliminate the error caused by position offset during the three-dimensional imaging process.
5. The lightweight unmanned forklift AI visual collision avoidance method based on a domestic embedded platform according to claim 1 is characterized in that: The stereo matching in S2 is based on a deep learning pyramid stereo matching network to perform pedestrian ranging, specifically including: Use convolutional layers for feature extraction; Use spatial pyramid pooling to unify the image tensors to form a matching cost volume; The matching cost volume is adjusted by 3D convolution, wherein the 3D convolution uses a stacked hourglass structure, and a disparity map output is formed behind each hourglass; Calculate the loss functions separately, and perform weighted summation on the loss functions to obtain the total loss function.
6. The lightweight unmanned forklift AI visual collision avoidance method based on a domestic embedded platform according to claim 1 is characterized in that: The compression method of the model compression in S2 is 3D convolution channel compression, which retains the feature extraction module of the original model and performs model compression on the cost aggregation module using 3D convolution.
7. The lightweight unmanned forklift AI visual collision avoidance method based on a domestic embedded platform according to claim 1 is characterized in that: The S3 specifically includes: S31, fusing the lightweight pedestrian detection feature matrix and the lightweight pedestrian distance estimation feature matrix to generate a comprehensive feature matrix F; S32, construct an adaptive sparse transformation model, use a dynamic adaptive sparse transformation network to perform feature analysis on the comprehensive feature matrix F, perform multi-level adaptive reconstruction, and generate a sparse feature matrix F ′ ; S33. Combining the multi-layer optimization of matrix tensor transformation and high-order feature coupling, define the adaptive sparse reconstruction objective function L: Where W represents the feature weight matrix, b represents the bias vector, α and β represent the adaptive weight coefficients, and tanh(·) represents the hyperbolic tangent function. represents the Kronecker product, represents the Frobenius norm, λ and η represent the sparse regularization and low-rank regularization coefficients respectively, ∑ i,j |W ij | p represents the sparse regularization term, p∈[0,1] represents the sparsity parameter, and SVD(W) represents the singular value decomposition of the weight matrix W; S34, for the sparse feature matrix F ′ Perform adaptive reconstruction to generate the reconstructed feature matrix F rec ; S35, using the reconstructed feature matrix F rec Perform multi-pedestrian detection and use deep convolutional neural networks to classify and label the features of different pedestrians; S36. Based on the detected pedestrian features, the relative distance of each pedestrian is estimated to generate a distance estimation value of the pedestrian, and combined with the anti-collision rules, it is determined whether to generate an alarm signal.
8. The lightweight unmanned forklift AI vision anti-collision method based on a domestic embedded platform according to claim 1 is characterized in that: The hardware of the S4 specifically includes a housing, an LCD touch screen, a video head, a voice intelligent switching button, an external speaker, and a data interface; The development board is ATK-DLRV1126, equipped with the domestic Rockchip RV1126 processor; the RV1126 processor integrates CPU and NPU, the CPU is a quad-core ARM Cortex-A7 processor with a main frequency of up to 1.5GHz, and a built-in 2TOPS NPU for accelerating deep learning reasoning tasks; the RV1126 processor provides communication interfaces, including USB, I2C and SPI; the development board ATK-DLRV1126 communicates with other functional modules through UART.
9. The lightweight unmanned forklift AI vision anti-collision method based on a domestic embedded platform according to claim 1 is characterized in that: The software in S4 is divided into four layers, including a kernel layer, a system service layer, an application framework layer and an application layer. The application layer includes a system basic capability subsystem set, a basic software service subsystem set, an enhanced software service subsystem set and a hardware service subsystem set.
Citation Information
Cited By
Forklift battery SOC prediction method and device based on neural network
CN120971980A
Large model construction method and device, equipment and medium
CN121168556A