Intelligent building material sorting method and system based on vision
Through the collaborative work of the multimodal vision module and the deep learning processing module, combined with precision robots and navigation systems, the full process automation of building materials sorting is achieved, solving the problems of low efficiency and insufficient accuracy of traditional building materials sorting systems, and improving identification accuracy and handling efficiency.
Patent Information
- Application Number
- CN202510613755.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-05-13
AI Technical Summary
Traditional building materials sorting systems rely on manual operation or simple mechanical equipment, which are inefficient and difficult to accurately sort building materials with similar appearance but different properties.
The coordinated work of multi-modal vision module, deep learning processing module, intelligent sorting execution module, fast calibration module and mobile bearing module is adopted to obtain three-dimensional material distribution data through a binocular stereo camera and ToF depth sensor, and identify and position it in combination with YOLO algorithm, CBAM optimization mechanism and data enhancement technology. It uses a six-axis robotic arm, vacuum suction cup and adaptive fixture for grabbing, and combines McNum wheel and laser SLAM navigation for precise movement.
The full process automation of building materials sorting has been realized, the identification accuracy and handling efficiency have been improved, and the adaptability and identification accuracy have been enhanced in complex environments.
Smart Images

Figure CN120502518A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of building material sorting, and in particular to a vision-based intelligent building material sorting method and system. Background Art
[0002] The rapid development of the construction industry has placed higher demands on the efficiency of building material sorting and handling. Traditional building material sorting systems rely primarily on manual operation or simple mechanical equipment to complete sorting and handling tasks. Manual sorting typically relies on the operator's vision and experience to identify and grasp items. This method is inefficient and prone to errors when dealing with large, complex, or mixed building materials.
[0003] In automated sorting systems, sorting and classification are commonly performed using robotic arms, conveyor belts, or simple sensors. However, these systems typically rely on simple features such as color, size, or weight for classification, making it difficult to accurately sort building materials that appear similar but have different properties. Summary of the Invention
[0004] In order to realize the automated grasping and handling of target building materials, the present application provides a vision-based intelligent sorting method and system for building materials.
[0005] The above-mentioned invention objective of this application is achieved through the following technical solutions: A vision-based intelligent sorting system for building materials, comprising a multimodal vision module, an intelligent sorting execution module, a deep learning processing module, a rapid calibration module, and a mobile carrying module, wherein: The multimodal vision module is used to communicate data with the deep learning processing module, and the multimodal vision module is connected to the deep learning processing module; The intelligent sorting execution module is used to communicate data with the deep learning processing module, and the intelligent sorting execution module is connected to the mobile carrying module; The deep learning processing module is used to communicate data with the multimodal vision module, and the deep learning processing module is used to communicate data with the intelligent sorting execution module and the mobile carrying module; The rapid calibration module is used to communicate data with the deep learning processing module and the intelligent sorting execution module; The communication interface of the mobile carrying module is used to communicate data with the intelligent sorting execution module and the deep learning processing module.
[0006] By implementing the above technical solution, the vision-based intelligent building materials sorting system achieves full automation from image acquisition to gripping and handling through the collaborative work of a multimodal vision module, a deep learning processing module, an intelligent sorting execution module, a rapid calibration module, and a mobile carrier module. The multimodal vision module's binocular stereo camera and Time of Flight (ToF) depth sensor jointly capture image and depth information, generating three-dimensional material distribution data that is transmitted to the deep learning processing module. This module uses the YOLO algorithm, CBAM optimization mechanism, and data augmentation technology to identify and locate building materials. Recognition results, after correction by the rapid calibration module, are then provided to the intelligent sorting execution module, which controls the six-axis robotic arm, vacuum suction cup, and adaptive gripper for gripping. The mobile carrier module uses Mecanum wheels and laser SLAM navigation for precise movement, dynamically adjusting based on path planning and positioning information. Communication interfaces efficiently transmit data and control commands between the system modules. A remote monitoring system is also integrated for result evaluation and feedback optimization, continuously adjusting the recognition algorithm and control strategy to improve recognition accuracy, handling efficiency, and environmental adaptability.
[0007] In a preferred example, the present application can be further configured as follows: the multimodal vision module includes a binocular stereo camera, a ToF depth sensor, a ring fill light and an anti-glare curtain assembly, the binocular stereo camera works in conjunction with the ToF depth sensor, and the multimodal vision module is connected to the deep learning processing module.
[0008] By implementing the aforementioned technical solution, the multimodal vision module integrates a binocular stereo camera and a Time-of-Flight (ToF) depth sensor, enabling the coordinated acquisition of image and depth data, effectively generating highly accurate three-dimensional material distribution data. The binocular camera captures spatial structure based on parallax, while the ToF sensor supplements depth information by measuring time of flight. The two complement each other, providing comprehensive and three-dimensional data. A ring-shaped fill light provides uniform illumination, reduces shadow interference, and enhances image clarity. The anti-glare curtain assembly effectively filters strong reflected light, improving image quality in complex or brightly lit environments.
[0009] In a preferred example, the present application can be further configured as follows: the intelligent sorting execution module includes a six-axis collaborative robotic arm, a vacuum suction cup and an adaptive clamp, the adaptive clamp includes a pressure sensor and a silicone damping layer, the intelligent sorting execution module is connected to the deep learning processing module, and the intelligent sorting execution module is connected to the mobile carrying module.
[0010] By adopting the above technical solutions, the intelligent sorting execution module consists of a six-axis collaborative robotic arm, a vacuum suction cup, and an adaptive clamp, which is used to achieve precise grasping and handling of different building materials. The robotic arm controls the precise movement of each joint based on the position and posture information provided by the deep learning processing module. The vacuum suction cup is suitable for building materials with smooth surfaces and ensures stable adsorption by adjusting the suction force. The adaptive clamp is suitable for building materials with irregular shapes. The built-in pressure sensor and silicone damping layer automatically adjust the clamping force and angle to prevent slippage and damage. The intelligent sorting execution module is connected to the mobile carrier module to achieve synchronous transmission of post-grab information and transportation path planning, improving the accuracy and efficiency of the overall operation.
[0011] In a preferred embodiment, the present application can be further configured as follows: the deep learning processing module includes a YOLO algorithm processing unit, a CBAM module, and a data enhancement unit; the deep learning processing module is connected to the multimodal vision module, the deep learning processing module is connected to the intelligent sorting execution module, and the deep learning processing module is connected to the mobile carrying module.
[0012] By implementing the above technical solution, the deep learning processing module, consisting of a YOLO algorithm processing unit, a CBAM module, and a data augmentation unit, collaborates to achieve building material identification and optimization. The YOLO algorithm rapidly detects and classifies three-dimensional material distribution data. The CBAM module utilizes an attention mechanism to optimize recognition results, improving precision and positioning accuracy. The data augmentation unit enhances the model's generalization capabilities through diversified training. This module connects with the multimodal vision module, the intelligent sorting execution module, and the mobile load module to provide stable data support, ensuring accurate sorting and handling operations.
[0013] In a preferred example, the present application can be further configured as follows: the rapid calibration module includes a nine-point nonlinear calibration unit and a calibration data output interface, and the calibration data output interface is respectively connected to the deep learning processing module and the intelligent sorting execution module.
[0014] By employing this technical solution, the rapid calibration module generates calibration data using a nine-point nonlinear calibration unit and transmits this data to the deep learning processing module and the intelligent sorting execution module via a calibration data output interface. The nine-point nonlinear calibration unit identifies and calculates preset calibration points to generate calibration data for error correction and optimized positioning accuracy. The connection to the calibration data output interface ensures that the deep learning processing module and the intelligent sorting execution module can make real-time adjustments and optimizations based on this calibration data.
[0015] In a preferred example, the present application can be further configured as follows: the mobile carrying module includes a Mecanum wheel module, a laser SLAM navigation module and a communication interface, the Mecanum wheel module and the laser SLAM navigation module work together to control the movement and positioning of the mobile carrying module, and the communication interface is respectively connected to the deep learning processing module and the intelligent sorting execution module.
[0016] By adopting the above technical solutions, the mobile carrying module realizes all-round movement and precise positioning through the collaborative work of the Mecanum wheel module and the laser SLAM navigation module. The Mecanum wheel module provides flexible multi-directional movement capabilities, including forward, backward, sideways and rotation. The laser SLAM navigation module realizes real-time scanning and positioning of the environment through the combination of laser radar and inertial measurement unit, thereby generating accurate location information and path planning. Data is realized through the connection between the communication interface and the deep learning processing module and the intelligent sorting execution module. A vision-based intelligent building material sorting method is applied to the vision-based intelligent building material sorting system. The vision-based intelligent building material sorting method includes: Obtaining first image data output by the binocular stereo camera and second depth data output by the ToF depth sensor, and inputting the first image data and the second depth data into the YOLO algorithm processing unit for preliminary recognition and positioning to obtain three-dimensional material distribution data; Acquiring calibration data output by the nine-point nonlinear calibration unit, and performing coordinate correction and error correction on the three-dimensional material distribution data based on the calibration data to obtain building material positioning data; Using the data enhancement unit to perform data set enhancement training on the building material positioning data to obtain an enhanced material distribution data set; Identify and classify the enhanced material distribution data set through the YOLO algorithm processing unit to obtain a material identification result; Input the material identification result into the CBAM module for optimization processing, and output the optimized material identification result; The optimized material identification result is input into the intelligent sorting execution module, and the intelligent sorting execution module grasps and carries the target building materials through the six-axis collaborative robot arm, the vacuum suction cup and the adaptive clamp; Acquire motion data output by the Mecanum wheel module and global positioning data output by the laser SLAM navigation module; Inputting the motion data and the global positioning data into the deep learning processing module to obtain path planning and positioning control information; The path planning and positioning control information is input into the intelligent sorting execution module to control the movement and positioning of the mobile carrying module.
[0017] By employing this technical solution, the coordinated acquisition of binocular stereo cameras and ToF depth sensors simultaneously captures image and depth information, generating high-precision three-dimensional material distribution data for subsequent identification and positioning. A nine-point nonlinear calibration unit provides calibration data for real-time correction of recognition errors and improved positioning accuracy. The deep learning processing module, consisting of the YOLO algorithm, CBAM module, and data augmentation unit, performs rapid recognition, feature optimization, and diversified training, enhancing the model's robustness and generalization capabilities and ensuring recognition accuracy in complex environments. The intelligent sorting execution module utilizes a six-axis robotic arm, vacuum suction cups, and adaptive grippers to achieve stable grasping and handling of various building materials. The Mecanum wheel and laser SLAM navigation module collaborate to support omnidirectional movement and precise positioning, enhancing adaptability in confined spaces. Each module efficiently exchanges data via communication interfaces. The deep learning processing module and the intelligent sorting execution module collaboratively optimize paths and control parameters. A remote monitoring system provides continuous feedback and model training, ensuring the accuracy, stability, and efficiency of the sorting process.
[0018] In a preferred example, the present application may be further configured as follows: using the data enhancement unit to perform dataset enhancement training on the building material positioning data to obtain an enhanced material distribution dataset, including: The data enhancement unit performs data set enhancement training on the building material positioning data based on the Mosaic data enhancement method to obtain the enhanced material distribution data set.
[0019] By adopting the above technical solution, the data enhancement unit performs data set enhancement training on building material positioning data based on the Mosaic data enhancement method, which can effectively improve the generalization ability and robustness of the recognition model.
[0020] In a preferred example, the present application may be further configured as follows: inputting the material identification result into the CBAM module for optimization processing, and outputting the optimized material identification result, including: The CBAM module optimizes the material recognition result based on the channel attention mechanism and the spatial attention mechanism, and outputs the optimized material recognition result.
[0021] By employing the above technical solutions, the CBAM module optimizes material recognition results through both channel-attention and spatial-attention mechanisms, improving both recognition and positioning accuracy. The channel-attention mechanism extracts the importance of each channel in the feature map through global pooling, weighting key channels and de-emphasizing irrelevant information. The spatial-attention mechanism further identifies key regions in the image and weights positional features. Combined, these two mechanisms optimize the feature map in both the channel and spatial dimensions, enhancing the model's recognition and feature extraction capabilities for target building materials. The optimized recognition results are clearer and more stable, providing accurate support for subsequent classification and positioning.
[0022] In a preferred example, the present application may be further configured as follows: the vision-based intelligent sorting method for building materials further includes: The remote monitoring system evaluates and provides feedback on the optimized material identification result and the path planning and positioning control information to obtain an evaluation feedback result; Based on the evaluation feedback results, optimization information is generated, and the optimization information is input into the deep learning processing module and the intelligent sorting execution module for optimization training to obtain an optimized training model and execution control strategy; Inputting the optimized training model into the YOLO algorithm processing unit and the CBAM module to update the recognition algorithm and optimization strategy, thereby obtaining an updated material recognition result, which is used to improve the recognition accuracy and target positioning accuracy of the YOLO algorithm processing unit and the CBAM module; The optimized execution control strategy is input into the intelligent sorting execution module to update the execution control parameters of the sorting and handling operations to obtain updated path control information. The updated path control information is used to optimize the motion path and grasping accuracy of the intelligent sorting execution module.
[0023] By adopting the above technical solution, the remote monitoring system can provide real-time evaluation and feedback on the optimized material recognition results and path planning information, comprehensively analyze recognition accuracy, path deviation, and control effectiveness, and generate optimization information that is input into the deep learning processing module and the intelligent sorting execution module for training and updating, continuously optimizing the recognition model and control strategy. The optimized model is loaded into the YOLO algorithm processing unit and CBAM module to improve recognition and positioning accuracy. At the same time, the updated control strategy is used to adjust the execution parameters of sorting and handling, improving path planning and grasping accuracy. Through this dynamic evaluation and optimization mechanism, the system achieves high-precision recognition and stable control in complex environments.
[0024] In summary, this application includes at least one of the following beneficial technical effects: 1. The vision-based intelligent building materials sorting system achieves full automation from image acquisition to gripping and handling through the collaborative work of a multimodal vision module, a deep learning processing module, an intelligent sorting execution module, a rapid calibration module, and a mobile carrying module. The binocular stereo camera and ToF depth sensor in the multimodal vision module jointly capture image and depth information, generate three-dimensional material distribution data, and transmit it to the deep learning processing module. This module uses the YOLO algorithm, CBAM optimization mechanism, and data enhancement technology to complete the identification and positioning of building materials. After correction by the rapid calibration module, the identification results are provided to the intelligent sorting execution module, which controls the six-axis robotic arm, vacuum suction cup, and adaptive gripper to perform the gripping operation. The mobile carrying module uses Mecanum wheels and laser SLAM navigation to achieve precise movement, and dynamically adjusts based on path planning and positioning information. The system modules efficiently transmit data and control instructions through communication interfaces. At the same time, a remote monitoring system is used to evaluate results and provide feedback optimization, continuously adjusting the recognition algorithm and control strategy to improve recognition accuracy, handling efficiency, and environmental adaptability. 2. Through the coordinated acquisition of binocular stereo cameras and ToF depth sensors, image and depth information can be obtained simultaneously to generate high-precision three-dimensional material distribution data for subsequent identification and positioning. The nine-point nonlinear calibration unit provides calibration data for real-time correction of recognition errors and improved positioning accuracy. The deep learning processing module consists of the YOLO algorithm, CBAM module and data enhancement unit, which completes rapid recognition, feature optimization and diversified training, enhances the robustness and generalization ability of the model, and ensures recognition accuracy in complex environments. The intelligent sorting execution module uses a six-axis robotic arm, vacuum suction cup and adaptive fixture to achieve stable grasping and handling of different building materials. The Mecanum wheel and laser SLAM navigation module work together to support omnidirectional movement and precise positioning, improving adaptability in confined spaces. Each module conducts efficient data exchange through the communication interface. The deep learning processing module and the intelligent sorting execution module work together to optimize the path and control parameters. The remote monitoring system continuously provides feedback and training models to ensure the accuracy, stability and execution efficiency of the sorting process. 3. The remote monitoring system provides real-time evaluation and feedback on optimized material recognition results and path planning information. It comprehensively analyzes recognition accuracy, path deviation, and control effectiveness, generating optimization information that is fed into the deep learning processing module and the intelligent sorting execution module for training and updating, continuously optimizing the recognition model and control strategy. The optimized model is loaded into the YOLO algorithm processing unit and CBAM module to improve recognition and positioning accuracy. The updated control strategy is used to adjust sorting and handling execution parameters, improving path planning and grasping accuracy. This dynamic evaluation and optimization mechanism enables high-precision recognition and stable control in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 This is a flowchart of a vision-based intelligent building material sorting system in one embodiment of the present application.
[0026] Figure 2 This is a flow chart of a vision-based intelligent sorting method for building materials in one embodiment of the present application; Figure 3 This is a flowchart for implementing step S30 in a vision-based intelligent building material sorting method in one embodiment of the present application; Figure 4 This is a flowchart for implementing step S50 in a vision-based intelligent building material sorting method in one embodiment of the present application; Figure 5 This is a flowchart for implementing step S90 in a vision-based intelligent building material sorting method in one embodiment of the present application.
[0027] Description of reference numerals: 1. Multimodal vision module; 2. Intelligent sorting execution module; 3. Deep learning processing module; 4. Rapid calibration module; 5. Mobile carrying module. DETAILED DESCRIPTION
[0028] The present application is further described in detail below with reference to the accompanying drawings.
[0029] In one embodiment, if Figure 1 As shown, the present application discloses a vision-based intelligent sorting system for building materials, which includes a multimodal vision module 1, an intelligent sorting execution module 2, a deep learning processing module 3, a fast calibration module 4 and a mobile carrying module 5, wherein: The multimodal vision module 1 is used to communicate data with the deep learning processing module 3, and the multimodal vision module 1 is connected to the deep learning processing module 3; The intelligent sorting execution module 2 is used to communicate data with the deep learning processing module 3, and the intelligent sorting execution module 2 is connected to the mobile carrying module 5; The deep learning processing module 3 is used to communicate data with the multimodal vision module 1, and the deep learning processing module 3 is used to communicate data with the intelligent sorting execution module 2 and the mobile carrying module 5; The rapid calibration module 4 is used to communicate data with the deep learning processing module 3 and the intelligent sorting execution module 2; The communication interface of the mobile carrying module 5 is used for data communication with the intelligent sorting execution module 2 and the deep learning processing module 3.
[0030] Specifically, the multimodal vision module 1 is used to communicate data with the deep learning processing module 3, the multimodal vision module 1 is connected to the deep learning processing module 3, the intelligent sorting execution module 2 is used to communicate data with the deep learning processing module 3, the intelligent sorting execution module 2 is connected to the mobile carrying module 5, the deep learning processing module 3 is used to communicate data with the multimodal vision module 1, the deep learning processing module 3 is used to communicate data with the intelligent sorting execution module 2 and the mobile carrying module 5, the fast calibration module 4 is used to communicate data with the deep learning processing module 3 and the intelligent sorting execution module 2, and the communication interface of the mobile carrying module 5 is used to communicate with the intelligent sorting execution module 2. The picking execution module 2 communicates data with the deep learning processing module 3; specifically, the multimodal vision module 1 is configured with a binocular stereo camera and a ToF depth sensor to simultaneously collect image data and depth data in the target area, and the image data obtained from different perspectives by the binocular stereo camera is fused with the depth data collected by the ToF depth sensor, and the spatial position and shape information of the target building materials are reconstructed and represented by constructing a three-dimensional point cloud model. During the reconstruction process, the calibration data generated by the nine-point nonlinear calibration unit 41 is used to perform coordinate correction and error correction on the three-dimensional point cloud data, thereby generating building material positioning data; the generated building material positioning data The data is input to the deep learning processing module 3 through the data interface. The deep learning processing module 3 includes a YOLO algorithm processing unit, a CBAM module and a data enhancement unit. The YOLO algorithm processing unit performs target detection and classification on the input building material positioning data, extracts the category, position and size information of the building materials through the recognition algorithm, and annotates the recognition results and positioning information. The preliminary recognition results are optimized by the CBAM module. The CBAM module performs weighted calculation on the input features through the channel attention mechanism and the spatial attention mechanism, enhances the extraction and filtering of key features, and outputs the optimized recognition results. The data enhancement unit is used by the Mo The SAIC data enhancement method expands and diversifies the training of building material positioning data, thereby improving the generalization ability of the deep learning model; the optimized recognition results are input into the intelligent sorting execution module 2 through the data interface. The intelligent sorting execution module 2 includes a six-axis collaborative robot arm, a vacuum suction cup and an adaptive clamp. The six-axis collaborative robot arm performs path planning and motion control based on the recognition results and positioning information output by the deep learning processing module 3. The vacuum suction cup adapts to the surface characteristics of different types of building materials by adjusting the suction force. The adaptive clamp detects the clamping force through the built-in pressure sensor and silicone damping layer and provides flexible protection to avoid damage to the surface of the building materials.A data transmission channel is established between the intelligent sorting execution module 2 and the mobile carrier module 5 via a communication interface. The mobile carrier module 5 includes a Mecanum wheel module and a laser SLAM navigation module. The Mecanum wheel module provides multi-directional mobility, while the laser SLAM navigation module is used to build an environmental map and perform path planning and positioning. The mobile carrier module 5 uses real-time path control information to adjust its own motion trajectory and positioning accuracy, ensuring that building materials accurately reach their target locations during transportation. Through the collaborative work of the deep learning processing module 3 and the intelligent sorting execution module 2, full process control is achieved, from building material identification, classification, grasping, to transportation.
[0031] Furthermore, the multimodal vision module 1 includes a binocular stereo camera 11, a ToF depth sensor 12, a ring fill light 13 and an anti-glare curtain assembly 14. The binocular stereo camera 11 and the ToF depth sensor 12 work together, and the multimodal vision module 1 is connected to the deep learning processing module 3.
[0032] Specifically, the multimodal vision module 1 includes a binocular stereo camera 11, a ToF depth sensor 12, a ring fill light 13 and an anti-glare curtain assembly 14. The binocular stereo camera 11 is used to collect image data in the target area from different perspectives, and the ToF depth sensor 12 is used to collect depth data in the target area. By synchronously collecting and fusing image data and depth data, the generation of three-dimensional material distribution data of building materials is realized. The binocular stereo camera 11 and the ToF depth sensor 12 collect data through a synchronous trigger mechanism. After the binocular stereo camera 11 obtains image data from the left and right perspectives, the image data is transmitted to the deep learning processing module 3 through a serial port or a parallel interface. When collecting depth data, the ToF depth sensor 12 emits a modulated light signal to the surface of the target object, and measures the reflected light signal. Phase difference or time delay is used to calculate the depth information of the object surface. The generated depth data is fused with the image data of the binocular stereo camera 11, and a three-dimensional point cloud model is constructed through a three-dimensional reconstruction algorithm. The generated three-dimensional material distribution data includes the spatial position, size, shape, surface texture and other information of the building materials. In order to ensure the stability and accuracy of data acquisition, a ring fill light 13 is set in the multimodal vision module 1 to provide a uniform lighting environment, and automatically adjust the light intensity and color temperature to adapt to the lighting conditions in different scenes. The anti-glare curtain assembly 14 is used to eliminate glare interference caused by strong light sources or reflected light in the environment, and improve the clarity and contrast of the image data by blocking or absorbing excess light. The multimodal vision module 1 transmits the generated three-dimensional material distribution data to the deep learning processing module 3 through the data interface.
[0033] Furthermore, the intelligent sorting execution module 2 includes a six-axis collaborative robotic arm 21, a vacuum suction cup 22 and an adaptive clamp 23, the adaptive clamp 23 includes a pressure sensor 231 and a silicone damping layer 232, the intelligent sorting execution module 2 is connected to the deep learning processing module 3, and the intelligent sorting execution module 2 is connected to the mobile carrying module 5.
[0034] Specifically, the intelligent sorting execution module 2 includes a six-axis collaborative robot arm 21, a vacuum suction cup 22 and an adaptive clamp 23. The six-axis collaborative robot arm 21 is used to realize multi-degree-of-freedom operation and precise control of the target building materials, and provides highly flexible movement capabilities by configuring six rotary joints and multiple servo motors. The end of the six-axis collaborative robot arm 21 is installed with a vacuum suction cup 22 and an adaptive clamp 23. The vacuum suction cup 22 is used to perform non-contact grasping when the surface of the target building material is relatively smooth or has good adsorption conditions. The adsorption force is adjusted by the built-in solenoid valve and negative pressure generator to adapt to building materials with different weights and surface characteristics. The adaptive clamp 23 is used to clamp building materials with irregular shapes or uneven surfaces. A buffering and shock absorption function is provided by arranging a flexible silicone damping layer 232 in the clamping structure. The silicone damping layer 232 can apply soft pressure to the surface of the building material during the clamping process, thereby avoiding damage or deformation of the building material surface due to excessive clamping force. The built-in pressure sensor of the adaptive clamp 23 The device 231 is used to detect the pressure changes in the clamping process in real time, and transmits the detected pressure data to the deep learning processing module 3, and adaptively adjusts the clamping force according to the pressure feedback information to ensure the stability and safety of the clamping. The intelligent sorting execution module 2 is connected to the deep learning processing module 3, and interacts with data transmission and control instructions through the communication protocol. When the deep learning processing module 3 outputs the optimized material identification results and target position coordinates, the intelligent sorting execution module 2 performs path planning and action execution based on the received data, calculates the motion trajectory of the robotic arm and the clamping position of the end tool through the motion control algorithm, and adjusts the clamping force and adsorption force in real time to adapt to building materials of different types and sizes. The intelligent sorting execution module 2 is connected to the mobile carrying module 5, and performs real-time data communication with the mobile carrying module 5 through the data interface. When the target building material is successfully grasped, the mobile carrying module 5 moves and navigates autonomously according to the received path planning information, thereby transporting the target building material to the designated location.
[0035] Furthermore, the deep learning processing module 3 includes a YOLO algorithm processing unit 31, a CBAM module 32 and a data enhancement unit 33. The deep learning processing module 3 is connected to the multimodal vision module 1, the deep learning processing module 3 is connected to the intelligent sorting execution module 2, and the deep learning processing module 3 is connected to the mobile carrying module 5.
[0036] Specifically, the deep learning processing module 3 includes a YOLO algorithm processing unit 31, a CBAM module 32 and a data enhancement unit 33. The YOLO algorithm processing unit 31 is used to identify and classify the three-dimensional material distribution data from the multimodal vision module 1, and perform feature extraction and target detection on the input three-dimensional material distribution data by configuring a convolutional neural network. First, the input three-dimensional point cloud data is segmented and preprocessed, and the data format is converted into a standard format suitable for neural network input. Then, the input data is subjected to layer-by-layer feature extraction and downsampling operations through a combination of convolutional layers, pooling layers and activation functions, thereby obtaining the target building materials at different scales and locations. Hierarchical feature representation, in the YOLO algorithm, by defining anchor boxes and loss functions to train the model for target detection and classification, the recognition results are output in the form of category labels and target position coordinates, and the preliminary recognition results output by the YOLO algorithm processing unit 31 are optimized and the accuracy is improved through the CBAM module 32. The CBAM module 32 enhances the selectivity and sensitivity of target features by introducing channel attention mechanism and spatial attention mechanism. First, the input feature map is subjected to global average pooling and global maximum pooling operations to extract the global information and maximum response value of different channels respectively, and the channel features are added through the full connection layer and activation function. The weight calculation is performed, and the calculation result is element-by-element multiplied with the original feature map to generate an optimized feature map; then the optimized feature map is convolved and pooled through the spatial attention mechanism to generate an attention weight matrix for positioning, and the dot multiplication operation is performed with the feature map to enhance the focusing ability of the target area. The optimized recognition result is transmitted to the data enhancement unit 33 through the data interface for data set enhancement training. The data enhancement unit 33 randomly crops, rotates, scales and mixes the input data through the Mosaic data enhancement method to generate a training sample set with diversity and robustness. The enhanced training of different data samples is used to enhance the recognition accuracy. In order to improve the generalization ability and robustness of the model, the output results of the data enhancement unit 33 are transmitted to the intelligent sorting execution module 2 and the mobile carrying module 5 through the data interface. The deep learning processing module 3 is connected with the multimodal vision module 1 to receive the three-dimensional material distribution data from the multimodal vision module 1; the deep learning processing module 3 is connected with the intelligent sorting execution module 2 to transmit the optimized recognition results and classification information to the intelligent sorting execution module 2 to guide the sorting and grasping operations of building materials; the deep learning processing module 3 is connected with the mobile carrying module 5 to transmit the path planning information and recognition results to the mobile carrying module 5 to guide the movement and positioning operations.
[0037] Furthermore, the fast calibration module 4 includes a nine-point nonlinear calibration unit 41 and a calibration data output interface 42 , and the calibration data output interface 42 is connected to the deep learning processing module 3 and the intelligent sorting execution module 2 respectively.
[0038] Specifically, the fast calibration module 4 includes a nine-point nonlinear calibration unit 41 and a calibration data output interface 42. The nine-point nonlinear calibration unit 41 is used to calibrate and correct the spatial coordinate system between the multimodal vision module 1, the deep learning processing module 3 and the intelligent sorting execution module 2. Nine calibration points with known positions and postures are set in the specified calibration area, and the calibration points are imaged and depth detected by the binocular stereo camera 11 and the ToF depth sensor 12 in the multimodal vision module 1 to generate three-dimensional point cloud data containing the calibration points and corresponding image data. The collected calibration data is input into the nine-point nonlinear calibration unit 41. The nine-point nonlinear calibration unit 41 fits and solves the data of the calibration points by introducing a nonlinear optimization algorithm and a least squares method, and calculates the spatial mapping relationship and error parameters between the visual system and the execution module, including parameter information such as the intrinsic parameter matrix, the extrinsic parameter matrix and the distortion coefficient. The calibration data is generated by comparing the actual measurement value with the theoretical calibration value and iteratively optimizing. According to the calibration data, the calibration data is used to perform coordinate correction and error correction on the three-dimensional material distribution data, so as to ensure that the building material positioning data obtained from the visual system can be accurately mapped to the operating space of the intelligent sorting execution module 2 and the mobile carrying module 5. The calibration data output interface 42 is used to transmit the generated calibration data to the deep learning processing module 3 and the intelligent sorting execution module 2 through the data interface. The deep learning processing module 3 is used to receive the calibration data, and correct and compensate the three-dimensional material distribution data according to the received data, and fuse and map the calibration parameters with the point cloud data through calculation to ensure that the generated building material positioning data has high accuracy and consistency; the intelligent sorting execution module 2 is used to receive the calibration data, and adjust and optimize the motion path and operation accuracy of the robot arm based on the received calibration information, and ensure that the end tool of the robot arm can accurately reach the target position and accurately grasp and carry the target building materials by comparing and calculating the calibration information with the actual operation data.
[0039] Furthermore, the mobile carrying module 5 includes a Mecanum wheel module 51, a laser SLAM navigation module 52 and a communication interface 501. The Mecanum wheel module 51 and the laser SLAM navigation module 52 work together to control the movement and positioning of the mobile carrying module 5. The communication interface 501 is connected to the deep learning processing module 3 and the intelligent sorting execution module 2 respectively.
[0040] Specifically, the mobile carrying module 5 includes a Mecanum wheel module 51, a laser SLAM navigation module 52 and a communication interface 501. The Mecanum wheel module 51 and the laser SLAM navigation module 52 work together to achieve omnidirectional movement and precise positioning of the mobile carrying module 5. The Mecanum wheel module 51 is formed by configuring multiple Mecanum wheels at the bottom of the mobile platform. Each Mecanum wheel is composed of multiple tilted rollers. The rotation direction and speed of each Mecanum wheel are driven by an independent motor. The friction difference between the rollers and the ground is used to achieve forward, backward, left and right translation and rotation movement of the mobile platform. During the movement, the speed and direction of each Mecanum wheel are adjusted in real time by the motion controller, thereby achieving high-precision omnidirectional movement and position control. The laser SLAM navigation module 52 is configured with a laser radar and an inertial measurement unit (IMU) for real-time scanning and map construction of the surrounding environment. The laser radar emits a laser beam in a high-speed rotating manner and receives the reflected light signal. The distance from the surrounding objects is calculated by measuring the flight time or phase difference of the laser pulse. The distance between the bodies is calculated, thereby generating two-dimensional or three-dimensional point cloud data. The inertial measurement unit is used to detect the acceleration and angular velocity information of the mobile carrying module 5, and is fused with the lidar data to improve the positioning accuracy and robustness. The laser SLAM navigation module 52 compares and matches the constructed environmental map with the current position information to generate path planning and navigation instructions, and transmits the navigation information to the deep learning processing module 3 and the intelligent sorting execution module 2 through the communication interface 501. The deep learning processing module 3 generates optimized path planning data based on the received navigation information, and transmits the optimization result to the mobile carrying module 5 through the data interface. The intelligent sorting execution module 2 receives the path planning data and position information through the execution control terminal 201, which is used to control the movement of the robotic arm and the grabbing and handling operations of the target building materials. When the mobile carrying module 5 performs autonomous navigation and movement, the environmental map and path planning information are updated in real time through the laser SLAM navigation module 52 to ensure accurate obstacle avoidance and precise positioning in complex or dynamic environments, thereby achieving efficient building material handling and path optimization.
[0041] like Figure 2 As shown, a vision-based intelligent sorting method for building materials is applied to a vision-based intelligent sorting system for building materials. The vision-based intelligent sorting method for building materials includes: S10: Obtain the first image data output by the binocular stereo camera 11 and the second depth data output by the ToF depth sensor 12, input the first image data and the second depth data into the YOLO algorithm processing unit (31) for preliminary recognition and positioning, and obtain three-dimensional material distribution data.
[0042] Specifically, the first image data output by the binocular stereo camera 11 and the second depth data output by the ToF depth sensor 12 are obtained, and the binocular stereo camera 11 and the ToF depth sensor 12 are configured to simultaneously collect image data and depth data in the target area. The binocular stereo camera 11 is installed at a certain baseline distance by setting two cameras at different positions to form two left and right perspectives. By capturing two images of the same scene, the spatial depth information of the target building materials in the scene is calculated using the parallax principle. By matching the feature points and performing stereo correction on the left and right views, matching pixel pairs are generated and their parallax values are calculated. The parallax information is input into the three-dimensional reconstruction algorithm for depth estimation and point cloud generation. The ToF depth sensor 12 transmits a modulated light signal to the target area and receives the reflected light signal. It calculates the depth information of the object surface by measuring the phase difference or time delay of the light signal during propagation. The collected depth data is aligned and fused with the point cloud data generated by the binocular stereo camera 11. The depth information and image data are uniformly encoded and processed through a coordinate registration algorithm to generate three-dimensional material distribution data including the location, shape, size and surface features of the building materials. To ensure the accuracy and completeness of data collection, the ring fill light 13 configured in the multimodal vision module 1 provides uniform lighting conditions, automatically adjusts the light intensity and color temperature to adapt to different lighting environments, and at the same time, the anti-glare curtain component 14 shields and absorbs the glare generated in the environment to reduce the impact of light reflection or strong light interference on the image quality. The acquired first image data and second depth data are input into the YOLO algorithm processing unit 31, and the input data are subjected to feature extraction, detection and classification by configuring a convolutional neural network. The input data are subjected to convolution and pooling operations by the constructed neural network model, and the high-dimensional features of the original data are extracted and compressed. The classification judgment and target positioning are performed through the fully connected layer and the activation function to generate three-dimensional material distribution data including the category information and spatial position of the target building material.
[0043] S20: Acquire the calibration data output by the nine-point nonlinear calibration unit 41, perform coordinate correction and error correction on the three-dimensional material distribution data based on the calibration data, and obtain building material positioning data.
[0044] Specifically, nine calibration points with known positions are arranged in a predetermined calibration area, and the position coordinates of each calibration point are known and accurate. The calibration area is imaged and depth detected by the binocular stereo camera 11 and the ToF depth sensor 12 in the multimodal vision module 1 to generate three-dimensional point cloud data containing the calibration points and corresponding image data. The collected calibration point data is input into the nine-point nonlinear calibration unit 41. The nine-point nonlinear calibration unit 41 processes and fits the calibration point data through a nonlinear optimization algorithm. Specifically, the calibration point data is first denoised and filtered to eliminate the influence of environmental noise and sensor error. The objective function is constructed by the least squares method, and the measured value of the calibration point is calculated and the error is analyzed by the theoretical value. The error parameters are adjusted and compensated by the iterative optimization algorithm to generate calibration data. The calibration data includes an intrinsic parameter matrix, an extrinsic parameter matrix, and an extrinsic parameter matrix. The internal parameter matrix and distortion coefficient and other parameter information, wherein the internal parameter matrix is used to describe the focal length, principal point coordinates and pixel scaling factor of the camera, the external parameter matrix is used to describe the spatial position relationship and rotation matrix between the camera and the calibration plane, and the distortion coefficient is used to correct the radial and tangential distortion caused by the lens. The generated calibration data is input into the deep learning processing module 3 for coordinate correction and error correction. The calibration data is coordinate-aligned and transformed with the three-dimensional material distribution data, and the three-dimensional point cloud data is aligned with the world coordinate system to eliminate the coordinate deviation caused by sensor error or system configuration. The error distribution of the calibration point is modeled and fitted by the error correction algorithm. The coordinate value of each point cloud data is corrected point by point to eliminate the influence of systematic error and random error, and obtain accurate building material positioning data. The building material positioning data includes the spatial position, size, shape and direction information of the building materials.
[0045] S30: Using the data enhancement unit 33 to perform data set enhancement training on the building material positioning data to obtain an enhanced material distribution data set.
[0046] Specifically, the input building material positioning data is first preprocessed and formatted, including standardizing and normalizing the point cloud data and position coordinates in the three-dimensional material distribution data to facilitate the input and training of the deep learning network. By normalizing the input data, the building material data of different scales are uniformly scaled, and the input data are segmented and resampled, thereby improving the uniformity of the data and the computational efficiency; the data enhancement unit 33 diversifies and expands the training of the building material positioning data through the configured Mosaic data enhancement method, and randomly crops, splices, scales, rotates and mixes multiple groups of different building material positioning data to generate a new training sample set. In the Mosaic data enhancement process, the building material positioning data of four different scenes are randomly selected and merged into a new scene. First, the data of each scene is The rows are randomly cropped and scaled, and then the cropped data is transformed by coordinate translation and rotation matrices. The transformed data are spliced and fused to generate training samples with multiple building material objects and different backgrounds. By mixing and reorganizing data from different sources, the adaptability and generalization ability of the model to complex scenes are improved. At the same time, random noise and brightness adjustment are introduced in the data enhancement process to increase the diversity and robustness of the data set. The generated enhanced sample set is input into the deep learning processing module 3 through the data interface for training and optimization. During the training process, the input data is feature extracted and classified by the constructed deep neural network, and the input data is processed and down-sampled layer by layer through the convolution layer and the pooling layer to extract high-level features with discrimination. Classification and prediction are performed through the fully connected layer and the activation function to generate an enhanced material distribution data set.
[0047] S40: The YOLO algorithm processing unit 31 identifies and classifies the enhanced material distribution data set to obtain a material identification result.
[0048] Specifically, the enhanced material distribution data set is first input into the YOLO algorithm processing unit 31 for feature extraction and target detection, and the configured convolutional neural network is used to perform multi-level feature extraction and classification judgment on the input three-dimensional point cloud data and building material positioning data. The enhanced data set is formatted and standardized in the input layer, including data normalization, coordinate alignment and batch processing, and data samples from different sources are uniformly standardized to facilitate the input and calculation of the neural network. The YOLO algorithm processing unit 31 performs layer-by-layer feature extraction and dimensionality reduction operations on the input data through the constructed convolution layer, pooling layer and activation function, thereby generating a feature map with a high-dimensional representation. During the feature extraction process, a sliding window operation is performed on the input data through the convolution kernel to extract feature information at different scales and resolutions, and the feature map is downsampled and compressed through the pooling operation to reduce the computational complexity and the number of parameters. In the target detection of the YOLO algorithm In the process, the category, position and size of building materials are predicted and classified by defining anchor boxes and loss functions. Multiple anchor boxes are generated on the feature map of the input data, and each anchor box is classified and regressed. The predicted results include category confidence, bounding box coordinates and relative scale information. The predicted results are compared with the true annotation values, and the error value is calculated by the defined loss function. The loss function includes classification error, positioning error and confidence error. The network parameters are updated and optimized through the back propagation algorithm to reduce the gap between the predicted results and the true values. During the training process, multi-scale detection mechanism and feature fusion strategy are introduced to improve the model's recognition ability for building materials of different scales and shapes. The trained model is applied to the enhanced material distribution data set for identification and classification. The input data is forward propagated and inferred through the constructed neural network to generate material recognition results containing the category information, position coordinates and size information of the target building materials.
[0049] S50: Input the material identification result to the CBAM module 32 for optimization processing, and output the optimized material identification result.
[0050] Specifically, by improving the accuracy and optimizing the features of the material recognition results generated by the YOLO algorithm processing unit 31, the CBAM module 32 is composed of a channel attention mechanism and a spatial attention mechanism. First, the feature map of the input material recognition result is decomposed and normalized, and the channels of the input feature map are separated and normalized to generate multiple independent channel feature matrices. In the processing of the channel attention mechanism, global information is extracted from the input feature map through global average pooling and global maximum pooling operations, and the feature information of each channel is compressed and aggregated. The channel features are weighted by the fully connected layer and the activation function to generate a channel weight matrix for optimization. The channel weight matrix is multiplied element by element with the original feature map, thereby enhancing important features and removing redundant or useless features. The key features are suppressed; in the processing of the spatial attention mechanism, the spatial dimension of the feature map is compressed and reconstructed by performing convolution and pooling operations on the input feature map, and the features at different positions are weighted by the constructed convolution kernel to generate a spatial weight matrix for optimization, and the spatial weight matrix is multiplied by the feature map after channel optimization to highlight and optimize the key areas; in the optimization process of the CBAM module 32, the useful information and useless information in the original feature map are effectively distinguished and extracted by performing multiple iterations and optimization calculations on the material recognition results, and the different dimensions of the feature map are weighted and fused to improve the feature discrimination ability and the accuracy of the model. The generated optimized material recognition result includes the category information, position coordinates, size information and confidence value of the target building material.
[0051] S60: The optimized material identification result is input into the intelligent sorting execution module 2, and the intelligent sorting execution module 2 grasps and carries the target building materials through the six-axis collaborative robot arm 21, the vacuum suction cup 22 and the adaptive clamp 23.
[0052] Specifically, after receiving the optimized material recognition result, the sorting and grasping operations are controlled and executed through the category information, position coordinates, size information and confidence value of the target building materials contained in the recognition result. The six-axis collaborative robot arm 21 calculates and plans the spatial position and grasping path of the target building materials in real time through the built-in motion control system and path planning algorithm, and realizes high-degree-of-freedom movement and precise positioning through the cooperation of multiple servo motors and rotary joints. Specifically, the six-axis collaborative robot arm 21 calculates the target position and grasping posture of the end tool based on the received optimized material recognition result, and solves and optimizes the angles and motion trajectories of each joint of the robot arm through inverse kinematics algorithm and interpolation calculation, and converts the calculation results into driving signals of each joint motor, so that the end tool of the robot arm accurately reaches the predetermined position of the target building material; the vacuum suction cup 22 adsorbs and grasps the surface of the target building material by adjusting the working state of the negative pressure generator and the solenoid valve. When the surface of the target building material is relatively smooth or has good adsorption conditions The vacuum suction cup 22 automatically adjusts its suction force and contact area to accommodate building materials of different materials and surface properties. A built-in pressure sensor provides real-time detection and feedback of the suction state to ensure stable and safe grasping. The adaptive gripper 23 is used to grip and carry building materials with irregular shapes or uneven surfaces. A flexible silicone damping layer 232 and a pressure sensor 231 are configured in the gripping structure. The flexible silicone damping layer 232 provides cushioning and shock absorption functions. The gripping force and gripping angle are adjusted based on the sensed pressure changes to avoid damage to the building material surface or slippage. The adaptive gripper 23 monitors pressure data in real time during the gripping process through a closed-loop control system and adaptively adjusts the gripping state. Feedback control and iterative optimization ensure accurate gripping and handling of the target building material. After gripping and gripping the target building material, the intelligent sorting execution module 2 transmits the gripping state and position information to the mobile carrier module 5 via the communication interface to facilitate subsequent handling and path planning operations.
[0053] S70: Acquire the motion data output by the Mecanum wheel module 51 and the global positioning data output by the laser SLAM navigation module 52.
[0054] Specifically, the configured Mecanum wheel module 51 works in conjunction with the laser SLAM navigation module 52 to provide precise motion control and global positioning capabilities. The Mecanum wheel module 51 is composed of multiple Mecanum wheels, each of which is driven by an independent motor to provide motion capabilities in different directions. Specifically, a number of tilted rollers are evenly distributed on the outer ring of the Mecanum wheel. By changing the speed and direction of each motor, the mobile carrier module 5 can achieve all-round movement, including forward, backward, lateral movement, oblique movement and rotation in place. The speed and relative motion direction of each motor are calculated and controlled by the motion controller to achieve the desired motion path and posture changes. During the movement, the Mecanum wheel module 51 detects the rotation speed, rotation direction and tilt angle of each Mecanum wheel in real time through the configured speed sensor and angle sensor, formats and encodes the detected motion data, and transmits it to the deep learning processing module through the data interface. Block 3 is processed and analyzed; at the same time, the laser SLAM navigation module 52 is used to provide global positioning data and environmental perception information through the configured laser radar and inertial measurement unit IMU. The laser radar scans and measures the surrounding environment through a high-speed rotating laser transmitter, and calculates the distance to the surrounding objects by measuring the flight time or phase difference of the laser pulse, thereby generating two-dimensional or three-dimensional point cloud data. The inertial measurement unit provides the posture and displacement information of the mobile carrier module 5 by detecting the acceleration and angular velocity information. The data of the laser radar and the inertial measurement unit are fused and matched, and the current position information is compared and updated with the environmental map through the laser SLAM algorithm, thereby generating accurate global positioning data, including the current position coordinates, movement direction and posture angle. The generated global positioning data is transmitted to the deep learning processing module 3 and the intelligent sorting execution module 2 through the data interface for further optimization and adjustment of path planning and motion control.
[0055] S80: Input the motion data and global positioning data into the deep learning processing module 3 to obtain path planning and positioning control information.
[0056] Specifically, the motion data and the global positioning data are input into the deep learning processing module 3, and the motion data output by the Mecanum wheel module 51 and the global positioning data output by the laser SLAM navigation module 52 are synchronously transmitted to the deep learning processing module 3 through the data interface. The deep learning processing module 3 includes a configured path planning unit and a positioning control unit for performing fusion calculation and real-time optimization of the input data and path planning. In the preprocessing stage of the input data, the motion data and the global positioning data are formatted and normalized, and the data from different sources are aligned and registered. The motion data of the Mecanum wheel module 51 is decomposed and analyzed to extract the current speed, acceleration, steering angle and posture information, and the raw data is converted into a standardized motion state matrix through a feature extraction algorithm. At the same time, the global positioning data output by the laser SLAM navigation module 52 is analyzed and reconstructed to match and compare the current position, target position and environmental map to generate a complete environmental model and a global path map. In the path planning unit, by constructing The deep neural network and reinforcement learning algorithm are used to process and optimize the input data. First, feature extraction and convolution operations are performed on the input motion state matrix and the global path graph. Multi-scale feature extraction and dimensionality reduction operations are performed on the input data through the configured convolution layer and pooling layer to generate a high-dimensional feature map for path planning. On the basis of the feature map, the target points and path nodes of the path planning are gradually inferred and predicted through the constructed full connection layer and activation function. The network parameters are adjusted and optimized through the back propagation algorithm and reinforcement learning strategy to improve the accuracy and efficiency of path planning. In the positioning control unit, the path planning results are compared and corrected with the input data, and the motion trajectory is corrected and adjusted in real time through the configured control algorithm and optimization module to ensure that the mobile carrier module 5 can accurately follow the predetermined trajectory during the path planning process, and perform real-time obstacle avoidance and path replanning when encountering dynamic obstacles or environmental changes. The input data is continuously optimized and iteratively calculated through the deep learning processing module 3 to generate information containing path planning information and positioning control information.
[0057] S90: Input the path planning and positioning control information into the intelligent sorting execution module 2 to control the movement and positioning of the mobile carrying module 5.
[0058] Specifically, the path planning and positioning control information generated by the deep learning processing module 3 is transmitted to the intelligent sorting execution module 2 through the data interface. The intelligent sorting execution module 2 accurately controls the motion state and position of the mobile carrier module 5 according to the received path planning and positioning control information. By analyzing the target position coordinates, motion trajectory and posture angle in the path planning information, the intelligent sorting execution module 2 first analyzes and decomposes the path planning information, divides the overall path into several continuous motion nodes, and extracts and calculates the motion parameters of each node, including information such as moving speed, acceleration, steering angle and motion direction. During the motion control process, the drive motor of the Mecanum wheel module 51 is adjusted and controlled in real time through the constructed control algorithm. By controlling the speed and direction of each Mecanum wheel, the full-dimensional movement and precise positioning of the mobile carrier module 5 are achieved. By calculating the error value between the current motion state and the target position, and dynamically adjusting and compensating the error based on the PID control algorithm, it is ensured that the mobile carrier module 5 can accurately reach the predetermined position. At the same time, through The global positioning data provided by the laser SLAM navigation module 52 is compared and matched, the deviation information between the current mobile position and the target position in the global map is detected, and the path planning information is updated and optimized in real time through the constructed adaptive control algorithm. When encountering dynamic obstacles or environmental changes, the current motion trajectory is adjusted and corrected by introducing obstacle avoidance strategies and path replanning mechanisms, and the priority of the target position and path nodes is recalculated to ensure that the mobile carrier module 5 can move and position according to the optimal path; in the execution process of positioning control, the real-time status information of the mobile carrier module 5 is collected and fed back, including parameters such as position coordinates, speed, acceleration and attitude angle, and compared and analyzed with the predetermined path planning information, and error detection and adjustment are performed on each motion node. The motion accuracy and positioning accuracy of the mobile carrier module 5 are ensured through iterative optimization and feedback control. Finally, the optimized position information and path planning results are transmitted to the intelligent sorting execution module 2 through the communication interface to achieve accurate grasping and handling operations of the target building materials.
[0059] In one embodiment, if Figure 3 As shown, in step S30, the data enhancement unit 33 is used to perform data set enhancement training on the building material positioning data to obtain an enhanced material distribution data set, including: S301: The data enhancement unit 33 performs data set enhancement training on the building material positioning data based on the Mosaic data enhancement method to obtain an enhanced material distribution data set.
[0060] In this embodiment, the Mosaic data enhancement method refers to randomly cropping, scaling, rotating, splicing and fusing building material positioning data of multiple different scenes to generate a training sample set with diversity and complexity.
[0061] Specifically, four groups of different building material positioning data samples are randomly selected from the original building material positioning data set. Each group of data samples includes category information, location coordinates, size information and depth data of the building materials. In order to ensure the diversity and uniformity of the data samples, the categories and sizes are randomized during the sample selection process; the four groups of selected data samples are standardized and preprocessed, including normalization, coordinate alignment and data format conversion for each group of data, so as to facilitate subsequent splicing and fusion operations, and then each data sample is randomly cropped, scaled and rotated. By performing transformation operations of different scales and angles on the input data, the diversity and robustness of the data samples are enhanced. In the cropping operation, the original data is intercepted by a randomly generated cropping window, and the intercepted data is scaled and rotated. The data is transformed by the constructed affine transformation matrix. time transformation and position adjustment; in the data splicing stage, by splicing and mixing the four groups of preprocessed data samples, each data sample is mapped to a quarter area, and the four areas are combined to form a complete enhanced sample image. By splicing and reorganizing data from different sources, training samples with multiple building material objects and complex backgrounds are generated. Subsequently, the spliced samples are subjected to brightness adjustment, noise introduction, color transformation and other operations to further increase the diversity and adaptability of the data samples; in the data set enhancement training process, the generated enhanced samples are input into the deep learning processing module 3 for training and optimization, and the input data is subjected to feature extraction, classification and recognition through the constructed deep neural network, and the network parameters are adjusted and updated through the back propagation algorithm and optimization strategy, thereby generating an enhanced material distribution data set.
[0062] In one embodiment, if Figure 4 As shown, in step S50, the material identification result is input to the CBAM module 32 for optimization processing, and the optimized material identification result is output, including: S501: The CBAM module 32 optimizes the material recognition results based on the channel attention mechanism and the spatial attention mechanism, and outputs the optimized material recognition results.
[0063] In this embodiment, the channel attention mechanism refers to the mechanism that improves the ability to extract and express key features by calculating and optimizing the weights of the input feature maps in the channel dimension. The spatial attention mechanism refers to the mechanism that improves the ability to focus on the target area and the accuracy by calculating and optimizing the weights of the input feature maps in the spatial dimension.
[0064] Specifically, the CBAM module 32 optimizes the material recognition results based on the channel attention mechanism and the spatial attention mechanism, and performs feature extraction and weighted optimization on the input material recognition results to improve the recognition accuracy and feature extraction ability of the model. In the specific implementation process, the input material recognition results are first preprocessed and standardized on the feature map, and the input multi-dimensional feature map is decomposed and formatted, including normalizing the feature map, channel separation and dimensionality adjustment to adapt to the subsequent channel attention and spatial attention calculations; in the channel attention mechanism, by performing global information extraction and weighted calculation on each channel of the input feature map, the feature map is first subjected to global average pooling and global maximum pooling operations to extract the global feature information and maximum response information of each channel, and the data of each channel is compressed and aggregated to generate a channel feature vector for optimization, which is input to the fully connected layer, and the feature vector is nonlinearly transformed and normalized through activation function and regularization operation to generate a channel feature vector for weighted optimization. The channel weight matrix is multiplied element-by-element by the original feature map to enhance the response to important features and suppress irrelevant or redundant features. In the spatial attention mechanism, the spatial dimension of the optimized channel feature map is extracted and optimized. The channel optimized feature map is first convolved and pooled to extract the spatial feature information of different positions. The constructed convolution kernel is used to perform weighted calculation and pixel-by-pixel optimization on each position of the input feature map to generate a weight matrix for spatial optimization. The generated spatial weight matrix is then multiplied by the channel optimized feature map to highlight and optimize important areas and suppress and filter irrelevant areas. In the optimization process, the useful information and useless information in the original feature map are effectively distinguished and extracted by performing multiple iterations and optimization calculations on the input material recognition results. The different dimensions of the feature map are weighted and fused to improve the feature discrimination ability and model accuracy. The optimized material recognition results include the category information, position coordinates, size information and confidence value of the building materials.
[0065] In one embodiment, if Figure 5 As shown, after step S90, a vision-based intelligent building material sorting method further includes: S901: The remote monitoring system evaluates and provides feedback on the optimized material identification results, path planning, and positioning control information to obtain evaluation feedback results.
[0066] In this embodiment, the remote monitoring system refers to a computer system or cloud platform used to monitor, evaluate and optimize the material identification and path planning process in real time.
[0067] Specifically, the optimized material recognition results and path planning and positioning control information are first uploaded to the remote monitoring system through the communication interface. The remote monitoring system parses and stores the received data through the configured network communication module and data processing unit, and classifies and records the category information, position coordinates, size information and confidence value in the material recognition results. At the same time, the target position, motion trajectory, speed, acceleration and posture information in the path planning and positioning control information are formatted and standardized for subsequent evaluation and analysis. During the evaluation process, the remote monitoring system uses the constructed data analysis algorithm to perform multi-dimensional comparison and calculation on the input data, and evaluates the accuracy, stability and consistency of the material recognition results and path planning information, compares the material recognition results with the pre-set standard value or target value, and calculates the recognition accuracy and error rate. At the same time, the path planning and positioning control information are evaluated. The deviation between the motion trajectory in the position control information and the target position is calculated and analyzed, and the rationality and accuracy of the path planning are judged by calculating the error value and cumulative error value of each node; in the feedback process, the remote monitoring system organizes and outputs the evaluation results through the constructed feedback control mechanism, and quantifies and models the evaluation results of the recognition results and the path planning information to generate evaluation feedback results including error distribution, accuracy statistics and optimization suggestions. The generated evaluation feedback results are transmitted to the deep learning processing module 3 and the intelligent sorting execution module 2 through the data interface for optimizing and adjusting the model parameters and control strategies. In this process, the remote monitoring system monitors and iteratively optimizes the recognition and control processes in real time through the constructed adaptive feedback mechanism, thereby improving the stability and accuracy of the system, and ensuring that the recognition and handling operations of the target building materials can be efficiently completed in complex or dynamic environments.
[0068] S902: Based on the evaluation feedback results, optimization information is generated and input into the deep learning processing module 3 and the intelligent sorting execution module 2 for optimization training to obtain the optimized training model and execution control strategy.
[0069] Specifically, by parsing and extracting the evaluation feedback results output by the remote monitoring system, including information such as recognition accuracy, path planning accuracy, positioning control accuracy and error distribution, the evaluation feedback results are decomposed and formatted through the constructed data processing unit, and the relevant data of recognition accuracy and path planning accuracy are extracted, and normalized and standardized for subsequent optimization training and parameter adjustment. In the process of generating optimization information, the error data in the evaluation feedback results are statistically analyzed and modeled to construct an error matrix containing recognition error and path planning error, and the error matrix is analyzed and optimized through the constructed optimization algorithm, including least squares optimization, gradient descent optimization and adaptive learning algorithm, etc., and the error distribution is fitted and solved to generate optimization information for optimization training; the generated optimization information is input into the deep learning processing module 3 and the intelligent sorting execution module 2, and the input optimization information is trained and optimized through the constructed deep learning model. In the process, the parameters of the YOLO algorithm processing unit 31, the CBAM module 32 and the data enhancement unit 33 are adjusted and optimized, the network weights are updated and corrected by the back propagation algorithm, and the errors in the optimization information are quantified and fed back by the constructed loss function, so as to gradually reduce the recognition error and positioning error. At the same time, in the intelligent sorting execution module 2, the control strategies of the six-axis collaborative robot arm 21, the vacuum suction cup 22 and the adaptive clamp 23 are optimized and adjusted, and the input optimization information is corrected and adjusted in real time by introducing the PID control algorithm, the adaptive control algorithm and the path replanning algorithm to generate an optimized execution control strategy and path planning parameters. The model and control strategy are continuously improved and optimized by iterative training and feedback optimization to generate an optimized training model and execution control strategy, and the optimization results are transmitted to the execution control end of the deep learning processing module 3 and the intelligent sorting execution module 2 through the data interface for subsequent recognition, classification and grasping operations.
[0070] S903: Input the optimized training model into the YOLO algorithm processing unit 31 and the CBAM module 32 to update the recognition algorithm and optimization strategy to obtain an updated material recognition result. The updated material recognition result is used to improve the recognition accuracy and target positioning accuracy of the YOLO algorithm processing unit 31 and the CBAM module 32.
[0071] Specifically, the optimized training model is input into the YOLO algorithm processing unit 31 and the CBAM module 32, and the optimized training model and the updated network parameters are transmitted to the YOLO algorithm processing unit 31 and the CBAM module 32 in the deep learning processing module 3 through the data interface, so as to update and improve the recognition algorithm and the optimization strategy. In the specific implementation process, the optimized training model is formatted and standardized, including updating and replacing the network weights, bias values, convolution kernel parameters and activation function parameters, and the optimized training model is compared and fused with the original model to ensure that the structure and logical integrity of the original model are maintained during the model update process; in the YOLO algorithm processing unit 31, the optimized training model is loaded into the deep neural network, and the input material distribution data set is subjected to feature extraction, detection and classification, the input data is subjected to multi-scale feature extraction and downsampling operations through the constructed convolution layer and pooling layer, the classification results are calculated and output through the fully connected layer and the activation function, and the network weights are updated and Adjustment is performed to improve the recognition accuracy and classification accuracy of the target building materials by the YOLO algorithm processing unit 31; in the CBAM module 32, by inputting the optimized training model into the channel attention mechanism and the spatial attention mechanism, the optimized material recognition results are subjected to multi-dimensional feature optimization and enhancement. In the channel attention mechanism, the input feature map is compressed and aggregated through global average pooling and global maximum pooling operations, and the feature vector is weightedly calculated and optimized through the fully connected layer and the activation function, and the generated weight matrix is multiplied element-by-element with the input feature map, thereby enhancing the response of important features and suppressing irrelevant features; in the spatial attention mechanism, the input feature map is subjected to spatial feature extraction and optimization through the constructed convolution kernel and pooling layer, and the extracted spatial feature information is fused and weighted with the original feature to generate the optimized material recognition result; in the optimization process, the input material recognition result is subjected to multiple iterations and feedback optimization, and the useful information and useless information in the original feature map are effectively distinguished and extracted, thereby improving the model's recognition accuracy and target positioning accuracy for the target building materials.
[0072] S904: The optimized execution control strategy is input into the intelligent sorting execution module 2, which updates the execution control parameters for sorting and handling operations and generates updated path control information. This updated path control information is then used to optimize the motion path and grasping accuracy of the intelligent sorting execution module 2. Specifically, the optimized execution control strategy is input into the intelligent sorting execution module 2, and the optimized execution control strategy and updated control parameters are transmitted to the intelligent sorting execution module 2 through the data interface, so as to update and optimize the control logic and execution parameters of the sorting and handling operations in real time. In the specific implementation process, the optimized execution control strategy is formatted and standardized, including the decomposition and reorganization of the control parameters, path planning information and grasping strategy, and the optimized execution control strategy is docked and integrated with the control algorithm of the intelligent sorting execution module 2 to ensure the stability and consistency of the system during the control strategy update process; in the sorting and handling operations, the execution control parameters of the six-axis collaborative robot 21, vacuum suction cup 22 and adaptive clamp 23 in the intelligent sorting execution module 2 are optimized and adjusted, and the motion path, speed, acceleration and posture angle of the robot are replanned and adjusted to improve the grasping accuracy and handling efficiency of the target building materials. Specifically, the optimized execution control strategy is analyzed and decomposed by the constructed path planning algorithm and motion control algorithm, and the entire sorting and handling operation process is divided into several The intelligent sorting execution module 2 is configured to perform step-by-step adjustment and optimization of the control parameters of each node, including the posture control of the end effector, path tracking, and adjustment of the clamping force. In the path planning process, the optimized execution control strategy is integrated with the motion data and global positioning data of the Mecanum wheel module 51 and the laser SLAM navigation module 52, and the path planning information is corrected and adjusted in real time through the PID control algorithm and the adaptive control algorithm, thereby generating optimized path control information. By gradually optimizing and iteratively updating the motion path and positioning accuracy, the intelligent sorting execution module 2 is able to accurately grasp and transport target building materials in a complex or dynamic environment. In the process of updating the execution control parameters, the error data is modeled and optimized by comparing and analyzing the feedback information with the real-time monitoring data, and the execution control strategy is iteratively trained and adjusted through the back propagation algorithm and the optimization strategy. The generated updated path control information includes comprehensive information of the optimized motion path, grasping accuracy, and transport efficiency, and is transmitted to the intelligent sorting execution module 2 through the data interface for guiding subsequent sorting and transport operations.
[0073] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0074] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0075] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A vision-based intelligent sorting system for building materials, characterized by: The vision-based intelligent building material sorting system comprises a multimodal vision module (1), an intelligent sorting execution module (2), a deep learning processing module (3), a fast calibration module (4) and a mobile carrying module (5), wherein: The multimodal vision module (1) is used to perform data communication with the deep learning processing module (3), and the multimodal vision module (1) is connected to the deep learning processing module (3); The intelligent sorting execution module (2) is used to perform data communication with the deep learning processing module (3), and the intelligent sorting execution module (2) is connected to the mobile carrying module (5); The deep learning processing module (3) is used to perform data communication with the multimodal vision module (1), and the deep learning processing module (3) is used to perform data communication with the intelligent sorting execution module (2) and the mobile carrying module (5); The rapid calibration module (4) is used to perform data communication with the deep learning processing module (3) and the intelligent sorting execution module (2); The communication interface of the mobile carrying module (5) is used for data communication with the intelligent sorting execution module (2) and the deep learning processing module (3).
2. The vision-based intelligent building material sorting system according to claim 1, characterized in that: The multimodal vision module (1) comprises a binocular stereo camera (11), a ToF depth sensor (12), a ring fill light (13) and an anti-glare curtain assembly (14); the binocular stereo camera (11) and the ToF depth sensor (12) work in coordination; and the multimodal vision module (1) is connected to the deep learning processing module (3).
3. The vision-based intelligent building material sorting system according to claim 1, characterized in that: The intelligent sorting execution module (2) comprises a six-axis collaborative robot arm (21), a vacuum suction cup (22) and an adaptive clamp (23), wherein the adaptive clamp (23) comprises a pressure sensor (231) and a silicone damping layer (232), the intelligent sorting execution module (2) is connected to the deep learning processing module (3), and the intelligent sorting execution module (2) is connected to the mobile carrying module (5).
4. The vision-based intelligent building material sorting system according to claim 1, characterized in that: The deep learning processing module (3) comprises a YOLO algorithm processing unit (31), a CBAM module (32) and a data enhancement unit (33); the deep learning processing module (3) is connected to the multimodal vision module (1); the deep learning processing module (3) is connected to the intelligent sorting execution module (2); and the deep learning processing module (3) is connected to the mobile carrying module (5).
5. The vision-based intelligent building material sorting system according to claim 1, characterized in that: The rapid calibration module (4) comprises a nine-point nonlinear calibration unit (41) and a calibration data output interface (42), and the calibration data output interface (42) is connected to the deep learning processing module (3) and the intelligent sorting execution module (2) respectively.
6. The vision-based intelligent building material sorting system according to claim 1, characterized in that: The mobile carrying module (5) comprises a Mecanum wheel module (51), a laser SLAM navigation module (52) and a communication interface (501). The Mecanum wheel module (51) and the laser SLAM navigation module (52) work in conjunction with each other to control the movement and positioning of the mobile carrying module (5). The communication interface (501) is connected to the deep learning processing module (3) and the intelligent sorting execution module (2) respectively.
7. A vision-based intelligent sorting method for building materials, applied to a vision-based intelligent sorting system for building materials as claimed in any one of claims 1 to 6, characterized in that: The vision-based intelligent building material sorting method includes: Acquire first image data output by the binocular stereo camera (11) and second depth data output by the ToF depth sensor (12), input the first image data and the second depth data into the YOLO algorithm processing unit (31) for preliminary recognition and positioning, and obtain three-dimensional material distribution data; Obtaining calibration data output by the nine-point nonlinear calibration unit (41), and performing coordinate correction and error correction on the three-dimensional material distribution data based on the calibration data to obtain building material positioning data; Using the data enhancement unit (33) to perform data set enhancement training on the building material positioning data to obtain an enhanced material distribution data set; Identifying and classifying the enhanced material distribution data set by the YOLO algorithm processing unit (31) to obtain a material identification result; Inputting the material identification result into the CBAM module (32) for optimization processing, and outputting the optimized material identification result; The optimized material identification result is input into the intelligent sorting execution module (2), and the intelligent sorting execution module (2) grasps and carries the target building material through the six-axis collaborative robot arm (21), the vacuum suction cup (22) and the adaptive clamp (23); Acquiring motion data output by the Mecanum wheel module (51) and global positioning data output by the laser SLAM navigation module (52); Inputting the motion data and the global positioning data into the deep learning processing module (3) to obtain path planning and positioning control information; The path planning and positioning control information is input into the intelligent sorting execution module (2) to control the movement and positioning of the mobile carrying module (5).
8. The method of intelligent building material sorting based on vision according to claim 7, characterized in that: The data enhancement unit (33) is used to perform data set enhancement training on the building material positioning data to obtain an enhanced material distribution data set, including: The data enhancement unit (33) performs data set enhancement training on the building material positioning data based on the Mosaic data enhancement method to obtain the enhanced material distribution data set.
9. The method of intelligent building material sorting based on vision according to claim 7, characterized in that: The material identification result is input into the CBAM module (32) for optimization processing, and the optimized material identification result is output, including: The CBAM module (32) optimizes the material recognition result based on the channel attention mechanism and the spatial attention mechanism, and outputs the optimized material recognition result.
10. The vision-based intelligent sorting method for building materials according to claim 7, characterized in that: The vision-based intelligent building material sorting method further includes: The remote monitoring system evaluates and provides feedback on the optimized material identification result and the path planning and positioning control information to obtain an evaluation feedback result; Based on the evaluation feedback results, optimization information is generated, and the optimization information is input into the deep learning processing module (3) and the intelligent sorting execution module (2) for optimization training to obtain an optimized training model and execution control strategy; Inputting the optimized training model into the YOLO algorithm processing unit (31) and the CBAM module (32) to update the recognition algorithm and optimization strategy, thereby obtaining an updated material recognition result, wherein the updated material recognition result is used to improve the recognition accuracy and target positioning accuracy of the YOLO algorithm processing unit (31) and the CBAM module (32); The optimized execution control strategy is input into the intelligent sorting execution module (2) to update the execution control parameters of the sorting and handling operations, thereby obtaining updated path control information. The updated path control information is used to optimize the motion path and grasping accuracy of the intelligent sorting execution module (2).
Citation Information
Patent Citations
Intelligent handling robotic arm system based on 3D vision and deep learning, and using method
CN111496770A
Material box stacking method, device and equipment based on AGV and storage medium
CN115123839A
Visual classification garbage automatic sorting system based on deep learning and self-adaptive grabbing
CN118385157A
Multi-article identification and distinguishing method, system, equipment and medium
CN119580017A
Garbage sorting robot based on deep learning
CN119795125A
Cited By
Kitchen service robot operation method based on VLN large model
CN120902023A