A mobile machine and its unmanned perception mapping system for unstructured scenes

By integrating a computing platform and a multi-sensor system onto mobile machinery, environmental data is preprocessed and dynamically detected to generate a global semantic map. This solves the problem of mapping mobile machinery in unstructured scenarios and enables high-precision unmanned driving control.

CN116168075BActive Publication Date: 2026-01-16HUAQIAO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310308010.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-27
Publication Date
2026-01-16
Estimated Expiration
2043-03-27

Smart Images

  • Figure CN116168075B_ABST
    Figure CN116168075B_ABST
Patent Text Reader

Abstract

The application provides a mobile machine and an unmanned driving perception mapping system for a non-structured scene, which introduces environmental elevation data to assist in accurate segmentation tasks between non-structured targets, improves the low edge discrimination ability and inaccurate recognition of traditional perception mapping systems in non-structured environments, and introduces instance segmentation mask information to assist in front-end mapping to solve the problems of low mapping robustness caused by low matching efficiency of traditional mapping SLAM in non-structured scenes, lack of human-computer interaction ability, and dynamic obstacle interference, thereby enhancing the perception ability of the mobile machine in the non-structured unmanned driving operation scene.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of unmanned perception mapping, in particular to a mobile machine and an unmanned perception mapping system for unstructured scenes thereof. BACKGROUND

[0002] The operation of various types of excavators, various types of loaders and other construction, water conservancy and mining operation machines in unstructured scenes on the market is often accompanied by high vibration, high temperature difference, high dust and other phenomena, which makes the driver long-term work in a high-risk operation environment. At the same time, the uneven technical skills of the drivers also make it difficult to guarantee the work efficiency. Therefore, the unmanned development of the current mobile machine has become a key problem.

[0003] With the development of electric mobility of mobile machines, great improvements have been made in the operation mechanism, actuating mechanism and driving force mode of mobile machines, which lays a good foundation for the unmanned development of mobile machines. In combination with the use of unmanned control planning algorithm on passenger cars, mobile machines have also achieved certain results. However, the deep learning perception and SLAM mapping technology used by the current unmanned control system of mobile machines in the environment perception is mostly applied to structured road scenes such as highways and urban streets in good weather conditions. The road environment structure of this type of road is obvious and easy to classify and extract target objects. However, the actual operation environment of mobile machines is mostly unstructured scenes such as muddy, wet and dusty scenes. At this time, the traditional perception technology cannot be well applied to the high dynamic, high dust and muddy chaotic unstructured operation scene of mobile machine unmanned driving. The traditional perception based on deep learning cannot accurately extract the edge features of the unstructured environment, and the traditional SLAM system cannot achieve good mapping effect in the high dynamic unstructured scene lacking regular point and line features.

[0004] Therefore, the present application is proposed. SUMMARY

[0005] Therefore, the present application is proposed.

[0006] The application discloses an unmanned driving perception mapping system for an unstructured scene.

[0007] The control platform comprises an electric proportional pressure reducing valve, left and right traveling motors, a multi-way valve and a vehicle controller, the output end of the computing platform is electrically connected with the input end of the vehicle controller, and the output end of the vehicle controller is electrically connected with the control end of the electric proportional pressure reducing valve, the control end of the left and right traveling motors and the control end of the multi-way valve.

[0008] The computing platform is configured to realize the following steps by executing a computer program stored therein.

[0009] The original environment perception acquisition component is subjected to time synchronization and space alignment processing, and surrounding environment data collected by the original environment perception acquisition component is acquired.

[0010] The surrounding environment data is preprocessed to generate preprocessed data, wherein the preprocessed data comprises segmentation mask information, preprocessed point cloud data, environment elevation information and accurate pose information of the mobile machine.

[0011] The segmentation mask information and the preprocessed point cloud data are subjected to dynamic detection and semantic mapping processing to generate dynamic detection results and point cloud data with semantic information.

[0012] The dynamic detection results, the point cloud data with semantic information and the accurate pose information of the mobile machine are subjected to optimization processing to generate a global semantic map.

[0013] The control platform is controlled according to the global semantic map to realize unmanned driving of the mobile machine.

[0014] Preferably, the original environment perception acquisition component comprises a camera, a laser radar sensor, an IMU pose sensor and a GNSS positioning sensor.

[0015] The output end of the camera, the output end of the laser radar sensor, the output end of the IMU pose sensor and the output end of the GNSS positioning sensor are electrically connected with the input end of the computing platform.

[0016] The camera is configured to collect image data of a working environment of the mobile machine.

[0017] The laser radar sensor is configured to collect point cloud data information and environment depth information of the working environment of the mobile machine.

[0018] The IMU pose sensor is configured to collect relative positioning and attitude information of the mobile machine in an initial state.

[0019] The GNSS positioning sensor is configured to collect centimeter-level positioning information of the mobile machine.

[0020] Preferably, the surrounding environment data collected by the original environment perception acquisition component is obtained, specifically:

[0021] Image data of the working environment of the mobile machine collected by the camera is obtained.

[0022] Point cloud data information and environment depth information of the working environment of the mobile machine collected by the laser radar sensor are obtained.

[0023] The IMU pose sensor is configured to collect relative positioning and attitude information of the mobile machine in an initial state.

[0024] The GNSS positioning sensor is configured to collect centimeter-level positioning information of the mobile machine.

[0025] Preferably, the surrounding environment data is preprocessed to generate preprocessed data, specifically:

[0026] The point cloud data information is filtered to remove noise and is subjected to clustering and segmentation processing to generate preprocessed point cloud data.

[0027] The environment depth information is preprocessed to generate environment elevation information.

[0028] Preferably, the surrounding environment data is preprocessed to generate preprocessed data, and further includes:

[0029] The image data is subjected to filtering processing, normalization, and information enhancement processing to generate image pixel data.

[0030] A corresponding relationship between the image pixel data and the environment elevation information is established, different elevation values between different target objects are used to pre-classify the image pixel data, and classified image pixel data is generated.

[0031] The classified image pixel data is subjected to segmentation network processing to generate segmentation mask information.

[0032] Preferably, the surrounding environment data is preprocessed to generate preprocessed data, and further includes:

[0033] The relative positioning, the attitude information, and the centimeter-level positioning information are subjected to tight coupling calculation and error correction to generate accurate pose information of the mobile machine.

[0034] Preferably, the segmentation mask information and the pre-processed point cloud data are subjected to dynamic detection and semantic mapping processing to generate dynamic detection results and point cloud data with semantic information, specifically:

[0035] The segmentation mask information and the pre-processed point cloud data are subjected to semantic mapping processing to generate point cloud data with semantic information.

[0036] The segmentation mask information is subjected to dynamic detection using an optical flow method to determine the state of the segmentation mask information.

[0037] When it is determined that the state of the segmentation mask information is static, dynamic detection results are generated.

[0038] When it is determined that the state of the segmentation mask information is dynamic, dynamic interference information in the segmentation mask information is removed to generate dynamic detection results.

[0039] Preferably, the dynamic detection results, the point cloud data with semantic information, and the precise pose information of the mobile machine are subjected to optimization processing to generate a global semantic map, specifically:

[0040] The point cloud data with semantic information and the precise pose information of the mobile machine are subjected to laser odometry processing to generate optimization data.

[0041] The optimization data and the dynamic detection results are subjected to loop detection update iteration processing to generate a global semantic map.

[0042] The application also discloses a mobile machine comprising a vehicle body and an unmanned driving perception mapping system for an unstructured scene according to any one of the above, wherein the unmanned driving perception mapping system for an unstructured scene is arranged on the vehicle body.

[0043] Preferably, it further comprises a horn and front and rear vehicle lights arranged on the vehicle body, wherein the control end of the horn and the control end of the front and rear vehicle lights are electrically connected to the output end of the vehicle controller.

[0044] In summary, the mobile machine and the unmanned driving perception mapping system for an unstructured scene thereof provided by the present embodiment use environmental elevation information to assist in segmenting the boundaries between unstructured environmental target objects, use segmentation mask information to assist in inter-frame matching at the front end of mapping, and use segmentation mask information for dynamic object detection, which can effectively improve the fusion perception mapping of the mobile machine in an unstructured scene with dynamic and irregular environmental scenes, ensure high-precision perception mapping of the machine in an unstructured scene, and improve the interaction capability with the environment. Thus, the problem of single adjustment of the cleaning operation state, poor effect, and waste of manpower and resources in the prior art automatic adjustment cleaning vehicle scheme is solved. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 is a structural schematic diagram of an unmanned driving perception mapping system for an unstructured scene provided by an embodiment of the present application.

[0046] Figure 2 is a flow schematic diagram of an unmanned driving perception mapping system for an unstructured scene provided by an embodiment of the present application.

[0047] Figure 3 is a network flow schematic diagram of an unmanned driving perception mapping system for an unstructured scene provided by an embodiment of the present application.

[0048] Figure 4 is an improved boundary enhanced segmentation network structure schematic diagram provided by an embodiment of the present application.

[0049] Figure 5 is a dynamic detection flow schematic diagram provided by an embodiment of the present application.

[0050] Figure 6 is a semantic map construction flow schematic diagram provided by an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0052] The specific embodiments of the present application will be described in detail below with reference to the drawings.

[0053] Please refer to Figures 1 to 3 The first embodiment of the present application provides an unmanned driving perception mapping system for an unstructured scene, comprising: a computing platform 5, a control platform 11, and a raw environment perception acquisition component;

[0054] The control platform 11 includes an electric proportional pressure reducing valve 8, left and right traveling motors 9, a multi-way valve 10, and a vehicle controller 12. The output end of the computing platform 5 is electrically connected to the input end of the vehicle controller 12. The output end of the vehicle controller 12 is electrically connected to the control end of the electric proportional pressure reducing valve 8, the control end of the left and right traveling motors 9, and the control end of the multi-way valve 10.

[0055] Specifically, in the embodiment, the original environment perception acquisition assembly includes a camera 2, a laser radar sensor 3, an IMU pose sensor 13, and a GNSS positioning sensor 4.

[0056] The output end of the camera 2, the output end of the laser radar sensor 3, the output end of the IMU pose sensor 13, and the output end of the GNSS positioning sensor 4 are electrically connected to the input end of the computing platform 5.

[0057] The camera 2 is configured to acquire image data of a mobile mechanical working environment.

[0058] The laser radar sensor 3 is configured to acquire point cloud data information and environmental depth information of the mobile mechanical working environment.

[0059] The IMU pose sensor 13 is configured to acquire relative positioning and attitude information of an initial state of the mobile machine.

[0060] The GNSS positioning sensor 4 is configured to acquire centimeter-level positioning information of the mobile machine.

[0061] The computing platform 5 is configured to execute a computer program stored therein to implement the following steps.

[0062] S101, time synchronization and space alignment processing are performed on the original environment perception acquisition assembly, and surrounding environment data collected by the original environment perception acquisition assembly is acquired.

[0063] Specifically, step S101 includes acquiring image data of a mobile mechanical working environment collected by the camera 2.

[0064] Point cloud data information and environmental depth information of a mobile mechanical working environment collected by the laser radar sensor 3 are acquired.

[0065] Relative positioning and attitude information of an initial state of the mobile machine collected by the IMU pose sensor 13 are acquired.

[0066] Centimeter-level positioning information of the mobile machine collected by the GNSS positioning sensor 4 is acquired.

[0067] In the embodiment, when the unmanned perception mapping system of the unstructured scene receives a start signal, the original environment perception acquisition component acquires unstructured working scene information through opening the environment information acquisition sensors arranged on the mobile machine after spatial alignment and time synchronization, including environment image pixel information, environment elevation information, three-dimensional point cloud data and the state and pose information of the mobile machine, and inputs the information into the data processing module of the computing platform 5 for processing. The original environment perception acquisition component uses a target-based calibration method for spatial alignment between sensors and uses an Ethernet protocol time synchronization generator for time synchronization between sensors. In simple terms, the time stamps of each sensor are obtained, and then the GNSS positioning sensor 4 is used as a reference time for time hard synchronization with the laser radar sensor 3, and the camera 2 and the laser radar sensor 3 use soft synchronization of the Ethernet protocol, so as to synchronize the time stamps between different sensors in a combination of soft and hard synchronization, calibrate the internal and external parameters of the sensors to obtain the internal parameters of the sensors and then obtain the rotation and translation matrix between the sensors, and align the space of different sensors according to the obtained external parameter, wherein the sensor platform includes a camera, a laser radar, an IMU and a GNSS.

[0068] S102, preprocessing the surrounding environment data to generate preprocessed data, wherein the preprocessed data includes segmentation mask information, preprocessed point cloud data, environment elevation information and accurate pose information of the mobile machine;

[0069] Specifically, step S102 includes filtering and removing noise from the point cloud data information, and generating preprocessed point cloud data through clustering and segmentation processing.

[0070] The environment depth information is preprocessed to generate environment elevation information.

[0071] Please refer to Figure 4, specifically, in the present embodiment, the computing platform 5 can be divided into a training unit and a real-time processing unit, wherein the training unit includes a data processing module, and the real-time processing unit includes a dynamic detection module and an environment mapping module; wherein data processing includes data preprocessing, such as image preprocessing, point cloud data preprocessing, environment elevation data acquisition, image pixel information acquisition, data association of elevation data and pixel data, IMU pre-integration solution, elevation image segmentation task and semantic information mapping of image point cloud; the elevation information is used to pre-classify the image pixel points according to the elevation information difference. The data processing module uses an embedded computing platform to perform image preprocessing on image pixel data; the environment depth map collected by the laser radar is converted and output as environment elevation information data, the image data is associated with the environment elevation information data, the image pixel points are pre-classified, the pre-classified image pixel point information is input into the improved boundary enhanced instance segmentation network for training, and the boundary segmentation accurate mask information is generated, that is, the image pixel information of the environment after preprocessing is associated with the elevation information; the elevation information in the unstructured environment is used to assist in segmenting the target object in the unstructured scene, the elevation image pixel associated data is segmented to generate the segmentation mask of the unstructured environment information, and the segmentation mask information of the environment information is mapped with the processed point cloud data to generate instance fusion data for environment mapping, that is, the point cloud data is preprocessed, the mask generated by the segmentation network and the preprocessed point cloud data are mapped to generate semantic point cloud data, so as to realize efficient use of the point cloud data; the IMU and GNSS are jointly solved by using extended Kalman filtering to realize high-precision self-position information.

[0072] The dynamic obstacle is identified and processed by the dynamic detection module to optimize the mapping, and the segmentation mask result information output by the elevation image training in the data processing module is input to the dynamic detection module. By comparing the segmentation mask category information, the dynamic object is identified, and the result is input to the back-end optimization part of the mapping module for loop detection to update and correct the map; that is, the optical flow method is used to calculate the optical flow field of the object to identify the dynamic target, including construction personnel, mobile machinery, etc. The environment mapping module introduces segmentation information to assist inter-frame matching in the front-end odometer, and introduces dynamic detection results for back-end optimization to obtain unstructured environment information and self-position information. The results are input to the subsequent planning and control module for transmission of control instructions, and the control instructions are sent to the vehicle execution unit, which includes an electric proportional pressure reducing valve, a multi-way valve, and left and right walking motors, which in turn drive the vehicle to walk, avoid obstacles, and warn, etc. The environment mapping module processes the environment data collection to achieve robust semantic map construction of unstructured scenes by using segmentation information in the instance point cloud data to assist inter-frame matching optimization; the segmentation mask data, semantic point cloud data, and pose information in the data processing module, and the detection results in the dynamic detection module are input to the front-end odometer, back-end optimization, and loop detection process in the environment mapping module. By using the correspondence between the segmentation mask information and the point cloud data to assist inter-frame matching, the semantic map of the unstructured scene is constructed to provide accurate unstructured scene maps for the subsequent planning and control module.

[0073] The laser radar sensor 3 acquires point cloud data, and pre-processes the synchronized point cloud data, which includes filtering to remove noise, clustering and segmentation, etc. to reduce the occupancy rate of the point cloud to improve the utilization efficiency of the point cloud, including generating environment elevation information using the depth information collected by the laser radar, and pre-classifying the image pixel data according to the generated elevation information. The pre-processed point cloud data is mapped with semantic information.

[0074] In this embodiment, the data processing module includes original data preprocessing, image pixel point pre-classification, image data segmentation task, and semantic mapping. The data processing module acquires point cloud data information and environment depth information of the working environment of the mobile machinery through the laser radar sensor 3, performs clustering filtering processing on the point cloud original data information, sets an effective range, retains the point cloud data information within the effective range, and further improves the utilization efficiency of the point cloud data. The depth information collected by the laser radar sensor 3 is vectorized to generate terrain elevation data information of the working environment; wherein the point cloud data generates depth information to establish an environment elevation map, and the point cloud data is processed by clustering, filtering, and point cloud pruning to improve the utilization efficiency of the point cloud data.

[0075] The mobile machine and its unmanned perception mapping system of unstructured scenes introduce environmental elevation data to assist in the accurate segmentation task of the boundary between non-structured targets, improve the low edge discrimination ability and inaccurate recognition of traditional perception mapping systems in unstructured environments, introduce instance segmentation mask information to assist in front-end mapping to solve the problems of low mapping robustness caused by low matching efficiency of traditional mapping SLAM in unstructured scenes, lack of human-computer interaction ability and dynamic obstacle interference, and further enhance the perception ability of mobile machines in unstructured unmanned operation scenes.

[0076] Filtering, normalizing, and information enhancing the image data to generate image pixel data;

[0077] Establishing a corresponding relationship between the image pixel data and the environmental elevation information, using different elevation values between different targets to pre-classify the image pixel data to generate classified image pixel data;

[0078] Segmentation network processing of the classified image pixel data to generate segmentation mask information.

[0079] Specifically, in the present embodiment, the data processing module acquires work environment image pixel information through the camera 2, and performs image preprocessing operations such as normalization, uniform cropping, noise filtering, and image segmentation annotation; that is, the camera 2 acquires image data, performs image preprocessing after time synchronization of the image data, establishes a corresponding relationship between elevation information and image pixel data according to the elevation information data, pre-classifies image pixels using different elevation values between different targets, inputs the pre-classified image pixel information into a designed lightweight instance segmentation network for instance segmentation to generate target class mask information, wherein the preprocessing includes filtering, normalization, and information enhancement to improve the feature information of the picture, and the obtained segmentation mask information is used for mapping and dynamic object detection.

[0080] The image pixel pre-classification uses a spatial nearest neighbor method to associate and match image pixel information with terrain elevation information, uses the boundary difference of terrain elevation information of different target objects to classify the target objects, and uses the classification result to pre-classify the image pixels; wherein the initial pose information of the mobile machine is used to associate the nearest neighbor data of the elevation information data and the image pixel data, the elevation data of the elevation information is used as a classification reference to pre-classify the corresponding image pixels, so that the division boundary between the image classes is more accurate. The image data segmentation task uses the associated information of the environmental elevation information data and the image pixel data as input to improve the boundary enhanced instance segmentation network, uses the elevation data to assist in training the image data, and uses the environmental elevation information to assist in accurately distinguishing the target objects in the irregular environment in the unstructured scene. The improved boundary enhanced instance segmentation network consists of a decoding part and an encoding part; that is, the image data segmentation task is to input the pre-classified image pixel data into the improved boundary enhanced segmentation network in the training unit of the computing platform 5 for segmentation mask model training, and input the trained segmentation model into the real-time processing unit of the computing platform 5 for real-time prediction of environmental segmentation mask information, and input the real-time segmentation prediction result into the semantic mapping module. The decoding part includes an input layer, a main network layer, a hollow convolution layer, a multi-inflation coefficient convolution layer, a cyclic convolution layer and a pooling layer; the encoding part includes a convolution layer, an up-sampling layer, a feature fusion layer and a feature output layer; and the classification mask information of the output pixel points.

[0081] In this embodiment, the image pixel pre-classification uses the point cloud depth information collected by the laser radar sensor 3 to restore the environmental terrain elevation information data through feature vectorization; the image pixel data is extracted from the environmental image information collected by the camera, and the coordinate center is the center of the vehicle body. The Euclidean distance between the camera 2 and the pixel point and the laser radar sensor 3 and the point cloud data is calculated, and they are associated together, and then the terrain elevation information of different target objects in the environment has a certain difference, and the terrain elevation information is classified according to the difference, and the corresponding image pixel data is associated, and the corresponding image pixel data is pre-classified according to the classification result of the elevation information, and the boundary segmentation effect of the image information in the segmentation task network is improved.

[0082] The improved boundary enhanced segmentation network model is composed of a feature input layer, a deep convolutional neural network layer, a multi-scale dilated convolution layer, a multi-head self-attention mechanism module, an LSTM network layer, a feature fusion layer and a pooling layer. After the pixel information of the preprocessed image is input into the feature input layer for feature extraction, the feature map is input into the deep convolutional neural network for convolution feature extraction, and high feature map and low feature map are output respectively. The high feature map is input into the multi-scale dilated convolution layer. The multi-scale dilated convolution layer adaptively adjusts the dilated factor size according to the pre-classification prior information of the input pixel points. In the position where the pixel feature difference is large, that is, the adjacent position of the target object, the dilated factor is increased to maintain the size of the original input feature map while expanding the receptive field at the boundary. At the same time, the multi-head self-attention mechanism module is introduced in the place where the pixel feature changes obviously. The multi-head self-attention module includes 4 convolution layers, 2 Dropout layers, a Softmax layer and a Concat layer pair. After the continuous pixel boundary sequence is extracted and trained by increasing the weight, a global average pooling operation is performed on the features, and then the features are fused and output. Then a 2*2 convolution layer is used to adjust the number of feature channels, and a feature map with high semantic information is output. After 4 times down-sampling operation is performed on the output feature map, the feature map is stacked and fused with the low feature map processed and output; the low feature map is input into a 2*2 ordinary convolution layer, and after convolution dimension increasing operation, it is input into the LSTM network layer. The LSTM network layer is composed of an input gate, a forget gate, an output gate and an internal memory unit. When the input image pixel point is in the input sequence feature change obvious area, it passes through the input gate to the internal memory storage unit, and then outputs through the output gate. When the input pixel point is the internal pixel point of the target object, it passes through the input gate and the forget gate to output, reducing the calculation intensity of network training. After the feature map output from the output gate is adjusted by a 2*2 convolution channel, it is stacked and fused with the high semantic information feature map after down-sampling. After the fused and output feature map is input into a 3*3 convolution network for feature extraction, 4 times up-sampling operation is performed for feature size adjustment, and the classification mask information of the pixel point is output.

[0083] The relative positioning, the attitude information and the centimeter-level positioning information are tightly coupled to solve errors and generate accurate pose information of the mobile machine.

[0084] Specifically, in the embodiment, the GNSS positioning sensor 4 outputs the position information of the mobile machine through satellite and differential joint positioning, that is, the GNSS positioning sensor 4 acquires centimeter-level positioning information, the IMU pose sensor 13 provides initial position information of the mobile machine and initial pose information of the vehicle body, that is, the IMU pose sensor 13 acquires relative positioning and attitude information in the initial state; the IMU pose sensor 13 performs pre-integration processing, and the GNSS positioning sensor 4 and the IMU pose sensor 13 jointly solve and output the pose information, which are tightly coupled and solved through Kalman filtering, continuously correct errors, and acquire accurate pose information, which is used as priori pose information of the front-end odometry to improve the accuracy of inter-frame matching; that is, the GNSS positioning information and the IMU pose information are tightly coupled and solved through extended Kalman filtering in the real-time processing module of the computing platform 5 to output the combined navigation information of the mobile machine, which is input into the dynamic detection module and the environment mapping module of the computing platform 5.

[0085] S103, performing dynamic detection and semantic mapping processing on the segmentation mask information and the preprocessed point cloud data to generate dynamic detection results and point cloud data with semantic information;

[0086] Specifically, step S103 includes: performing semantic mapping processing on the segmentation mask information and the preprocessed point cloud data to generate point cloud data with semantic information.

[0087] Specifically, in the embodiment, the preprocessed point cloud data and the classification information of the pixel points are classified using a hash table to map the pixel point classification information to the preprocessed point cloud data to output point cloud semantic data, and the Bayesian is used for incremental updating. The semantic mapping module uses a hash table to establish a mapping relationship between image pixel points and point cloud data according to the point pair relationship in the same position range according to the priori pose information of the sensor by corresponding matching the spatial position relationship between the segmentation mask information and the point cloud data, and the mapping result is input into the dynamic detection module and the environment mapping module; that is, after the image data segmentation mask and the preprocessed point cloud data are time-synchronized and spatially aligned, the mapping is performed through the form of the hash table, the image semantic pixel points are mapped to the point cloud data for fusion, the point cloud data with semantic information is obtained, and the inter-frame matching of the front-end odometry is performed.

[0088] The segmentation mask information is detected dynamically by using an optical flow method to judge the state of the segmentation mask information.

[0089] When it is judged that the state of the segmentation mask information is static, the dynamic detection result is generated.

[0090] When it is judged that the state of the segmentation mask information is dynamic, the dynamic interference information in the segmentation mask information is removed, and the dynamic detection result is generated.

[0091] Referring to Figure 5 , specifically, in the embodiment, the dynamic detection module uses the optical flow method to calculate the light flow field intensity of the pixel point category on each image frame, sets a condition threshold, if it is a dynamic pixel point category, the category information is marked, then input to the environment mapping for back-end optimization update and elimination of the last dynamic object key frame, if it is a static object, directly input to the back-end optimization update; that is, according to the segmentation mask information, the target object is dynamically detected using the optical flow method, the detection corresponding mask information is directly used as the loop detection input for the back-end pose state iterative optimization solution when the dynamic detection result is a static object, if the dynamic detection result is a dynamic object, the matching information of the dynamic interference object in the last frame is input to the back-end for elimination and the tail shadow caused is eliminated, and the result after the back-end optimization iterative solution is input to the mapping module to generate a global semantic map. The environment mapping module includes a front-end odometer, a back-end optimization, a loop detection, and a semantic mapping, the front-end odometer combines the point cloud data with semantics as the key frame selection standard for inter-frame matching, improves the robustness of inter-frame matching, and according to the IMU data as the initial position information for mapping, the position information is jointly solved with GNSS as the position correction input to the back-end optimization in the mapping process.

[0092] In the embodiment, the dynamic detection module of the computing platform 5 inputs the image acquisition information of the camera 2 and the real-time output segmentation prior information of the real-time computing unit of the computing platform 5; a two-dimensional instantaneous velocity field of the target object pixel points in the collected image is calculated using a sparse optical flow method, the target feature point information in the segmentation prior information is classified through the two-dimensional velocity vector of the pixel points, a two-dimensional velocity vector sequence is established for the pixel points of the image sequence, and two frames are taken as a sequence segment, for example, x1, x2, …, xn, xn+1. Taking one of the segments xi as an example, the moving mechanical current speed is V, the two-dimensional lost amount velocity of the current frame pixel point is V1, the two-dimensional lost amount velocity of the same pixel point in the next frame is V2, the two-dimensional velocity vector of the pixel point is xi, that is, V-(V2-V1)=xi, when xi, xi+1>0 or xi, xi+1<0, the segmentation mask category to which the target feature pixel point belongs is a dynamic target object; when xi, xi+1=0 or xi, xi+1≈0, the segmentation mask category to which the target feature pixel point belongs is a static target object; when xi>0, xi+1<0 or xi=0, xi+1≠0, the segmentation mask category to which the target feature pixel point belongs is a dynamic target object static state. The static target object is input to the front-end odometer of the environment mapping module for auxiliary inter-frame matching, the dynamic target object and the dynamic target object static state are input to the back-end optimization and loop detection of the environment mapping module, and the map is updated and optimized.

[0093] S104, optimizing the dynamic detection result, the point cloud data with semantic information and the accurate pose information of the mobile machine to generate a global semantic map;

[0094] Specifically, step S104 includes: performing laser odometry processing on the point cloud data with semantic information and the accurate pose information of the mobile machine to generate optimized data.

[0095] Performing loop detection update iteration processing on the optimized data and the dynamic detection result to generate a global semantic map.

[0096] Please refer to Figure 6 Specifically, in the embodiment, after the semantic point cloud data and the combined navigation pose information are time-synchronized and spatially aligned, the semantic pixel information is used for point cloud inter-frame matching in the front-end odometry to obtain an inter-frame pose transformation relationship, the IMU information of the combined navigation is input to the front-end odometry as prior information of the front-end odometry to initialize the inter-frame matching, the drift generated by the front-end matching process is corrected and iteratively optimized using the combined navigation information, the dynamic detection result is input to the back-end for loop detection update, and the Levenberg-Marquardt nonlinear optimization method is used for global state estimation through twice confidence domain search, and the solution is iteratively solved until convergence, and a semantic map is established according to the update result; the generated global semantic map is input to a planning control module to realize real-time obstacle avoidance, path planning and decision-making and subsequent tasks of the mobile machine.

[0097] In the embodiment, the environment mapping module of the computing platform 5 is similar to a conventional SLAM structure, and is composed of a front-end odometry, a back-end optimization, loop detection and a global state estimation. Figure FourThe unmanned driving perception mapping system of the unstructured scene adopts a new front-end odometer, uses the moving mechanical combination navigation information of the data processing module in the real-time processing unit of the computing platform 5 and the point cloud image pixel corresponding mapping result output by the semantic mapping module as the mapping input data, the new rear-end optimization uses the segmentation mask information of the static objects in the unstructured scene output by the dynamic detection module to replace the irregular point line features in the unstructured scene to assist the inter-frame matching, and uses the segmentation mask information of the dynamic target objects and the static state of the dynamic target objects in the unstructured scene to perform the rear-end optimization loop detection. The IMU data in the moving mechanical combination navigation information provides the pose information at the initial moment in the unstructured scene, and the GNSS data acquires the position information of the moving machine through the satellite to continuously correct the cumulative error generated by the IMU and the trajectory deviation of the moving machine. The new front-end odometer combines the image pixel input pre-classified by the environment elevation terrain data to improve the boundary enhanced segmentation network to obtain the segmentation task of the accurate target object boundary segmentation mask information in the unstructured scene, uses the static target objects detected by the dynamic detection module as the inter-frame matching objects, and uses the segmentation mask information thereof as the inter-frame matching feature information, thereby improving the extraction accuracy and robustness of the inter-frame feature information of the system in the unstructured scene. The new rear-end optimization inputs the segmentation mask information of the dynamic target objects and the static state of the target objects of the dynamic target objects detected by the dynamic detection module as the object information of the loop detection, updates and corrects the map according to the optical flow velocity vector relationship of the dynamic target object segmentation masks of different frames, updates the pose of the dynamic target object in time, and at the same time, uses the segmentation mask map of the static target object to correct once every 10 frames, divides the whole data into two parts, improves the utilization efficiency of the data, avoids the calculation delay burden of the system caused by a large amount of object data, improves the mapping robustness of the perception mapping system in the unstructured scene, and finally outputs the environment interactive map of the unstructured scene.

[0098] S105, control the control platform 11 according to the global semantic map to realize unmanned driving of the moving machine.

[0099] Specifically, in the embodiment, after the control platform 11 receives the task operation result of the environment interactive map of the unstructured scene established by the computing platform 5, the vehicle controller 12 controls the electric proportional pressure reducing valve 8, the multi-way valve 10, drives the left and right walking motors 9 to work and the work of other auxiliary devices, realizes autonomous planning movement of the moving machine, obstacle avoidance and real-time interaction with the target object.

[0100] In summary, the unmanned driving perception mapping system of the unstructured scene adopts the elevation terrain information to pre-classify the image pixel points, collects the elevation terrain information of the unstructured scene by the laser radar, establishes the associated information of the elevation terrain information and the image pixel points, pre-classifies the image pixel points by the difference value of the elevation terrain information at the target object boundary, and indirectly improves the boundary segmentation accuracy of the image segmentation task in the unstructured scene. Furthermore, the unmanned driving perception mapping system of the unstructured scene adopts the improved boundary enhanced segmentation network for model training, introduces the attention mechanism and the LSTM for the environmental features before multiple regressions, makes the model more focused in the training process, and directly improves the segmentation accuracy of the segmentation task. In addition, the unmanned driving perception mapping system of the unstructured scene adopts the optical flow method for dynamic detection, calculates the corresponding optical flow field intensity of the speed vector of the prior category information and performs state judgment, and uses different state information to assist in mapping, avoids the interference of dynamic target objects in mapping, and solves the problems of frame matching failure, mapping ghosting and the like. In addition, the environmental mapping module of the unmanned driving perception mapping system of the unstructured scene uses the segmentation information to replace the point and line features for frame matching, and uses the dynamic detection result for residual calculation and back-end optimization, avoids the mapping failure in the unstructured scene lacking of regular point and line features, and improves the mapping performance of the perception mapping system in the unstructured scene.

[0101] In summary, the unmanned driving perception mapping system of the unstructured scene adopts the elevation terrain information to pre-classify the image pixel points, collects the elevation terrain information of the unstructured scene by the laser radar, establishes the associated information of the elevation terrain information and the image pixel points, pre-classifies the image pixel points by the difference value of the elevation terrain information at the target object boundary, and indirectly improves the boundary segmentation accuracy of the image segmentation task in the unstructured scene. Furthermore, the unmanned driving perception mapping system of the unstructured scene adopts the improved boundary enhanced segmentation network for model training, introduces the attention mechanism and the LSTM for the environmental features before multiple regressions, makes the model more focused in the training process, and directly improves the segmentation accuracy of the segmentation task. In addition, the unmanned driving perception mapping system of the unstructured scene adopts the optical flow method for dynamic detection, calculates the corresponding optical flow field intensity of the speed vector of the prior category information and performs state judgment, and uses different state information to assist in mapping, avoids the interference of dynamic target objects in mapping, and solves the problems of frame matching failure, mapping ghosting and the like. In addition, the environmental mapping module of the unmanned driving perception mapping system of the unstructured scene uses the segmentation information to replace the point and line features for frame matching, and uses the dynamic detection result for residual calculation and back-end optimization, avoids the mapping failure in the unstructured scene lacking of regular point and line features, and improves the mapping performance of the perception mapping system in the unstructured scene.

[0102] Referring to Figure 1 The second embodiment of the present application provides a mobile machine, comprising a vehicle body 1 and a non-structured scene unmanned driving perception mapping system according to any one of the above, wherein the non-structured scene unmanned driving perception mapping system is arranged on the vehicle body 1.

[0103] In a possible embodiment of the present application, a horn 6 and front and rear vehicle lights 7 arranged on the vehicle body 1 are further included, wherein the control end of the horn 6 and the control end of the front and rear vehicle lights 7 are electrically connected with the output end of the vehicle controller 12.

[0104] The above merely describes the preferred embodiments of the present application, and the protection scope of the present application is not limited to the above-mentioned embodiments.

Claims

1. An unmanned perception mapping system for unstructured scenes, the system comprising: The application relates to a computing platform, a control platform and an original environment perception acquisition component. The control platform comprises an electric proportional pressure reducing valve, left and right traveling motors, a multi-way valve and a vehicle controller, the output end of the computing platform is electrically connected with the input end of the vehicle controller, and the output end of the vehicle controller is electrically connected with the control end of the electric proportional pressure reducing valve, the control end of the left and right traveling motors and the control end of the multi-way valve. The computing platform is configured to realize the following steps by executing a computer program stored therein: The original environment perception acquisition component is subjected to time synchronization and space alignment processing, and surrounding environment data collected by the original environment perception acquisition component is obtained; The surrounding environment data is preprocessed to generate preprocessed data, wherein the preprocessed data comprises segmentation mask information, preprocessed point cloud data, environment elevation information and accurate pose information of a mobile machine; The segmentation mask information and the preprocessed point cloud data are subjected to dynamic detection and semantic mapping processing to generate dynamic detection results and point cloud data with semantic information; The dynamic detection results, the point cloud data with semantic information and the accurate pose information of the mobile machine are subjected to optimization processing to generate a global semantic map; The control platform is controlled according to the global semantic map to realize unmanned driving of the mobile machine; The surrounding environment data is preprocessed to generate preprocessed data, and the preprocessed data further comprises: Image data is subjected to filtering processing, normalization and information enhancement processing to generate image pixel data; A corresponding relationship between the image pixel data and the environment elevation information is established, different elevation values between different target objects are used to pre-classify the image pixel data to generate classified image pixel data; The segmentation mask information and the preprocessed point cloud data are subjected to dynamic detection and semantic mapping processing to generate dynamic detection results and point cloud data with semantic information, specifically: The classified image pixel data is processed by a segmentation network. The segmentation network stacks the feature map output from the output gate with a feature map with semantic information after convolution channel adjustment and down-sampling of the classified image pixel data. The feature map output after the fusion is input into a convolution network of 3 The classification mask information of the pixel points is output after the feature extraction by the convolution network of 3, the size adjustment of the feature by 4 times up-sampling operation, and the size adjustment of the feature. The segmentation mask information and the preprocessed point cloud data are subjected to semantic mapping processing to generate point cloud data with semantic information; An optical flow method is used to detect the state of the segmentation mask information; When it is judged that the state of the segmentation mask information is static, dynamic detection results are generated; When it is judged that the state of the segmentation mask information is dynamic, dynamic interference information in the segmentation mask information is removed to generate dynamic detection results. The original environment perception acquisition component comprises a camera, a laser radar sensor, an IMU pose sensor and a GNSS positioning sensor; 2. The unmanned perception mapping system for unstructured scenes of claim 1, wherein, The output end of the camera, the output end of the laser radar sensor, the output end of the IMU pose sensor and the output end of the GNSS positioning sensor are electrically connected with the input end of the computing platform; The camera is configured to collect image data of a mobile machine working environment; The laser radar sensor is configured to collect point cloud data information and environment depth information of the mobile machine working environment; The IMU pose sensor is configured to collect relative positioning and attitude information of an initial state of the mobile machine. ​ The GNSS positioning sensor is configured to collect the centimeter-level positioning information of the mobile machine.

3. The unmanned perception mapping system for unstructured scenes of claim 2, wherein, The surrounding environment data is preprocessed to generate preprocessed data, specifically: The point cloud data information is filtered to remove noise and is clustered and segmented to generate preprocessed point cloud data. The environment depth information is preprocessed to generate environment elevation information.

4. The unmanned perception mapping system for unstructured scenes of claim 3, wherein, The surrounding environment data is preprocessed to generate preprocessed data, and the preprocessing further includes: The relative positioning, the attitude information, and the centimeter-level positioning information are tightly coupled to be solved, and errors are corrected to generate accurate pose information of the mobile machine.

5. The unmanned perception mapping system for unstructured scenes of claim 1, wherein, The dynamic detection result, the point cloud data with semantic information, and the accurate pose information of the mobile machine are optimized to generate a global semantic map, specifically: The point cloud data with semantic information and the accurate pose information of the mobile machine are processed by laser odometry to generate optimized data. The optimized data and the dynamic detection result are processed by loop detection update iteration to generate a global semantic map.

6. A mobile machine characterized by, The non-structured scene unmanned driving perception mapping system includes a vehicle body and the non-structured scene unmanned driving perception mapping system according to any one of claims 1 to 5, wherein the non-structured scene unmanned driving perception mapping system is arranged on the vehicle body.

7. A mobile machine according to claim 6 wherein, The non-structured scene unmanned driving perception mapping system further includes a horn and front and rear vehicle lights arranged on the vehicle body, wherein control ends of the horn and the front and rear vehicle lights are electrically connected to an output end of the vehicle controller.

Citation Information

Patent Citations

  • Unstructured road-oriented point cloud map construction and maintenance method

    CN114659513A

  • Semantic map construction method based on laser and vision fusion

    CN115187737A