Unmanned cleaning equipment data acquisition method and system based on big data mining
By configuring multiple sensors on unmanned cleaning equipment and conducting big data mining, the problems of incomplete data acquisition and neglecting time and space consistency in traditional technology are solved, and high-precision environmental model construction and data analysis are realized, which improves the equipment's autonomous navigation and cleaning efficiency.
Patent Information
- Application Number
- CN202510040028.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional unmanned driving cleaning equipment data acquisition technology is difficult to obtain high-precision environmental images and three-dimensional spatial data, and ignores the spatial and temporal consistency of data, resulting in insufficient accuracy and reliability of autonomous navigation and data analysis.
Using a method based on big data mining, a high-definition camera, lidar and weight sensor are configured, combined with image processing and spatial reconstruction algorithms, comprehensive data acquisition and accurate spatio-temporal label annotation are performed, and key features are extracted through deep data preprocessing and multi-dimensional data fusion.
It realizes high-precision environmental model construction and data spatiotemporal consistency, and improves the accuracy and efficiency of autonomous navigation and data analysis of unmanned cleaning equipment.
Smart Images

Figure CN119989261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field, and in particular to a data collection method and system for unmanned cleaning equipment based on big data mining. Background Art
[0002] With the acceleration of urbanization and the improvement of environmental awareness, unmanned cleaning equipment is increasingly used in urban cleaning operations. This type of equipment achieves efficient and intelligent cleaning operations by integrating sensor technology, big data processing capabilities and autonomous navigation algorithms.
[0003] There are obvious deficiencies in traditional technologies. On the one hand, due to limited data collection methods, traditional technologies find it difficult to obtain high-precision images and three-dimensional spatial data of the environment around the equipment, resulting in inaccurate construction of environmental models, affecting the equipment's autonomous navigation and obstacle avoidance capabilities. On the other hand, traditional technologies often ignore the temporal and spatial consistency of data when processing data, that is, the correlation and changing trends of data at different times and spaces. This neglect makes the data analysis results lack accuracy and reliability, and it is difficult to use them to guide the intelligent decision-making of equipment and cleaning task planning. In addition, traditional technologies also have deficiencies in data preprocessing and feature extraction, and are easily affected by noise interference and abnormal data, further reducing the accuracy and efficiency of data analysis.
[0004] To sum up, traditional technologies have obvious shortcomings. In order to overcome these shortcomings, it is particularly important to develop data collection methods and systems for unmanned cleaning equipment based on big data mining. Summary of the invention
[0005] The purpose of the present invention is to make up for the deficiencies of the prior art and to provide a data collection method and system for unmanned cleaning equipment based on big data mining, which can be achieved through.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a data collection method for unmanned cleaning equipment based on big data mining, the specific steps of the method are:
[0007] S1. Comprehensive data collection strategy steps
[0008] Various types of sensors are configured on the unmanned cleaning equipment, including high-definition cameras, laser radars, and weight sensors. The high-definition cameras collect image information of the surrounding environment at a frequency of f1 frames per second, and process the images through image processing algorithm A1. Algorithm A1 is as follows:
[0009] I new =α·Filter(I old )+β·Enhance(I old )
[0010] Among them, I old is the original image, I new is the processed image, Filter is the filtering function, which uses a custom hybrid filtering algorithm, combines Gaussian filtering and median filtering, and dynamically adjusts the weights of the two according to the image noise type. Enhance is the image enhancement function, which is based on adaptive histogram equalization improvement and adjusts the contrast by statistical information of the local area of the image. α and β are weight parameters. The laser radar scans at a frequency of f2 times / second to obtain the three-dimensional space data around the device, and uses the space reconstruction algorithm A2 to build the environment model. Algorithm A2 is as follows:
[0011]
[0012] Among them, S is the reconstructed spatial model, PointCloud(r i ,θ i ) represents the point cloud data obtained by radar scanning, r i is the distance, θ i is the angle, w i To weight the point cloud data, the weight sensor monitors the weight change of the garbage collection box in real time to obtain the data of garbage generation. At the same time, the equipment is connected to the urban geographic information system (GIS) to obtain the geographical location information of the area where the equipment is located. Various sensors on the equipment continuously collect the equipment's own operating parameters, including operating speed v, remaining power e, cleaning time t, and cleaning area s;
[0013] S2. Accurate spatiotemporal labeling steps
[0014] Using the built-in high-precision clock system and GPS positioning module of the equipment, each piece of collected data is given an accurate time and space tag. The time tag is accurate to the millisecond level and is realized through the clock system and GPS time synchronization mechanism. The space tag consists of the longitude and latitude coordinates (x, y) obtained by GPS positioning and the area division information provided by the GIS system. According to the equipment operation trajectory and the cleaning task area, the space tag is further refined, and the space coding algorithm A3 is used to encode the space position. Algorithm A3 is as follows:
[0015] C = Hash(x,y,z,AreaType)
[0016] Where C is the spatial code, z is the altitude, AreaType is the area type identifier, and Hash is a custom hash function. By uniquely encoding different spatial locations, the efficiency of data storage and query is improved. The hash function parameters are determined by analyzing the distribution characteristics of geographic spatial data.
[0017] S3. Deep data preprocessing steps
[0018] For image data, a multi-stage process is adopted. First, the image noise type is identified by the noise detection algorithm A4. The algorithm A4 is as follows: N = Classify (Statistical Features (I)) where N is the noise type, Statistical Features (I) is the statistical feature of the image, and Classify is the classification function. According to the noise type, the corresponding filtering algorithm is selected for denoising. Then, the image enhancement algorithm A5 is used to improve the image quality. The algorithm A5 is as follows: I enhanced =γ·AdaptiveContrast(I denoised )+(1-γ)·EdgeEnhance(I denoised ), where I denoised is the denoised image, I enhanced is the enhanced image, AdaptiveContrast is the adaptive contrast adjustment function, EdgeEnhance is the edge enhancement function, γ is the weight parameter, and for the numerical data collected by the sensor, the outlier detection algorithm A6 is used to identify and remove outliers. Algorithm A6 is as follows:
[0019] O=Outlier(DataSet,μ,σ)
[0020] Among them, O is the outlier set, DataSet is the data set collected by the sensor, μ is the data mean, σ is the data standard deviation, Outlier is the outlier detection function, based on the improved 3σ criterion, the threshold is adjusted in combination with the data distribution characteristics, and is determined by statistical analysis of historical data. Then, the data normalization algorithm A7 is used to unify the data into the same dimension and numerical range. Algorithm A7 is as follows:
[0021]
[0022] Among them, D is the original data, D norm is the normalized data;
[0023] S4. Data fusion and feature extraction steps
[0024] The preprocessed image data, sensor numerical data and spatiotemporal label data are fused. First, a multidimensional data fusion framework is constructed to arrange and combine different types of data according to their spatiotemporal attributes and feature dimensions. For image data, the key features after enhancement are extracted, and the improved SIFT algorithm A8 is used for feature extraction. The algorithm A8 is as follows:
[0025] F=Keypoints(I enhanced )·Descriptor(I enhanced )
[0026] Among them, F is the extracted feature vector, Keypoints(I enhanced ) is the key point detected on the enhanced image, Descriptor (I enhanced ) is the descriptor of the key point. By optimizing the key point detection and descriptor generation process of the traditional SIFT algorithm, for numerical data, its statistical features and trend features are extracted, and the extracted features are fused with the spatiotemporal label data to form a data set with spatiotemporal features and multi-dimensional attributes.
[0027] Furthermore, the high-definition cameras are configured in multiple locations distributed in different positions of the unmanned cleaning equipment to achieve 360-degree all-round environmental monitoring. The image data collected by each high-definition camera is processed by an independent image processing module. During the processing, each image processing module adaptively adjusts the weight parameters α and β in the algorithm A1 according to the different image collection positions and angles. The specific adjustment method is: for image samples at different positions and angles, first analyze their image features, and quantize these features into feature vectors By training a large number of image samples with known accurate weight parameters α and β, a feature vector A neural network model is used as input and weight parameters α and β as output. When processing image data in real time, the feature vector of the current image is input into the neural network model, and the dynamically adjusted weight parameters α and β are output, so as to ensure that images at different positions and angles can get the best processing effect and improve the accuracy and completeness of image information.
[0028] Furthermore, the laser radar adopts multi-line scanning technology with a scanning angle range of [0,360°]. In the spatial reconstruction algorithm A2, for different distances r i and angle θ i Point cloud data, weight w i In addition, the influence of environmental factors is fully considered to build an environmental feature database, which stores the point cloud data features under different environmental types and the corresponding weight adjustment strategies. During the operation of the equipment, the current environment type is determined through real-time perception and analysis of the surrounding environment, and then the weight adjustment strategy corresponding to the environment type is matched from the environmental feature database. The weights w of different point cloud data are adjusted according to the strategy. i Adjustments are made. For example, in areas with dense buildings, the point cloud data may be occluded and become sparse. At this time, the weight of the point cloud data close to the device and the occluded area is increased, thereby improving the accuracy of spatial reconstruction.
[0029] Furthermore, when the geographic location information is obtained by connecting with the urban geographic information system, detailed information on topography, road slope, and traffic flow is also obtained. The topography information is obtained by deep matching and fusion with the high-precision terrain database. The specific method is as follows: the longitude and latitude information of the area where the equipment is located is used as an index to query the corresponding terrain data in the high-precision terrain database, including the altitude and terrain undulation curve. For the road slope information, the road network data and elevation data in the GIS system are used for calculation. First, the road node and section information within a certain range around the location of the equipment and its surroundings is obtained. Combined with the elevation data of these nodes, the road is calculated through trigonometric function relationships. The slope of each road section and traffic flow information are obtained through real-time interaction with the data interface of the traffic management department. The interface adopts a standardized data transmission protocol. The device sends a request to the server of the traffic management department, and the server returns real-time traffic flow data in the area where the device is located and its surroundings. These detailed information are used in subsequent data processing and analysis to further optimize the collection of equipment operating parameters and the planning of cleaning tasks. For example, the power output parameters of the equipment are adjusted according to the slope of the road so that the equipment can maintain a stable operating slope on roads with different slopes. The time and route of the cleaning operation are adjusted according to the traffic flow to avoid cleaning operations during peak traffic hours, improve cleaning efficiency and reduce the impact on traffic.
[0030] Furthermore, in the process of precise spatiotemporal labeling, the accuracy of the time label is ensured by regularly calibrating it with the GPS satellite time. The built-in clock system of the device adopts a high-precision crystal oscillator, and its frequency stability ensures the stable operation of time. At the same time, in order to cope with the abnormal situation of GPS signal interruption, the device is equipped with a backup clock system. The backup clock system adopts atomic clock technology and can maintain high-precision time recording for a long time. Under normal circumstances, the built-in clock system of the device is synchronized with the GPS satellite time according to a preset calibration period. The calibration period is T hours. During the calibration process, the precise time signal sent by the GPS satellite is obtained and compared with the time of the built-in clock system of the device to calculate the time deviation Δt, and then adjust the time of the built-in clock system of the device according to the time deviation. When the GPS signal is interrupted, the backup clock system automatically starts and continues to record time. After the GPS signal is restored, the device automatically synchronizes and calibrates the time recorded by the backup clock system with the GPS time. The specific synchronization and calibration method is: calculate the time t recorded by the backup clock system during the GPS signal interruption backup , and the duration of GPS signal interruption t interruptAccording to these two time information and the time signal after GPS recovery, the time record of the device is adjusted to ensure the continuity and accuracy of the time tag. In the process of refining the spatial tag, the spatial coding algorithm A3 is optimized in combination with the movement speed v and direction information of the device. The specific optimization method is: when calculating the spatial code C, the movement speed v and direction angle φ of the device are introduced as parameters, and they are input into the spatial coding algorithm together with the longitude and latitude coordinates (x, y), altitude z and area type identifier AreaType. By adjusting the algorithm, the spatial coding can more accurately reflect the position changes of the device at different times, thereby improving the accuracy of the spatial tag.
[0031] Furthermore, in the depth data preprocessing, for the noise detection algorithm A4 of the image data, in the classification function improved based on the support vector machine (SVM), the kernel function K(x, y) is used to process the nonlinear classification problem, and the kernel function K(x, y) selects the Gaussian kernel function:
[0032]
[0033] Among them, x and y are image feature vectors, σ is the kernel function parameter, and the optimal σ value is determined by cross-validation of a large number of noisy image samples. The specific cross-validation process is: divide the noisy image samples into k non-overlapping subsets, select one of the subsets as the validation set each time, and the remaining k-1 subsets as the training set. During the training process, for different σ values, the classification model based on the support vector machine is trained and evaluated on the validation set. The evaluation indicators include accuracy and recall rate. After multiple iterations, the σ value that makes the evaluation indicators optimal is selected as the final kernel function parameter. In the image enhancement algorithm A5, the adaptive contrast adjustment function AdaptiveContrast is based on The local histogram statistical information is specifically implemented as follows: the image is divided into multiple sub-blocks of equal size, and for each sub-block, its grayscale histogram is calculated. According to the distribution of the histogram, the contrast within the sub-block is adjusted by using a histogram equalization or an adaptive histogram equalization method. Specifically, for sub-blocks with a relatively concentrated histogram distribution, an adaptive histogram equalization method is used. By stretching the histogram of the local area, the contrast within the sub-block is enhanced. For sub-blocks with a relatively uniform histogram distribution, a simple histogram equalization method is used to make the grayscale distribution within the sub-block more uniform. In this way, the local contrast of the image is enhanced, and the visual effect and information recognizability of the image are improved.
[0034] Furthermore, in the numerical data outlier detection algorithm A6, based on the improved 3σ criterion, the threshold is adjusted in combination with the autocorrelation and trend of the data. The specific method is as follows: firstly, the numerical data D collected by the sensor is analyzed in time series, and the autoregressive model AR(p) of the data is constructed:
[0035]
[0036] Among them, X t is the data at time t, φ i is the autoregressive coefficient, p is the autoregressive order, ∈ t is a white noise sequence. By fitting and analyzing historical data, the autoregressive order p and autoregressive coefficient φ are determined. i , thereby determining the autocorrelation of the data, and then using the polynomial fitting method to perform trend analysis on the data, assuming that the fitting polynomial is y = a n x n +a n-1 x n-1 +…+a1x+a0, determine the coefficient a of the polynomial by the least squares method i According to the fitted polynomial, the trend line of the data is calculated, and the data points that deviate from the trend line to a certain extent are taken as candidate points of outliers. This degree is measured by a dynamic threshold τ. The determination of τ combines the 3σ criterion and the autocorrelation and trend of the data. Specifically, according to the standard deviation σ of the data and the residual analysis of the autoregressive model, the threshold τ is dynamically adjusted. For example, when the autocorrelation of the data is strong and the trend is relatively stable, the threshold τ is appropriately lowered to more strictly detect outliers. When the autocorrelation of the data is weak and the trend fluctuates greatly, the threshold τ is appropriately increased to avoid misjudging normal data as outliers. In this way, the accuracy of outlier detection is improved, providing a more reliable data basis for subsequent data processing.
[0037] On the other hand, the data collection system for unmanned cleaning equipment based on big data mining is characterized in that the system includes a comprehensive data collection strategy module, a precise spatiotemporal labeling module, a deep data preprocessing module, and a data fusion and feature extraction module:
[0038] The omnidirectional data collection strategy module: multiple types of sensors are configured on the unmanned cleaning equipment, including high-definition cameras, laser radars, and weight sensors. The high-definition cameras collect image information of the surrounding environment at a frequency of f1 frames / second, and the images are processed by the image processing algorithm A1. The algorithm A1 is as follows:
[0039] I new =α·Filter(I old )+β·Enhance(I old )
[0040] Among them, I old is the original image, I new is the processed image, Filter is the filtering function, which uses a custom hybrid filtering algorithm, combines Gaussian filtering and median filtering, and dynamically adjusts the weights of the two according to the image noise type. Enhance is the image enhancement function, which is based on adaptive histogram equalization improvement and adjusts the contrast by statistical information of the local area of the image. α and β are weight parameters. The laser radar scans at a frequency of f2 times / second to obtain the three-dimensional space data around the device, and uses the space reconstruction algorithm A2 to build the environment model. Algorithm A2 is as follows:
[0041]
[0042] Among them, S is the reconstructed spatial model, PointCloud(r i ,θ i ) represents the point cloud data obtained by radar scanning, r i is the distance, θ i is the angle, w i To weight the point cloud data, the weight sensor monitors the weight change of the garbage collection box in real time to obtain the data of garbage generation. At the same time, the equipment is connected to the urban geographic information system (GIS) to obtain the geographical location information of the area where the equipment is located. Various sensors on the equipment continuously collect the equipment's own operating parameters, including operating speed v, remaining power e, cleaning time t, and cleaning area s;
[0043] The precise spatiotemporal labeling module uses the built-in high-precision clock system and GPS positioning module of the device to give precise time and space labels to each piece of collected data. The time label is accurate to the millisecond level and is realized through the clock system and GPS time synchronization mechanism. The space label is composed of the longitude and latitude coordinates (x, y) obtained by GPS positioning and the area division information provided by the GIS system. According to the equipment operation trajectory and the cleaning task area, the space label is further refined, and the space coding algorithm A3 is used to encode the space position. Algorithm A3 is as follows:
[0044] C = Hash(x,y,z,AreaType)
[0045] Where C is the spatial code, z is the altitude, AreaType is the area type identifier, and Hash is a custom hash function. By uniquely encoding different spatial locations, the efficiency of data storage and query is improved. The hash function parameters are determined by analyzing the distribution characteristics of geographic spatial data.
[0046] The depth data preprocessing module: for image data, a multi-stage process is adopted. First, the image noise type is identified by the noise detection algorithm A4. The algorithm A4 is as follows: N = Classify (Statistical Features (I)) where N is the noise type, Statistical Features (I) is the statistical feature of the image, and Classify is the classification function. According to the noise type, the corresponding filtering algorithm is selected for denoising. Then, the image enhancement algorithm A5 is used to improve the image quality. The algorithm A5 is as follows: I enhanced =γ·AdaptiveContrast(I denoised )+(1-γ)·EdgeEnhance(I denoised ), where I denoised is the denoised image, I enhanced is the enhanced image, AdaptiveContrast is the adaptive contrast adjustment function, EdgeEnhance is the edge enhancement function, γ is the weight parameter, and for the numerical data collected by the sensor, the outlier detection algorithm A6 is used to identify and remove outliers. Algorithm A6 is as follows:
[0047] O=Outlier(DataSet,μ,σ)
[0048] Among them, O is the outlier set, DataSet is the data set collected by the sensor, μ is the data mean, σ is the data standard deviation, Outlier is the outlier detection function, based on the improved 3σ criterion, the threshold is adjusted in combination with the data distribution characteristics, and is determined by statistical analysis of historical data. Then, the data normalization algorithm A7 is used to unify the data into the same dimension and numerical range. Algorithm A7 is as follows:
[0049]
[0050] Among them, D is the original data, D norm is the normalized data;
[0051] The data fusion and feature extraction module: fuses the preprocessed image data, sensor numerical data and spatiotemporal label data. First, a multidimensional data fusion framework is constructed to arrange and combine different types of data according to their spatiotemporal attributes and feature dimensions. For image data, the key features after enhancement processing are extracted, and the improved SIFT algorithm A8 is used for feature extraction. The algorithm A8 is as follows:
[0052] F=Keypoints(I enhanced )·Descriptor(I enhanced )
[0053] Among them, F is the extracted feature vector, Keypoints(I enhanced ) is the key point detected on the enhanced image, Descriptor (I enhanced ) is the descriptor of the key point. By optimizing the key point detection and descriptor generation process of the traditional SIFT algorithm, for numerical data, its statistical features and trend features are extracted, and the extracted features are fused with the spatiotemporal label data to form a data set with spatiotemporal features and multi-dimensional attributes.
[0054] Compared with the prior art, the data collection method and system for unmanned cleaning equipment based on big data mining has the following beneficial effects:
[0055] 1. The present invention provides strong support for big data mining through precise spatiotemporal labeling of data and efficient fusion feature extraction. By using the built-in high-precision clock system and GPS positioning module of the equipment, each piece of collected data is given a time label accurate to the millisecond level and a detailed spatial label, ensuring the spatiotemporal consistency of the data. In addition, by constructing a multidimensional data fusion framework, image data, sensor numerical data, and spatiotemporal label data are effectively fused, and key features are extracted, providing rich, multi-dimensional data input for big data mining algorithms, which not only improves the accuracy and efficiency of data mining, but also provides strong data support for the intelligent decision-making and autonomous navigation of unmanned cleaning equipment.
[0056] 2. The present invention significantly improves the data collection quality and processing efficiency of unmanned cleaning equipment through all-round data collection strategies and deep data preprocessing technology. The comprehensive use of multiple sensors such as high-definition cameras, lidars and weight sensors, combined with customized hybrid filtering algorithms and image enhancement functions, can capture the equipment's surrounding environment information and equipment operating status in real time and accurately, providing a solid foundation for subsequent data analysis and cleaning task planning. At the same time, the deep data preprocessing module further ensures the accuracy and reliability of the data through multi-stage processing of image data and outlier detection of numerical data, reducing the impact of noise interference and abnormal data on system performance.
[0057] Other advantages, objectives and features of the present invention will be set forth in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be taught from the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0059] Figure 1 This is a process operation diagram of the data collection method for unmanned cleaning equipment based on big data mining;
[0060] Figure 2 This is the process operation diagram of the data collection system for unmanned cleaning equipment based on big data mining. DETAILED DESCRIPTION
[0061] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.
[0062] Embodiment 1
[0063] This embodiment describes the deployment of unmanned cleaning equipment in a large urban park.
[0064] The device is equipped with 4 high-definition cameras, located at the front, back, left and right of the device, and the frequency f1 = 10 frames / second to collect images. For example, when collecting images of the company's park road, the original image I old In the case of a mixture of salt and pepper noise and Gaussian noise, a custom hybrid filtering algorithm Filter is used for processing. Due to the complexity of the noise type, the Gaussian filter weight is set to 0.4 and the median filter weight is set to 0.6 according to the weights determined by the previous analysis of a large number of similar noise image samples. The image enhancement function Enhance is improved based on adaptive histogram equalization. For the image of the vegetation area beside the road, the contrast is adjusted by adjusting the statistical information of the local area, and α is set to 0.3, β is set to 0.7. The processed image I new The laser radar uses multi-line scanning technology, with a scanning frequency of f2 = 20 times / second and a scanning angle range of [0,360°]. When constructing the park environment model, for the distance r i Close (e.g. less than 5 meters) and angle θ i The point cloud data in front of the device (such as 0°-30°) is more critical for device obstacle avoidance and path planning. According to the strategy corresponding to the park environment type in the environmental feature database, the point cloud data weight w is set. i =0.8 For point cloud data that is farther away (e.g., greater than 15 meters) and on the side of the device, set the weight w i =0.2, through the spatial reconstruction algorithm Build an environment model.
[0065] The weight sensor monitors the weight changes of the garbage collection box in real time. During time periods with more visitors to the park, such as weekend afternoons, the garbage generation data increases rapidly, and the cleaning route and frequency can be adjusted accordingly. At the same time, the equipment is connected to the city's geographic information system (GIS) to obtain the equipment's geographical location information in the park, as well as the park's topography (such as grasslands, hillsides, and around lakes), road slope, and traffic flow (mainly pedestrian flow). The equipment's own operating parameters such as operating speed v = 0.5 m / s, remaining power e = 80%, cleaning time t = 2 hours, and cleaning area s = 500 square meters are also continuously collected.
[0066] The device's built-in high-precision clock system uses a high-precision crystal oscillator with a calibration period of T = 1 hour. During normal operation, it is synchronized with the GPS satellite time. For example, when synchronized at 10 a.m., the time deviation Δt = 0.005 seconds is calculated, and the device's built-in clock system time is adjusted. In terms of spatial labels, when the device is cleaning the lakeside path in the park, the longitude and latitude coordinates are (x, y), the altitude z = 50 meters, and the area type identifier AreaType is "park lakeside path". It is encoded using the spatial coding algorithm C = Hash (x, y, z, AreaType). The hash function parameters are determined based on an analysis of the distribution characteristics of the park's geographic spatial data. For example, the parameters are determined after comprehensive consideration of the area and shape factors of different regions, so that the encoding can accurately distinguish different locations in the park.
[0067] For image data, the noise detection algorithm identifies the noise type through the classification function improved by the support vector machine (SVM). For example, when collecting images of trees by the lake, the image feature vectors x and y represent the texture and grayscale features of the image, and the Gaussian kernel function is used. σ=0.5 to deal with nonlinear classification problems, determine the noise type as a mixture of Gaussian noise and salt and pepper noise, then select the corresponding filtering algorithm for denoising according to the noise type, and then use the image enhancement algorithm, set γ=0.4, and use the adaptive contrast adjustment function (divide the image into multiple sub-blocks, such as using the adaptive histogram equalization method to adjust the contrast of the sub-blocks in the shadow part of the tree) and the edge enhancement function to obtain the enhanced image I enhanced For numerical data, such as garbage collection bin weight data, first perform time series analysis on it and build an autoregressive model By fitting and analyzing historical data, the autoregressive order p=3 and the autoregressive coefficient are determined. Then use the polynomial fitting method to analyze the trend, assuming that the fitting polynomial is y=a3x 3 +a2x 2 +a1x+a0, determine the coefficient a by the least squares method i, calculate the trend line based on the fitted polynomial, determine the dynamic threshold τ based on the 3σ criterion and the autocorrelation and trend of the data, identify and remove outliers, and finally use the data normalization algorithm to unify the data into the [0, 1] interval. For example, the original data of garbage weight at a certain moment is D = 50 kg. After normalization, (Assuming the minimum historical garbage weight is 10 kg and the maximum is 100 kg).
[0068] A multidimensional data fusion framework is constructed to fuse the preprocessed image data, sensor numerical data, and spatiotemporal label data. For image data, an improved scale-invariant feature transform (SIFT) algorithm is used to extract features. In the image at the entrance of the park, key points are detected on the enhanced image through the optimized key point detection and descriptor generation process, and their descriptors are calculated to obtain the feature vector F = Keypoints (I enhanced )·Descriptor(I enhanced ), for numerical data, its statistical characteristics (such as mean, standard deviation) and trend characteristics are extracted and fused with the spatiotemporal label data to form a data set with spatiotemporal characteristics and multi-dimensional attributes, which is used for subsequent data analysis and equipment operation optimization, such as adjusting the cleaning strategy and path planning of cleaning equipment in different areas of the park according to the fused data.
[0069] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modifications to the technical contents disclosed above without departing from the scope of the technical solution of the present invention. Any simple modification, different changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. The data collection method of unmanned cleaning equipment based on big data mining is characterized by: The specific steps of this method are: S1. Comprehensive data collection strategy steps Various types of sensors are configured on the unmanned cleaning equipment, including high-definition cameras, laser radars, and weight sensors. The high-definition cameras collect image information of the surrounding environment at a frequency of f1 frames per second, and process the images through image processing algorithm A1. Algorithm A1 is as follows: I new =α·Filter(I old )+β·Enhance(I old ) Among them, I old is the original image, I new is the processed image, Filter is the filtering function, which uses a custom hybrid filtering algorithm, combines Gaussian filtering and median filtering, and dynamically adjusts the weights of the two according to the image noise type. Enhance is the image enhancement function, which is based on adaptive histogram equalization improvement and adjusts the contrast by statistical information of the local area of the image. α and β are weight parameters. The laser radar scans at a frequency of f2 times / second to obtain the three-dimensional space data around the device, and uses the space reconstruction algorithm A2 to build the environment model. Algorithm A2 is as follows: Among them, S is the reconstructed spatial model, PointCloud(r i ,θ i ) represents the point cloud data obtained by radar scanning, r i is the distance, θ i is the angle, w i To weight the point cloud data, the weight sensor monitors the weight change of the garbage collection box in real time to obtain the data of garbage generation. At the same time, the equipment is connected to the urban geographic information system (GIS) to obtain the geographical location information of the area where the equipment is located. Various sensors on the equipment continuously collect the equipment's own operating parameters, including operating speed v, remaining power e, cleaning time t, and cleaning area s; S2. Accurate spatiotemporal labeling steps Using the built-in high-precision clock system and GPS positioning module of the equipment, each piece of collected data is given an accurate time and space tag. The time tag is accurate to the millisecond level and is realized through the clock system and GPS time synchronization mechanism. The space tag consists of the longitude and latitude coordinates (x, y) obtained by GPS positioning and the area division information provided by the GIS system. According to the equipment operation trajectory and the cleaning task area, the space tag is further refined, and the space coding algorithm A3 is used to encode the space position. Algorithm A3 is as follows: C = Hash(x,y,z,AreaType) Where C is the spatial code, z is the altitude, AreaType is the area type identifier, and Hash is a custom hash function. By uniquely encoding different spatial locations, the efficiency of data storage and query is improved. The hash function parameters are determined by analyzing the distribution characteristics of geographic spatial data. S3. Deep data preprocessing steps For image data, a multi-stage process is used. First, the image noise type is identified by the noise detection algorithm A4. The algorithm A4 is as follows: N = Classify (Statistical Features (I) where N is the noise type, Statistical Features (I) is the statistical feature of the image, and Classify is the classification function. According to the noise type, the corresponding filtering algorithm is selected for denoising. Then, the image enhancement algorithm A5 is used to improve the image quality. The algorithm A5 is as follows: I enhanced =γ·AdaptiveContrast(I denoised )+(1-γ)EdgeEnhance(I denoised ), where I denoised is the denoised image, I enhanced is the enhanced image, AdaptiveContrast is the adaptive contrast adjustment function, EdgeEnhance is the edge enhancement function, γ is the weight parameter, and for the numerical data collected by the sensor, the outlier detection algorithm A6 is used to identify and remove outliers. Algorithm A6 is as follows: O=Outlier(DataSet,μ,σ) Among them, O is the outlier set, DataSet is the data set collected by the sensor, μ is the data mean, σ is the data standard deviation, Outlier is the outlier detection function, based on the improved 3σ criterion, the threshold is adjusted in combination with the data distribution characteristics, and is determined by statistical analysis of historical data. Then, the data normalization algorithm A7 is used to unify the data into the same dimension and numerical range. Algorithm A7 is as follows: Among them, D is the original data, D norm is the normalized data; S4. Data fusion and feature extraction steps The preprocessed image data, sensor numerical data and spatiotemporal label data are fused. First, a multidimensional data fusion framework is constructed to arrange and combine different types of data according to their spatiotemporal attributes and feature dimensions. For image data, the key features after enhancement are extracted, and the improved SIFT algorithm A8 is used for feature extraction. The algorithm A8 is as follows: F=Keypoints(I enhanced )·Descriptor(I enhanced ) Among them, F is the extracted feature vector, Keypoints(I enhanced ) is the key point detected on the enhanced image, Descriptor (I enhanced ) is the descriptor of the key point. By optimizing the key point detection and descriptor generation process of the traditional SIFT algorithm, for numerical data, its statistical features and trend features are extracted, and the extracted features are fused with the spatiotemporal label data to form a data set with spatiotemporal features and multi-dimensional attributes.
2. The data collection method for unmanned cleaning equipment based on big data mining according to claim 1 is characterized in that: The high-definition cameras are configured in multiple locations at different positions of the unmanned cleaning equipment. The image data collected by each high-definition camera is processed by an independent image processing module. During the processing, each image processing module adaptively adjusts the weight parameters α and β in the algorithm A1 according to the different image collection positions and angles. The specific adjustment method is: for image samples at different positions and angles, the image features are first analyzed, and these features are quantified into feature vectors By training a large number of image samples with known accurate weight parameters α and β, a feature vector A neural network model is used as input and weight parameters α and β are output. When processing image data in real time, the feature vector of the current image is input into the neural network model, and the dynamically adjusted weight parameters α and β are output.
3. The data collection method for unmanned cleaning equipment based on big data mining according to claim 1 is characterized in that: The laser radar adopts multi-line scanning technology with a scanning angle range of [0,360°]. In the spatial reconstruction algorithm A2, for different distances r i and angle θ i Point cloud data, weight w i In addition, the influence of environmental factors is fully considered to build an environmental feature database, which stores the point cloud data features under different environmental types and the corresponding weight adjustment strategies. During the operation of the equipment, the current environment type is determined through real-time perception and analysis of the surrounding environment, and then the weight adjustment strategy corresponding to the environment type is matched from the environmental feature database. The weights w of different point cloud data are adjusted according to the strategy. i Make adjustments.
4. The data collection method for unmanned cleaning equipment based on big data mining according to claim 1 is characterized in that: When connecting with the urban geographic information system to obtain geographic location information, detailed information on topography, road slope, and traffic flow is also obtained. The topography information is obtained through deep matching and fusion with a high-precision terrain database. The interface uses a standardized data transmission protocol. The device sends a request to the server of the traffic management department, and the server returns real-time traffic flow data in the area where the device is located and its surroundings. These detailed information is used in subsequent data processing and analysis to further optimize the collection of equipment operating parameters and the planning of cleaning tasks.
5. The data collection method for unmanned cleaning equipment based on big data mining according to claim 1 is characterized in that: In the process of precise spatiotemporal labeling, the accuracy of the time label is ensured by regularly calibrating it with the GPS satellite time. The built-in clock system of the device adopts a high-precision crystal oscillator, and its frequency stability ensures the stable operation of time. At the same time, in order to cope with the abnormal situation of GPS signal interruption, the device is equipped with a backup clock system. The backup clock system adopts atomic clock technology and can maintain high-precision time recording for a long time. Under normal circumstances, the built-in clock system of the device is synchronized with the GPS satellite time according to a preset calibration cycle. The calibration cycle is T hours. During the calibration process, the precise time signal sent by the GPS satellite is obtained and compared with the time of the built-in clock system of the device to calculate the time deviation Δt, and then adjust the time of the built-in clock system of the device according to the time deviation. When the GPS signal is interrupted, the backup clock system automatically starts and continues to record time. After the GPS signal is restored, the device automatically synchronizes and calibrates the time recorded by the backup clock system with the GPS time. The specific synchronization and calibration method is: calculate the time t recorded by the backup clock system during the GPS signal interruption backup , and the duration of GPS signal interruption t interrupt According to these two time information and the time signal after GPS recovery, the time record of the device is adjusted. In the process of refining the spatial label, the movement speed v and direction information of the device are combined to optimize the spatial coding algorithm A3. The specific optimization method is: when calculating the spatial code C, the movement speed v and direction angle φ of the device are introduced as parameters, and they are input into the spatial coding algorithm together with the longitude and latitude coordinates (x, y), altitude z and area type identifier AreaType. By adjusting the algorithm, the spatial coding can more accurately reflect the position changes of the device at different times.
6. The data collection method for unmanned cleaning equipment based on big data mining according to claim 1 is characterized in that: In the depth data preprocessing, for the noise detection algorithm A4 of the image data, in the classification function improved based on the support vector machine (SVM), the kernel function K(x, y) is used to process the nonlinear classification problem, and the kernel function K(x, y) selects the Gaussian kernel function: Among them, x and y are image feature vectors, σ is a kernel function parameter. In the image enhancement algorithm A5, the adaptive contrast adjustment function AdaptiveContrast is based on local histogram statistical information. The specific implementation method is: divide the image into multiple sub-blocks of equal size, and for each sub-block, calculate its grayscale histogram. According to the distribution of the histogram, use histogram equalization or adaptive histogram equalization to adjust the contrast within the sub-block. Specifically, for sub-blocks with more concentrated histogram distribution, the adaptive histogram equalization method is used. By stretching the histogram of the local area, the contrast within the sub-block is enhanced. For sub-blocks with more uniform histogram distribution, a simple histogram equalization method is used to make the grayscale distribution within the sub-block more uniform.
7. The data collection method for unmanned cleaning equipment based on big data mining according to claim 1 is characterized in that: In the numerical data outlier detection algorithm A6, based on the improved 3σ criterion, the threshold is adjusted in combination with the autocorrelation and trend of the data. The specific method is as follows: firstly, the numerical data D collected by the sensor is analyzed in time series to construct the autoregressive model AR(p) of the data: Among them, X t is the data at time t, φ i is the autoregressive coefficient, p is the autoregressive order, ∈ t is a white noise sequence. By fitting and analyzing historical data, the autoregressive order p and autoregressive coefficient φ are determined. i , thereby determining the autocorrelation of the data, and then using the polynomial fitting method to perform trend analysis on the data, assuming that the fitting polynomial is y = a n x n +a n-1 x n-1 +…+a1x+a0, determine the coefficient a of the polynomial by the least squares method i , according to the fitted polynomial, the trend line of the data is calculated, and the data points that deviate from the trend line to a certain extent are regarded as candidate points of outliers. This degree is measured by a dynamic threshold τ. The determination of τ combines the 3σ criterion and the autocorrelation and trend of the data. Specifically, the threshold τ is dynamically adjusted according to the standard deviation σ of the data and the residual analysis of the autoregressive model.
8. The data collection system for unmanned cleaning equipment based on big data mining is characterized by: The system includes a comprehensive data collection strategy module, a precise spatiotemporal labeling module, a deep data preprocessing module, and a data fusion and feature extraction module: The omnidirectional data collection strategy module: multiple types of sensors are configured on the unmanned cleaning equipment, including high-definition cameras, laser radars, and weight sensors. The high-definition cameras collect image information of the surrounding environment at a frequency of f1 frames / second, and the images are processed by the image processing algorithm A1. The algorithm A1 is as follows: I new =α·Filter(I old )+β·Enhance(I old ) Among them, I old is the original image, I new is the processed image, Filter is the filtering function, which uses a custom hybrid filtering algorithm, combines Gaussian filtering and median filtering, and dynamically adjusts the weights of the two according to the image noise type. Enhance is the image enhancement function, which is based on adaptive histogram equalization improvement and adjusts the contrast by statistical information of the local area of the image. α and β are weight parameters. The laser radar scans at a frequency of f2 times / second to obtain the three-dimensional space data around the device, and uses the space reconstruction algorithm A2 to build the environment model. Algorithm A2 is as follows: Among them, S is the reconstructed spatial model, PointCloud(r i ,θ i ) represents the point cloud data obtained by radar scanning, r i is the distance, θ i is the angle, w i To weight the point cloud data, the weight sensor monitors the weight change of the garbage collection box in real time to obtain the data of garbage generation. At the same time, the equipment is connected to the urban geographic information system (GIS) to obtain the geographical location information of the area where the equipment is located. Various sensors on the equipment continuously collect the equipment's own operating parameters, including operating speed v, remaining power e, cleaning time t, and cleaning area s; The precise spatiotemporal labeling module uses the built-in high-precision clock system and GPS positioning module of the device to give precise time and space labels to each piece of collected data. The time label is accurate to the millisecond level and is realized through the clock system and GPS time synchronization mechanism. The space label is composed of the longitude and latitude coordinates (x, y) obtained by GPS positioning and the area division information provided by the GIS system. According to the equipment operation trajectory and the cleaning task area, the space label is further refined, and the space coding algorithm A3 is used to encode the space position. Algorithm A3 is as follows: C = Hash(x,y,z,AreaType) Where C is the spatial code, z is the altitude, AreaType is the area type identifier, and Hash is a custom hash function. By uniquely encoding different spatial locations, the efficiency of data storage and query is improved. The hash function parameters are determined by analyzing the distribution characteristics of geographic spatial data. The depth data preprocessing module: for image data, a multi-stage process is adopted. First, the image noise type is identified by the noise detection algorithm A4. The algorithm A4 is as follows: N = Classify (Statistical Features (I)) where N is the noise type, Statistical Features (I) is the statistical feature of the image, and Classify is the classification function. According to the noise type, the corresponding filtering algorithm is selected for denoising. Then, the image enhancement algorithm A5 is used to improve the image quality. The algorithm A5 is as follows: I enhanced =γ·AdaptiveContrast(I denoised )+(1-γ)·EdgeEnhance(I denoised ), where I denoised is the denoised image, I enh anced is the enhanced image, AdaptiveContrast is the adaptive contrast adjustment function, EdgeEnhance is the edge enhancement function, γ is the weight parameter, and for the numerical data collected by the sensor, the outlier detection algorithm A6 is used to identify and remove outliers. Algorithm A6 is as follows: O=Outlier(DataSet,μ,σ) Among them, O is the outlier set, DataSet is the data set collected by the sensor, μ is the data mean, σ is the data standard deviation, Outlier is the outlier detection function, based on the improved 3σ criterion, the threshold is adjusted in combination with the data distribution characteristics, and is determined by statistical analysis of historical data. Then, the data normalization algorithm A7 is used to unify the data into the same dimension and numerical range. Algorithm A7 is as follows: Among them, D is the original data, D norm is the normalized data; The data fusion and feature extraction module: fuses the preprocessed image data, sensor numerical data and spatiotemporal label data. First, a multidimensional data fusion framework is constructed to arrange and combine different types of data according to their spatiotemporal attributes and feature dimensions. For image data, the key features after enhancement processing are extracted, and the improved SIFT algorithm A8 is used for feature extraction. The algorithm A8 is as follows: F=Keypoints(I enh anced )·Descriptor(I enh anced ) Among them, F is the extracted feature vector, Keypoints(I enh anced ) is the key point detected on the enhanced image, Descriptor (I enh anced ) is the descriptor of the key point. By optimizing the key point detection and descriptor generation process of the traditional SIFT algorithm, for numerical data, its statistical features and trend features are extracted, and the extracted features are fused with the spatiotemporal label data to form a data set with spatiotemporal features and multi-dimensional attributes.
Citation Information
Cited By
Ophthalmic nerve scanning monitoring method and system based on eye movement gating
CN120240989A