Real-time evaluation method of street cleanliness based on multimodal data fusion

Through the multimodal data fusion method, the street cleanliness evaluation method is extracted, which solves the problems of high manual inspection costs and inaccurate evaluation of single sensors, real-time and accurate assessment of urban street cleanliness, adapts to complex environmental changes, and improves the reliability and efficiency of the evaluation.

CN120449111BActive Publication Date: 2025-09-02SHANGHAI BODLE ENVIRONMENTAL TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510954369.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-02
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

The existing street cleanliness assessment method relies on high cost and in real time for manual inspections. A single sensor cannot fully and accurately reflect the cleanliness of urban streets. There are heterogeneous characteristics differences, data space-time alignment and information redundancy in multimodal data fusion, which affects the accuracy and real-timeness of the assessment.

Method used

By obtaining the initial acquisition parameters and original monitoring data of the multimodal perception device, performing multimodal preprocessing to extract heterogeneous feature differences points, generating an association relationship map, adjusting the acquisition parameters and performing spatiotemporal alignment, establishing a data fusion model for feature fusion, conflict elimination and redundancy filtering, and generating real-time cleanliness evaluation results.

Benefits of technology

It has achieved a comprehensive and accurate assessment of street cleanliness, improved data consistency and accuracy, adapted to changes in urban environments, reduced the cost of manual inspections, and improved the reliability and efficiency of assessments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449111B_ABST
    Figure CN120449111B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of street cleanliness assessment, and discloses a real-time street cleanliness assessment method based on multimodal data fusion. The method first obtains the initial acquisition parameters and original monitoring data of the multimodal sensing device. These data contain overlapping area environmental information and are acquired in a preset direction. The original data is then subjected to multimodal preprocessing to extract heterogeneous feature difference points and generate a correlation relationship map. The effective monitoring range is then divided based on the map, and the acquisition parameters are adjusted to form an optimized parameter set. The adjusted data is then aligned in time and space to generate a multi-source data sequence. Finally, the overlapping areas are fused to eliminate feature conflicts and information redundancy, and a real-time assessment result is generated. In addition, the correlation relationship map can be corrected in combination with manual inspection records. This method can effectively fuse multimodal data, realize real-time and accurate assessment of street cleanliness, and improve the reliability and adaptability of the assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of street cleanliness assessment, and in particular to a real-time street cleanliness assessment method based on multimodal data fusion. Background Art

[0002] With the rapid advancement of urbanization, urban environmental governance, particularly the management of street cleanliness, faces numerous challenges. Traditional street cleanliness assessment methods rely primarily on manual inspections, which have significant limitations. Manual inspections require significant manpower and material resources, resulting in high costs. Furthermore, the scope and frequency of inspections are severely limited, making it difficult to achieve comprehensive, real-time monitoring of an entire city's streets. Furthermore, manual assessments are susceptible to subjective factors, making it difficult to ensure accuracy and consistency.

[0003] In addition, some single-sensor monitoring technologies have been applied to street cleanliness assessments. However, a single sensor can only capture a specific aspect of environmental information, such as the visual characteristics of garbage or a single indicator like air quality, and cannot comprehensively and accurately reflect street cleanliness. Furthermore, a single sensor has a limited monitoring range and is susceptible to occlusion and interference in complex urban environments, resulting in insufficient data reliability and integrity.

[0004] In real-world urban street environments, cleanliness is influenced by many factors, including the type, quantity, and distribution of garbage, as well as the degree of road pollution and water accumulation. These factors are interconnected and mutually influential, requiring multi-dimensional monitoring and analysis. Therefore, a single data source and simple assessment methods are no longer sufficient to meet the needs of real-time assessment of urban street cleanliness.

[0005] With the continuous development of sensor technology and information technology, multimodal sensing devices are increasingly being used in urban environmental monitoring. Multimodal sensing devices can simultaneously acquire multiple types of environmental data, such as visual images, infrared data, and sensor data, enabling comprehensive assessments of street cleanliness. However, multimodal data fusion presents numerous technical challenges, such as heterogeneous feature differences between modal data, spatial and temporal alignment issues, and information conflicts and redundancy in overlapping areas. If these issues are not effectively addressed, they will seriously impact the accuracy and real-time nature of street cleanliness assessments.

[0006] Existing multimodal data fusion methods for street cleanliness assessment often fail to fully account for the complexity and dynamic nature of urban street environments. For example, during data preprocessing, heterogeneous feature differences between modal data cannot be effectively extracted, resulting in inaccurate subsequent construction of correlation maps. Furthermore, during the data fusion process, feature conflicts and information redundancy at the assessment boundary cannot be effectively eliminated, impacting the reliability of the assessment results. Therefore, there is an urgent need for a multimodal data fusion method that can adapt to the complex environment of urban streets and achieve real-time, accurate assessment of street cleanliness. Summary of the Invention

[0007] The purpose of the present invention is to provide a real-time street cleanliness assessment method based on multimodal data fusion to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides a real-time street cleanliness assessment method based on multimodal data fusion, the system comprising:

[0009] Acquiring initial acquisition parameters of a multimodal sensing device and raw monitoring data corresponding to the multimodal sensing device; the raw monitoring data includes environmental information of overlapping areas and is acquired according to a preset acquisition direction;

[0010] Performing multimodal preprocessing on the original monitoring data, extracting heterogeneous feature difference points between the data, and generating a correlation relationship map based on the distribution pattern of the heterogeneous feature difference points;

[0011] Divide the data area of ​​each sensing device into an effective monitoring range based on the association relationship map, and adjust the collection parameters of each sensing device according to the division result to form an optimized parameter set;

[0012] Performing spatiotemporal alignment on the adjusted monitoring data of each sensing device according to the optimized parameter set to generate an aligned multi-source data sequence;

[0013] The overlapping regions of the multi-source data sequences are fused to eliminate feature conflicts and information redundancy at the evaluation boundaries, thereby generating real-time cleanliness evaluation results.

[0014] Preferably, performing multimodal preprocessing on the original monitoring data includes:

[0015] Calculating initial difference parameters of each data according to the deployment position and sensing characteristics of the multimodal sensing device;

[0016] According to the distribution density of the heterogeneous feature difference points, points with feature rates exceeding a threshold on both sides are selected from the overlapping area as difference dividing points;

[0017] Calculate the spatiotemporal transformation matrix of the difference demarcation point based on the initial difference parameters, and use the data center point as the reference transformation point;

[0018] Feature data interpolation is performed according to the positions of the spatiotemporal transformation matrix, the reference transformation points, and the difference demarcation points to generate multimodal preprocessed monitoring data.

[0019] Preferably, performing feature data interpolation according to the positions of the spatiotemporal transformation matrix, the reference transformation points, and the difference demarcation points includes:

[0020] The area between the reference transformation point and the difference demarcation point is used as the main processing area, and the other areas are used as auxiliary processing areas;

[0021] Constructing a data feature compensation function based on the initial difference parameter; the data feature compensation function is a correspondence between a feature position in the monitoring data and an actual environment position;

[0022] interpolating the main processing area according to the data feature compensation function to generate a first processing coordinate, and interpolating the auxiliary processing area to generate a second processing coordinate;

[0023] In combination with the characteristic difference point distribution of the second processing coordinate, the second processing coordinate is corrected by combining the characteristic gradient function and the data characteristic compensation function to generate a third processing coordinate;

[0024] The first processing coordinates and the third processing coordinates are combined to form multimodal preprocessed monitoring data.

[0025] Preferably, dividing the effective monitoring range for the data area of ​​each sensing device based on the association relationship map includes:

[0026] Calculating the monitoring coverage range corresponding to the reference transformation point in a spatiotemporal coordinate system as a reference range according to the acquisition angle of the multimodal sensing device;

[0027] Determining the boundary coordinates of each sensing device data in the spatiotemporal coordinate system based on the reference range and the spatiotemporal transformation matrix;

[0028] The data resolution and acquisition angle of each sensing device are adjusted according to the boundary coordinates to form an optimized parameter set.

[0029] Preferably, performing overlapping region fusion processing on the multi-source data sequences includes:

[0030] The multi-source data sequence is input into a data fusion model, and the output result of the data fusion model is used as a real-time cleanliness evaluation result.

[0031] Preferably, the establishment of the data fusion model includes:

[0032] Collect monitoring data samples under various environmental conditions and divide the samples into feature fusion samples, conflict elimination samples and redundant filtering samples;

[0033] Establish fusion criteria for feature fusion, conflict elimination and redundant filtering in HSV feature space;

[0034] Constructing a fusion weight distribution model according to the fusion standard;

[0035] Training a deep neural network based on the classified monitoring data samples to form a feature fusion sub-model, a conflict elimination sub-model, and a redundancy filtering sub-model;

[0036] The fusion weight distribution model, feature fusion sub-model, conflict elimination sub-model and redundant filtering sub-model are combined into a data fusion model.

[0037] Preferably, the multi-source data sequence is input into a data fusion model, and the output result of the data fusion model is used as a real-time cleanliness assessment result, including:

[0038] Assigning feature weights, conflict weights, and redundancy weights to each data point in the overlapping area using the fusion weight assignment model;

[0039] Performing feature enhancement processing on the data region assigned to the feature weight through the feature fusion sub-model;

[0040] Performing conflict resolution processing on the data regions assigned to the conflict weights through the conflict resolution sub-model;

[0041] The data area assigned to the redundant weight is subjected to redundancy elimination processing by the redundant filtering sub-model;

[0042] The output results of the fusion weight allocation model, feature fusion sub-model, and conflict elimination sub-model are combined into real-time cleanliness assessment results.

[0043] Preferably, the fusion criteria for feature fusion, conflict elimination, and redundancy filtering are established in the HSV feature space respectively, including:

[0044] Calculating an average value of characteristic values ​​in the monitoring data sample as a baseline characteristic, an average value of saturation as a baseline saturation, and an average value of brightness as a baseline brightness;

[0045] The feature fusion standard is that the fluctuation range of the feature value in the overlapping area does not exceed a preset ratio of the reference feature, and the saturation fluctuation range is less than a first threshold;

[0046] The conflict elimination standard is that the distribution variance of the feature values ​​in the overlapping area is less than the second threshold, and the difference between the brightness value and the reference brightness is less than the third threshold;

[0047] The redundancy filtering standard is that the characteristic gradient change rate of adjacent data points in the overlapping area is less than a fourth threshold, and the saturation gradient change rate is less than a fifth threshold.

[0048] Preferably, after generating the real-time cleanliness assessment result, the method further includes: obtaining the ground-truth record of manual inspection, performing feature matching on the real-time cleanliness assessment result and the ground-truth record, and extracting the difference feature points between the assessment result and the ground-truth record; and correcting the feature mapping rules of the association relationship map according to the difference feature points.

[0049] Preferably, the multimodal preprocessing of the original monitoring data also includes: detecting abnormal data points in the original monitoring data, the abnormal data points being discrete points whose characteristic values ​​deviate from the baseline characteristics by more than a preset range; marking the abnormal data points as invalid data and eliminating them to generate cleaned multimodal monitoring data.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] The real-time street cleanliness assessment method based on multimodal data fusion proposed in the present invention obtains the initial acquisition parameters and original monitoring data of the multimodal sensing device. These data contain environmental information of overlapping areas and are acquired in a preset direction, providing a rich data source for subsequent processing. The original monitoring data is subjected to multimodal preprocessing, heterogeneous feature difference points are extracted, and a correlation relationship map is generated, which can deeply analyze the differences and connections between different modal data. Based on the correlation relationship map, the effective monitoring range is divided and the acquisition parameters are adjusted to make the data collection of each sensing device more targeted and effective. The adjusted monitoring data is aligned in time and space to generate a multi-source data sequence, which lays a good foundation for data fusion. The multi-source data sequence is fused in overlapping areas to eliminate feature conflicts and information redundancy, thereby generating accurate real-time cleanliness assessment results.

[0052] During multimodal preprocessing, initial difference parameters are calculated based on the deployment location and sensor characteristics. Difference cutoff points are selected from overlapping areas. A spatiotemporal transformation matrix is ​​calculated, and feature data interpolation is performed based on the central data point. This effectively addresses data discrepancies and improves data consistency and accuracy. By dividing the area into primary and secondary processing areas, constructing a data feature compensation function for interpolation, and correcting the coordinates of the secondary processing area, the quality of the preprocessed data is further improved.

[0053] When dividing the effective monitoring range, the reference range is calculated according to the acquisition angle, the boundary coordinates are determined based on the space-time transformation matrix, and the data resolution and acquisition angle are adjusted to form an optimized parameter set, making the monitoring range of the sensing device more reasonable and improving the efficiency and accuracy of data acquisition.

[0054] In the process of establishing the data fusion model, monitoring data samples under various environmental conditions were collected and classified. Fusion criteria were established in the HSV feature space, a fusion weight distribution model was constructed, and a deep neural network was trained to form sub-models, which were then combined into a data fusion model. This model can process different types of data accordingly. By assigning feature weights, conflict weights, and redundancy weights, it performs feature enhancement, conflict resolution, and redundancy elimination, effectively solving the data fusion problem in overlapping areas and improving the reliability and accuracy of the evaluation results.

[0055] After generating the real-time cleanliness assessment results, the ground truth records of manual inspections are obtained, feature matching is performed and difference feature points are extracted, and the feature mapping rules of the association relationship map are corrected, thereby achieving continuous optimization and improvement of the assessment method, enabling it to better adapt to changes in the actual environment.

[0056] Furthermore, the preprocessing process detects and removes outliers, ensuring the quality of the input data and further improving the stability and reliability of the entire assessment method. This method fully leverages the advantages of multimodal data to comprehensively and accurately assess street cleanliness, providing strong technical support for urban environmental governance and possessing broad application prospects and practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 This is a working principle diagram of the real-time street cleanliness assessment method based on multimodal data fusion according to the present invention;

[0058] Figure 2 Flowchart for multimodal preprocessing;

[0059] Figure 3 Flowchart for interpolating feature data;

[0060] Figure 4 Flowchart established for the data fusion model;

[0061] Figure 5 Flowchart of data fusion model processing. DETAILED DESCRIPTION

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0063] See also Figure 1-Figure 5 The present invention provides a real-time street cleanliness assessment method based on multimodal data fusion, and the specific implementation steps are as follows:

[0064] Step 1: Obtain the initial acquisition parameters of the multimodal sensing device and the raw monitoring data corresponding to the multimodal sensing device; the raw monitoring data includes environmental information of overlapping areas and is acquired according to a preset acquisition direction. Multimodal sensing devices may include, but are not limited to, high-definition cameras, lidars, infrared sensors, etc., and the initial acquisition parameters include acquisition frequency, resolution, angular range, etc. The raw monitoring data covers street image information, object distance information, temperature information, etc. Due to overlapping coverage in device deployment, the data will contain environmental information of overlapping areas, and all devices collect data in a pre-set direction (such as along the street or perpendicular to the street).

[0065] Step 2: Perform multimodal preprocessing on the raw monitoring data to extract heterogeneous feature differences between the data. Based on the distribution patterns of these heterogeneous feature differences, a correlation map is generated. During preprocessing, differences in feature representation are identified for data collected by different types of equipment (e.g., pixel features in images and point cloud features from lidar). For example, inconsistencies between color features in images and reflectivity features from lidar are detected. By analyzing the distribution density and positional relationships of these differences, a correlation map is constructed to reflect the inherent connections between data from different modalities.

[0066] Step 3: Based on the association graph, the data area of ​​each sensing device is divided into an effective monitoring range. The acquisition parameters of each sensing device are adjusted based on the division results to form an optimized parameter set. Based on the correlation strength and characteristic distribution of the data in the association graph, the boundaries of the area that each device can effectively monitor are determined to avoid ineffective or duplicate monitoring. Based on the divided effective range, the device's acquisition parameters are adjusted, such as reducing the acquisition angle in duplicate areas and increasing the resolution in key areas, to form an optimized parameter set.

[0067] Step 4: Temporally and spatially align the adjusted monitoring data from each sensing device according to the optimized parameter set to generate an aligned multi-source data sequence. Based on the time synchronization requirements and spatial coordinate mapping relationships in the optimized parameter set, the monitoring data from different devices in the same time period and spatial region are matched to ensure data consistency in both temporal and spatial dimensions, forming an ordered multi-source data sequence.

[0068] Step 5: Fusion is performed on overlapping regions of the multi-source data series to eliminate feature conflicts and information redundancy within the assessment boundaries, generating real-time cleanliness assessment results. A fusion algorithm is used to address feature conflicts (e.g., conflicting cleanliness assessments of the same location by different devices) and information redundancy (e.g., duplicate recordings of the same environmental features) within the overlapping regions of the multi-source data series. Ultimately, real-time cleanliness assessment results for each street area are generated.

[0069] Example 1: When performing multimodal preprocessing on the original monitoring data, it is necessary to calculate the initial difference parameters of each data based on the deployment location and sensing characteristics of the multimodal sensing device. The deployment location of the multimodal sensing device covers its specific installation coordinates on the street, including longitude and latitude, altitude, and relative distance from surrounding fixed reference objects (such as street lights and curbs). This location information directly determines the spatial reference for data collection. Sensing characteristics vary depending on the type of device. For example, the sensing characteristics of a high-definition camera include lens focal length, photosensitive element size, and color channel response curve. The sensing characteristics of a lidar include laser wavelength, scanning frequency, and point cloud density. The sensing characteristics of an infrared sensor include detection band, temperature sensitivity, etc. When calculating the initial difference parameters, these factors need to be considered comprehensively. For spatial position differences, the spatial offset of data collected by different devices can be calculated through coordinate transformation. For differences in sensing characteristics, the output results of different devices in the same environment can be compared to determine the inherent deviation in feature expression, such as the deviation in the correspondence between the camera color features and the lidar reflectivity features. Ultimately, a set of initial difference parameters including spatial offset coefficients, feature response deviation values, and data scale difference rates are formed.

[0070] Based on the distribution density of heterogeneous feature difference points, points with feature ratios exceeding a threshold on both sides of the overlapping area are selected as difference points. Overlapping areas refer to street areas where the monitoring ranges of multiple sensing devices intersect. Environmental information within this area is recorded simultaneously by at least two devices, but due to differences in device characteristics, the recorded features may be inconsistent. These inconsistent points are heterogeneous feature difference points. Distribution density is calculated using street spatial grids as units, counting the ratio of the number of heterogeneous feature difference points within each grid to the total number of data points in that grid. Feature ratio refers to the probability that a feature appears in the data on both sides of a difference point. For example, the proportion of the feature "white garbage present" in the data on the left side of a point is the left feature ratio, while the proportion of the same feature in the data on the right side is the right feature ratio. When both the left and right feature ratios of a data point exceed a preset threshold (this threshold can be set based on the complexity of the street environment and the required assessment accuracy), it indicates a significant demarcation between the feature distributions on both sides of the point. This point is marked as a difference point, used to distinguish the feature boundaries of data from different devices within the overlapping area.

[0071] The spatiotemporal transformation matrix of the difference demarcation points is calculated based on the initial difference parameters, with the center point serving as the reference transformation point. The spatiotemporal transformation matrix describes the transformation relationship between the difference demarcation points in both time and space. Spatial transformations include coordinate translation, rotation, and scaling to eliminate differences in spatial coordinate systems caused by device installation angles and positions. Temporal transformations primarily involve timestamp calibration to address data collection asynchrony between devices. During the calculation process, the spatial offset coefficients in the initial difference parameters are used to determine the coordinate translation. The rotation matrix is ​​calculated based on the device installation angle. The scaling factor is set based on the data scale difference rate. The time offset is calculated based on the device's time synchronization error. Ultimately, a spatiotemporal transformation matrix consisting of both spatial and temporal transformation submatrices is constructed. The center point is the geometric center of all raw monitoring data in space. This point is selected as the reference transformation point, and its coordinates are set as the reference origin for the spatiotemporal transformation. This ensures that all transformations of the difference demarcation points are based on this point, ensuring uniformity of the coordinate system after transformation.

[0072] Feature data interpolation is performed based on the spatiotemporal transformation matrix, the locations of the benchmark transformation points, and the difference demarcation points to generate multimodal preprocessed monitoring data. The purpose of feature data interpolation is to fill gaps between data from different devices, ensuring that the preprocessed monitoring data fully covers the entire street monitoring area and maintains consistent feature representation. During the interpolation process, the benchmark transformation points are used as the center and the locations of the difference demarcation points are used to divide the interpolation region. The interpolation algorithm within each region references the transformation relationships in the spatiotemporal transformation matrix to ensure that the interpolation results match the spatiotemporal characteristics of the surrounding data. For the spatial dimension, a distance-weighted interpolation method is used. Data points closer to the benchmark transformation points or difference demarcation points have a greater influence in the interpolation calculation. For the temporal dimension, data at different times are mapped onto a unified time axis based on the temporal transformation submatrix, and linear or spline interpolation methods are used to fill gaps in the time series. This interpolation process transforms the originally scattered, feature-variable raw monitoring data into a continuous, consistent dataset. This dataset preserves the valid environmental features collected by each device while eliminating feature conflicts caused by device differences.

[0073] Example 2: When interpolating feature data based on the positions of the spatiotemporal transformation matrix, the reference transformation points, and the difference demarcation points, the area between the reference transformation points and the difference demarcation points needs to be defined as the main processing area, and the remaining areas are used as auxiliary processing areas. The scope of the main processing area is defined by the coordinates of the reference transformation points and the coordinates of each difference demarcation point, forming a closed polygonal area. Most of the heterogeneous feature difference points in the original monitoring data are concentrated in this area, and the data density is relatively high. It is a key area that affects the accuracy of subsequent evaluation. The auxiliary processing area includes all monitoring areas outside the main processing area. The data points in this area are relatively sparse, and the heterogeneous feature difference points are less distributed, so the impact on the overall evaluation results is relatively low.

[0074] A data feature compensation function is constructed based on the initial difference parameters. This function is used to describe the correspondence between the feature position in the monitoring data and the actual environmental position. The initial difference parameters contain information about feature position deviations caused by factors such as equipment installation errors and measurement accuracy limitations. For example, camera pixel position offsets caused by lens distortion and LiDAR point cloud coordinate deviations caused by ranging errors are examples. The construction of the data feature compensation function requires combining this deviation information. By analyzing the position differences of the same environmental feature in data from different devices, a mathematical mapping relationship is established so that the feature position in the monitoring data can be accurately mapped to the actual physical location of the street. For example, for a fixed curb, there is a deviation between the pixel coordinates displayed in the camera image and the three-dimensional coordinates in the LiDAR point cloud. The compensation function needs to be able to unify these two coordinates to the same actual position.

[0075] According to the data feature compensation function, the main processing area is interpolated to generate the first processing coordinates, and the auxiliary processing area is interpolated to generate the second processing coordinates. The interpolation process of the main processing area requires a high-precision algorithm. First, the existing data point coordinates in the main processing area are corrected to the actual environment coordinates according to the data feature compensation function. Then, based on the corrected coordinates, interpolation operations are performed between adjacent data points according to the preset spatial resolution. When interpolating, the feature similarity of the data points needs to be considered. The more similar the features of the adjacent points, the more balanced the weight distribution of the interpolation results, to ensure that the generated first processing coordinates can accurately reflect the distribution of environmental features in the main processing area. The interpolation of the auxiliary processing area can adopt a relatively simplified algorithm. There is no need to perform complex feature similarity calculations. Only uniform interpolation is performed based on the spatial distribution density of the data points to generate the second processing coordinates to ensure the integrity of the auxiliary processing area data.

[0076] Combined with the distribution of characteristic difference points in the second processing coordinates, the second processing coordinates are corrected by the joint characteristic gradient function and the data characteristic compensation function to generate the third processing coordinates. The characteristic gradient function is used to describe the rate of change of environmental characteristics in space. By calculating the amplitude and direction of the change of the characteristic value in the auxiliary processing area, the gradient trend of the characteristic distribution is determined. Analysis of the distribution of characteristic difference points in the second processing coordinates shows that these difference points are mainly concentrated in areas where the characteristic gradient changes are large, such as the junction of trash cans on the street and the surrounding road surface. During the correction process, the gradient direction of the area where the difference points are located is first determined according to the characteristic gradient function, and then the coordinates of these difference points are corrected twice using the data characteristic compensation function so that the corrected coordinates can fit the characteristic distribution trend of the actual environment. For areas with small characteristic gradient changes, only slight adjustments are required to the second processing coordinates, and finally the third processing coordinates are generated to improve the accuracy of the auxiliary processing area data.

[0077] The first processing coordinates and the third processing coordinates are merged to form the monitoring data after multimodal preprocessing. During the merging process, the two types of coordinates need to be uniformly verified in the spatial coordinate system to ensure that the coordinates of the main processing area and the auxiliary processing area can be seamlessly connected to avoid coordinate overlap or gaps. For the coordinate points at the intersection, if there is a slight difference between the first processing coordinates and the third processing coordinates, the first processing coordinates of the main processing area shall be used as the standard to ensure the accuracy of the data in the key area. The merged monitoring data covers the environmental characteristic coordinates of the entire street monitoring area. These coordinates have been corrected by the data feature compensation function and have taken into account the changing trend of the feature gradient, which can accurately reflect the actual environmental conditions of the street.

[0078] Example 3: When dividing the effective monitoring range for the data area of ​​each sensing device based on the association relationship map, it is necessary to calculate the monitoring coverage range corresponding to the reference transformation point in the space-time coordinate system according to the acquisition angle of the multimodal sensing device as the reference range. The acquisition angle of the multimodal sensing device includes horizontal viewing angle and vertical viewing angle. The acquisition angles of different devices are different. For example, the horizontal viewing angle of a high-definition camera may be 90 degrees, and the horizontal viewing angle of a lidar may be 120 degrees. The space-time coordinate system takes a fixed point on the street as the origin, the horizontal axis represents the length direction of the street, the vertical axis represents the width direction of the street, and the vertical axis represents the time dimension. When calculating the reference range, take the reference transformation point as the vertex, and combine the acquisition angle of the device to draw a conical coverage area in the space-time coordinate system. The boundary of the area is determined by the edge ray of the acquisition angle. All spatial points and time points within the coverage area belong to the initial monitoring range of the device.

[0079] Based on the reference range and the space-time transformation matrix, the boundary coordinates of each sensing device data in the space-time coordinate system are determined. The space-time transformation matrix contains the transformation parameters of the device data in time and space. By substituting the boundary rays of the reference range into the space-time transformation matrix for calculation, the converted boundary rays can be obtained. The polygon vertices formed by the intersection of these rays in the space-time coordinate system are the boundary coordinates. The determination of the boundary coordinates needs to take into account the time response delay of the device. For example, if there is a fixed deviation between the data collection time of a certain device and the standard time, the space-time transformation matrix will include this deviation in the calculation, so that the boundary coordinates are shifted accordingly in the time dimension. For the boundary coordinates of the overlapping area, it is necessary to determine the coordinate value of the intersection by comparing the boundary rays of different devices to ensure that the boundary coordinates can accurately distinguish the monitoring ranges of different devices.

[0080] According to the boundary coordinates, the collection parameters of each sensing device are adjusted to form an optimized parameter set. The collection parameters include collection frequency, resolution, exposure time, etc. The adjustment process needs to be combined with the effective range defined by the boundary coordinates. If the boundary coordinates show that the monitoring range of a certain device is redundant in a certain area of ​​the street, that is, the area has been completely covered by other devices, the corresponding collection frequency of the area can be reduced or the collection angle can be narrowed; if the boundary coordinates show that the monitoring data resolution of a certain area is insufficient, that is, the spacing between adjacent boundary coordinates is too large, the collection resolution of the area can be increased to increase the data point density. The optimized parameter set needs to be stored in the form of a file, containing the identification information of each device, the adjusted parameter values ​​and the corresponding boundary coordinate range, so that the device can call the corresponding parameters according to its own position and monitoring range during operation.

[0081] When fusing overlapping regions of multi-source data sequences, the multi-source data sequences must be input into a data fusion model, and the output of the data fusion model serves as the real-time cleanliness assessment result. Multi-source data sequences are device data that has been aligned in time and space, encompassing the environmental characteristics of various street areas at different times. The input to the data fusion model is the feature matrix of the multi-source data sequence, and the output is the cleanliness assessment value for each street grid. During the fusion process, the model first extracts features from the input data, identifying characteristics related to cleanliness, such as the color, shape, and distribution density of garbage. For data in overlapping regions, the model performs a weighted fusion based on the degree of feature consistency. Regions with more consistent features receive a more balanced weight distribution. For regions with differing features, the model uses pre-set rules to determine the more reliable feature value as the fusion result.

[0082] In the above process, the boundary coordinates can be calculated using the following formula:

[0083]

[0084] in, Represents the transformed boundary coordinates; represents the space-time transformation matrix; Represents the device's rotation matrix, which is used to describe the coordinate rotation caused by the device's installation angle; The coordinates of the initial boundary points representing the benchmark range; This represents the device's position offset, used to correct for deviations between the device's installed position and its theoretical position. This formula can be used to convert the device's initial boundary coordinates into a unified space-time coordinate system to obtain accurate boundary coordinates.

[0085] When processing multi-source data sequences, the data fusion model performs feature matching, conflict detection, and information integration in a pre-set process. Feature matching involves associating identical environmental features across different devices, such as associating a "black stain" in a camera image with a "high reflectivity area" in a LiDAR point cloud. Conflict detection involves identifying inconsistent descriptions of the same area by different devices, such as an area that appears "clean" in camera data but "contains garbage" in infrared sensor data. Information integration integrates all valid features to generate a final cleanliness assessment, expressed as a numerical value or grade, covering all monitored areas of the street.

[0086] Example 4: Building a data fusion model requires collecting monitoring data samples under various environmental conditions. These samples are then divided into feature fusion samples, conflict resolution samples, and redundancy filtering samples. This diversity of environmental conditions is reflected in varying weather conditions. For example, on a sunny day, a street under direct sunlight at noon offers ample light, allowing high-definition cameras to clearly capture road surface details, while lidar reflectivity data is less affected by shadows. On rainy days, water accumulates on the road, causing temperature fluctuations in infrared sensors, and the camera image may be blurred by raindrops. The collected samples must cover these diverse scenarios. Each sample contains the raw monitoring data from each sensing device in the corresponding environment and a record of the street's actual cleanliness. During classification, feature fusion samples select data from different devices that consistently describe the same area, such as data from an area where the camera captures "no obvious garbage" and the lidar detects no unusual objects. Conflict resolution samples select data with conflicting information, such as data from an area where the camera indicates "fallen leaves" but the infrared sensor fails to detect a corresponding temperature anomaly. Redundancy filtering samples select data that repeatedly records the same information, such as data from the same location in multiple consecutive camera frames that show "clean road surface."

[0087] Fusion standards for feature fusion, conflict resolution, and redundant filtering are established in the HSV feature space. The HSV feature space describes color features using three dimensions: hue, saturation, and brightness. It is suitable for processing image data collected by devices such as cameras. It can also correlate data such as the reflectivity of lidar and the temperature of infrared sensors through feature mapping. Hue reflects the type of color, such as the distinct difference in hue between red plastic bags in garbage and the gray pavement. Saturation reflects the purity of the color, such as the lower saturation of wet pavement than dry pavement. Brightness indicates the degree of brightness, with the brightness of the pavement under streetlights at night higher than in shadowed areas. When establishing standards, it is necessary to consider the association between these features and street cleanliness. For example, unusual hues with high saturation may correspond to colored garbage, and areas with sudden brightness changes may indicate the accumulation of objects.

[0088] A fusion weight assignment model is constructed based on the fusion criteria. The core of this model is to assign weights to different data points, and the size of the weight depends on the degree to which the data point meets the various fusion criteria. For data points in the feature fusion sample, if their features meet the fusion criteria, a higher feature weight is assigned; for data points in the conflict elimination sample, corresponding conflict weights are assigned based on the severity of the conflict; for data points in the redundant filtering sample, redundant weights are assigned based on the degree of information duplication. The model sets weight calculation rules. For example, when the hue of a data point is close to a known garbage color and the saturation meets the fusion criteria, its feature weight is automatically increased; when the frequency of data points describing conflicts in different devices is high, the conflict weight is increased accordingly.

[0089] A deep neural network is trained based on classified monitoring data samples, forming a feature fusion sub-model, a conflict resolution sub-model, and a redundancy filtering sub-model. The feature fusion sub-model is trained using the feature fusion samples. The neural network learns the correspondence between features from different modal data, for example, correlating the point cloud density features of a lidar with the image texture features of a camera, and outputs a fused composite feature. The conflict resolution sub-model uses conflict resolution samples as input and learns to identify conflict types and resolve them. For example, when the camera and lidar conflict in their judgments of the same object, the model prioritizes the lidar data based on the object's size (due to lidar's greater accuracy in ranging and size determination). The redundancy filtering sub-model is trained using redundancy filtering samples to learn to identify patterns of repeated information. For example, blocks of data with unchanged position and features across consecutive frames are classified as redundant information. During training, the neural network's inter-layer parameters are adjusted to gradually align the outputs of each sub-model with the actual labels of the samples.

[0090] The fusion weight assignment model, feature fusion sub-model, conflict resolution sub-model, and redundancy filtering sub-model are combined into a data fusion model. The model's operational flow must be defined during this combination. First, the fusion weight assignment model assigns weights to the input multi-source data sequence. Then, data is distributed to the corresponding sub-models based on the weights: data with high feature weights enters the feature fusion sub-model, data with high conflict weights enters the conflict resolution sub-model, and data with high redundancy weights enters the redundancy filtering sub-model. After processing, each sub-model outputs intermediate results, which are then merged through the model's integration module to form complete street cleanliness assessment data. During the combination process, data interface compatibility between the models must be ensured. For example, the fused feature format output by the feature fusion sub-model must match the input requirements of the conflict resolution sub-model to avoid data conversion errors.

[0091] Example 5: When a multi-source data sequence is input into a data fusion model and the output of the data fusion model is used as a real-time cleanliness assessment result, a feature weight, a conflict weight, and a redundancy weight are first assigned to each data point in the overlapping area through a fusion weight allocation model. The fusion weight allocation model analyzes the performance of each data point in different modal data. The feature weight reflects the contribution of the data point to the cleanliness assessment. For example, if a data point clearly shows the shape and color of garbage, its feature weight will be higher. The conflict weight is used to mark the possibility of contradictions between data points from different devices. If the camera data at the same location shows garbage but the lidar data does not detect it, the conflict weight of the point will increase. The redundancy weight is related to the degree of repetition of the data point. When multiple consecutive data points describe the same cleanliness state, the redundancy weight of the subsequent data points will increase.

[0092] Data regions assigned to feature weights are enhanced through a feature fusion sub-model. This sub-model focuses on regions with high feature weights, extracting complementary information from different modal data. For example, it extracts the color and texture features of garbage from camera data and the three-dimensional size features of garbage from lidar data, and then overlays and integrates these features. For a plastic bottle on the street, the model combines the blue hue captured by the camera, the cylindrical outline detected by the lidar, and the temperature features recorded by the infrared sensor to form a more comprehensive feature description, making the object's characteristics more prominent in the evaluation data.

[0093] The data areas assigned conflict weights are processed through the conflict resolution sub-model. The conflict resolution sub-model first identifies the type of conflict. If the conflict is caused by a difference in device perspective, such as when the camera misses trash behind a trash can due to an angle issue but the lidar detects it, the model prioritizes the more comprehensive device data based on the device's location and perspective parameters. If the conflict is caused by environmental interference, such as when a blurry camera image on a rainy day is mistakenly identified as trash, the model references historical data on device performance in similar environments and reduces the weight of the data from the affected device. This process eliminates inconsistencies between different device data and forms a consistent feature description.

[0094] Data areas assigned redundant weights are processed through a redundant filtering sub-model for redundancy elimination. This sub-model scans areas with high redundant weights and identifies duplicate feature information. For example, if the same street location appears clean in 10 consecutive frames, the model will retain the data from one frame as a representative and eliminate the remaining nine frames to reduce the data volume. For features that differ slightly but are essentially the same, such as consecutive data points showing a small amount of dust on the road, the model will merge these data points and use a single composite feature to represent the cleanliness of the area, avoiding duplicate storage and processing of information.

[0095] The outputs of the fusion weight allocation model, feature fusion sub-model, and conflict resolution sub-model are combined to produce a real-time cleanliness assessment. During this merging process, the features output by each model are aggregated according to the street's spatial grid. The assessment result for each grid combines the fusion features of that area, resolved conflict information, and filtered valid data. The final assessment results are labeled on a grid-by-grid basis, indicating the cleanliness level of each area, such as "clean," "lightly polluted," or "heavily polluted," along with descriptions of key features, such as "white plastic bag on the northeast corner."

[0096] After generating the real-time cleanliness assessment results, the ground-truth records of manual inspections are obtained, and feature matching is performed between the real-time cleanliness assessment results and the ground-truth records to extract the difference feature points between the assessment results and the actual records. The manual inspection records contain information such as photos taken on-site by the inspectors, the recorded location and type of garbage, etc. By comparing the grid features in the assessment results with the corresponding area features in the manual records, inconsistencies are identified. For example, if the manual records show an accumulation of fallen leaves in an area marked as "clean" in the assessment results, this area is a difference feature point. Or if the garbage type determined by the assessment results does not match the manual records, such as if the model identifies paper but it is actually a plastic bottle, this will also be marked as a difference feature point.

[0097] Based on the difference feature points, the feature mapping rules of the association graph are modified. The cause of the difference feature points is analyzed. If the mapping of a certain feature type in the association graph is inaccurate, such as mistakenly mapping dark shadows as garbage, the association rules for that feature type in the graph need to be adjusted to reduce the correlation between shadow features and garbage features. If the mapping ratio of different modal data is inappropriate, such as over-reliance on camera data and ignoring LiDAR depth information, the weights of different modal data in the graph need to be redistributed. Through such corrections, the association graph more accurately reflects the correspondence between data features and actual cleaning status.

[0098] Multimodal preprocessing of raw monitoring data also includes detecting anomalous data points within the data. Anomalous data points are discrete points whose characteristic values ​​deviate from baseline characteristics by more than a preset range. Baseline characteristics are defined by analyzing a large amount of normal data, such as the tonal range of a normal road surface or the laser reflectivity range. Data points identified as anomalous are those whose characteristic values ​​exceed this range, such as the sudden appearance of completely black pixels in camera data or points with unusually large distance values ​​in lidar data.

[0099] Abnormal data points are marked as invalid and removed, generating cleaned multimodal monitoring data. During this removal process, the system records the location and characteristics of the abnormal data points for subsequent analysis to determine if the equipment is faulty. After cleaning, only valid feature points are retained in the monitoring data, preventing abnormal data from interfering with preprocessing and evaluation results, ensuring that subsequent feature extraction and fusion processes are based on more reliable data.

[0100] When establishing fusion standards for feature fusion, conflict resolution, and redundancy filtering in the HSV feature space, the average value of the feature values ​​in the monitoring data samples is calculated as the baseline feature, the average saturation value is used as the baseline saturation, and the average brightness value is used as the baseline brightness. By statistically analyzing a large amount of sample data, the average hue, saturation, and brightness values ​​of the street environment under normal conditions are obtained. For example, the baseline hue of a street pavement at noon on a sunny day may be close to gray, with low baseline saturation and high baseline brightness.

[0101] The feature fusion criteria are that the fluctuation range of the feature values ​​in the overlapping area does not exceed a preset ratio of the baseline feature, and the saturation fluctuation range is less than a first threshold. When the data features in the overlapping area vary within a certain ratio of the baseline feature and the saturation fluctuation is small, it indicates that the data features are highly consistent and suitable for fusion. For example, if the hue descriptions of the same area by different devices are within ±10% of the baseline hue and the saturation fluctuation does not exceed 5%, the feature fusion conditions are met.

[0102] The conflict resolution criterion is that the distribution variance of the eigenvalues ​​within the overlapping region is less than the second threshold, and the difference between the brightness value and the reference brightness is less than the third threshold. A small eigenvalue distribution variance indicates concentrated data features, while a brightness close to the reference brightness indicates consistent lighting conditions. In this case, if data conflicts still exist, they are more likely caused by device error, making it easier for the model to resolve them. For example, in areas where the eigenvalue distribution variance is less than 0.02 and the brightness difference is within ±8%, the conflict resolution criterion applies.

[0103] The redundancy filtering criteria require that the feature gradient change rate of adjacent data points within the overlapping region be less than the fourth threshold, and the saturation gradient change rate be less than the fifth threshold. A small feature gradient change rate indicates that the features of adjacent data points change slowly, while a small saturation gradient change rate indicates a smooth color transition. Information in such areas is prone to duplication. For example, in an area with a smooth and debris-free road surface, the feature gradient change rate of adjacent data points is less than 0.01, and the saturation gradient change rate is less than 0.03, meeting the redundancy filtering criteria.

[0104] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0105] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A real-time street cleanliness assessment method based on multimodal data fusion, characterized in that: include: Obtaining initial acquisition parameters of a multimodal sensing device and raw monitoring data corresponding to the multimodal sensing device; The original monitoring data contains environmental information of overlapping areas and is acquired according to a preset collection direction; Performing multimodal preprocessing on the original monitoring data, extracting heterogeneous feature difference points between the data, and generating a correlation relationship map based on the distribution pattern of the heterogeneous feature difference points; Divide the data area of ​​each sensing device into an effective monitoring range based on the association relationship map, and adjust the collection parameters of each sensing device according to the division result to form an optimized parameter set; Performing spatiotemporal alignment on the adjusted monitoring data of each sensing device according to the optimized parameter set to generate an aligned multi-source data sequence; Performing overlapping region fusion processing on the multi-source data sequences to eliminate feature conflicts and information redundancy at the assessment boundaries and generate real-time cleanliness assessment results; The multimodal preprocessing of the original monitoring data includes: Calculating initial difference parameters of each data according to the deployment position and sensing characteristics of the multimodal sensing device; According to the distribution density of the heterogeneous feature difference points, points with feature rates exceeding a threshold on both sides are selected from the overlapping area as difference dividing points; Calculate the spatiotemporal transformation matrix of the difference demarcation point based on the initial difference parameters, and use the data center point as the reference transformation point; Performing feature data interpolation according to the positions of the spatiotemporal transformation matrix, the reference transformation points, and the difference demarcation points to generate multimodal preprocessed monitoring data; Performing feature data interpolation according to the spatiotemporal transformation matrix, the reference transformation point, and the position of the difference demarcation point includes: The area between the reference transformation point and the difference demarcation point is used as the main processing area, and the other areas are used as auxiliary processing areas; Constructing a data feature compensation function based on the initial difference parameter; the data feature compensation function is a correspondence between a feature position in the monitoring data and an actual environment position; interpolating the main processing area according to the data feature compensation function to generate a first processing coordinate, and interpolating the auxiliary processing area to generate a second processing coordinate; In combination with the characteristic difference point distribution of the second processing coordinate, the second processing coordinate is corrected by combining the characteristic gradient function and the data characteristic compensation function to generate a third processing coordinate; The first processing coordinates and the third processing coordinates are combined to form multimodal preprocessed monitoring data.

2. The real-time street cleanliness assessment method based on multimodal data fusion according to claim 1 is characterized in that: The effective monitoring range of each sensing device's data area based on the association relationship graph includes: Calculating the monitoring coverage range corresponding to the reference transformation point in a spatiotemporal coordinate system as a reference range according to the acquisition angle of the multimodal sensing device; Determining the boundary coordinates of each sensing device data in the spatiotemporal coordinate system based on the reference range and the spatiotemporal transformation matrix; The data resolution and acquisition angle of each sensing device are adjusted according to the boundary coordinates to form an optimized parameter set.

3. The real-time street cleanliness assessment method based on multimodal data fusion according to claim 1 is characterized in that: Performing overlapping region fusion processing on the multi-source data sequences includes: The multi-source data sequence is input into a data fusion model, and the output result of the data fusion model is used as a real-time cleanliness evaluation result.

4. The method for real-time street cleanliness assessment based on multimodal data fusion according to claim 3 is characterized in that: The establishment of the data fusion model includes: Collect monitoring data samples under various environmental conditions and divide the samples into feature fusion samples, conflict elimination samples and redundant filtering samples; Establish fusion criteria for feature fusion, conflict elimination and redundant filtering in HSV feature space; Constructing a fusion weight distribution model according to the fusion standard; Training a deep neural network based on the classified monitoring data samples to form a feature fusion sub-model, a conflict elimination sub-model, and a redundancy filtering sub-model; The fusion weight distribution model, feature fusion sub-model, conflict elimination sub-model and redundant filtering sub-model are combined into a data fusion model.

5. The real-time street cleanliness assessment method based on multimodal data fusion according to claim 4 is characterized in that: Inputting the multi-source data sequence into a data fusion model, and using the output result of the data fusion model as a real-time cleanliness assessment result, including: Assigning feature weights, conflict weights, and redundancy weights to each data point in the overlapping area using the fusion weight assignment model; Performing feature enhancement processing on the data region assigned to the feature weight through the feature fusion sub-model; Performing conflict resolution processing on the data regions assigned to the conflict weights through the conflict resolution sub-model; The data area assigned to the redundant weight is subjected to redundancy elimination processing by the redundant filtering sub-model; The output results of the fusion weight allocation model, feature fusion sub-model, and conflict elimination sub-model are combined into real-time cleanliness assessment results.

6. The method for real-time street cleanliness assessment based on multimodal data fusion according to claim 4 is characterized in that: The fusion criteria for feature fusion, conflict elimination, and redundant filtering are established in the HSV feature space, including: Calculating an average value of characteristic values ​​in the monitoring data sample as a baseline characteristic, an average value of saturation as a baseline saturation, and an average value of brightness as a baseline brightness; The feature fusion standard is that the fluctuation range of the feature value in the overlapping area does not exceed a preset ratio of the reference feature, and the saturation fluctuation range is less than a first threshold; The conflict elimination standard is that the distribution variance of the feature values ​​in the overlapping area is less than the second threshold, and the difference between the brightness value and the reference brightness is less than the third threshold; The redundancy filtering standard is that the characteristic gradient change rate of adjacent data points in the overlapping area is less than a fourth threshold, and the saturation gradient change rate is less than a fifth threshold.

7. The real-time street cleanliness assessment method based on multimodal data fusion according to claim 1 is characterized in that: After generating the real-time cleanliness assessment result, the method also includes: obtaining the ground-truth record of manual inspection, performing feature matching between the real-time cleanliness assessment result and the ground-truth record, extracting the difference feature points between the assessment result and the ground-truth record; and correcting the feature mapping rules of the association relationship map according to the difference feature points.

8. The real-time street cleanliness assessment method based on multimodal data fusion according to claim 1 is characterized in that: The multimodal preprocessing of the original monitoring data also includes: detecting abnormal data points in the original monitoring data, wherein the abnormal data points are discrete points whose characteristic values ​​deviate from the baseline characteristics by more than a preset range; marking the abnormal data points as invalid data and eliminating them to generate cleaned multimodal monitoring data.

Citation Information

Patent Citations

  • Theft crime risk assessment method based on multi-modal data fusion

    CN118428732A

  • Smart city environmental sanitation unmanned vehicle path planning and real-time monitoring method

    CN120087580A