Road monitoring multi-mode sensing method and system adapting to dynamic environment
Through multimodal perception method and edge cloud collaborative processing technology, the shortcomings of traditional road monitoring systems in dynamic environments and complex road conditions are solved, and efficient and reliable road monitoring and perception effects are achieved.
Patent Information
- Application Number
- CN202510189836.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional road monitoring systems have problems such as insufficient environmental adaptability, difficult sensor fusion, low computing efficiency and poor compatibility when dealing with dynamic environments and complex road conditions.
A multimodal perception method for road monitoring adapted to a dynamic environment is proposed. By acquiring multiple sensor data, feature extraction and weighted fusion are performed, road conditions and weather classification are used for road conditions and weather classification, sensor configuration is dynamically adjusted, and data collaborative processing is adopted for edge computing and cloud computing to achieve spatial and temporal alignment and data fusion.
It improves the overall performance and reliability of the system in complex and variable environments, achieves the optimal perception effect in dynamic environments, and enhances the scalability and compatibility of the system.
Smart Images

Figure CN120047899A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of road monitoring, and particularly relates to a multi-modal perception method and system for road monitoring adapted to a dynamic environment. Background Art
[0002] With the rapid development of intelligent transportation systems (ITS) and autonomous driving technologies, the application requirements of road monitoring and perception systems in the fields of traffic safety, autonomous driving, automated logistics, etc. have gradually increased. Traditional road monitoring systems mostly rely on single sensors or simple data fusion technologies, mainly including video surveillance, radar sensors, lidar (LiDAR), etc. However, these traditional technologies often face many challenges when dealing with dynamic environments and complex road conditions.
[0003] First of all, traditional road monitoring systems often lack environmental adaptability. Autonomous driving or intelligent transportation systems usually need to conduct real-time monitoring and perception on different road conditions such as highways, urban roads, and rural roads. However, due to significant differences in road environments in different scenarios, such as traffic flow, vehicle speed, obstacle types, weather changes, etc., existing systems are difficult to achieve rapid switching and precise adaptation in different environments. Especially in environments such as highways, urban roads, and rural roads, the road conditions change very quickly, and sensors need to be dynamically adjusted according to the current road condition state to maintain efficient operation. Existing systems cannot flexibly adjust the sensor configuration according to environmental changes and often rely on fixed sensor modes, resulting in suboptimal perception effects in specific environments.
[0004] Secondly, the sensor fusion technologies of existing technologies often have problems of spatio-temporal alignment and data heterogeneity. Road monitoring systems usually rely on multiple types of sensors, such as vision sensors, lidar, millimeter-wave radar, etc. The output data formats and acquisition times of these sensors are different. For example, vision sensors collect image data, lidar outputs point cloud data, and millimeter-wave radar provides spectrum data. Due to the different working principles and data characteristics of each sensor, how to effectively align and fuse these heterogeneous data in different times and spaces has become an urgent problem to be solved. Most existing sensor fusion methods adopt static fusion models and cannot cope with the rapid changes in dynamic environments. Especially in complex traffic and bad weather conditions, the accuracy and real-time performance of sensor data may be affected, resulting in a decline in system performance.
[0005] Furthermore, existing road monitoring systems also have bottlenecks in terms of computational efficiency and real-time performance. With the increase in the types of sensors and the amount of data, the system needs to process a large amount of real-time data, and traditional computing platforms often struggle to ensure sufficient processing speed and response capabilities. In extreme environments (such as rainy or snowy weather, night driving, etc.), sensors may face problems such as signal attenuation and noise interference. How to complete complex data processing and make a quick response within a limited time has become a major challenge in system design.
[0006] Existing technologies also lack comprehensive consideration of system compatibility. With the continuous development of hardware and algorithms, traditional systems often lack flexible modular design, resulting in poor scalability of the system. Under the requirements of sensor upgrades, algorithm optimizations, etc., it is difficult for traditional systems to achieve seamless integration and compatibility. This makes many existing systems unable to maintain high efficiency in new environments or new technologies, and have poor compatibility between the hardware and software of different manufacturers, increasing the difficulty of maintenance and upgrade.
[0007] Therefore, how to overcome these technical deficiencies and design a road monitoring system with dynamic environment adaptability, efficient data fusion capabilities, fast computing response, and good compatibility has become an urgent problem to be solved. Summary of the Invention
[0008] The object of the present invention is to propose a multi-modal perception method and system for road monitoring that adapts to dynamic environments, which not only overcomes the deficiencies of traditional road monitoring systems in terms of environmental adaptability, sensor fusion, computational efficiency, and compatibility, but also significantly improves the overall performance and reliability of the system in complex and changing environments, and has broad application prospects, especially in the fields of autonomous driving and intelligent transportation.
[0009] To achieve the above object, in the first aspect of the present invention, a multi-modal perception method for road monitoring that adapts to dynamic environments is provided, and the method includes:
[0010] Obtain data collected by infrared, camera arrays, and radar sensors, and perform feature extraction;
[0011] Based on the extracted features, perform multi-modal weighted fusion to obtain a weighted fusion feature set;
[0012] Based on the weighted fusion feature set, use a multi-layer perceptron network to classify the fusion features and output classification results of road conditions and weather conditions;
[0013] Based on the classification results of road conditions and weather conditions, conduct environmental assessment to generate a comprehensive assessment result;
[0014] Dynamically adjust the working priorities of each sensor according to the comprehensive assessment result, and configure and adjust the data collection of the sensors according to the working priorities;
[0015] For the heterogeneous data collected by each sensor, align them in terms of time and space to avoid perception distortion caused by the asynchronization of data from different sensors, and perform data fusion on the aligned data to output the fused data after alignment;
[0016] Based on the fused data after alignment, automatically adjust the data fusion strategy according to the real-time road conditions and environmental conditions, and output the high-level feature data after fusion;
[0017] Using the high-level feature data as input, perform data collaborative processing by means of edge computing and cloud computing to generate a local model generated by edge computing, a global model generated by cloud computing, and an optimization result; among them, the edge computing performs real-time data preprocessing and feature extraction, and hands over the complex deep learning model training and global data optimization tasks to cloud computing for analysis.
[0018] Furthermore, the infrared sensor data is processed through a convolutional neural network, and a spatial adaptive weight mechanism is introduced to enhance the recognition ability of targets with large temperature differences; the camera array data extracts local features through the convolutional neural networks of multiple cameras, and then performs weighted fusion on the features from different perspectives through a multi-view depth convolutional fusion layer to enhance the global perception ability of the environment; the radar data extracts the spatial features of targets through an enhanced model based on dynamic convolutional kernels, especially accurately capturing the distance and speed of targets under adverse weather conditions.
[0019] Furthermore, the environmental assessment integrates the impacts of road conditions and weather through a weighting function to generate a comprehensive assessment result;
[0020] Dynamically adjust the working priorities of each sensor based on the comprehensive assessment result, and configure and adjust the data collection of the sensors according to the working priorities, specifically including:
[0021] Define that the priority weights of each sensor are adjusted through a weighting function, where the weighting function depends on the environmental assessment result;
[0022] Generate the final sensor configuration plan according to the priority weights and working modes of the sensors, including enabling, standby, and adjusting the frequency;
[0023] For the configuration function, introduce an environment adaptive coefficient to adjust the sensitivity of sensor configuration.
[0024] Furthermore, for the heterogeneous data collected by each sensor, adopt an adaptive interpolation method for time alignment, and through a spatial transformation matrix, convert the data of the sensor from the sensor coordinate system to the global coordinate system for spatial alignment.
[0025] Furthermore, based on the time-aligned and space-aligned data, weighted fusion is performed according to the reliability of the sensors, where the reliability is a reliability function determined according to the reliability evaluation value of the sensor at the corresponding moment.
[0026] Furthermore, in the weighted fusion according to the reliability of the sensors, the data fusion strategy will be automatically adjusted according to the real-time road conditions and environmental conditions, specifically including:
[0027] Based on the characteristics of the reliability of each sensor changing over time, it is necessary to dynamically assign weights to different sensors, where the weight assignment adjustment is realized by an adaptive weighting factor, and the adaptive weighting factor is dynamically adjusted according to the historical performance of the sensor and the current environmental state;
[0028] Introduce temporal consistency constraints and spatial consistency constraints to avoid errors caused by inconsistent sensor data;
[0029] Finally, the joint optimization of dynamic weighted fusion and spatio-temporal consistency constraints is achieved through a fusion optimization algorithm.
[0030] Furthermore, the temporal consistency constraint ensures the temporal synchronization of the data of each sensor, while the spatial consistency constraint ensures that the spatial position information of the sensor data does not undergo unreasonable offsets. The functions of the temporal consistency constraint and the spatial consistency constraint are:
[0031]
[0032] where C sync (t r ) is the constraint function, C sync (t r ) is the temporal and spatial consistency constraint, is the space-aligned data of sensor i at time t r , is the reference data, and the standard data with spatial consistency at time t r ;
[0033] The fusion optimization algorithm is:
[0034]
[0035] where is the objective function, including the error term and the consistency constraint term of data fusion, is the data of sensor i at time t r after dynamic weighting, X true (t r ) is the real target data, and α, β are weight coefficients that adjust the relative importance of the data error term and the consistency constraint term.
[0036] Furthermore, the data collaborative processing using edge computing and cloud computing specifically includes:
[0037] According to the characteristics and real-time requirements of the current data, the computing tasks are divided into two categories: edge computing tasks and cloud computing tasks;
[0038] For edge computing tasks, when performing edge computing tasks, it is necessary to compress the input data and only upload the necessary features and updates of the local model. During the update process of the local model, compression technology is used to convert the original data into a low-dimensional feature representation, and at the same time, the update of the local model is optimized, and the gradient descent algorithm is used for parameter update;
[0039] For cloud computing tasks, after collecting the local model parameters from multiple edge devices, they are uploaded to the cloud for cloud computing tasks. Cross-device collaborative optimization is performed through the cloud, and a global loss function is introduced. The global loss function combines the loss functions of the local models of each edge device and the optimization objectives of the global tasks in the cloud. The cloud updates the global model parameters by minimizing the global loss function, and these global model parameters will be used as new guidance, and the optimized model will be returned to the edge devices for continued use.
[0040] Furthermore, in the data collaborative processing based on edge computing and cloud computing, constraint conditions for transmission bandwidth and computing power are established to reduce unnecessary data transmission, specifically including:
[0041] Design an intelligent data transmission strategy:
[0042] Introduce an optimization model for transmission delay. The optimization model for transmission delay optimizes the transmission path according to the transmission bandwidth limit, real-time requirements, and computing load of each device. By defining a data transmission delay loss function to guide the collaborative scheduling between edge devices and the cloud:
[0043]
[0044] where B i (t r ) is the data transmission bandwidth of edge device i; C i (t r ) is the computing power of edge device i; δ i is the transmission delay term of device i, which is used to adjust the impact of network delay on the system;
[0045] By minimizing the data transmission delay loss function schedule the data upload and processing process to improve the response speed of edge computing and cloud computing.
[0046] In the second aspect of the present invention, a road monitoring multi-modal perception system adapted to a dynamic environment is provided. Based on any of the above-mentioned road monitoring multi-modal perception methods adapted to a dynamic environment, the system includes:
[0047] The edge computing module is responsible for data preprocessing, local model updating, and real-time response;
[0048] The cloud computing module is used to handle global model updating and cross-device collaborative computing;
[0049] The communication module is used to ensure data transmission between the edge device and the cloud, and achieve load balancing;
[0050] And / or, the system further includes:
[0051] A task acceleration module, which is used to accelerate specific computing tasks through a dedicated hardware accelerator, and dynamically select an acceleration strategy according to the characteristics of data input and computing complexity; wherein, the selection of the acceleration strategy is based on a dynamic computing optimization function, and the optimization function automatically adjusts the use of hardware resources according to the computing load of the current task;
[0052] A layer data compression and transmission scheduling mechanism is designed in the data transmission between the edge computing module and the cloud computing module. This mechanism dynamically determines the timing and method of data uploading in combination with data transmission delay, network bandwidth, and computing load;
[0053] The system also adopts a dynamic feedback mechanism. Each time the system state changes, the system will readjust its computing and communication loads based on an optimization strategy; wherein, the dynamic feedback mechanism evaluates the performance state of the hardware platform through a state evaluation function, and adjusts the hardware resource allocation strategy based on the evaluation result. The state evaluation function is calculated from computing load, computing power, data transmission delay, and available bandwidth.
[0054] The beneficial technical effects of the present invention are at least as follows:
[0055] The present invention proposes a dynamic sensor configuration strategy based on road condition types and real-time states. This strategy dynamically adjusts the use and configuration of sensors according to different road condition types (such as highways, urban roads, etc.) and real-time road condition states (such as traffic density, vehicle speed, etc.). For example, on highways or rural roads, lidar and millimeter-wave radars are preferentially used, while on urban roads or in bad weather, the system will preferentially enable vision sensors and infrared sensors. Through this adaptive configuration, the system can efficiently cope with different environmental changes and avoid the limitations of traditional systems that rely on fixed configurations.
[0056] To overcome the problems of spatio-temporal alignment and data heterogeneity in the existing technology, the present invention adopts a dynamic data fusion model based on deep learning, which can perform spatio-temporal alignment on data of different modalities in real time and automatically optimize the fusion strategy according to the changes in the current road conditions. Especially in complex road conditions (such as traffic peaks, bad weather), the model can dynamically adjust the fusion weights, improve the perception accuracy and real-time performance of the system, and solve the problem of poor adaptability of traditional fusion models to dynamic environments.
[0057] The present invention proposes an innovative edge computing and cloud collaboration mechanism. Through preliminary data processing by edge devices, the data transmission delay is reduced, and the cloud platform is used for global optimization and training of deep learning models. In extreme environments, edge devices can quickly respond to and process real-time data from sensors, while the cloud is responsible for the optimization and feedback of the overall system. This architecture not only ensures the high efficiency and real-time performance of the system, but also has good scalability and compatibility, and can support the flexible integration of different types of sensors and algorithm models, solving the problems of computational bottlenecks and poor modular compatibility of traditional systems.
[0058] To solve the deficiencies of the existing system in terms of hardware compatibility and scalability, the present invention proposes a modular hardware platform design. Through standardized interfaces, sensors and computing modules produced by different manufacturers can be seamlessly docked, and the entire system does not need to be replaced when the hardware is upgraded. The modular design enables the system to be flexibly expanded according to requirements, supports the rapid integration of new sensors and algorithms, and thus ensures the long-term stable and efficient operation of the system.
[0059] Through these innovations, the present invention not only overcomes the deficiencies of traditional road monitoring systems in terms of environmental adaptability, sensor fusion, computing efficiency and compatibility, but also significantly improves the overall performance and reliability of the system in complex and changing environments, and has broad application prospects, especially in the fields of autonomous driving and intelligent transportation. Brief Description of the Drawings
[0060] The present invention is further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation to the present invention. For those of ordinary skill in the art, other drawings can also be obtained according to the following drawings without creative efforts.
[0061] Figure 1 It is a flowchart of a multi-modal perception method for road monitoring adapting to dynamic environments according to the present invention.
[0062] Figure 2 It is a framework diagram of a multi-modal perception system for road monitoring adapting to dynamic environments according to the present invention. Detailed Embodiments
[0063] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.
[0064] like Figure 1 As shown, an embodiment of the present invention provides a road monitoring multimodal perception method adapted to a dynamic environment, the method comprising:
[0065] S1. Obtain data collected by infrared, camera array, and radar sensors, and perform feature extraction; among them, firstly, perform spatiotemporal alignment processing on the data obtained from the infrared sensor, camera array, and radar. Since the data of different sensors deviate in time and space, the present invention adopts a timestamp alignment method to integrate the data of different sensors into a unified time step through interpolation and time synchronization mechanisms. This is the premise for ensuring multimodal data fusion.
[0066] Input: Data collected by infrared, camera array, radar sensors.
[0067] Output: A unified dataset after spatiotemporal alignment, providing a basis for subsequent data processing and analysis.
[0068] S2. Perform multimodal weighted fusion based on the extracted features to obtain a weighted fused feature set.
[0069] Specifically, feature extraction is performed on the spatiotemporal aligned data of each sensor.
[0070] The infrared sensor data is processed by a convolutional neural network (CNN), and a spatial adaptive weight mechanism is introduced to enhance the recognition ability of targets with large temperature differences (such as pedestrians, vehicles, etc.).
[0071] The camera array data is extracted from local features through the convolutional neural network of multiple cameras, and then the features from different perspectives are weightedly fused through the multi-perspective deep convolutional fusion layer to enhance the global perception of the environment.
[0072] Radar data extracts the spatial features of the target through an enhanced model based on dynamic convolution kernels, accurately capturing information such as the target's distance and speed, especially in severe weather conditions.
[0073] formula:
[0074] Infrared sensor feature extraction:
[0075] F infrared =CNN infrared (I infrared ; W infrared )
[0076] Camera array feature fusion:
[0077]
[0078] Radar data feature extraction:
[0079] F radar = CNN radar (D radar ; K radar )
[0080] Output: Infrared, camera array, and radar features after extraction and fusion, providing rich information for multi-modal perception.
[0081] Furthermore, the features processed by infrared, camera array, and radar sensors are fused. Since different sensor data have different signal-to-noise ratios and environmental adaptability, this step uses a dynamic weighted fusion method to assign a weighting coefficient to the data of each sensor according to its performance under specific environmental conditions. This weight coefficient is dynamically adjusted according to real-time environmental information to ensure optimal fusion effects under different weather, lighting, and road conditions.
[0082] Formula:
[0083] F fuse = w infrared ·F infrared + w camera ·F camera + w radar ·F radar
[0084] Input: Processed infrared, camera array, and radar features.
[0085] Output: Weighted fusion feature set, providing support for subsequent environmental classification and road condition recognition.
[0086] S3. Using the multi-layer perceptron network based on the weighted fusion feature set, classify the fusion features and output the classification results of road condition types and weather conditions.
[0087] Specifically, using the fused multi-modal features, perform environmental classification to identify the current road condition type (such as urban road, highway, mountain road, etc.) and weather condition (such as sunny, rainy, snowy, etc.). This step uses a multi-layer perceptron (MLP) network to classify the fusion features and output the classification results of road condition types and weather conditions.
[0088] Formula:
[0089] C road ,C weather= MLP(F fuse )
[0090] Input: Multimodal features after weighted fusion.
[0091] Output: Current road condition type C road and weather condition C weather .
[0092] S4. Conduct environmental assessment based on the classification results of road condition types and weather conditions, and generate a comprehensive assessment result.
[0093] Specifically, according to the provided road condition classification results, the system will intelligently select and configure the working modes of sensors. Under different road conditions, the system will preferentially select suitable sensors (such as using high-precision vision and millimeter-wave radar on urban roads, and using radar and infrared sensors in bad weather). By automatically adjusting the sensor configuration, ensure that data collection can best adapt to environmental changes. The output of this step is a sensor configuration scheme generated according to the road condition classification results.
[0094] Furthermore, the system first conducts environmental assessment according to the current road condition information C road and weather condition C weather . The environmental assessment integrates the impacts of road conditions and weather through a weighting function to generate a comprehensive assessment result C eval :
[0095] C eval = f(C road , C weather ) = α 1 ·C road + α 2 ·C weather
[0096] where C road represents the road condition, C weather represents the weather condition, and α 1 and α 2 are the weighting coefficients of road conditions and weather.
[0097] S5. Dynamically adjust the working priorities of each sensor according to the comprehensive assessment result, and configure and adjust the data collection of the sensors according to the working priorities.
[0098] Specifically, based on the environmental assessment result C eval , the system dynamically adjusts the working priorities of each sensor to ensure the best performance.
[0099] Furthermore, the priority of each sensor is adjusted through a weighting function α i (C eval ), and its weight w i is determined by the following formula:
[0100] w i = α i (C eval )
[0101] where w i is the working priority of sensor i, and α i (C eval ) is the dynamic adjustment function of sensor i, which depends on the environmental assessment result C eval .
[0102] Furthermore, the system generates the final sensor configuration scheme C i based on the priority weight w i of the sensor and the working mode S sensor , which is represented by the following formula:
[0103]
[0104] where C sensor is the final sensor configuration scheme, w i is the working priority of sensor i, and S i is the working mode of sensor i (such as "enabled", "standby", "adjust frequency", etc.).
[0105] Furthermore, to further refine the sensor configuration, the system introduces the environmental adaptability coefficient λ adapt to adjust the sensitivity of the sensor configuration. The final sensor configuration takes into account the adaptive adjustment of the environment and is calculated by the following formula
[0106]
[0107] where λ adapt is the environmental adaptability coefficient, and w i ·S i is the combination of the sensor priority and the working mode.
[0108] For example, in different environments, the working mode of the sensor will be dynamically adjusted according to the priority weight.
[0109] Under urban road conditions, the weight of the camera will be increased to ensure high-resolution image data acquisition; while under harsh weather conditions, the priority of the radar will be increased to ensure the accuracy of the sensing distance.
[0110] Furthermore, by dynamically adjusting the sensor configuration, the system can optimize the use of resources, avoid unnecessary computational loads, and at the same time ensure the sensing ability in different environments. The system can respond in real time to changes in complex environments, improving the intelligence and robustness of the overall sensing system.
[0111] This solution realizes the optimization of sensor configuration in different environments through real-time road condition and weather assessment, sensor priority adjustment, and environmental adaptation mechanisms. This method has important application value in the fields of autonomous driving, intelligent transportation, etc., can effectively improve the accuracy and efficiency of environmental perception, and at the same time reduce system resource consumption.
[0112] S6. For the heterogeneous data collected by each sensor, align them in terms of time and time to avoid perception distortion caused by the asynchronization of different sensor data, and perform data fusion on the aligned data to output the fused data after alignment.
[0113] Specifically, due to the sampling frequency differences of different sensors, their timestamps are inconsistent. At this time, an adaptive interpolation method is used for time alignment. Set the reference time t r , and through the interpolation function T align (t r ), map the data of each sensor to the same time axis:
[0114]
[0115] Among them, represents the data of sensor i at time t. represents the data of sensor i after alignment, at the reference time t r . T align (t r ) represents the interpolation function for time alignment.
[0116] Furthermore, due to the different installation positions and directions of each sensor, spatial alignment is required. Through the spatial transformation matrix Ri, the present invention transforms the data of sensor i from the sensor coordinate system to the global coordinate system. The specific transformation formula is:
[0117]
[0118] Among them, represents the data after spatial alignment. R i represents the rotation matrix of sensor i, which transforms the data from the sensor coordinate system to the global coordinate system. represents the position vector of sensor i. represents the data after time alignment.
[0119] Furthermore, for multi-modal sensor data, the present invention needs to perform weighted fusion according to the reliability of the sensors. An adaptive weighted algorithm is adopted, and the weight of each sensor is dynamically adjusted through the reliability evaluation function α i (t r ). The weighted fusion formula is as follows:
[0120]
[0121] Among them, X fusion (t r ) represents the fused data. w i (t r ) represents the weighting coefficient of sensor i at time t r , which is dynamically adjusted based on the reliability of the sensor.
[0122] The reliability weighting function w i (t r ) is as follows:
[0123]
[0124] Among them, α i (t r ) represents the reliability evaluation value of sensor i at time t r .
[0125] Finally, after spatio-temporal alignment and multi-modal data synchronization processing, the obtained fused data X fusion (t r ) will be used as the input for downstream tasks (such as object detection, path planning, etc.). These data have undergone spatio-temporal alignment, spatial coordinate transformation, and weighted fusion, and can provide more accurate environmental perception information.
[0126] S7. Based on the aligned fused data, automatically adjust the data fusion strategy according to the real-time road conditions and environmental conditions, and output the fused high-level feature data.
[0127] Specifically, based on the synchronized data, the system will automatically adjust the data fusion strategy according to the real-time road conditions and environmental conditions. For example, when the line of sight is limited or the weather is bad, increase the weights of radar and lidar data, while under good weather and lighting conditions, increase the weight of visual data. Data fusion uses deep learning algorithms for multi-level data fusion, so as to generate high-precision and highly reliable perception results. The output is the fused high-level feature data, which is used for subsequent road state analysis and judgment. In this step, the input comes from the output of the previous step, that is, the fused data X fusion (t r ) after spatio-temporal alignment and multi-modal data synchronization. These data have already undergone spatio-temporal alignment, spatial coordinate transformation, and weighted fusion, providing a unified data set from multiple sensors. The present invention will further perform dynamic fusion and integration processing based on this data.
[0128] Furthermore, in the dynamic fusion strategy, the present invention takes into account the time-varying nature and importance of sensor data. The reliability of each sensor changes over time, so the present invention needs to dynamically assign weights to different sensors. This process is achieved through an adaptive weighting factor w i (t r ) and the weighting factor is dynamically adjusted according to the historical performance of the sensor and the current environmental state. The specific formula is:
[0129]
[0130] Wherein, represents the final output after dynamic weighted fusion. w i (t r ) represents the dynamic weighting factor of sensor i at time t r , which depends on the reliability and historical performance of the sensor. represents the fused data of sensor i at time t r .
[0131] Wherein, w i (t r ) is calculated by the following reliability evaluation function:
[0132]
[0133] Wherein, R i (t r ) represents the reliability evaluation value of sensor i at time t r , which is calculated based on sensor data quality, sensor stability, and environmental factors (such as noise, interference, etc.).
[0134] Furthermore, in the dynamic fusion process, in order to avoid errors caused by inconsistent sensor data, the present invention introduces temporal consistency and spatial consistency constraints. The temporal consistency constraint ensures the temporal synchronization of the data of each sensor, while the spatial consistency constraint ensures that the spatial position information of the sensor data does not undergo unreasonable offsets. To achieve this goal, the present invention defines a constraint function C sync (t r ) and adds this constraint in the optimization process:
[0135]
[0136] Wherein, C sync (t r ) represents the temporal and spatial consistency constraint. represents the spatially aligned data of sensor i at time t r . represents the reference data at time t rStandard data with spatial consistency (such as global positioning system or prior map data).
[0137] During the optimization process, C sync (t r ) is added to the loss function to ensure the spatial and temporal consistency of sensor data.
[0138] Furthermore, finally, the present invention realizes the joint optimization of dynamic weighted fusion and spatio-temporal consistency constraint through an optimization algorithm. The present invention defines the objective function as:
[0139]
[0140] wherein, represents the objective function, including the error term of data fusion and the consistency constraint term. represents the data of sensor i after dynamic weighting at time t r X true (t r ) represents the true target data (e.g., ground truth or high-precision map data). α, β represent weight coefficients to adjust the relative importance of the data error term and the consistency constraint term. By minimizing the objective function the present invention can obtain the final fusion result, which satisfies the spatio-temporal consistency constraint and maximally improves the accuracy of data fusion.
[0141] Finally, the data after dynamic weighting and spatio-temporal consistency optimization will be used as the input for the next decision-making, prediction or control system. These data will be used for tasks such as path planning, target detection, behavior prediction, etc., and can provide a more accurate and stable environmental perception model.
[0142] S8. Using high-level feature data as input, performing data collaborative processing by using edge computing and cloud computing, and generating a local model generated by edge computing, a global model generated by cloud computing and an optimization result; wherein, the edge computing performs real-time data preprocessing and feature extraction, and hands over the complex deep learning model training and global data optimization tasks to cloud computing analysis.
[0143] Specifically, the system uses an edge computing platform to perform real-time data preprocessing and feature extraction, and hands over the complex deep learning model training and global data optimization tasks to the cloud for processing. The edge computing module is responsible for fast response and local data processing, while the cloud platform performs large-scale global optimization and calculation. Through edge computing and cloud collaborative processing, it is ensured that the system can quickly respond and process a large amount of data, while avoiding the burden of a single computing platform.
[0144] Furthermore, according to the characteristics and real-time requirements of the current data, the present invention divides the computing tasks into two major categories: edge computing tasks (with strong real-time performance and relatively low computing requirements) and cloud computing tasks (with high computing complexity and global optimization). This task division method is determined by a dynamic task division function T split to decide how to allocate data to the edge and the cloud. The goal of task division is to maximize the response speed and accuracy of the system. The specific task division formula is as follows:
[0145]
[0146] where X edge represents the subset of data processed by the edge device. X cloud represents the global data processed by the cloud.
[0147] The tasks of edge computing devices usually include preliminary data cleaning, feature extraction, real-time monitoring, and local model updates, while cloud computing devices are responsible for performing complex global modeling, optimization, and cross-device collaboration.
[0148] Furthermore, edge devices usually have limitations in computing and storage resources. Therefore, it is necessary to compress the input data and only upload the necessary features and updates of the local model. During the update process of the local model, the present invention uses compression technology to transform the original data into a low-dimensional feature representation. To ensure efficient computing and transmission, the present invention optimizes the update of local model parameters and uses the gradient descent algorithm to update the parameters. The update formula is as follows:
[0149]
[0150] where θ local (t r ) represents the local model parameters updated by the edge device at time t r . η represents the learning rate. represents the gradient of the loss function in the edge computing task, which is calculated based on the local data X edge (t r ) input by the edge device.
[0151] The core goal of this formula is to achieve real-time local model updates on the edge device, so as to make a quick response without increasing the computing burden. To reduce the amount of data transmission, only the model update results will be uploaded to the cloud instead of the complete data set.
[0152] Furthermore, in the cloud, after collecting the local model parameters from multiple edge devices, the present invention performs cross-device collaborative optimization through the high computing power of the cloud. To enhance the effect of the global optimization process, the present invention introduces a global loss function This function combines the loss functions of the local models of each edge device and the optimization objectives of the global tasks in the cloud. Specifically, the present invention globally optimizes the weighted loss function of the local models, and the optimization formula is as follows:
[0153]
[0154] Among them, represents the global optimization objective in the cloud, which combines the local tasks of the edge devices and the global tasks in the cloud. wi(t r ) represents the weighting factor of edge device i at time t r , reflecting the contribution degree of each device in the global optimization. represents the loss function of the edge computing task of device i. Y represents the influence coefficient of controlling the cloud task on the global optimization. represents the loss function of the global optimization task in the cloud, which involves cross-device collaboration and global modeling.
[0155] The cloud updates the global model parameter θglobal(t ) by minimizing r . This global model parameter will be used as a new guidance, and the optimized model will be returned to the edge devices for continued use. The update formula of the global model is as follows:
[0156]
[0157] Among them, θ global (t r ) represents the global model parameter in the cloud. λ represents the learning rate of the global optimization process in the cloud.
[0158] Ensures that the cloud can adjust the global strategy according to the models of all edge devices, enabling the entire system to flexibly adjust and optimize for changing data and task requirements.
[0159] Furthermore, in order to further improve the system efficiency, the present invention proposes a cross-layer data transmission optimization scheme based on edge-cloud collaboration. The present invention designs an intelligent data transmission strategy by establishing constraint conditions for transmission bandwidth and computing power, reducing unnecessary data transmission. For example, for redundant data and historical data that are no longer needed, the bandwidth consumption of data transmission is reduced through compression and selective uploading.
[0160] For this reason, the present invention introduces an optimization model for transmission delay, which considers the transmission bandwidth limit, real-time requirement, and computing load of each device to optimize the transmission path. By defining the data transmission delay loss function to guide the collaborative scheduling between edge devices and the cloud:
[0161]
[0162] Among them, B i (t r ) represents the data transmission bandwidth of edge device i. C i (t r ) represents the computing power of edge device i. δ i represents the transmission delay term of device i, which is used to adjust the impact of network delay on the system. By minimizing the data upload and processing processes can be effectively scheduled, thereby reducing the burden on the system and improving the response speed of the entire system.
[0163] Finally, the model and optimization results after collaborative computing between the edge and the cloud will be output to the final application. Each edge device will continue to process data based on the updated global model after receiving it and participate in the next round of model updates.
[0164] To better implement the above method, as Figure 2 shown, the present invention also proposes a road monitoring multimodal perception system adapted to a dynamic environment. Based on any one of the above-mentioned road monitoring multimodal perception methods adapted to a dynamic environment, the system includes:
[0165] The edge computing module is responsible for data preprocessing, local model update, and real-time response;
[0166] The cloud computing module is used to process global model updates and cross-device collaborative computing;
[0167] The communication module is used to ensure data transmission between the edge device and the cloud and achieve load balancing;
[0168] And / or, the system further includes:
[0169] The task acceleration module is used to accelerate specific computing tasks through a dedicated hardware accelerator and dynamically select an acceleration strategy according to the characteristics of the data input and the computing complexity; among them, the selection of the acceleration strategy is based on a dynamic computing optimization function, and the optimization function automatically adjusts the use of hardware resources according to the computing load of the current task;
[0170] Design a layer data compression and transmission scheduling mechanism in the data transmission between the edge computing module and the cloud computing module. This mechanism dynamically determines the timing and method of data upload in combination with data transmission delay, network bandwidth, and computing load conditions;
[0171] The system also adopts a dynamic feedback mechanism. Each time the system state changes, the system will readjust its computing and communication loads based on the optimization strategy. Among them, the dynamic feedback mechanism evaluates the performance state of the hardware platform through a state evaluation function, and adjusts the allocation strategy of hardware resources based on the evaluation result. The state evaluation function is calculated from the computing load, computing power, data transmission delay, and available bandwidth.
[0172] Among them, according to the edge and cloud computing task division completed in step 5, the system hardware design will adopt a modular architecture, which is divided into three core parts:
[0173] Edge computing module: responsible for data preprocessing, local model update, and real-time response.
[0174] Cloud computing module: processes global model updates and cross-device collaborative computing.
[0175] Communication module: ensures data transmission between edge devices and the cloud, and realizes load balancing.
[0176] Specifically, the computing power and storage capacity of each module will be adjusted according to actual needs. Specifically, the hardware of edge devices should support fast data cleaning and local feature extraction (such as data compression), while cloud devices need to have strong parallel computing capabilities, especially when performing global optimization. To ensure the flexibility and scalability of the system, the hardware platform should have pluggable characteristics between modules, so that each module can be independently upgraded and optimized.
[0177] Furthermore, for the application scenario of the patent, the hardware platform needs to be customized, especially in the task collaboration between the edge computing module and the cloud computing module. To reduce the computing load and latency, the present invention designs a task acceleration module (TaskAccelerationModule, TAM), which accelerates specific computing tasks (such as local model update and feature extraction) through a dedicated hardware accelerator (such as FPGA or ASIC). The design goal of TAM is to dynamically select an acceleration strategy according to the characteristics of data input and computing complexity. The selection of the acceleration strategy is based on a dynamic computing optimization function C opt , which automatically adjusts the use of hardware resources according to the computing load of the current task.
[0178] Optimization function C opt is calculated by considering the load of each computing module and the response time of the system. The specific calculation formula is as follows:
[0179]
[0180] Among them, L i (t r ) represents the moment tr The computing load of device i, D i (t r ) represents the data transmission delay of device i. α and β represent weighting coefficients used to balance the weights between computing load and transmission delay.
[0181] Through this optimization function, TAM can dynamically select whether to enable the hardware accelerator according to the load conditions of the edge and the cloud, thus significantly reducing the computing delay and improving the response speed of the system.
[0182] Furthermore, the communication module of the hardware platform needs to solve the problem of data transmission efficiency between edge devices and the cloud. To improve the throughput of the system, the present invention designs a multi-layer data compression and transmission scheduling mechanism, which combines data transmission delay, network bandwidth, and computing load conditions to dynamically determine the timing and method of data upload.
[0183] In this mechanism, the present invention introduces a communication delay optimization function L comm , which comprehensively considers the computing load of each edge device and the urgency of data, and optimizes the data upload strategy. The specific calculation formula is as follows:
[0184]
[0185] where γ i represents the data transmission priority weight of edge device i. B i (t r ) represents the bandwidth of edge device i. δ i represents the computing load weight of edge device i. R i (t r ) represents the remaining processing capacity of device i. This function realizes transmission scheduling based on the current device status and computing requirements in the hardware platform. By dynamically adjusting the timing of data upload, it avoids data transmission in case of network congestion or device overload, thus improving the efficiency and response speed of the entire system.
[0186] Furthermore, after completing the design of the hardware platform, the present invention integrates the modular hardware platform with the collaborative computing framework of the system. During the integration process, each module of the hardware platform will cooperate efficiently with the upper-layer cloud system and the lower-layer edge devices according to the task allocation and computing load conditions. To ensure the optimization of the hardware platform, the system adopts a dynamic feedback mechanism. Each time the system state changes, the hardware platform will readjust its computing and communication loads based on the optimization strategy.
[0187] The feedback mechanism evaluates the performance state of the hardware platform through a simple state evaluation function S eval and adjusts the allocation strategy of hardware resources based on the evaluation result:
[0188]
[0189] Among them, L i (t r ) represents the computing load of device i. C i (t r ) represents the computing power of device i. D i (t r ) represents the data transmission delay of device i. T i (t r ) represents the available bandwidth of device i. By dynamically adjusting the hardware resources and task allocation, the system can adaptively optimize according to environmental changes to ensure high-efficiency and stable performance under different workloads and operating conditions.
[0190] Unless otherwise specifically stated, the relative steps, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present application.
[0191] If the described functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0192] In the description of the present application, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is habitually placed during use. It is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present application.
[0193] In the description of the present application, it should also be noted that, unless otherwise clearly specified and defined, the terms "arranged", "installed", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0194] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A multimodal perception method for road monitoring adapted to dynamic environments, characterized in that: The method comprises: Obtain data collected by infrared, camera array, and radar sensors, and perform feature extraction; Based on the extracted features, multi-modal weighted fusion is performed to obtain a weighted fused feature set; Based on the weighted fused feature set, a multi-layer perceptron network is used to classify the fused features and output the classification results of road condition type and weather condition; Conduct environmental assessment based on road condition type and weather condition classification results to generate comprehensive assessment results; Dynamically adjust the working priority of each sensor according to the comprehensive evaluation results, and adjust the configuration of sensor data collection according to the working priority; The heterogeneous data collected by each sensor are aligned in time and space to avoid perception distortion caused by the asynchrony of data from different sensors. The aligned data are then fused and the aligned fused data are output. Based on the aligned fusion data, the data fusion strategy is automatically adjusted according to the real-time road conditions and environmental conditions, and the fused high-level feature data is output; Taking high-level feature data as input, edge computing and cloud computing are used for collaborative data processing to generate local models generated by edge computing and global models and optimization results generated by cloud computing; wherein, the edge computing performs real-time data preprocessing and feature extraction, and entrusts complex deep learning model training and global data optimization tasks to cloud computing analysis.
2. According to claim 1, a road monitoring multimodal perception method adapted to dynamic environments is characterized in that: The infrared sensor data is processed through a convolutional neural network, and a spatial adaptive weight mechanism is introduced to enhance the recognition ability of targets with large temperature differences. The camera array data extracts local features through the convolutional neural network of multiple cameras, and then the features of different perspectives are weighted and fused through the multi-perspective deep convolution fusion layer to enhance the global perception of the environment. The radar data extracts the spatial features of the target through an enhanced model based on dynamic convolution kernels, especially in severe weather conditions, to accurately capture the distance and speed of the target.
3. According to claim 1, a road monitoring multimodal perception method adapted to dynamic environments is characterized in that: The environmental assessment integrates the impact of road conditions and weather through a weighted function to generate a comprehensive assessment result; Based on the comprehensive evaluation results, the working priority of each sensor is dynamically adjusted, and the data collection of the sensor is configured and adjusted according to the working priority, including: The priority weight of each sensor is defined and adjusted by a weighting function, wherein the weighting function depends on the environmental assessment result; Generate the final sensor configuration scheme according to the sensor's priority weight and working mode, including enabling, standby, and frequency adjustment; For the configuration function, an environmental adaptive coefficient is introduced to adjust the sensitivity of the sensor configuration.
4. The multimodal perception method for road monitoring adapted to dynamic environments according to claim 1 is characterized in that: For the heterogeneous data collected by each sensor, an adaptive interpolation method is used for time alignment, and the sensor data is transformed from the sensor coordinate system to the global coordinate system through the spatial transformation matrix for spatial alignment.
5. The multimodal perception method for road monitoring adapted to dynamic environments according to claim 4 is characterized in that: Based on the time-aligned and space-aligned data, weighted fusion is performed according to the reliability of the sensor, wherein the reliability is a reliability function determined according to the reliability evaluation value of the sensor at the corresponding moment.
6. The multimodal perception method for road monitoring adapted to dynamic environments according to claim 5 is characterized in that: In weighted fusion based on sensor reliability, the data fusion strategy is automatically adjusted according to real-time road conditions and environmental conditions, including: Based on the characteristic that the reliability of each sensor changes over time, it is necessary to dynamically assign weights to different sensors, wherein the weight adjustment is implemented by an adaptive weighting factor, and the adaptive weighting factor is dynamically adjusted according to the historical performance of the sensor and the current environmental status; Introducing temporal consistency constraints and spatial consistency constraints to avoid errors caused by inconsistent sensor data; Finally, the joint optimization of dynamic weighted fusion and spatiotemporal consistency constraints is achieved through a fusion optimization algorithm.
7. The multimodal perception method for road monitoring adapted to dynamic environments according to claim 6 is characterized in that: The temporal consistency constraint ensures the temporal synchronization of the data of each sensor, while the spatial consistency constraint ensures that the spatial position information of the sensor data does not have unreasonable offset. The temporal consistency constraint and the spatial consistency constraint constitute the function: Among them, C sync (t r ) is the constraint function, C sync (t r ) are temporal and spatial consistency constraints, is sensor i at time t r The spatial alignment data, is the reference data, at time t r Spatially consistent standard data; The fusion optimization algorithm is: in, is the objective function, including the error term and consistency constraint term of data fusion, is sensor i at time t r Dynamically weighted data, X true (t r ) is the real target data, α and β are weight coefficients, which adjust the relative importance of data error term and consistency constraint term.
8. The multimodal perception method for road monitoring adapted to dynamic environments according to claim 5 is characterized in that: The collaborative data processing using edge computing and cloud computing specifically includes: According to the characteristics of current data and real-time requirements, computing tasks are divided into two categories: edge computing tasks and cloud computing tasks; For edge computing tasks, when performing edge computing tasks, the input data needs to be compressed, and only the necessary features and local model updates are uploaded. In the process of updating the local model, the compression technology is used to convert the original data into a low-dimensional feature representation. At the same time, the update of the local model is optimized, and the gradient descent algorithm is used to update the parameters; For cloud computing tasks, local model parameters from multiple edge devices are collected and uploaded to the cloud for cloud computing tasks. Cross-device collaborative optimization is performed through the cloud, and a global loss function is introduced. The global loss function combines the loss function of the local model of each edge device and the optimization goal of the global task on the cloud. The cloud updates the global model parameters by minimizing the global loss function. The global model parameters will serve as new guidance, and the optimized model will be returned to the edge device for continued use.
9. The multimodal perception method for road monitoring adapted to dynamic environments according to claim 8, characterized in that: In data collaborative processing based on edge computing and cloud computing, constraints on transmission bandwidth and computing power are established to reduce unnecessary data transmission, including: Design smart data transfer strategies: The transmission delay optimization model is introduced. The transmission delay optimization model optimizes the transmission path according to the transmission bandwidth limit, real-time requirements and computing load of each device. By defining the data transmission delay loss function To guide the coordinated scheduling between edge devices and the cloud: Among them, B i (t r ) is the data transmission bandwidth of edge device i; C i (t r ) is the computing power of edge device i; δ i is the transmission delay term of device i, which is used to adjust the impact of network delay on the system; By minimizing the data transmission delay loss function Schedule data upload and processing to improve the response speed of edge computing and cloud computing.
10. A road monitoring multimodal perception system adapted to a dynamic environment, based on the road monitoring multimodal perception method adapted to a dynamic environment according to any one of claims 1 to 9, characterized in that: The system comprises: The edge computing module is responsible for data preprocessing, local model updating and real-time response; The cloud computing module is used to handle global model updates and cross-device collaborative computing; The communication module is used to ensure data transmission between edge devices and the cloud and achieve load balancing; And / or, the system further comprises: A task acceleration module is used to accelerate specific computing tasks through a dedicated hardware accelerator, and dynamically select an acceleration strategy based on the characteristics of data input and the computational complexity; wherein the selection of the acceleration strategy is based on a dynamic computation optimization function, and the optimization function automatically adjusts the use of hardware resources according to the computational load of the current task; Designing a layer data compression and transmission scheduling mechanism in the data transmission between the edge computing module and the cloud computing module, which dynamically determines the timing and method of data upload based on data transmission delay, network bandwidth and computing load; The system also adopts a dynamic feedback mechanism. Each time the system state changes, the system will readjust its computing and communication loads based on the optimization strategy. The dynamic feedback mechanism evaluates the performance state of the hardware platform through a state evaluation function, and adjusts the allocation strategy of hardware resources based on the evaluation result. The state evaluation function is calculated by computing load, computing power, data transmission delay, and available bandwidth.
Citation Information
Cited By
Internet of vehicles channel prediction method based on multi-modal fusion and related equipment
CN120342527A
Network security threat intelligent detection method based on big data analysis
CN120498872A
High-precision image measurement system based on multi-feature fusion and measurement method thereof
CN120673396A
Automatic driving system reliability self-diagnosis method and platform based on multi-modal sensor fusion
CN120708310A
Automatic driving equipment monitoring system on expressway
CN120913386A