A real-time scene analysis method for the interaction between intelligent agents and urban roads

By collecting three-dimensional point cloud data and constructing a voxel grid model through intelligent body sensors, the problem of ignoring the three-dimensional features of objects in existing technologies is solved, and efficient decision-making and behavior execution of intelligent bodies in complex environments are achieved.

CN120340262BActive Publication Date: 2025-09-16XIANGJIANG LAB
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510824864.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-16
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing real-time scene analysis methods ignore the three-dimensional features of objects, making it difficult for intelligent agents to accurately understand the three-dimensional structure and spatial position of objects in complex or highly dynamic environments, affecting decision-making and operation accuracy.

Method used

Three-dimensional point cloud data is collected through intelligent body sensors, registered and converted into voxel grids using the ICP algorithm, and combined with multi-level voxel grid storage and deep learning, a three-dimensional environment model is constructed to analyze object distribution and movement trends, generate decisions, and update in real time.

Benefits of technology

It improves the intelligent agent's ability to understand the space of complex environments, enhances the accuracy of decision support and behavior execution, and reduces the incidence of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340262B_ABST
    Figure CN120340262B_ABST
Patent Text Reader

Abstract

The present application relates to a real-time scene analysis method for the interaction between an intelligent agent and an urban road, which collects three-dimensional point cloud data of the road environment and updates it in real time; aligns the three-dimensional point cloud data collected by different sensors through the ICP algorithm, converts the real-time updated three-dimensional point cloud data into voxels, and uses a multi-level voxel grid to store the voxels to construct a three-dimensional environment model; in the three-dimensional environment model, analyzes the road space occupancy based on the distribution and attribute information of objects in the voxel grid, and predicts the future movement trend of objects based on the position changes of objects in the voxel grid; combines historical traffic data and real-time environmental information to analyze the changing patterns of traffic flow and provide a driving path reference for the intelligent agent; generates a decision based on the analysis results including the road space occupancy, the future movement trend of objects and the driving path reference of the intelligent agent, and the intelligent agent executes the decision and provides real-time feedback on the execution results to update the three-dimensional environment model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of scene analysis, and in particular to a real-time scene analysis method for the interaction between an intelligent agent and an urban road. Background Art

[0002] With the continuous development of intelligent transportation systems, real-time interaction between intelligent agents (such as autonomous vehicles and driverless buses) and urban roads has become an important research direction for improving the efficiency, safety, and intelligence of road traffic. Multimodal perception technology, as one of the key technologies, can provide intelligent agents with more comprehensive and accurate road environment information by fusing data from different sensors (such as vision, radar, lidar, infrared sensors, etc.).

[0003] Existing real-time scene analysis methods use traditional two-dimensional grid storage methods. By projecting scene data onto a plane, the two-dimensional grid storage method can usually only process information on the plane (such as position, speed, etc.), while ignoring the three-dimensional features of the object such as height, shape, and material. This makes it impossible for the intelligent agent to accurately understand the three-dimensional structure or relative position of the object in space when perceiving and analyzing the environment, thereby limiting the intelligent agent's ability to understand complex three-dimensional environments. The two-dimensional grid storage method is difficult to effectively represent the three-dimensional form and changes of these objects, resulting in reduced accuracy or information loss when the intelligent agent handles complex or highly dynamic environments, affecting its decision-making and operation. Summary of the Invention

[0004] Based on this, it is necessary to provide a real-time scene analysis method for the interaction between an intelligent agent and urban roads, which includes:

[0005] S1: The intelligent agent uses sensors installed on the intelligent agent to perceive the road environment in real time, collects 3D point cloud data of the road environment, and updates it in real time;

[0006] S2: registering the three-dimensional point cloud data collected by different sensors using an ICP algorithm, converting the three-dimensional point cloud data updated in real time into voxels, and storing the voxels using a multi-level voxel grid to construct a three-dimensional environment model;

[0007] S3: In the three-dimensional environment model, analyzing the road space occupancy according to the distribution and attribute information of the objects in the voxel grid, and predicting the future movement trend of the objects according to the position changes of the objects in the voxel grid;

[0008] Combine historical traffic data with real-time environmental information to analyze traffic flow changes and provide driving path references for intelligent agents;

[0009] S4: Generate a decision based on the analysis results including road space occupancy, future movement trends of objects, and reference to the agent's driving path. The agent executes the decision and provides real-time feedback on the execution results to update the three-dimensional environment model.

[0010] Beneficial effects: This method combines three-dimensional point cloud data and voxel grid technology, and the intelligent agent's spatial understanding ability is significantly enhanced. It can more accurately grasp the spatial relationship between objects and their motion trajectories, providing stronger support for subsequent decision-making and behavior execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 This is a flowchart of the real-time scene analysis method for the interaction between an intelligent agent and urban roads in an embodiment of the present application. DETAILED DESCRIPTION

[0013] To make the above-mentioned objects, features, and advantages of the present application more clearly understood, the specific embodiments of the present application are described in detail below with reference to the accompanying drawings. The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0014] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0015] like Figure 1 As shown, this embodiment provides a real-time scene analysis method for interaction between an intelligent agent and a city road, the method comprising:

[0016] S1: The road environment is perceived in real time through sensors installed on the intelligent body, and three-dimensional point cloud data of the road environment is collected and updated in real time.

[0017] Specifically, the sensor includes:

[0018] Cameras are used to collect image information including traffic signs, road markings, pedestrians, vehicles, and road status data;

[0019] LiDAR, used to collect the first three-dimensional point cloud data of the obstacle's position, speed, and distance from the intelligent body;

[0020] Millimeter-wave radar, used to collect second and third-dimensional point cloud data of the speed of surrounding objects and their distance from the intelligent body;

[0021] Ultrasonic sensors for detecting close-range objects, especially effective when parking or driving at low speeds;

[0022] Noise sensors are used to monitor environmental noise to determine traffic flow, road conditions, and special events (such as traffic accidents, sudden obstacles, etc.).

[0023] Furthermore, real-time updates of 3D point cloud data include:

[0024] Status update, the update formula is:

[0025] ;

[0026] in, represents the predicted state at time k; A represents the state transfer matrix; represents the predicted state at time k-1 after correction; B represents the control input matrix; represents the control input at time k; the state update formula is used to predict the state of the system at the next moment, which depends on the system's motion model and current state information.

[0027] Measurement update, the update formula is:

[0028] ;

[0029] ;

[0030] ;

[0031] in, represents the Kalman gain; represents the prediction error covariance at time k; H represents the measurement matrix; R represents the measurement noise covariance; T represents the transpose; represents the predicted state at time k after correction; represents the measurement value at time k; represents the covariance of the prediction error at time k after correction. The measurement update formula is used to correct the predicted state based on the sensor data.

[0032] State updates provide predicted states, and measurement updates correct predicted values ​​based on observed data. Together, they form a closed-loop cycle of data fusion, continuously optimizing the system state. By combining model predictions with real-time observations, the two improve the robustness and accuracy of data fusion and form a complete state estimation framework.

[0033] S2: aligning the three-dimensional point cloud data collected by different sensors through the ICP algorithm, converting the three-dimensional point cloud data updated in real time into voxels, and storing the voxels using a multi-level voxel grid to construct a three-dimensional environment model.

[0034] Currently, the traditional storage method is a two-dimensional grid storage method. This solution introduces the perception and analysis of three-dimensional space, considering factors such as the height, shape, and material of objects. By obtaining three-dimensional point cloud data of the scene and converting this information into a three-dimensional model that the intelligent agent can understand through deep learning algorithms, it improves the comprehensive understanding of the road environment. The three-dimensional point cloud data, image information, and road conditions (such as road signs and lane lines) from the perception layer are integrated to construct a three-dimensional environment model that includes information such as height, object shape, and material. This model can not only help the intelligent agent capture objects, but also provide spatial relationships and depth information between objects. Through a multi-level voxel grid (Octree structure) combined with real-time data fusion, incremental updates, and deep learning methods, it can achieve efficient updates and iterations of three-dimensional storage, thereby providing the intelligent agent with stronger spatial understanding and decision support capabilities, ensuring higher adaptability and safety.

[0035] Specifically, the process of converting 3D point cloud data into voxels includes:

[0036] Calculate the position of the 3D point cloud data in 3D space and map it to a voxel grid with a fixed resolution. The voxel allocation formula is:

[0037] ;

[0038] in, represents a voxel in a voxel grid; Represents a point set in 3D point cloud data; Indicates the p The location attributes of each point within the voxel;

[0039] The mapping formula from spatial position to voxel coordinates is:

[0040] ;

[0041] in, Indicates the position of 3D point cloud data in 3D space; Indicates the resolution of the voxel grid; Indicates rounding down.

[0042] Furthermore, the storage format is:

[0043] Voxel(x,y,z)→{Object Type,Material,Height,Velocity,Reflection Data};

[0044] Among them, Voxel(x,y,z) represents the voxel value at position (x,y,z);

[0045] Object Type indicates the type of object (such as vehicle, sidewalk, etc.);

[0046] Material\text{Material}Material represents the material of an object (such as metal, glass, etc.);

[0047] Height\text{Height}Height represents the height or position of an object;

[0048] Velocity\text{Velocity} Velocity represents the speed of an object;

[0049] Reflection Data\text{Reflection Data}Reflection Data represents the reflection information of an object and is used for combining sensor data.

[0050] Three-dimensional point cloud data is a data structure used for three-dimensional object capture and scene reconstruction. Each point contains three-dimensional coordinates and additional attributes (such as reflection intensity). These points are aggregated through point cloud processing algorithms to form a complete three-dimensional environment model.

[0051] Furthermore, the three-dimensional point cloud data collected by different sensors are registered using an ICP algorithm, including:

[0052] Minimize the Euclidean distance between the midpoints of two 3D point cloud data. The calculation formula is:

[0053] ;

[0054] ;

[0055] Where n represents the number of points in the 3D point cloud data; Represents the i-th point in the three-dimensional point cloud data p; Represents the i-th point in the three-dimensional point cloud data q; represents the Euclidean distance; Indicates a point The three-dimensional coordinates of Indicates a point The three-dimensional coordinates of .

[0056] Furthermore, when converting three-dimensional point cloud data into voxel grids, this embodiment adopts an innovative mixed-resolution voxel grid construction method. The resolution of the voxel grid is automatically adjusted according to the distribution density and importance of objects in the road environment. In areas with dense vehicles or near key transportation facilities (such as intersections, bridges, etc.), high-resolution voxel grids are used to accurately represent the shape and spatial relationship of objects; in relatively open areas, lower resolutions are used to save storage space and computing resources. At the same time, combined with the prediction of the driving path of the intelligent agent, the construction of the voxel grid is optimized in advance to provide more efficient and accurate spatial information support for the decision-making of the intelligent agent. The resolution of the voxel grid can be dynamically adjusted according to the distribution density of the object and the predicted distance of the driving path of the intelligent agent. The adjustment formula is:

[0057] ;

[0058] in, Represents a voxel grid at a point in space resolution; Represents a spatial point The density of objects in the area; Indicates the distribution density threshold of high-density areas; Indicates the highest resolution value; The attenuation coefficient that indicates the effect of the predicted distance on the resolution; Represents a spatial point The predicted distance to the agent's travel path.

[0059] The mixed-resolution voxel grid construction method proposed in this embodiment combines a dynamic adjustment strategy for the distribution characteristics of objects in the road environment to maximize the accuracy and efficiency of data representation. This method generates a voxel grid model with regional differentiation characteristics by analyzing the density and importance of objects and the predicted driving path of the intelligent agent. In areas with high object density or key traffic nodes such as intersections and bridges, the system uses a high-resolution voxel grid to ensure that the detailed morphology and accurate spatial distribution of objects can be captured; in open areas, to reduce storage redundancy and computing resource consumption, the system chooses a low-resolution voxel grid for coarse representation.

[0060] In this embodiment, the ICP algorithm is used to register 3D point cloud data acquired by multiple sensors. The 3D point cloud to voxel conversion converts these point cloud data into a voxel grid that is easy to store and process. ICP aligns multiple point clouds by minimizing the Euclidean distance between point clouds and iteratively calculating the optimal rotation matrix and displacement vector.

[0061] The output of the ICP algorithm (aligned point cloud data) provides accurate spatial position data for the conversion of 3D point clouds to voxels. After ICP alignment, the errors in the point cloud data are minimized, and the generated point cloud can be more accurately converted to a voxel grid. In the voxel grid, the data is assigned to the position of each voxel. Therefore, these two steps are related. The ICP algorithm provides an accurate position reference for each voxel in 3D space.

[0062] Voxel grids are used to discretize 3D space and map objects to space. Each voxel represents a small unit in space and can be filled and updated based on point cloud data. In incremental updates, new point cloud data needs to be merged into the existing voxel grid. If the point cloud is updated frequently, the voxel values ​​need to be adjusted based on the new information.

[0063] ;

[0064] in, is the updated voxel grid value; is the weight coefficient, which is selected based on the credibility of the sensor data; is the new voxel value;

[0065] Through the combination of the above algorithmic formulas, including voxel grid generation, Octree data structure, ICP point cloud registration, Kalman filtering and incremental update and deep learning methods, the above methods work together to process, store and update real-time data, thereby ensuring the accuracy and real-time performance of the three-dimensional environment model, while ensuring efficient storage and update iteration. The voxel grid divides the three-dimensional space through voxel grids of uniform size, which saves more memory than point cloud storage, especially when storing large environmental models. In the real-time scene analysis of the interaction between multimodal perception intelligent agents and road environments, point cloud data provides detailed spatial information, while the voxel grid can better handle the balance between spatial resolution and computing requirements.

[0066] The expressive capabilities of the above models are enhanced by combining hybrid volumetric scene representations of voxels and point clouds, as well as technologies such as 3D Gaussian reconstruction. Specifically, by mapping point cloud data into a voxel grid, the regularity of voxels and the high precision of point clouds are utilized to enhance spatial representation capabilities. For example, point cloud data can be used as a basis, and voxel structures can be used for efficient indexing and searching. Newly collected point cloud data can be processed in real time through incremental updates. The surface of the point cloud can be reconstructed using a Gaussian mixture model (GMM), especially for complex object surfaces. In sparse areas of the point cloud, Gaussian distribution is used to simulate and fill in missing spatial information, providing a more continuous and accurate spatial representation. This method can be used to process point cloud data with noise and irregular gaps, helping intelligent agents make more accurate decisions in incomplete or irregular data environments.

[0067] To increase the creativity of this solution, this solution has been further optimized and innovated based on existing technologies. A deep learning-based adaptive fusion strategy is adopted when fusing 3D point cloud data, image information, and road status data at the data fusion layer. By learning from a large amount of actual road scene data, the system can automatically capture the importance weights of different data types under different traffic conditions and dynamically fuse them accordingly. For example, in a traffic congestion scenario, the length of vehicle queues and road sign information in the image information may be more critical. The system will increase their weights accordingly to ensure that the fused data can more accurately reflect the actual road conditions. This is a unique fusion method that distinguishes it from other technical solutions.

[0068] In this solution, the optimization of the data fusion layer realizes the dynamic weight adjustment of three-dimensional point cloud data, image information and road status data through the adaptive fusion strategy of deep learning, so as to more accurately characterize complex traffic conditions. Specifically, the system first uses a deep neural network to extract high-dimensional feature vectors of different data types. For example, the features of three-dimensional point cloud data reflect the elevation changes and obstacle distribution of the road surface, the features of image data capture vehicle density, queue length and road sign information, and the features of road status data describe vehicle speed, flow and congestion index. The extracted features are input into the fusion model through an adaptive weight allocation module, which is optimized based on the actual distribution of different traffic scenarios in the training samples. The system defines a dynamic weight calculation formula, which is as follows:

[0069] ;

[0070] in, represents the dynamic weight of the i-th category data, is the feature importance score of the i-th type of data in the current traffic scene, extracted by the scene capture module, is the feature importance score of the j-th type of data in the current traffic scene; is a trainable parameter used to adjust the sensitivity of the weight of the i-th type of data, is a trainable parameter that adjusts the sensitivity of the weight of the jth data type; N is the total number of data types. This formula achieves a normalized distribution of weights through a soft maximization function, ensuring that the sum of all weights is 1, thus making the contribution ratio of different data types have clear physical meaning.

[0071] In the specific implementation process, the system uses a trained deep learning model to classify traffic scenes, such as normal traffic, mild congestion and severe congestion, and adjusts the feature scores of various types of data based on the classification results. Taking a severely congested scene as an example, the system will significantly increase the weight of features in the image data, such as vehicle queue length and road signs, based on key features captured in the training data, while reducing its reliance on road elevation changes in the 3D point cloud data. After completing the weight assignment, features of different data types are linearly superimposed according to the weights to form a fused feature vector:

[0072] ;

[0073] Among them, F is the final fused feature vector, is the feature vector extracted from the i-th category of data. This dynamic fusion method not only allows the system to adjust data importance in real time based on varying traffic conditions, but also significantly improves the model's adaptability and prediction accuracy in complex traffic environments. Experimental results show that in congested traffic scenarios, this adaptive strategy can reduce traffic flow prediction errors by approximately 15% compared to traditional fixed-weight fusion methods. This deep learning-based dynamic fusion approach is one of the method's core innovations, fully embodying the integration of technical optimization and practical application scenarios.

[0074] In the application of the ICP algorithm, this solution develops a real-time dynamic ICP algorithm optimization mechanism tailored to the characteristics of intelligent agent-road interaction scenarios. This mechanism adjusts the ICP algorithm's iteration parameters in real time based on the agent's motion state and changes in the road environment. When the agent is traveling at high speeds or encountering complex road conditions (such as road construction or emergency scenes), the algorithm converges quickly, improving the real-time and accuracy of registration and reducing the error accumulation caused by environmental changes. This optimization, building on existing technologies, further enhances the practicality and effectiveness of the ICP algorithm in road scenarios.

[0075] Specifically, the number of iterations of the ICP algorithm and the convergence threshold of the registration error are dynamically adjusted according to the speed of the agent and the complexity index of the environment. The dynamic adjustment formula is:

[0076] ;

[0077] ;

[0078] in, Indicates the maximum number of iterations; Indicates the default number of iterations; The adjustment coefficient representing the complexity of the environment; Indicates the complexity of the environment, which is calculated based on the rate of change and local curvature of the 3D point cloud data; represents the adjusted registration error convergence threshold; Indicates the default registration error convergence threshold; Indicates the speed adjustment coefficient; Indicates the instantaneous speed of the agent. When the agent is in a complex environment, such as road construction or emergency scene, the complexity of the environment Improvement, the algorithm will appropriately increase the maximum number of iterations , to ensure that the registration process fully optimizes more feature points; and when driving at high speed, in order to avoid the algorithm calculation being too large and affecting real-time performance, the convergence threshold Appropriate relaxation allows the algorithm to complete the matching in fewer iterations.

[0079] Furthermore, to further enhance robustness, this mechanism incorporates a deep learning-based initial pose estimation module. This coarse registration of the source point cloud reduces the initial error range, thereby reducing the iterative burden of the ICP algorithm. Experiments have shown that this dynamic optimization mechanism significantly improves registration accuracy and efficiency in complex road scenarios.

[0080] S3: In the three-dimensional environment model, the road space occupancy is analyzed based on the distribution and attribute information of objects in the voxel grid (for example, in a vehicle-dense area, the traffic congestion level of the area is evaluated by counting the number and distribution range of vehicle-related voxels in the voxel grid), and the future movement trend of the object is predicted based on the position change of the object in the voxel grid.

[0081] Specifically, within the 3D environmental model, the voxel grid is first traversed to count the number of voxels associated with various types of objects within different areas. For example, in areas with dense traffic, the number of voxels associated with vehicles is counted. Combined with their distribution, the system analyzes road space occupancy and determines the area's traffic congestion level based on pre-set congestion assessment criteria. Simultaneously, the position information of the corresponding voxels in the voxel grid is continuously tracked, recording position data at different times. By calculating parameters such as position change, speed, and direction, the system uses kinematic models or machine learning algorithms, such as Kalman filtering, to predict the object's future motion trends.

[0082] Combine historical traffic data and real-time environmental information to analyze the changing patterns of traffic flow and provide driving path references for intelligent agents.

[0083] Specifically, historical traffic data is collected over a specific time period, including information on traffic volume, speed, and congestion across different time periods and road sections. Sensors are also used to obtain real-time information about the road environment, such as weather conditions, road construction, and current traffic volume. This historical data is categorized and statistically analyzed by time and road section, identifying patterns and trends. These patterns are then modified and adjusted based on real-time environmental information. For example, if road construction occurs, traffic flow patterns on that section will change. Finally, the comprehensive analysis results, taking into account factors such as real-time road conditions, estimated travel times, and congestion risks for different routes, allow the agent to plan a more reasonable driving route as a reference.

[0084] S4: Generate a decision based on the analysis results including road space occupancy, future movement trends of objects, and reference to the agent's driving path. The agent executes the decision and provides real-time feedback on the execution results to update the three-dimensional environment model.

[0085] Specifically, the decision-making process is as follows:

[0086] First, the system considers road space occupancy to determine which areas are passable and which are congested or obstructed. It also predicts potential dynamic risks based on the future movement of objects, such as whether a vehicle will suddenly change lanes or a pedestrian will cross the road. It also considers the rationality and safety of the agent's travel path, taking into account the agent's previous travel path. This information is fed into a decision-making algorithm or model, which weighs various factors and generates specific decision instructions, such as acceleration, deceleration, and turning. The agent then executes the corresponding action, while its sensors continuously collect environmental data during the execution process, such as changes in position and the new state of surrounding objects. This real-time feedback is used to update the 3D environmental model, revising the position and state of objects in the model to ensure a high degree of consistency between the model and the actual environment, providing a more accurate basis for the next decision.

[0087] The real-time scene analysis method for interaction between an intelligent agent and a city road provided in this embodiment has the following beneficial effects:

[0088] 1. This method incorporates multimodal perception technology, enabling the intelligent agent to obtain comprehensive information from multiple sensing devices, including cameras, lidar, radar sensors, and infrared sensors. Through real-time data fusion, the intelligent agent can fully perceive the road environment, capturing not only static objects (such as road signs and lane markings) but also dynamic objects (such as vehicles and pedestrians) in real time. By combining 3D point cloud data with voxel grid technology, the intelligent agent's spatial understanding is significantly enhanced, enabling it to more accurately grasp the spatial relationships between objects and their motion trajectories, providing stronger support for subsequent decision-making and action execution.

[0089] 2. By constructing a three-dimensional environmental model that incorporates information such as height, object shape, and material, this method enables the agent to make more accurate and intelligent decisions in complex traffic scenarios. The model accounts for dynamically changing road conditions, real-time traffic flow, and various environmental changes (such as weather and road conditions), enabling the agent to flexibly respond to diverse scenarios. Assisted by deep learning algorithms, the model is continuously optimized and iterated, improving the agent's adaptability and safety in various complex environments and reducing the incidence of traffic accidents.

[0090] 3. This method introduces a voxel grid (Octree structure) and an incremental update mechanism, enabling efficient spatial storage of 3D environment models. This hierarchical spatial partitioning and data compression not only significantly saves storage space but also improves data update efficiency. In real-time scenarios, where the environment is constantly changing (such as objects appearing, disappearing, or shifting), the incremental update mechanism allows new environmental information to be quickly incorporated into the existing model, ensuring that the agent maintains real-time updates and an accurate understanding of the environment.

[0091] 4. Use the ICP algorithm to register multi-sensor data, effectively aligning 3D point cloud data collected by different sensors and minimizing errors, making the point cloud data more accurate. After processing by the ICP algorithm, the point cloud data can be accurately converted into a voxel grid, enabling efficient 3D environment reconstruction. This process accurately extracts and stores information such as the height, shape, and material of objects, providing the intelligent agent with more detailed and precise object capture capabilities and improving the quality of scene reconstruction.

[0092] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0093] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A real-time scene analysis method for interaction between an intelligent agent and a city road, characterized by: include: S1: The intelligent agent uses sensors installed on the intelligent agent to perceive the road environment in real time, collects 3D point cloud data of the road environment, and updates it in real time; S2: registering the three-dimensional point cloud data collected by different sensors using an ICP algorithm, converting the three-dimensional point cloud data updated in real time into voxels, and storing the voxels using a multi-level voxel grid to construct a three-dimensional environment model; The number of iterations of the ICP algorithm and the convergence threshold of the registration error are dynamically adjusted according to the speed of the agent and the complexity index of the environment. The dynamic adjustment formula is: ; ; in, Indicates the maximum number of iterations; Indicates the default number of iterations; The adjustment coefficient representing the complexity of the environment; Indicates the complexity of the environment, which is calculated based on the rate of change and local curvature of the 3D point cloud data; represents the adjusted registration error convergence threshold; Indicates the default registration error convergence threshold; Indicates the speed adjustment coefficient; represents the instantaneous speed of the agent; S3: In the three-dimensional environment model, analyzing the road space occupancy according to the distribution and attribute information of the objects in the voxel grid, and predicting the future movement trend of the objects according to the position changes of the objects in the voxel grid; Combine historical traffic data with real-time environmental information to analyze traffic flow changes and provide driving path references for intelligent agents; S4: Generate a decision based on the analysis results including road space occupancy, future movement trends of objects, and reference to the agent's driving path. The agent executes the decision and provides real-time feedback on the execution results to update the three-dimensional environment model.

2. The method for real-time scene analysis of interaction between an intelligent agent and a city road according to claim 1, characterized in that: The sensor comprises: Cameras are used to collect image information including traffic signs, road markings, pedestrians, vehicles, and road status data; LiDAR, used to collect the first three-dimensional point cloud data of the obstacle's position, speed, and distance from the intelligent body; Millimeter-wave radar, used to collect second and third-dimensional point cloud data of the speed of surrounding objects and their distance from the intelligent body; Ultrasonic sensors for detecting objects at close range; Noise sensor, used to monitor environmental noise.

3. The real-time scene analysis method for interaction between an intelligent agent and a city road according to claim 1 is characterized in that: Real-time updates of 3D point cloud data include: Status update, the update formula is: ; in, represents the predicted state at time k; A represents the state transfer matrix; represents the predicted state at time k-1 after correction; B represents the control input matrix; represents the control input at time k; Measurement update, the update formula is: ; ; ; in, represents the Kalman gain; represents the prediction error covariance at time k; H represents the measurement matrix; R represents the measurement noise covariance; T represents the transpose; represents the predicted state at time k after correction; represents the measurement value at time k; Represents the covariance of the forecast error at time k after correction.

4. The method for real-time scene analysis of interaction between an intelligent agent and a city road according to claim 1, characterized in that: The process of converting 3D point cloud data into voxels includes: Calculate the position of the 3D point cloud data in 3D space and map it to a voxel grid with a fixed resolution. The voxel allocation formula is: ; in, represents a voxel in a voxel grid; Represents a point set in 3D point cloud data; Indicates the p The location attributes of each point within the voxel; The mapping formula from spatial position to voxel coordinates is: ; in, Indicates the position of 3D point cloud data in 3D space; Indicates the resolution of the voxel grid; Indicates rounding down.

5. The method for real-time scene analysis of interaction between an intelligent agent and a city road according to claim 4 is characterized in that: The storage format is: Voxel(x,y,z)→{Object Type,Material,Height,Velocity,Reflection Data}; Among them, Voxel(x,y,z) represents the voxel value at position (x,y,z); Object Type indicates the type of object; Material\text{Material}Material represents the material of an object; Height\text{Height}Height represents the height or position of an object; Velocity\text{Velocity} Velocity represents the speed of an object; Reflection Data\text{Reflection Data}Reflection Data represents the reflection information of an object and is used for combining sensor data.

6. The method for real-time scene analysis of interaction between an intelligent agent and a city road according to claim 1, characterized in that: The three-dimensional point cloud data collected by different sensors is registered using an ICP algorithm, including: Minimize the Euclidean distance between the midpoints of two 3D point cloud data. The calculation formula is: ; ; Where n represents the number of points in the 3D point cloud data; Represents the i-th point in the three-dimensional point cloud data p; Represents the i-th point in the three-dimensional point cloud data q; represents the Euclidean distance; Indicates a point The three-dimensional coordinates of Indicates a point The three-dimensional coordinates of .

7. The method for real-time scene analysis of interaction between an intelligent agent and a city road according to claim 1, characterized in that: The resolution of the voxel grid is dynamically adjusted based on the object distribution density and the predicted distance of the agent's travel path. The adjustment formula is: ; in, Represents a voxel grid at a point in space resolution; Represents a spatial point The density of objects in the area; Indicates the distribution density threshold of high-density areas; Indicates the highest resolution value; The attenuation coefficient that indicates the effect of the predicted distance on the resolution; Represents a spatial point The predicted distance to the agent's travel path.

Citation Information

Patent Citations

  • Intelligent inspection device and alarm method

    CN118658126A

  • Intelligent vehicle dynamic target three-dimensional sensing method in complex environment

    CN119152490A

  • Road network operation situation research method based on multi-source data fusion

    CN120126306A