Intelligent agent and urban road interaction real-time scene analysis method
By collecting three-dimensional point cloud data and constructing voxel grid model, the problem that agents in the prior art cannot accurately understand the three-dimensional structure of objects is solved, and more efficient spatial understanding and decision-making support are achieved.
Patent Information
- Application Number
- CN202510824864.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Existing real-time scenario analysis methods use two-dimensional grid storage to fail to accurately understand the three-dimensional structure and spatial location of objects, resulting in reduced decision-making accuracy of agents in complex or dynamic environments.
By collecting three-dimensional point cloud data, registering and converting it into voxel mesh using ICP algorithm, combining multi-level voxel mesh storage and deep learning, a three-dimensional environmental model is built, road space occupation and object movement trends are analyzed, decisions are generated and real-time updates are updated.
It improves the agent's spatial understanding of complex environments, enhances the accuracy and adaptability of decisions, and reduces the incidence of traffic accidents.
Smart Images

Figure CN120340262A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of scene analysis, and particularly to a real-time scene analysis method for the interaction between an agent and an urban road. Background Art
[0002] With the continuous development of intelligent transportation systems, the real-time interaction between agents (such as autonomous vehicles, driverless buses, etc.) and urban roads has become an important research direction for improving road traffic efficiency, safety, and intelligence. As one of the key technologies, multi-modal perception technology can provide more comprehensive and accurate road environment information for agents by fusing data from different sensors (such as vision, radar, lidar, infrared sensors, etc.).
[0003] Existing real-time scene analysis methods use traditional two-dimensional grid storage methods. By projecting scene data onto a plane, two-dimensional grid storage methods can usually only process information on the plane (such as position, speed, etc.), while ignoring three-dimensional features of objects such as height, shape, and material. This makes it impossible for agents to accurately understand the three-dimensional structure of objects or their relative positions in space when perceiving and analyzing the environment, thus limiting the agent's ability to understand complex three-dimensional environments. Two-dimensional grid storage methods are difficult to effectively represent the three-dimensional forms and changes of these objects, resulting in a situation where agents are prone to reduced accuracy or information loss when dealing with complex or highly dynamic environments, affecting their decision-making and operations. Summary of the Invention
[0004] Based on this, it is necessary to provide a real-time scene analysis method for the interaction between an agent and an urban road, which includes: S1: Real-time sense the road environment through sensors installed on the agent, collect three-dimensional point cloud data of the road environment, and update it in real time; S2: Register the three-dimensional point cloud data collected by different sensors through the ICP algorithm, convert the real-time updated three-dimensional point cloud data into voxels, and store the voxels using a multi-level voxel grid to construct a three-dimensional environment model; S3: In the three-dimensional environment model, analyze the road space occupancy according to the distribution and attribute information of objects in the voxel grid, and predict the future movement trend of objects according to the position changes of objects in the voxel grid; Analyze the change law of traffic flow by combining historical traffic data and real-time environment information, and provide a driving path reference for the agent; S4: Generate a decision based on the analysis results including road space occupancy, future movement trend of objects, and driving path reference for the agent. The agent executes the decision and feeds back the execution result in real time to update the three-dimensional environment model.
[0005] Beneficial effects: By combining 3D point cloud data and voxel grid technology, the spatial understanding ability of the intelligent agent is significantly enhanced, enabling it to more accurately grasp the spatial relationships between objects and their movement trajectories, providing stronger support for subsequent decision-making and behavior execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0007] Figure 1 It is a flowchart of the method for real-time scene analysis of the interaction between an intelligent agent and an urban road in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0008] To make the above objects, features, and advantages of the present application more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present application with reference to the drawings. Many specific details are set forth in the following description to fully understand the present application. However, the present application can be implemented in many other ways different from those described herein. Those skilled in the art can make similar improvements without departing from the connotation of the present application. Therefore, the present application is not limited by the specific embodiments disclosed below.
[0009] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0010] As Figure 1 shown, this embodiment provides a method for real-time scene analysis of the interaction between an intelligent agent and an urban road. The method includes: S1: Real-time sense the road environment through the sensors set on the intelligent agent, collect the 3D point cloud data of the road environment, and update it in real time.
[0011] Specifically, the sensors include: A camera for collecting image information including traffic signs, road markings, pedestrians, vehicles, and road status data; A lidar for collecting the first 3D point cloud data of the positions, speeds, and distances of obstacles from the intelligent agent; A millimeter-wave radar for collecting second three-dimensional point cloud data on the speeds of surrounding objects and their distances from the agent; An ultrasonic sensor for monitoring nearby objects, especially effective when parking or driving at low speeds; A noise sensor for monitoring ambient noise to judge traffic flow, road conditions, and special events (such as traffic accidents, sudden obstacles, etc.).
[0012] Furthermore, the real-time update of the three-dimensional point cloud data includes: State update, with the update formula: ; where represents the predicted state at time k; A represents the state transition matrix; represents the predicted state at time k-1 after correction; B represents the control input matrix; represents the control input at time k; The state update formula is used to predict the state of the system at the next moment, depending on the motion model of the system and the current state information.
[0013] Measurement update, with the update formula: ; ; ; where represents the Kalman gain; represents the predicted error covariance at time k; H represents the measurement matrix; R represents the measurement noise covariance; T represents the transpose; represents the predicted state at time k after correction; represents the measurement value at time k; represents the predicted error covariance at time k after correction. The measurement update formula is used to correct the predicted state based on sensor data.
[0014] The state update provides the predicted state, and the measurement update corrects the predicted value based on the observed data. The two together form a closed-loop cycle of data fusion, continuously optimizing the system state. The two cooperate to improve the robustness and accuracy of data fusion by combining model prediction and real-time observation, forming a complete state estimation framework.
[0015] S2: Register the three-dimensional point cloud data collected by different sensors through the ICP algorithm, convert the real-time updated three-dimensional point cloud data into voxels, and store the voxels using a multi-level voxel grid to construct a three-dimensional environmental model.
[0016] Currently, the traditional storage method is a two-dimensional grid storage method. This solution introduces the perception and analysis of three-dimensional space, considering factors such as the height, shape, and material of objects. By obtaining the three-dimensional point cloud data of the scene and using deep learning algorithms to convert this information into a three-dimensional model that can be understood by the agent, the comprehensive understanding of the road environment is enhanced. The three-dimensional point cloud data, image information, and road status (such as road signs, lane lines, etc.) from the perception layer are fused to construct a three-dimensional environmental model that includes information such as height, object shape, and material. This model can not only help the agent capture objects but also provide the spatial relationship and depth information between objects. Through a multi-level voxel grid (Octree structure) combined with real-time data fusion, incremental update, and deep learning methods, efficient update and iteration of three-dimensional storage can be achieved, thereby providing the agent with stronger spatial understanding and decision support capabilities, ensuring higher adaptability and security.
[0017] Specifically, the process of converting three-dimensional point cloud data into voxels includes: Calculate the position of the three-dimensional point cloud data in three-dimensional space and map it to a voxel grid with a fixed resolution. The voxel assignment formula is: ; Where, represents a voxel in the voxel grid; represents a point set in the three-dimensional point cloud data; represents the position attribute of the p th point within the voxel; The mapping formula from spatial position to voxel coordinates is: ; Where, represents the position of the three-dimensional point cloud data in three-dimensional space; represents the resolution of the voxel grid; represents rounding down.
[0018] Furthermore, the storage format is: Voxel(x,y,z)→{Object Type,Material,Height,Velocity,Reflection Data}; Where, Voxel(x,y,z) represents the voxel value at position (x,y,z); Object Type represents the type of the object (such as vehicle, sidewalk, etc.); Material\text{Material}Material represents the material of the object (such as metal, glass, etc.); Height represents the height or position of an object; Velocity represents the moving speed of an object; Reflection Data represents the reflection information of an object and is used for the combination of sensor data.
[0019] Three-dimensional point cloud data is a data structure used for three-dimensional object capture and scene reconstruction. Each point contains three-dimensional coordinates and additional attributes (such as reflection intensity). These points are aggregated through point cloud processing algorithms to form a complete three-dimensional environment model.
[0020] Furthermore, the three-dimensional point cloud data collected by different sensors is registered through the ICP algorithm, including: Minimize the Euclidean distance between points in two three-dimensional point cloud data, and the calculation formula is: ; ; where n represents the number of points in the three-dimensional point cloud data; represents the i-th point in the three-dimensional point cloud data p; represents the i-th point in the three-dimensional point cloud data q; represents the Euclidean distance; represents the point 's three-dimensional coordinates; represents the point 's three-dimensional coordinates.
[0021] Furthermore, when converting the three-dimensional point cloud data into a voxel grid, this embodiment adopts an innovative hybrid-resolution voxel grid construction method. According to the distribution density and importance degree of objects in the road environment, the resolution of the voxel grid is automatically adjusted. In areas with dense vehicles or near key traffic facilities (such as intersections, bridges, etc.), a high-resolution voxel grid is used to accurately represent the object shape and spatial relationship; in relatively empty areas, a lower resolution is adopted to save storage space and computing resources. At the same time, combined with the prediction of the driving path of the intelligent agent, the construction of the voxel grid is optimized in advance to provide more efficient and accurate spatial information support for the decision-making of the intelligent agent. The resolution of the voxel grid can be dynamically adjusted according to the object distribution density and the predicted distance of the intelligent agent's driving path, and the adjustment formula is: ; where, represents the resolution of the voxel grid at the spatial point ; Represents the object distribution density in the area where the spatial point is located ; Represents the distribution density threshold of the high-density area Represents the highest resolution value Represents the attenuation coefficient of the influence of the predicted distance on the resolution Represents the spatial point The predicted distance to the agent's driving path
[0022] The hybrid resolution voxel grid construction method proposed in this embodiment combines a dynamic adjustment strategy for the object distribution characteristics in the road environment to maximize the accuracy and efficiency of data representation. This method generates a voxel grid model with regional differentiation characteristics by analyzing object density, importance level, and the agent's driving path prediction. In high object density areas or key traffic nodes, such as intersections and bridges, the system uses high-resolution voxel grids to ensure that the detailed shape and precise spatial distribution of objects can be captured; while in open areas, to reduce storage redundancy and computational resource consumption, the system selects low-resolution voxel grids for rough representation.
[0023] In this embodiment, the ICP algorithm is used to register the three-dimensional point cloud data obtained by multiple sensors, and the conversion from three-dimensional point cloud to voxel is to convert this point cloud data into a voxel grid that is convenient for storage and processing. ICP continuously iteratively calculates the optimal rotation matrix and displacement vector by minimizing the Euclidean distance between point clouds, so as to align multiple point clouds; The output of the ICP algorithm (the aligned point cloud data) provides accurate spatial position data for the conversion from three-dimensional point cloud to voxel. After ICP registration, the error in the point cloud data is minimized, and the generated point cloud can be more accurately converted into a voxel grid. In the voxel grid, the data is assigned to the position of each voxel. Therefore, these two steps are related in sequence, and the ICP algorithm provides an accurate position reference for the assignment of each voxel in three-dimensional space; The voxel grid is used to discretize the three-dimensional space and perform spatial mapping on objects. Each voxel represents a small unit in space and can be filled and updated according to the point cloud data. In incremental updates, new point cloud data needs to be merged into the existing voxel grid. If there are many point cloud updates, the voxel values need to be adjusted according to the new information; ; Among them, is the updated voxel grid value; is the weight coefficient, which is selected according to the credibility of the sensor data; is the new voxel value; Through the combination of the above algorithm formulas, including voxel grid generation, Octree data structure, ICP point cloud registration, Kalman filtering and incremental update, and deep learning methods, the above methods work together on the processing, storage and update of real-time data, thus ensuring the accuracy and real-time performance of the 3D environmental model, while ensuring efficient storage and update iteration. The voxel grid divides the 3D space by voxel cells of uniform size, which saves more memory than point cloud storage, especially when storing large environmental models. In the real-time scene analysis of the interaction between multi-modal perception agents and the road environment, point cloud data provides detailed spatial information, while the voxel grid can better handle the balance between spatial resolution and computational requirements.
[0024] Combined with hybrid volume scene representations that integrate voxels and point clouds, as well as techniques such as 3D Gaussian reconstruction, etc., to enhance the performance of the above model. Specifically, by mapping point cloud data into the voxel grid, the regularity of voxels and the high precision of point clouds are used to enhance the spatial representation ability. For example, point cloud data can be used as the basis, and the voxel structure can be used for efficient indexing and searching. Newly acquired point cloud data can be processed in real time through incremental update. The surface reconstruction of point clouds can be carried out through the Gaussian mixture model (GMM), especially for complex object surfaces. In the sparse areas of point clouds, Gaussian distribution is used to simulate and fill in the missing spatial information, providing a more continuous and accurate spatial representation. By processing point cloud data with noise and irregular gaps in this way, it helps the agent to make more accurate decisions in an incomplete or irregular data environment.
[0025] To increase the creativity of this solution, this solution has been further optimized and innovated on the basis of existing technologies. When fusing 3D point cloud data, image information and road state data at the data fusion layer, an adaptive fusion strategy based on deep learning is adopted. Through learning a large amount of actual road scene data, the system can automatically capture the importance weights of different data types under different traffic conditions and perform dynamic fusion accordingly. For example, in a traffic congestion scenario, the vehicle queue length and road sign information in the image information may be more critical, and the system will increase their weights accordingly to ensure that the fused data can more accurately reflect the actual road conditions. This is a unique fusion method different from other technical solutions.
[0026] In this solution, the optimization of the data fusion layer realizes the dynamic weight adjustment of 3D point cloud data, image information, and road status data through the adaptive fusion strategy of deep learning, thereby more accurately representing complex traffic conditions. Specifically, the system first uses a deep neural network to extract high-dimensional feature vectors of different data types. For example, the features of 3D point cloud data reflect the elevation changes of the road surface and the distribution of obstacles, the features of image data capture vehicle density, queue length, and road sign information, and the features of road status data describe vehicle speed, traffic flow, and congestion index. The extracted features are input into the fusion model through an adaptive weight allocation module, which is optimized based on the actual distribution of different traffic scenarios in the training samples. The system defines a dynamic weight calculation formula, specifically: ; where, represents the dynamic weight of the i-th type of data, is the feature importance score of the i-th type of data in the current traffic scenario, extracted by the scenario capture module, is the feature importance score of the j-th type of data in the current traffic scenario; is a trainable parameter for adjusting the weight sensitivity of the i-th type of data, is a trainable parameter for adjusting the weight sensitivity of the j-th type of data; N is the total number of data categories. This formula realizes the normalized distribution of weights through the softmax function, ensuring that the sum of all weights is 1, so that the contribution ratio of different data types has a clear physical meaning.
[0027] In the specific implementation process, the system uses a trained deep learning model to classify traffic scenarios, such as normal traffic, mild congestion, and severe congestion, and adjusts the feature scores of various types of data according to the classification results . Taking the severe congestion scenario as an example, the system will significantly increase the weights of features such as vehicle queue length and road signs in the image data according to the key features captured in the training data, while reducing the dependence on the road elevation changes in the 3D point cloud data. After completing the weight allocation, the features of different data types are linearly superimposed according to the weights to form a fused feature vector: ; where, F is the finally fused feature vector, The feature vector extracted for the i-th type of data. Through this dynamic fusion method, the system can not only adjust the data importance in real time according to different traffic conditions, but also significantly improve the adaptability and prediction accuracy of the model to complex traffic environments. Experimental results show that in traffic congestion scenarios, compared with the traditional fixed-weight fusion method, this adaptive strategy can reduce the error of traffic flow prediction by about 15%. This dynamic fusion method based on deep learning is one of the core innovations of this method, fully reflecting the combination of technology optimization and actual application scenarios.
[0028] In the application of the ICP algorithm, this solution develops a real-time dynamic ICP algorithm optimization mechanism for the characteristics of the intelligent agent-road interaction scenario. This mechanism can adjust the iterative parameters of the ICP algorithm in real time according to the motion state of the intelligent agent and the changes in the road environment. When the intelligent agent is driving at high speed or encountering complex road conditions (such as road construction, emergency sites, etc.), the algorithm can converge quickly, improve the real-time performance and accuracy of registration, and reduce the error accumulation caused by environmental changes. This optimization further improves the practicality and effectiveness of the ICP algorithm in road scenarios on the basis of existing technologies.
[0029] Specifically, the number of iterations of the ICP algorithm and the registration error convergence threshold are dynamically adjusted according to the speed of the intelligent agent and the environmental complexity index. The dynamic adjustment formula is: ; ; where represents the maximum number of iterations; represents the default number of iterations; represents the adjustment coefficient of environmental complexity; represents the environmental complexity, calculated based on the change rate and local curvature of the three-dimensional point cloud data; represents the adjusted registration error convergence threshold; represents the default registration error convergence threshold; represents the adjustment coefficient of speed; represents the instantaneous speed of the intelligent agent. When the intelligent agent is in a complex environment, such as road construction or an emergency site, the environmental complexity increases, and the algorithm will appropriately increase the maximum number of iterations to ensure that the registration process fully optimizes more feature points; while when driving at high speed, to avoid excessive computational complexity of the algorithm affecting real-time performance, the convergence threshold is appropriately relaxed, enabling the algorithm to complete the matching in fewer iterations.
[0030] In addition, to further enhance robustness, this mechanism combines an initial pose estimation module based on deep learning to reduce the initial error range through rough registration of the source point cloud, thereby reducing the iterative burden of the ICP algorithm. Experiments show that this dynamic optimization mechanism can significantly improve the registration accuracy and efficiency in complex road scenarios.
[0031] S3: In the three-dimensional environment model, analyze the road space occupancy according to the distribution and attribute information of objects in the voxel grid (for example, in a vehicle-dense area, evaluate the traffic congestion level in this area by counting the number and distribution range of vehicle-related voxels in the voxel grid), and predict the future movement trend of objects according to the position changes of objects in the voxel grid.
[0032] Specifically, in the three-dimensional environment model, first traverse the voxel grid, count the number of voxels related to various objects in different regions. For example, in a vehicle-dense area, count the number of vehicle-related voxels, and combine its distribution range. According to the preset congestion level evaluation criteria, analyze and obtain the road space occupancy and determine the traffic congestion level in this area. At the same time, continuously track the position information of the voxels corresponding to the objects in the voxel grid, record the position data at different times, and calculate parameters such as the position change amount, speed, and direction. Then use a kinematic model or machine learning algorithm, such as Kalman filter, to predict the future movement trend of objects.
[0033] Combine historical traffic data and real-time environmental information to analyze the change law of traffic flow and provide a driving path reference for the intelligent agent.
[0034] Specifically, first collect historical traffic data within a certain period of time, including information such as traffic flow, vehicle speed, and congestion conditions at different times and on different road sections. At the same time, use sensors to obtain the current road environment information in real time, such as weather conditions, road construction conditions, and current traffic flow. Classify, organize, and statistically analyze the historical data according to dimensions such as time and road sections to find out features such as periodic laws and trend changes. Then, combine the real-time environmental information to correct and adjust these laws. For example, when there is road construction, the traffic flow law of the corresponding road section will change. Finally, comprehensively analyze the results, consider factors such as the real-time road conditions, estimated travel time, and congestion risk of different paths, and plan a relatively reasonable driving path for the intelligent agent as a reference.
[0035] S4: Generate a decision based on the analysis results including road space occupancy, future movement trend of objects, and driving path reference for the intelligent agent. The intelligent agent executes the decision and feeds back the execution result in real time to update the three-dimensional environment model.
[0036] Specifically, the process of generating a decision is as follows: First, comprehensively consider the road space occupancy to determine which areas are passable, which areas are congested or have obstacles; based on the future movement trends of objects, predict possible dynamic risks, such as whether a vehicle will suddenly change lanes or a pedestrian will cross the road. At the same time, refer to the driving path of the agent and consider the rationality and safety of the path. Input this information into a decision-making generation algorithm or model, which will weigh various factors and generate specific decision instructions such as accelerating, decelerating, and turning. Then, after receiving the instructions, the agent executes the corresponding actions, and its sensors continuously collect environmental data during the execution process, such as position changes and new states of surrounding objects. These real-time feedback data are used to update the three-dimensional environmental model, correct the position, state, etc. of the objects in the model, ensure that the model always maintains a high degree of fit with the actual environment, and provide a more accurate basis for the next decision.
[0037] The intelligent agent and urban road interaction real-time scene analysis method provided in this embodiment has the following beneficial effects: 1. By introducing multi-modal perception technology, the agent can obtain comprehensive information from various perception devices such as cameras, lidar, radar sensors, and infrared sensors. Through real-time data fusion, the agent can comprehensively perceive the road environment, not only capture static objects (such as road signs, lane lines, etc.), but also track dynamic objects (such as vehicles, pedestrians, etc.) in real time. Combining three-dimensional point cloud data and voxel grid technology, the agent's spatial understanding ability is significantly enhanced, and it can more accurately grasp the spatial relationships between objects and their movement trajectories, providing stronger support for subsequent decision-making and behavior execution.
[0038] 2. By constructing a three-dimensional environmental model that includes information such as height, object shape, and material, the agent can make more accurate and intelligent decisions in complex traffic scenarios. The model takes into account dynamic road conditions, real-time traffic flow, and various changes in the environment (such as weather conditions, road surface conditions, etc.), enabling the agent to flexibly respond to different scenarios. With the assistance of deep learning algorithms, the model is continuously optimized and iterated, thereby improving the adaptability and safety of the agent in various complex environments and reducing the incidence of traffic accidents.
[0039] 3. The method introduces a voxel grid (Octree structure) and an incremental update mechanism, enabling the three-dimensional environmental model to be efficiently stored in space. Through hierarchical spatial partitioning and data compression methods, not only is the storage space significantly saved, but the data update efficiency is also improved. In a real-time scene, the environment is constantly changing (such as the appearance, disappearance, or displacement of objects), and the incremental update mechanism enables new environmental information to be quickly incorporated into the existing model, ensuring that the agent can update in real time and maintain an accurate understanding of the environment.
[0040] 4. Use the ICP algorithm to register multi-sensor data, effectively align the 3D point cloud data collected by different sensors, and minimize errors to make the point cloud data more accurate. After being processed by the ICP algorithm, the point cloud data can be accurately converted into a voxel grid to achieve efficient 3D environment reconstruction. During this process, information such as the height, shape, and material of the object is accurately extracted and stored, providing the intelligent agent with more detailed and accurate object capture capabilities while improving the quality of scene reconstruction.
[0041] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0042] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for real-time scene analysis of the interaction between an intelligent agent and an urban road, characterized in that, Including: S1: The road environment is perceived in real time through sensors set on the agent, three-dimensional point cloud data of the road environment is collected and updated in real time; S2: The three-dimensional point cloud data collected by different sensors is registered through the ICP algorithm, the real-time updated three-dimensional point cloud data is converted into voxels, and the voxels are stored using a multi-level voxel grid to construct a three-dimensional environment model; S3: In the three-dimensional environment model, the road space occupancy is analyzed according to the distribution and attribute information of objects in the voxel grid, and the future movement trend of objects is predicted according to the position changes of objects in the voxel grid; Combined with historical traffic data and real-time environment information, analyze the change law of traffic flow, and provide a driving path reference for the agent; S4: Generate a decision based on the analysis results including road space occupancy, future movement trend of objects, and driving path reference of the agent, the agent executes the decision, and the execution result is fed back in real time to update the three-dimensional environment model.
2. The method for real-time scenario analysis of the interaction between an intelligent agent and an urban road according to claim 1, wherein The sensors include: A camera for collecting image information including traffic signs, road markings, pedestrians, vehicles, and road status data; A lidar for collecting first three-dimensional point cloud data of the position, speed, and distance from the agent of obstacles; A millimeter-wave radar for collecting second three-dimensional point cloud data of the speed and distance from the agent of surrounding objects; An ultrasonic sensor for monitoring nearby objects; A noise sensor for monitoring environmental noise.
3. The real-time scene analysis method for the interaction between an intelligent agent and an urban road according to claim 1, wherein, The real-time update of the three-dimensional point cloud data includes: State update, and the update formula is: ; Among them, represents the predicted state at time k; A represents the state transition matrix; represents the predicted state at time k-1 after correction; B represents the control input matrix; represents the control input at time k; Measurement update, and the update formula is: ; ; ; Among them, represents the Kalman gain; represents the predicted error covariance at time k; H represents the measurement matrix; R represents the measurement noise covariance; T represents the transpose; represents the predicted state at time k after correction; represents the measurement value at time k; represents the predicted error covariance at time k after correction.
4. The real-time scene analysis method for the interaction between an intelligent agent and an urban road according to claim 1, wherein The process of converting the three-dimensional point cloud data into voxels includes: Calculate the position of the three-dimensional point cloud data in three-dimensional space and map it to a voxel grid with a fixed resolution. The voxel assignment formula is: ; Among them, represents a voxel in the voxel grid; represents a point set in the three-dimensional point cloud data; represents the p th position attribute of a point within the voxel. The mapping formula from spatial position to voxel coordinates is: ; Among them, represents the position of the three-dimensional point cloud data in three-dimensional space; represents the resolution of the voxel grid; represents rounding down.
5. The method for real-time scenario analysis of the interaction between an intelligent agent and an urban road according to claim 4, wherein The storage format is: Voxel(x,y,z)→{Object Type,Material,Height,Velocity,Reflection Data}; Among them, Voxel(x,y,z) represents the voxel value at the position (x,y,z); Object Type represents the type of the object; Material\text{Material}Material represents the material of the object; Height\text{Height}Height represents the height or position of the object; Velocity\text{Velocity}Velocity represents the movement speed of the object; Reflection Data\text{Reflection Data}Reflection Data represents the reflection information of the object and is used for the combination of sensor data.
6. The real-time scene analysis method for the interaction between an intelligent agent and an urban road according to claim 1, wherein Registering the three-dimensional point cloud data collected by different sensors through the ICP algorithm includes: Minimize the Euclidean distance between points in the two three-dimensional point cloud data, and the calculation formula is: ; ; Among them, n represents the number of points in the three-dimensional point cloud data; represents the i-th point in the three-dimensional point cloud data p; represents the i-th point in the three-dimensional point cloud data q; represents the Euclidean distance; represents the point whose three-dimensional coordinates are; represents the point whose three-dimensional coordinates are.
7. The real-time scene analysis method for the interaction between an intelligent agent and an urban road according to claim 1, characterized in that, It also includes dynamically adjusting the iteration times of the ICP algorithm and the registration error convergence threshold according to the speed of the agent and the environmental complexity index. The dynamic adjustment formula is: ; ; Among them, represents the maximum number of iterations; represents the default number of iterations; represents the adjustment coefficient of the environmental complexity; represents the environmental complexity, which is calculated based on the change rate and local curvature of the 3D point cloud data; represents the adjusted registration error convergence threshold; represents the default registration error convergence threshold; represents the adjustment coefficient of the speed; represents the instantaneous speed of the agent.
8. The method for real-time scenario analysis of the interaction between an intelligent agent and an urban road according to claim 1, characterized in that, The resolution of the voxel grid is dynamically adjusted according to the object distribution density and the predicted distance of the agent's driving path. The adjustment formula is as follows: ; Among them, represents the resolution of the voxel grid at the spatial point ; represents the object distribution density in the region where the spatial point is located; represents the distribution density threshold of the high-density region; represents the highest resolution value; represents the attenuation coefficient of the influence of the predicted distance on the resolution; represents the spatial point to the predicted distance of the agent's driving path.
Citation Information
Patent Citations
A live-line work scene electric power part three-dimensional reconstruction method based on point cloud
CN109934855A
Multi-solid-state-laser-radar external parameter calibration method based on SIFT-SHOT characteristics
CN113470090A
Intelligent inspection device and alarm method
CN118658126A
Intelligent vehicle dynamic target three-dimensional sensing method in complex environment
CN119152490A
Power transmission line hazard source distance measurement method and system based on point cloud space information fusion
CN119478844A