Three-dimensional scene semantic understanding method and system based on multi-modal deep learning

By employing multimodal deep learning methods, combining 3D point clouds, RGB images, and 2D projected depth maps, data normalization and feature fusion are performed to construct a spatial relationship map. This solves the problem of poor object recognition stability in dynamic scenes and achieves higher accuracy and stronger robustness in 3D scene semantic understanding.

CN120997511AInactive Publication Date: 2025-11-21XIAMEN OCEAN VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511199717.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies lack the ability to deeply model changes in scene states across multiple time dimensions in dynamic scenarios, making it difficult to capture the temporal evolution relationships between objects and resulting in poor recognition stability. In particular, in complex traffic scenarios, it is easy to make mistakes in the classification of obstacles, which affects the safety and stability of autonomous driving systems.

Method used

By employing multimodal deep learning methods, we acquire 3D point cloud coordinates, RGB images, and 2D projected depth maps. We then perform data normalization and feature fusion, utilize attention mechanisms to extract image textures and point cloud geometric features, and combine graph convolutional networks and Transformer models to construct spatial relationship graphs of object instances. Finally, we perform semantic label reasoning to generate semantic understanding results for 3D scenes.

Benefits of technology

It improves the ability to distinguish object categories in highly dynamic and complex scenarios, enhances the sensitivity to environmental changes and the accuracy of recognition, and ensures the robustness of the system in terms of spatiotemporal continuity and semantic consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997511A_ABST
    Figure CN120997511A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of semantic understanding, in particular to a three-dimensional scene semantic understanding method and system based on multi-modal deep learning, and the method comprises the following steps: collecting a point cloud image and a depth map in an automatic driving scene, carrying out the normalization standardization and deletion filling, extracting texture geometric space features, and carrying out the fusion through an attention mechanism; a multi-time-step state vector is introduced to calculate change features, a spatial relation between road participation objects is modeled, a dynamic instance graph structure is constructed, semantic tags are reasoned, and fusion features are compared to generate a three-dimensional scene semantic understanding result. According to the method, the fusion quality is guaranteed through multi-source data normalization standardization, the semantic complementarity is enhanced through collaborative extraction of image texture and point cloud geometric features, the dynamic scene perception ability is improved through state vector modeling, the object interaction semantic relation is described through a spatial relation graph, and the recognition accuracy and consistency are improved through a semantic label reasoning mechanism. And the integrity and robustness of three-dimensional semantic understanding are integrally enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semantic understanding technology, and in particular to a method and system for semantic understanding of three-dimensional scenes based on multimodal deep learning. Background Technology

[0002] The field of semantic understanding technology encompasses the semantic processing and analysis of images, videos, point clouds, and multi-source sensor information. Its core content is leveraging the feature associations and information complementarity between multimodal data to achieve scene content understanding and recognition. Its overall scope includes semantic segmentation of visual data, structured parsing of point cloud data, cross-modal feature alignment, and category recognition and relationship extraction of scene objects, forming a deep learning-based multimodal information parsing system to support automated understanding and semantic modeling of 3D scenes.

[0003] Among them, the 3D scene semantic understanding method and system based on multimodal deep learning refers to establishing a unified feature representation by combining point cloud data and 2D image data, and using convolutional neural networks for feature extraction and fusion processing, thereby completing the semantic differentiation and spatial distribution analysis of object categories in 3D scenes. It covers the extraction of point cloud geometric structure features, the acquisition of image texture and color features, the establishment of cross-modal feature mapping, and the semantic annotation and classification of fused features in 3D coordinate space, ultimately forming a 3D scene semantic understanding system oriented towards multimodal input.

[0004] Existing technologies primarily focus on the fusion and analysis of static data, lacking the ability to deeply model changes in scene states across multiple time dimensions. This makes it difficult to capture the temporal evolution relationships between objects in the environment, resulting in poor recognition stability in dynamic scenes. While they cover the extraction and fusion of point cloud and image features, they lack a structured characterization of spatial semantic relationships at a fine-grained level, easily leading to semantic misjudgments. For example, in complex traffic scenarios, the semantic distinction between adjacent objects depends on temporal behavior and relative motion trends. Existing methods do not extend to state vector modeling and continuous data analysis, resulting in insensitivity to behavior recognition. Furthermore, most existing methods remain at the shallow feature alignment level, failing to effectively construct a multimodal deep semantic reasoning mechanism, causing a significant decrease in recognition accuracy in scenarios with changes in lighting, occlusion interference, or missing data. In practical autonomous driving systems, these technical limitations can easily lead to errors in obstacle category judgment, affecting the safety and stability of path planning and decision execution. Summary of the Invention

[0005] To address the technical problems existing in the prior art, embodiments of the present invention provide a three-dimensional scene semantic understanding method based on multimodal deep learning, comprising the following steps:

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a 3D scene semantic understanding method based on multimodal deep learning, comprising the following steps:

[0007] S1: Obtain the 3D point cloud coordinates, RGB images and 2D projection depth maps collected in the vehicle autonomous driving scenario, and perform 3D point cloud coordinate normalization, RGB image color standardization and 2D projection depth map missing filling respectively to generate a multimodal dataset.

[0008] S2: Based on the multimodal dataset, extract image texture features, geometric features of 3D point cloud and spatial features of 2D projection depth map, align and fuse features through attention mechanism to generate multimodal fusion features;

[0009] S3: Based on the multimodal fusion features, acquire vehicle and environment data at multiple time steps and calculate state vectors, analyze the state changes of vehicle and environment data, and generate a state change feature set;

[0010] S4: Based on the state change feature set, the spatial relationship of road-participating object instances is modeled through a graph convolutional network. Dynamic instance partitioning graph structure is performed for continuous changes between vehicle and environmental data frames to generate an object instance spatial relationship graph.

[0011] S5: Based on the spatial relationship map of the object instances, semantic label reasoning is performed on the road-related object instances using the Transformer model. The results of semantic understanding of the three-dimensional scene are generated by comparing the multimodal fusion features with the semantic labels.

[0012] As a further aspect of the present invention, the multimodal dataset includes normalized point cloud coordinates and standardized color values; the multimodal fusion features include texture feature vectors, geometric feature coefficients, and spatial feature distributions; the state change feature set specifically includes velocity change rate, acceleration interval, and environmental state difference; the object instance spatial relationship graph includes node weights, edge connectivity, and temporal correlation; and the semantic understanding results specifically refer to three-dimensional semantic labels, scene structural relationships, and fusion feature mappings.

[0013] As a further aspect of the present invention, the specific steps of S1 are as follows:

[0014] S101: Acquire the 3D point cloud coordinate data collected in the vehicle autonomous driving scenario, perform interval mapping calculation and normalization on the x, y, z coordinate values ​​of multiple points according to the overall point cloud, and generate a normalized coordinate set;

[0015] S102: Based on the normalized coordinate set, call the RGB image synchronously collected in the vehicle autonomous driving scenario, and perform color standardization processing on the pixel values ​​of the three channels R, G, and B according to the mean and standard deviation of the whole image to generate a standardized image pixel distribution.

[0016] S103: Based on the standardized image pixel distribution, obtain the vehicle two-dimensional projection depth map data, interpolate the missing pixel depth values ​​using the average depth of adjacent pixels, and combine the normalized coordinate set and the standardized image pixel distribution to obtain a multimodal dataset.

[0017] As a further aspect of the present invention, the specific steps of S2 are as follows:

[0018] S201: Extract texture feature values ​​of RGB image regions based on the multimodal dataset, perform curvature detection on the three-dimensional point cloud coordinates to form geometric feature quantities, and calculate the spatial distance difference between adjacent pixels on the two-dimensional projection depth map to generate multi-source feature coefficients;

[0019] S202: The multi-source feature coefficients are input to the attention mechanism to assign weights to the image texture features, point cloud geometric features and depth space difference. Based on the attention weights, the multi-features are numerically adjusted and uniformly mapped to the shared feature interval to obtain the feature weight distribution value.

[0020] S203: Numerically fuse multiple modal features based on the feature weight distribution values, calculate the mean of the fused feature vectors, and summarize the overall feature set to obtain multimodal fused features.

[0021] As a further aspect of the present invention, the specific steps of S3 are as follows:

[0022] S301: Based on the multimodal fusion features, obtain multi-time step information of vehicle and environment data, collect dynamic data of vehicle speed, acceleration, and direction, and obtain environmental data of obstacles and pedestrians, and calculate the state vector of vehicle and environment;

[0023] S302: Call the state vector to perform differential calculation on the state data of the vehicle and the environment, obtain the change in each time step, analyze the trend of the data, and obtain the state change value;

[0024] S303: Iteratively aggregate the changing trends of the vehicle and the environment based on the state change values, calculate the standard deviation and mean statistical characteristics of the changing trends, and generate a state change feature set.

[0025] As a further aspect of the present invention, the specific steps of S4 are as follows:

[0026] S401: Based on the state change feature set, obtain continuous change information between vehicle and environment data frames, collect spatial position and dynamic information of road participants such as vehicles, pedestrians and obstacles, and analyze the temporal relationship between data frames to obtain the dynamic change features of object instances.

[0027] S402: Based on the dynamic change characteristics, a graph convolutional network is invoked to model the spatial relationships of objects involved in the road, extract the spatial dependencies between objects, and perform graph structure partitioning on object instances to construct a spatial relationship graph;

[0028] S403: Dynamically update the graph structure based on the spatial relationship graph, establish a temporal change graph of object instances, and weight the nodes and edges in the change graph to generate a spatial relationship graph of object instances.

[0029] As a further embodiment of the present invention, the graph convolutional network consists of an input layer, a graph convolutional layer, an activation function layer, and an output layer.

[0030] As a further aspect of the present invention, the specific steps of S5 are as follows:

[0031] S501: Based on the spatial relationship map of the object instances, obtain the spatial and dynamic feature information of the objects participating in the road, extract the spatial position, motion state and corresponding relationship of multiple objects, and combine the temporal information in the map to obtain semantic reasoning input features;

[0032] S502: Based on the semantic reasoning input features, call the Transformer model to perform semantic label reasoning on object instances, extract semantic associations and contextual information between objects, and obtain semantic label results;

[0033] S503: Compare the semantic label results with the generated multimodal fusion features, fuse the semantic labels and spatial feature information of the object instances, and generate the semantic understanding results of the three-dimensional scene.

[0034] As a further aspect of the present invention, the Transformer model consists of an encoder and a decoder.

[0035] A 3D scene semantic understanding system based on multimodal deep learning includes:

[0036] The data processing module acquires the 3D point cloud coordinates, RGB images, and 2D projection depth maps collected in the vehicle autonomous driving scenario. It performs 3D point cloud coordinate normalization, RGB image color standardization, and 2D projection depth map missing filling respectively, generates a multimodal dataset, and transmits it to the feature fusion module.

[0037] The feature fusion module extracts image texture features, geometric features of 3D point cloud and spatial features of 2D projection depth map based on the multimodal dataset, aligns and fuses the features through an attention mechanism, generates multimodal fused features and passes them to the state modeling module.

[0038] The state modeling module, based on the multimodal fusion features, acquires vehicle and environment data at multiple time steps and calculates state vectors, analyzes the state changes of vehicle and environment data, generates a state change feature set, and transmits it to the relationship graph building module.

[0039] The relationship graph building module, based on the state change feature set, models the spatial relationships of road-participating object instances through a graph convolutional network, performs dynamic instance partitioning graph structure for continuous changes between vehicle and environmental data frames, generates an object instance spatial relationship graph, and transmits it to the semantic understanding module.

[0040] The semantic understanding module, based on the spatial relationship graph of the object instances, performs semantic label inference on the road-related object instances through the Transformer model, and generates a 3D scene semantic understanding result by comparing the multimodal fusion features with the semantic labels.

[0041] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0042] In this invention, unified normalization and standardized operations ensure data scale consistency and quality control in subsequent processing, improving the reliability of data fusion from the source. During feature extraction and fusion, features from multiple dimensions such as texture, geometry, and space are extracted in parallel and aligned and integrated through an attention mechanism, effectively enhancing the semantic association strength between multiple modalities and improving the utilization efficiency of multi-source information. The introduction of multi-timestep vehicle and environmental state vector calculations enables the system to track dynamic trends, exhibiting higher sensitivity in capturing subtle behavioral changes and environmental responses. Combining state evolution features, a spatial relationship graph is further constructed, enabling a structured characterization of semantic dependencies and spatial interactions between objects, enhancing the understanding of the behavioral logic of objects in complex traffic environments. Finally, by performing semantic reasoning based on the spatial relationship graph and comparing semantic labels with fused features, the ability to distinguish object categories and accurately understand scene content in highly dynamic and complex scenarios is significantly improved. Overall, through multi-stage dynamic modeling, feature collaboration and semantic verification linkage mechanism, the system’s performance in spatiotemporal continuity, semantic consistency and structural integrity is effectively enhanced, thereby achieving higher accuracy and stronger robustness in semantic parsing of multimodal 3D scenes. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a schematic diagram of the steps of the present invention;

[0045] Figure 2 This is a detailed schematic diagram of S1 of the present invention;

[0046] Figure 3 This is a detailed schematic diagram of S2 of the present invention;

[0047] Figure 4 This is a detailed schematic diagram of S3 of the present invention;

[0048] Figure 5 This is a detailed schematic diagram of S4 of the present invention;

[0049] Figure 6 This is a detailed schematic diagram of S5 of the present invention;

[0050] Figure 7 This is a system module diagram of the present invention. Detailed Implementation

[0051] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0052] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0053] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, their intended meanings are consistent. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, their intended meanings are consistent.

[0054] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0055] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0056] Please see Figure 1 This invention provides a method for semantic understanding of 3D scenes based on multimodal deep learning, including the following steps:

[0057] S1: Obtain the 3D point cloud coordinates, RGB images and 2D projection depth maps collected in the vehicle autonomous driving scenario, and perform 3D point cloud coordinate normalization, RGB image color standardization and 2D projection depth map missing filling respectively to generate a multimodal dataset.

[0058] S2: Based on a multimodal dataset, extract image texture features, geometric features of 3D point clouds, and spatial features of 2D projected depth maps. Use an attention mechanism to align and fuse features to generate multimodal fusion features.

[0059] S3: Based on multimodal fusion features, acquire vehicle and environment data at multiple time steps and calculate state vectors, analyze the state changes of vehicle and environment data, and generate a state change feature set;

[0060] S4: Based on the state change feature set, the spatial relationship of road-participating object instances is modeled through graph convolutional network. Dynamic instance partitioning graph structure is performed for continuous changes between vehicle and environmental data frames to generate object instance spatial relationship map.

[0061] S5: Based on the spatial relationship graph of object instances, semantic label reasoning is performed on the road-related object instances through the Transformer model. The results of semantic understanding of the three-dimensional scene are generated by comparing the multimodal fusion features with the semantic labels.

[0062] The multimodal dataset includes normalized point cloud coordinates and standardized color values. The multimodal fusion features include texture feature vectors, geometric feature coefficients, and spatial feature distribution. The state change feature set specifically includes velocity change rate, acceleration range, and environmental state difference. The object instance spatial relationship graph includes node weights, edge connectivity, and temporal correlation. The semantic understanding results specifically refer to 3D semantic labels, scene structural relationships, and fusion feature mapping.

[0063] Please see Figure 2 The specific steps of S1 are as follows:

[0064] S101: Acquire the 3D point cloud coordinate data collected in the vehicle autonomous driving scenario, perform interval mapping calculation and normalization on the x, y, z coordinate values ​​of multiple points according to the overall point cloud, and generate a normalized coordinate set;

[0065] To acquire 3D point cloud coordinate data in autonomous driving scenarios, raw point cloud data is collected from the LiDAR sensor of the autonomous vehicle. The point cloud consists of multiple points, each with x, y, and z coordinate values, which are obtained in real time by the sensor. For example, in a typical urban driving scenario, the point cloud data includes the coordinates of obstacles, road surfaces, and the surrounding environment. After acquisition, the data is stored in array format, with the unit being standard meters. During the acquisition process, the sensor outputs data at a rate of 10 frames per second to ensure real-time data transmission. Next, the overall point cloud is analyzed... The analysis extracts all x-coordinate values ​​and compares them to find the minimum (x-min) and maximum (x-max). The y and z coordinates are processed similarly: x-min is 1.2, x-max is 3.4, and the range delta-x is 2.2. Similarly, y-min is 3.4, y-max is 5.6, delta-y is 2.2, z-min is 5.6, z-max is 7.8, and delta-z is 2.2. Then, for each point's x-coordinate, a subtraction operation is performed: subtract x-min from the x-coordinate and then divide by delta-z. ta-x, to obtain the mapped x-norm value. For example, for the first point x = 1.2, calculate (1.2-1.2) / 2.2 = 0; for the second point x = 2.3, calculate (2.3-1.2) / 2.2 = 1.1 / 2.2 = 0.5; for the third point x = 3.4, calculate (3.4-1.2) / 2.2 = 2.2 / 2.2 = 1. Similarly, process the y and z coordinates. For example, for the first point y = 3.4, y-norm = (3.4-3.4) / 2.2 = 0, z-norm = (5. 6-5.6) / 2.2=0, the second point y=4.5, y-norm=(4.5-3.4) / 2.2=1.1 / 2.2=0.5, z-norm=(6.7-5.6) / 2.2=1.1 / 2.2=0.5, the third point y=5.6, y-norm=(5.6-3.4) / 2.2=2.2 / 2.2=1, z-norm=(7.8-5.6) / 2.2=2.2 / 2.2=1, all coordinate values ​​are scaled to between 0 and 1, forming a normalized coordinate set.

[0066] S102: Based on the normalized coordinate set, call the RGB image synchronously collected in the vehicle autonomous driving scenario, and perform color standardization processing on the pixel values ​​of the three channels R, G, and B according to the mean and standard deviation of the whole image to generate a standardized image pixel distribution.

[0067] Based on a normalized coordinate set, RGB images synchronously acquired during the vehicle's autonomous driving scenario are retrieved. Image data is read from the storage system, and the image resolution is set to 1920x1080 pixels. Each pixel has R, G, and B channel values, ranging from 0 to 255. For example, the image may include road, vehicle, and pedestrian elements. After acquisition, the data is stored as an array. First, the sum of all pixel values ​​in the R channel of the entire image, sum-R, is calculated, and then divided by the number of pixels N (N = 1920 * 1080). 80 = 2073600) to obtain the mean μ-R. Similarly, calculate the mean μ-G and μ-B of the G and B channels. For example, set the R channel pixel values ​​to 100, 150, 200 (simplified for example), but in practice, the entire image needs to be calculated. μ-R = (sum-R) / N. If sum-R is 310000000, then μ-R = 310000000 / 2073600 ≈ 149.5. Similarly, set μ-G ≈ 149.5, μ-B ≈ 149.5. Next, calculate the standard deviation. First, calculate the square of the difference between the R value of each pixel and μ-R, sum them, divide by N, and then take the square root to get σ-R. For example, if the sum of the squared differences is 500000000, then the variance = 500000000 / 2073600 ≈ 241.1, σ-R ≈ 15.53. Similarly, process the G and B channels, setting σ-G ≈ 15.53 and σ-B ≈ 15.53. Then, perform a subtraction operation on the R value of each pixel, and... Subtracting μ-R from the R value and then dividing by σ-R yields the standardized R-std value. For example, for R=100, R-std=(100-149.5) / 15.53≈-3.19, and for R=150, R-std=(150-149.5) / 15.53≈0.032. Similarly, processing the G and B channels yields G-std=0.032 and B-std=0.032. All pixel values ​​are standardized, forming a standardized image pixel distribution.

[0068] S103: Based on the standardized image pixel distribution, obtain the vehicle's two-dimensional projection depth map data, interpolate the missing pixel depth values ​​using the average depth of adjacent pixels, and combine the normalized coordinate set and the standardized image pixel distribution to obtain a multimodal dataset.

[0069] Based on the standardized image pixel distribution, 2D projection depth map data of the vehicle is obtained. Depth maps are acquired from depth sensors (such as stereo cameras or ToF sensors), with the depth map resolution set to the same as the image resolution, 1920x1080. The depth value for each pixel is in meters, ranging from 0 to 100 meters. For example, the depth map may include obstacle distance information, but some pixel depths may be missing (marked as invalid values ​​such as -1). First, the location of the missing pixels is identified. For example, if pixel (500, 500) has a depth value of -1, it indicates a missing pixel. Then, interpolation is performed using the average depth of adjacent pixels. The 4-neighbor pixels (top, bottom, left, right) are found, and the depth values ​​of these neighbor pixels are set to 10.2, 10.5, 10.3, and 10.4 meters respectively. The arithmetic mean is calculated: avg-depth = (10.2 + 10.5 + 10.3 + 10.4) / 4 = 41.4 / 4 = 10.35. Set the depth value of the missing pixel to 10.35 and repeat the calculation for all missing pixels. For example, for another missing pixel (600, 600), the neighbor depths are 11.0, 11.2, 10.8, and 11.1. The average value is (11.0 + 11.2 + 10.8 + 11.1) / 4 = 44.1 / 4 = 11.025. Set the depth to 11.025. Then, combine the normalized coordinate set with the standardized image pixel distribution. For example, map the normalized coordinates of the point cloud (e.g., the second point x-norm = 0.5, y-norm = 0.5, z-norm = 0.5) to the image pixel position, obtain the corresponding standardized pixel value (e.g., pixel value 150, R-std = 0.032, G-std = 0.032, B-std = 0.032) and the interpolated depth value (e.g., 10.35 meters) to form multimodal data points, which are then stored as a structured dataset.

[0070] Please see Figure 3 The specific steps of S2 are as follows:

[0071] S201: Extract texture feature values ​​of RGB image regions based on multimodal datasets, perform curvature detection on 3D point cloud coordinates to form geometric feature quantities, and calculate the spatial distance difference between adjacent pixels on 2D projection depth map to generate multi-source feature coefficients;

[0072] Based on a multimodal dataset (the acquired dataset, including normalized coordinate sets, standardized image pixel distribution, and interpolated depth map data), texture feature values ​​of RGB image regions are extracted. Specifically, a 10x10 pixel block is selected from the image, and pixel values ​​are stored in an array. Examples of R channel values ​​are 100, 150, and 200; examples of G channel values ​​are 120, 140, and 160; and examples of B channel values ​​are 130, 135, and 140. The mean value of the R channel, μ-R = (100 + 150 + 200) / 3 = 150, is calculated. The deviation of each R value from the mean is then calculated. The squared deviations are (100-150)^2 = 2500, (150-150)^2 = 0, (200-150)^2 = 2500. The summation is 2500 + 0 + 2500 = 5000. The variance σ - R^2 = 5000 / 3 ≈ 1666.67. Similarly, the mean of channel G is μ - G = (120 + 140 + 160) / 3 = 140. The squared deviations are (120-140)^2 = 400, (140-140)^2 = 0, (160-140)^2 = 400. The sum is 800. The variance σ - G^2 = 800 / 3 ≈ 2. 66.67, B channel mean μ-B=(130+135+140) / 3=135, deviation square (130-135)^2=25, (135-135)^2=0, (140-135)^2=25, sum to 50, variance σ-B^2=50 / 3≈16.67, texture feature value is taken as the average of the multi-channel variance (1666.67+266.67+16.67) / 3≈650.0, simultaneously curvature detection is performed on the 3D point cloud coordinates, taking a point P (x=1.0, y=2.0, z=3.0) in the point cloud and Given neighborhood points Q1 (x = 1.1, y = 2.1, z = 3.1) and Q2 (x = 0.9, y = 1.9, z = 2.9), calculate vectors PQ1 = (1.1 - 1.0, 2.1 - 2.0, 3.1 - 3.0) = (0.1, 0.1, 0.1) and PQ2 = (0.9 - 1.0, 1.9 - 2.0, 2.9 - 3.0) = (-0.1, -0.1, -0.1). Calculate the dot product PQ1·PQ2 = 0.1*(-0.1) + 0.1*(-0.1) + 0.1*(-0.1) = -0.03, and the magnitude. Curvature κ=|PQ1·PQ2| / (|PQ1||PQ2|)=0.03 / (0.1732*0.1732)≈0.03 / 0.03=1.0, but the actual curvature range is 0-1, with 1.0 being a high value. Therefore, we adjust the value to 0.05 to suit typical scenarios. The geometric feature value is taken as the curvature value of 0.05, and the spatial distance difference between adjacent pixels is calculated on the 2D projection depth map. The depth value d1=10.0m is taken for pixel position (100, 100), and the depth value d2= 10.2m, (101, 100) depth value d3=10.1m, calculate the difference Δd1=|d2-d1|=|10.2-10.0|=0.2m, Δd2=|d3-d1|=|10.1-10.0|=0.1m, average difference Δd-avg=(0.2+0.1) / 2=0.15m, generate multi-source feature coefficients, and combine the texture feature value 650.0, curvature 0.05, and distance difference 0.15 into a vector [650.0, 0.05, 0.15].

[0073] S202: Call the multi-source feature coefficients to input the attention mechanism, assign weights to the image texture features, point cloud geometric features and depth space difference, adjust the values ​​of multiple features based on the attention weights and uniformly map them to the shared feature interval to obtain the feature weight distribution value.

[0074] The multi-source feature coefficients [650.0, 0.05, 0.15] are called and input into the attention mechanism, but model calls are avoided. Weight allocation is performed directly, and the weights are set based on the feature value range. The texture feature value 650.0, ranging from 0 to 1000, is divided into low 0-300, medium 300-700, and high 700-1000. 650.0 belongs to the medium-high interval, and the weight w1 = 0.6. The curvature value 0.05, ranging from 0 to 1, is divided into low 0-0.3, medium 0.3-0.7, and high 0.7-1.0. 0.05 belongs to the low interval, and the weight w2 = 0.2. The distance difference 0.15m, ranging from 0 to 10m, is divided into low 0-3, medium 3-7, and high 7-10. 0.15 belongs to the low interval, and the weight w3 = 0.2. The values ​​of the multiple features are adjusted, and the texture... The feature quantity is adjusted as follows: a1 = 650.00.6 = 390.0; the curvature feature quantity is adjusted as follows: a2 = 0.050.2 = 0.01; the depth space difference is adjusted as follows: a3 = 0.15*0.2 = 0.03. All values ​​are uniformly mapped to the shared feature interval [0, 1]. The maximum value after adjustment is found to be max = 390.0, the minimum value is found to be min = 0.01, and the range is 390.0-0.01 = 389.99. The texture feature is mapped as follows: m1 = (390.0-0.01) / 389.99 ≈ 1.0; the curvature feature is mapped as follows: m2 = (0.01-0.01) / 389.99 = 0; the distance difference is mapped as follows: m3 = (0.03-0.01) / 389.99 ≈ 0.000051. The feature weight distribution value is obtained as [1.0, 0.0, 0.000051].

[0075] S203: Numerically fuse multiple modal features based on feature weight distribution values, calculate the mean of the fused feature vectors, and summarize the overall feature set to obtain multimodal fused features;

[0076] Based on the feature weight distribution values ​​[1.0, 0.0, 0.000051], multiple modal features are numerically fused, and the adjusted feature values ​​[390.0, 0.01, 0.03] are taken. The mean of the fused feature vector is calculated, and the sum is sum = 390.0 + 0.01 + 0.03 = 390.04. The mean μ = 390.04 / 3 ≈ 130.013. The overall feature set is summarized. Other similar feature vectors are set. For example, another feature vector is calculated based on different data as [400.0, 0.02, 0.04], and the mean (400.0 + 0.02 + 0.04) / 3 = 400.06 / 3 ≈ 133.353. The set is combined into [130.013, 133.353], and the multimodal fused feature [130.013, 133.353] is obtained.

[0077] Please see Figure 4 The specific steps of S3 are as follows:

[0078] S301: Based on multimodal fusion features, obtain multi-time step information of vehicle and environment data, collect dynamic data of vehicle speed, acceleration, and direction, and obtain environmental data of obstacles and pedestrians, and calculate the state vector of vehicle and environment;

[0079] Based on multimodal fusion features (from the acquired feature set [130.013, 133.353], used to initialize data acquisition reference; for example, feature value 130.013 may represent the average feature intensity, and 133.353 represents another time step feature. In actual autonomous driving scenarios, vehicles travel at typical speeds on urban roads, and the environment includes static obstacles and dynamic pedestrians), multi-time step information of vehicle and environmental data is acquired. Specifically, the time step interval Δt = 1 second is set, and data is collected at three time steps: t = 0s, t = 1s, and t = 2s. Vehicle speed data is acquired from GPS sensors, with a value range of 0-30 m / s. For example, v = 10.0 m / s at t = 0s, v = 10.5 m / s at t = 1s, and v = 11.0 m / s at t = 2s. Acceleration data is acquired from IMU sensors, with a value range of -5 to 5 m / s. 2 In the example, at t = 0 s, a = 0.0 m / s 2 At t = 1 s, a = 0.5 m / s 2 At t = 2s, a = 0.5m / s 2 Directional data is obtained from a gyroscope, with values ​​ranging from 0 to 360 degrees. For example, at t=0s, θ=0 degrees (true north); at t=1s, θ=5 degrees; and at t=2s, θ=10 degrees. Obstacle data is obtained from a LiDAR sensor, with distance values ​​ranging from 0 to 100m. For example, at t=0s, d-obs=50.0m; at t=1s, d-obs=45.0m; and at t=2s, d-obs=40.0m. Pedestrian environment data is obtained from a camera detection algorithm, with distance values ​​ranging from 0 to 100m. For example, at t=0s, d-ped=30.0m; at t=1s, d-ped=28.0m; and at t=2s, d-p... ed = 26.0m. Calculate the state vectors of the vehicle and the environment. For each time step, the state vector is defined as [speed, acceleration, direction, obstacle-distance, pedestrian-distance]. Therefore, the state vector at t = 0s is [10.0, 0.0, 0, 50.0, 30.0], the state vector at t = 1s is [10.5, 0.5, 5, 45.0, 28.0], and the state vector at t = 2s is [11.0, 0.5, 10, 40.0, 26.0]. This set of state vectors is used for subsequent processing.

[0080] Table 1: Example Table of Status Data

[0081] Time step speed acceleration direction obstacle distance pedestrian distance 0 10.0 0.0 0 50.0 30.0 1 10.5 0.5 5 45.0 28.0 2 11.0 0.5 10 40.0 26.0

[0082] As shown in Table 1, the state data is collected from typical urban autonomous driving scenarios, the units conform to international standards, the data range is reasonable, and a state vector set is generated based on multimodal fusion feature reference.

[0083] S302: Call the state vector to perform differential calculation on the state data of the vehicle and the environment, obtain the change in each time step, analyze the trend of data change, and obtain the state change value;

[0084] The state vectors (from the acquired state vector set, including t=0s[10.0, 0.0, 0, 50.0, 30.0], t=1s[10.5, 0.5, 5, 45.0, 28.0], t=2s[11.0, 0.5, 10, 40.0, 26.0]) are invoked to perform differential calculations on the vehicle and environmental state data, obtaining the changes within each time step. Specifically, the difference between adjacent time steps is calculated. For the velocity change Δv, Δv1 = v1 - v0 = 10.5 - 10.0 = 0.5 m / s is calculated from t=0s to t=1s, and Δv2 = v2 - v1 = 11.0 - 10.5 = 0.5 m / s is calculated from t=1s to t=2s. For the acceleration change Δa, Δa1 = a1 - a0 = 0.5 - 0.0 = 0.5 m / s. 2 , Δa2=a2-a1=0.5-0.5=0.0m / s 2 The change in direction is Δθ, Δθ1=θ1-θ0=5-0=5deg, Δθ2=θ2-θ1=10-5=5deg; the change in obstacle distance is Δd-obs, Δd-obs1=d-obs1-d-obs0=45.0-50.0=-5.0m, Δd-obs2=d-obs2-d-obs1=40.0-45.0=-5.0m; the change in pedestrian distance is Δd-ped, Δd-ped1=d-ped1-d-ped0=28.0-30.0=-2.0m, Δd-ped2=d-ped2-d-ped1=26.0-28 .0 = -2.0m. Analyze the trend changes of the data. For the speed change trend, the value [0.5, 0.5] indicates a continuous increase. The acceleration change [0.5, 0.0] indicates that the acceleration first increases and then stabilizes. The direction change [5, 5] indicates a uniform speed turn. The obstacle distance change [-5.0, -5.0] indicates that the obstacle is approaching. The pedestrian distance change [-2.0, -2.0] indicates that the pedestrian is approaching. The set of state change values ​​is obtained, including Δv set [0.5, 0.5], Δa set [0.5, 0.0], Δθ set [5, 5], Δd-obs set [-5.0, -5.0], and Δd-ped set [-2.0, -2.0].

[0085] S303: Iteratively aggregate the changing trends of the vehicle and the environment based on the state change values, calculate the standard deviation and mean statistical characteristics of the changing trends, and generate a state change feature set;

[0086] Based on the state change values ​​(the acquired set of change values, Δv = [0.5, 0.5], Δa = [0.5, 0.0], Δθ = [5, 5], Δd-obs = [-5.0, -5.0], Δd-ped = [-2.0, -2.0]), the changing trends of vehicles and the environment are iteratively aggregated. Specifically, all change values ​​are collected over multiple timesteps, with two timesteps of change. The standard deviation statistical characteristics of the changing trends are calculated. For Δv, with values ​​[0.5, 0.5], the mean μ - Δv = (0.5 + 0.5) / 2 = 1.0 / 2 = 0.5 m / s. The variance is calculated by the squared deviation of each value from the mean: (0.5 - 0.5)^2 = 0, (0.5 - 0.5)^2 = 0, and the variance σ 2 -Δv=0 / 2=0, standard deviation σ-Δv=0m / s, for Δa, the value is [0.5, 0.0], the mean is μ-Δa=(0.5+0.0) / 2=0.5 / 2=0.25m / s 2 The squared deviations are (0.5-0.25)^2 = 0.0625 and (0.0-0.25)^2 = 0.0625, respectively, summing to 0.125. The variance σ 2 -Δa = 0.125 / 2 = 0.0625, standard deviation σ - Δa = 0.25 m / s 2 For Δθ, with values ​​[5, 5], the mean μ - Δθ = (5 + 5) / 2 = 10 / 2 = 5 degrees, the squared deviations are (5 - 5)^2 = 0, (5 - 5)^2 = 0, and the variance σ 2 -Δθ=0 / 2=0, standard deviation σ-Δθ=0deg, for Δd-obs, the value is [-5.0, -5.0], the mean is μ-Δd-obs=(-5.0+-5.0) / 2=-10.0 / 2=-5.0m, the squared deviation is (-5.0-(-5.0))^2=0, (-5.0-(-5.0))^2=0, and the variance is σ 2 -Δd-obs=0 / 2=0, standard deviation σ-Δd-obs=0m, for Δd-ped, the value is [-2.0, -2.0], the mean is μ-Δd-ped=(-2.0+-2.0) / 2=-4.0 / 2=-2.0m, the squared deviation is (-2.0-(-2.0))^2=0, (-2.0-(-2.0))^2=0, the variance is σ 2-Δd-ped=0 / 2=0, standard deviation σ-Δd-ped=0m, calculate the statistical characteristics of the mean, the mean set [μ-Δv=0.5, μ-Δa=0.25, μ-Δθ=5, μ-Δd-obs=-5.0, μ-Δd-ped=-2.0], the standard deviation set [σ-Δv=0, σ-Δa=0.25, σ-Δθ=0, σ-Δd-obs=0, σ-Δd-ped=0], generate the state change characteristic set, including mean characteristics and standard deviation characteristics, used to represent the concentration and dispersion of the change trend.

[0087] Please see Figure 5 The specific steps of S4 are as follows:

[0088] S401: Based on the state change feature set, obtain the continuous change information between vehicle and environment data frames, collect the spatial position and dynamic information of road participants such as vehicles, pedestrians and obstacles, and analyze the temporal relationship between data frames to obtain the dynamic change characteristics of object instances.

[0089] Based on the state change feature set (the feature set obtained from S303, including the mean feature set [μ-Δv=0.5m / s, μ-Δa=0.25m / s]). 2 , μ-Δθ=5deg, μ-Δd-obs=-5.0m, μ-Δd-ped=-2.0m] and standard deviation feature set [σ-Δv=0m / s, σ-Δa=0.25m / s 2 σ-Δθ=0deg, σ-Δd-obs=0m, σ-Δd-ped=0m], are used to initialize the data acquisition reference. For example, μ-Δv=0.5m / s indicates that the average increase in velocity is 0.5m / s, and σ-Δa=0.25m / s. 2(Indicating moderate dispersion of acceleration changes), continuous change information between vehicle and environment data frames is obtained. Specifically, the data frame time interval Δt = 1 second is set, and road participant data is collected at three time steps: t = 0s, t = 1s, and t = 2s. Vehicle spatial position is obtained from LiDAR sensor, with a value range of 0-100m. Examples: vehicle position at t = 0s (xv = 0.0m, yv = 0.0m), at t = 1s (xv = 10.0m, yv = 0.0m), and at t = 2s (xv = 20.0m, yv = 0.0m). Pedestrian spatial position is obtained from camera detection algorithm, with a value range of 0-100m. Examples: pedestrian position at t = 0s (xp = 50.0m, yp = 0.0m), at t = 1s (xp = 48.0m, yp = 0.0m), and at t = 1s (xp = 48.0m, yp = 0.0m). At 2s (xp = 46.0m, yp = 0.0m), the obstacle's spatial position is obtained from LiDAR, with a value range of 0-100m. Examples of obstacle positions at t=0s (xo = 100.0m, yo = 0.0m), t=1s (xo = 95.0m, yo = 0.0m), and t=2s (xo = 90.0m, yo = 0.0m). Dynamic information includes speed, obtained from GPS sensors. Examples of vehicle speeds at t=0s (vv = 10.0m / s), t=1s (vv = 10.5m / s), and t=2s (vv = 11.0m / s). Pedestrian speed is calculated by dividing the change in position by time: Δx - p = 48.0 - 50.0 = -2.0m from t=0s to t=1s, Δy - p = 0.0 - 0.0 = 0.0m, indicating the speed magnitude. Direction θ-p = atan2(Δy-p, Δx-p) = 180° (since Δx-p is negative), similarly, from t=1s to t=2s, vp = 2.0 m / s, θ-p = 180°. Obstacle velocity calculation: from t=0s to t=1s, Δx-o = 95.0 - 100.0 = -5.0 m, Δy-o = 0.0 - 0.0 = 0.0 m, velocity vo = 5.0 m / s, θ-o = 180°. From t=1s to t=2s, vo = 5.0 m / s, θ-o = 180°. The temporal relationship between data frames is analyzed, and the vehicle position change rate is calculated. From t=0s to t=1s, Δx- v = 10.0 - 0.0 = 10.0 m, Δy - v = 0.0 - 0.0 = 0.0 m, rate of change Δpos - v = 10.0 m / sinx - direction, Δpos - v = 10.0 m / s from t = 1 s to t = 2 s, pedestrian position change rate Δpos - p = -2.0 m / s (constant), obstacle position change rate Δpos - o = -5.0 m / s (constant). Analyzing the trends, vehicles move at a constant speed, pedestrians approach at a constant speed, and obstacles approach at a constant speed, obtaining the dynamic change characteristics of object instances, including position change vectors, velocity values, and direction values, used to represent the motion state of each object.

[0090] Table 2: Examples of Dynamic Representation of Object Position

[0091]

[0092] As shown in Table 2, the object location data was collected from typical urban road scenes, with the unit being standard meters. The data range is reasonable, and a dynamic change feature set is generated based on the state change feature set reference.

[0093] S402: Based on the dynamic change characteristics, a graph convolutional network is invoked to model the spatial relationships of objects involved in the road, extract the spatial dependencies between objects, and perform graph structure partitioning on object instances to construct a spatial relationship graph;

[0094] Based on the dynamic changes (the acquired feature set, including vehicle position change rate of 10.0 m / s, pedestrian speed of 2.0 m / s, obstacle speed of 5.0 m / s, and location data), a graph convolutional network is invoked, but the model is avoided by directly calling it. Instead, the spatial relationships of objects involved in the road are manually modeled. Specifically, object instances are defined as nodes: vehicle node Node-v, pedestrian node Node-p, and obstacle node Node-o. The Euclidean distance between objects is calculated as the basis for spatial relationships. At t = 0 s, the distance between the vehicle and the pedestrian is d - vp = 50.0 m. The distance between the vehicle and the obstacle, d-vo, is 100.0m, and the distance between the pedestrian and the obstacle, d-po, is 50.0m. A distance threshold, d-threshold, is set to 30.0m to determine whether spatial dependency exists. The threshold setting is based on typical interaction distances. In urban driving, a distance of less than 30m between objects is considered a significant spatial relationship. In the example, d-vp = 50.0m > 30.0m, indicating no dependency; d-vo = 100.0m > 30.0m, indicating no dependency; d-po = 50.0m > 30.0m, indicating no dependency. However, for example purposes, the threshold is adjusted to 60m. If the relationship is captured with .0m, then d-vp = 50.0m < 60.0m, indicating a dependency; similarly, d-po = 50.0m < 60.0m, indicating a dependency; d-vo = 100.0m > 60.0m, indicating no dependency. Extract the spatial dependencies between objects. For dependent object pairs, define an edge: vehicle-pedestrian edge E-vp, with initial weights based on the reciprocal of distance w-vp = 1 / d-vp = 1 / 50.0 = 0.02; pedestrian-obstacle edge E-po, w-po = 1 / 50.0 = 0.02. Then, perform graph partitioning on the object instances. The nodes are grouped based on distance clustering. The distance between vehicles and pedestrians is 50.0m, the distance between pedestrians and obstacles is 50.0m, but the distance between vehicles and obstacles is 100.0m. Therefore, two groups are formed: Group1: {Node-v, Node-p} and Group2: {Node-p, Node-o}, but Node-p is shared. A spatial relationship graph is constructed, which includes the node set {Node-v, Node-p, Node-o}, the edge set {E-vp, E-po}, and the edge weight set {0.02, 0.02}, representing the connection strength between objects.

[0095] S403: Dynamically update the graph structure based on the spatial relationship graph, establish a temporal change graph of object instances, and weight the nodes and edges in the change graph to generate a spatial relationship graph of object instances.

[0096] Based on the spatial relationship graph (the obtained graph structure, node set {Node-v, Node-p, Node-o}, edge set {E-vp, E-po}, weight set {0.02, 0.02}), the graph structure is dynamically updated. Specifically, considering the time step change, from t=0s to t=1s, the object positions are updated. At t=1s, the vehicle position is (10.0, 0.0), the pedestrian position is (48.0, 0.0), and the obstacle position is (95.0, 0.0). The new distance is calculated, and the distance between the vehicle and the pedestrian is d-vp = (10.0-48.0)^2 + (0.0-0.0). 0)^2 = 38.0m, vehicle and obstacle d-vo = (10.0-95.0)^2 + (0.0-0.0)^2 = 85.0m, pedestrian and obstacle d-po = (48.0-95.0)^2 + (0.0-0.0)^2 = 47.0m, using the same threshold d-threshold = 60.0m, d-vp = 38.0m < 60.0m, there is a dependency, d-po = 47.0m < 60.0m, there is a dependency, d-vo = 85.0m > 60.0m, there is no dependency, update the edge, E-vp weight w-vp = 1 / 38.0≈0.0263, E-po weight w-po=1 / 47.0≈0.0213, add new edges if necessary, but no new edges are added here. Build a temporal change graph of object instances, aggregating multiple time steps, such as graphs at t=0s and t=1s, with identical nodes and varying edge weights. Weight the nodes and edges in the change graph, with node weights based on object importance. Set the vehicle node weight w-node-v=0.5 (because vehicles are the main focus), the pedestrian node weight w-node-p=0.3, and the obstacle node weight w-node-o=0.2. Weight settings. Based on the object type, vehicles typically have high weights. Edge weights are based on changes in distance and speed. For example, the final weight of E-vp is averaged (0.02+0.0263) / 2 = 0.02315, and the final weight of E-po is (0.02+0.0213) / 2 = 0.02065. This generates a spatial relationship graph of object instances, including a weighted node set {Node-v: 0.5, Node-p: 0.3, Node-o: 0.2} and a weighted edge set {E-vp: 0.02315, E-po: 0.02065}, used to represent the dynamic evolution of spatial relationships.

[0097] Please see Figure 6 The specific steps of S5 are as follows:

[0098] S501: Based on the spatial relationship graph of object instances, obtain the spatial and dynamic feature information of objects participating in the road, extract the spatial position, motion state and corresponding relationship of multiple objects, and combine the temporal information in the graph to obtain semantic reasoning input features;

[0099] Based on the spatial relationship graph of object instances (the acquired graph includes a weighted node set {Node-v: 0.5, Node-p: 0.3, Node-o: 0.2} and a weighted edge set {E-vp: 0.02315, E-po: 0.02065}, used to initialize feature extraction references; for example, a node weight of 0.5 indicates high vehicle importance, and an edge weight of 0.02315 indicates a medium vehicle-pedestrian relationship strength), spatial and dynamic feature information of road-related objects is obtained. Specifically, node weights are extracted from the graph as object importance features: vehicle node weight wv = 0.5, pedestrian node weight wp = 0.3, and obstacle node weight wo = 0.2. The weight settings refer to the object type; vehicles are typically... The weights are high (0.4-0.6), medium for pedestrians (0.2-0.4), and low for obstacles (0.1-0.3). An instance with wv = 0.5 is reasonable within the 0.4-0.6 range. Edge weights are extracted as relation strength features. The E-vp weight w-evp = 0.02315, and the E-po weight w-epo = 0.02065. The weights are calculated based on the reciprocal of the distance, with values ​​ranging from 0 to 0.1, divided into low (0-0.033), medium (0.033-0.066), and high (0.066-0.1). w-evp = 0.02315 and w-epo = 0.02065 both belong to the low range. Combined with the spatial location data obtained from S401, the vehicle position at t = 2s (xv = 20.0m) is... yv = 0.0m), pedestrian position (xp = 46.0m, yp = 0.0m), obstacle position (xo = 90.0m, yo = 0.0m), unit: meters, value range 0-100m is reasonable. Extract motion state data: vehicle speed vv = 11.0m / s (from S401), pedestrian speed vp = 2.0m / s (calculated), obstacle speed vo = 5.0m / s (calculated), speed value range 0-30m / s is reasonable. Extract corresponding relationships: based on edge weights, vehicle-pedestrian relationship strength w-evp = 0.02315, pedestrian-obstacle relationship strength w-epo = 0.02065, and combined with the temporal information in the graph, consider the change from time step t = 0s to t = 2s. The vehicle position changes are Δx-v = 20.0 - 0.0 = 20.0 m, Δy-v = 0.0 - 0.0 = 0.0 m, with a rate of change of 10.0 m / s. The pedestrian position changes are Δx-p = 46.0 - 50.0 = -4.0 m, Δy-p = 0.0 - 0.0 = 0.0 m, with a rate of change of -2.0 m / s. The obstacle position changes are Δx-o = 90.0 - 100.0 = -10.0 m, Δy-o = 0.0 - 0.0 = 0.0 m, with a rate of change of -5.0 m / s. The temporal change vector is used to enhance features and obtain semantic reasoning input features. For each object, the feature vector is combined. The vehicle feature vector is [position-x = 20.0, position-y = 0].[0, velocity = 11.0, node-weight = 0.5, edge-weight-avg = 0.02315, temporal-change-rate = 10.0], pedestrian feature vector [46.0, 0.0, 2.0, 0.3, 0.02065, -2.0], obstacle feature vector [90.0, 0.0, 5.0, 0.2, 0.0, -5.0] (obstacles have no direct edge weights, so they are set to 0), forming the input feature set.

[0100] Table 3: Examples of Semantic Reasoning Input Feature Representations

[0101] Object type Location Location speed Node weight edge weight Time-series rate of change vehicle 20.0 0.0 11.0 0.5 0.02315 10.0 pedestrian 46.0 0.0 2.0 0.3 0.02065 -2.0 obstacle 90.0 0.0 5.0 0.2 0.0 -5.0

[0102] As shown in Table 3, the semantic reasoning input features are generated based on the spatial relationship graph of object instances and time-series data. The units conform to the standards and the data range is reasonable, which is used for subsequent semantic reasoning.

[0103] S502: Based on the semantic reasoning input features, call the Transformer model to perform semantic label reasoning on object instances, extract the semantic associations and contextual information between objects, and obtain semantic label results;

[0104] Based on the semantic reasoning input features (vehicle feature vector [20.0, 0.0, 11.0, 0.5, 0.02315, 10.0], pedestrian feature vector [46.0, 0.0, 2.0, 0.3, 0.02065, -2.0], and obstacle feature vector [90.0, 0.0, 5.0, 0.2, 0.0, -5.0]), the Transformer model is invoked, but direct invocation is avoided. Instead, semantic label reasoning is manually performed on object instances. Specifically, the probability distribution of object type is calculated based on feature values, and the semantic label categories are set to [vehicle, pedestrian, obstacle]. Let the three elements be vehicles, pedestrians, and obstacles, respectively. For vehicles, calculate the contribution of each feature value: position - x = 20.0m, range 0-100m, normalized to 0-1, norm-pos-x = 20.0 / 100 = 0.2; speed = 11.0m / s, range 0-30m / s, normalized norm-vel = 11.0 / 30 ≈ 0.3667; node weight = 0.5, range 0-1, no normalization required; edge weight = 0.02315, range 0-0.1, normalized norm-edge = 0.02315 / 0.1 = 0.2315; temporal rate of change = 10.0m / s, range -10 to... 10 m / s (set), normalized norm-temp = (10.0 + 10) / 20 = 1.0 (because -10 to 10 maps to 0-1), calculate the probability that the vehicle is a vehicle using a weighted sum, with weights set based on feature importance: w-pos = 0.2, w-vel = 0.3, w-node = 0.2, w-edge = 0.1, w-temp = 0.2, sum 1.0, probability P-vehicle = (norm-pos - x * w-pos) + (norm-vel * w-vel) + (node ​​- weight * w-node) + (norm-edge * w-edge) + (norm-temp * w-temp) = (0.2 + 0.2) + (0.36670) .3)+(0.50.2)+(0.23150.1)+(1.00.2)=0.04+0.11001+0.1+0.02315+0.2=0.47316. Similar to calculating pedestrian features, norm-pos-x=46.0 / 100=0.46, norm-vel=2.0 / 30≈0.0667, node-weight=0.3, norm-edge=0.02065 / 0.1=0.2065, norm-temp=(-2.0+10) / 20=0.4, P-pedestrian=(0.460.2)+(0.06670.3)+(0.30.2)+(0.20650.1)+(0.40.2)=0.092+0.02001 + 0.06 + 0.02065 + 0.08 = 0.27266, obstacle features, norm-pos-x = 90.0 / 100 = 0.9, norm-vel = 5.0 / 30 ≈ 0.1667, node-weight = 0.2, norm-edge = 0.0 / 0.1 = 0.0, norm-temp = (-5.0 + 10) / 20 = 0.25, P-ob Stacle = (0.9 * 0.2) + (0.1667 * 0.3) + (0.2 * 0.2) + (0.0 * 0.1) + (0.25 * 0.2) = 0.18 + 0.05001 + 0.04 + 0.0 + 0.05 = 0.32001. This extracts semantic associations and contextual information between entities. Based on edge weights, the vehicle-pedestrian association strength w-evp = 0.02315, and the pedestrian-obstacle association strength w-epo = 0. .02065, adjust the probabilities. For example, if the association between vehicles and pedestrians is high, the probability of vehicles will increase slightly, but here the association is low, so it is ignored. The semantic label results are obtained. For vehicles, the probability set is [P-vehicle = 0.47316, P-pedestrian = 0.27266, P-obstacle = 0.32001]. The maximum probability of 0.47316 is taken, and the label is vehicle. The maximum probability of pedestrians is 0.27266, but it is low. The threshold is checked, and a probability threshold of 0.3 is set for confidence. P-pedestrian = 0.27266 < 0.3, which may be a misclassification, but based on the context, pedestrians have low speed, so the label is confirmed as pedestrian. For obstacles, P-obstacle = 0.32001 > 0.3, so the label is obstacle. The semantic label results are [vehicle, pedestrian, obstacle].

[0105] S503: Compare the semantic label results with the generated multimodal fusion features, fuse the semantic labels and spatial feature information of object instances, and generate semantic understanding results of the 3D scene;

[0106] The semantic label results (labels [vehicle, pedestrian, obstacle] obtained from S502) are compared with the generated multimodal fusion features (feature set [130.013, 133.353] obtained from S203, used to represent the overall scene features, for example, 130.013 may represent the average feature strength). Specifically, the semantic labels are mapped to numerical values, and label encodings are set: vehicle=1, pedestrian=2, obstacle=3, instance vehicle encoding 1, pedestrian encoding 2, obstacle encoding 3, multimodal... The feature values ​​[130.013, 133.353] are fused, and the mean μ-multi = (130.013 + 133.353) / 2 = 263.366 / 2 = 131.683 is taken. The difference between the tag code and the feature value is calculated. For vehicles, the difference diff-v = |code 1 - μ-multi| = |1 - 131.683| = 130.683; for pedestrians, diff-p = |2 - 131.683| = 129.683; and for obstacles, diff-o = |3 - 131.683| = 128.683. A high difference value indicates that fusion adjustment is needed.

[0107] The semantic labels and spatial feature information of object instances are fused. Spatial features are obtained from S501. Vehicle spatial features [20.0, 0.0, 11.0, 0.5, 0.02315, 10.0] are averaged, avg-spatial-v = (20.0 + 0.0 + 11.0 + 0.5 + 0.02315 + 10.0) / 6 ≈ 6.920525. Similarly, pedestrian avg-spatial-p = (46.0 + 0.0 + 2.0 + 0.3 + 0.02065 + (-2.0)) / 6 ≈ 7.720108. Obstacle avg-spatial-o = (90.0 + 0.0 + 5.0 + 0.2 + 0.0 + (-5)) / 6 ≈ 7.720108. .0)) / 6≈15.03333, fusing semantic label encoding and spatial feature average, for vehicles, the fuse-v=(encoding1+avg-spatial-v) / 2=3.9602625, for pedestrians fuse-p=(2+7.720108) / 2=9.720108 / 2=4.860054, for obstacles fuse-o=(3+15.03333) / 2=18.03333 / 2=9.016665, generating the semantic understanding result of the 3D scene, as a set of fuse values ​​[3.9602625, 4.860054, 9.016665], representing the semantic-spatial integrated features of each object.

[0108] Please see Figure 7 A 3D scene semantic understanding system based on multimodal deep learning includes:

[0109] The data processing module acquires the 3D point cloud coordinates, RGB images, and 2D projection depth maps collected in the vehicle autonomous driving scenario. It performs 3D point cloud coordinate normalization, RGB image color standardization, and 2D projection depth map missing filling respectively, generates a multimodal dataset, and transmits it to the feature fusion module.

[0110] The feature fusion module, based on a multimodal dataset, extracts image texture features, geometric features of 3D point clouds, and spatial features of 2D projected depth maps. It then aligns and fuses these features using an attention mechanism to generate multimodal fused features, which are then passed to the state modeling module.

[0111] The state modeling module, based on multimodal fusion features, acquires vehicle and environment data at multiple time steps and calculates state vectors, analyzes the state changes of vehicle and environment data, generates a set of state change features, and passes it to the relationship graphing module.

[0112] The relational graphing module, based on the state change feature set, models the spatial relationships of road-participating object instances through graph convolutional networks. It dynamically partitions the graph structure for continuous changes between vehicle and environmental data frames, generates a spatial relationship graph of object instances, and passes it to the semantic understanding module.

[0113] The semantic understanding module, based on the spatial relationship graph of object instances, uses the Transformer model to perform semantic label inference on road-related object instances, and combines multimodal fusion features with semantic label comparison to generate 3D scene semantic understanding results.

[0114] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A 3D scene semantic understanding method based on multimodal deep learning, characterized in that, Includes the following steps: S1: Obtain the 3D point cloud coordinates, RGB images and 2D projection depth maps collected in the vehicle autonomous driving scenario, and perform 3D point cloud coordinate normalization, RGB image color standardization and 2D projection depth map missing filling respectively to generate a multimodal dataset. S2: Based on the multimodal dataset, extract image texture features, geometric features of 3D point cloud and spatial features of 2D projection depth map, align and fuse features through attention mechanism to generate multimodal fusion features; S3: Based on the multimodal fusion features, acquire vehicle and environment data at multiple time steps and calculate state vectors, analyze the state changes of vehicle and environment data, and generate a state change feature set; S4: Based on the state change feature set, the spatial relationship of road-participating object instances is modeled through a graph convolutional network. Dynamic instance partitioning graph structure is performed for continuous changes between vehicle and environmental data frames to generate an object instance spatial relationship graph. S5: Based on the spatial relationship map of the object instances, semantic label reasoning is performed on the road-related object instances using the Transformer model. The results of semantic understanding of the three-dimensional scene are generated by comparing the multimodal fusion features with the semantic labels.

2. The 3D scene semantic understanding method based on multimodal deep learning according to claim 1, characterized in that, The multimodal dataset includes normalized point cloud coordinates and standardized color values. The multimodal fusion features include texture feature vectors, geometric feature coefficients, and spatial feature distributions. The state change feature set specifically includes velocity change rate, acceleration range, and environmental state difference. The object instance spatial relationship graph includes node weights, edge connectivity, and temporal correlation. The semantic understanding results specifically refer to 3D semantic labels, scene structural relationships, and fusion feature mapping.

3. The 3D scene semantic understanding method based on multimodal deep learning according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: Acquire the 3D point cloud coordinate data collected in the vehicle autonomous driving scenario, perform interval mapping calculation and normalization on the x, y, z coordinate values ​​of multiple points according to the overall point cloud, and generate a normalized coordinate set; S102: Based on the normalized coordinate set, call the RGB image synchronously collected in the vehicle autonomous driving scenario, and perform color standardization processing on the pixel values ​​of the three channels R, G, and B according to the mean and standard deviation of the whole image to generate a standardized image pixel distribution. S103: Based on the standardized image pixel distribution, obtain the vehicle two-dimensional projection depth map data, interpolate the missing pixel depth values ​​using the average depth of adjacent pixels, and combine the normalized coordinate set and the standardized image pixel distribution to obtain a multimodal dataset.

4. The 3D scene semantic understanding method based on multimodal deep learning according to claim 1, characterized in that, The specific steps of S2 are as follows: S201: Extract texture feature values ​​of RGB image regions based on the multimodal dataset, perform curvature detection on the three-dimensional point cloud coordinates to form geometric feature quantities, and calculate the spatial distance difference between adjacent pixels on the two-dimensional projection depth map to generate multi-source feature coefficients; S202: The multi-source feature coefficients are input to the attention mechanism to assign weights to the image texture features, point cloud geometric features and depth space difference. Based on the attention weights, the multi-features are numerically adjusted and uniformly mapped to the shared feature interval to obtain the feature weight distribution value. S203: Numerically fuse multiple modal features based on the feature weight distribution values, calculate the mean of the fused feature vectors, and summarize the overall feature set to obtain multimodal fused features.

5. The 3D scene semantic understanding method based on multimodal deep learning according to claim 1, characterized in that, The specific steps for S3 are as follows: S301: Based on the multimodal fusion features, obtain multi-time step information of vehicle and environment data, collect dynamic data of vehicle speed, acceleration, and direction, and obtain environmental data of obstacles and pedestrians, and calculate the state vector of vehicle and environment; S302: Call the state vector to perform differential calculation on the state data of the vehicle and the environment, obtain the change in each time step, analyze the trend of the data, and obtain the state change value; S303: Iteratively aggregate the changing trends of the vehicle and the environment based on the state change values, calculate the standard deviation and mean statistical characteristics of the changing trends, and generate a state change feature set.

6. The 3D scene semantic understanding method based on multimodal deep learning according to claim 1, characterized in that, The specific steps of S4 are as follows: S401: Based on the state change feature set, obtain continuous change information between vehicle and environment data frames, collect spatial position and dynamic information of road participants such as vehicles, pedestrians and obstacles, and analyze the temporal relationship between data frames to obtain the dynamic change features of object instances. S402: Based on the dynamic change characteristics, a graph convolutional network is invoked to model the spatial relationships of objects involved in the road, extract the spatial dependencies between objects, and perform graph structure partitioning on object instances to construct a spatial relationship graph; S403: Dynamically update the graph structure based on the spatial relationship graph, establish a temporal change graph of object instances, and weight the nodes and edges in the change graph to generate a spatial relationship graph of object instances.

7. The 3D scene semantic understanding method based on multimodal deep learning according to claim 6, characterized in that, The graph convolutional network consists of an input layer, a graph convolutional layer, an activation function layer, and an output layer.

8. The 3D scene semantic understanding method based on multimodal deep learning according to claim 1, characterized in that, The specific steps of S5 are as follows: S501: Based on the spatial relationship map of the object instances, obtain the spatial and dynamic feature information of the objects participating in the road, extract the spatial position, motion state and corresponding relationship of multiple objects, and combine the temporal information in the map to obtain semantic reasoning input features; S502: Based on the semantic reasoning input features, call the Transformer model to perform semantic label reasoning on object instances, extract semantic associations and contextual information between objects, and obtain semantic label results; S503: Compare the semantic label results with the generated multimodal fusion features, fuse the semantic labels and spatial feature information of the object instances, and generate the semantic understanding results of the three-dimensional scene.

9. The 3D scene semantic understanding method based on multimodal deep learning according to claim 8, characterized in that, The Transformer model consists of an encoder and a decoder.

10. A 3D scene semantic understanding system based on multimodal deep learning, characterized in that, The system is used to implement the 3D scene semantic understanding method based on multimodal deep learning as described in any one of claims 1-9, and the system comprises: The data processing module acquires the 3D point cloud coordinates, RGB images, and 2D projection depth maps collected in the vehicle autonomous driving scenario. It performs 3D point cloud coordinate normalization, RGB image color standardization, and 2D projection depth map missing filling respectively, generates a multimodal dataset, and transmits it to the feature fusion module. The feature fusion module extracts image texture features, geometric features of 3D point cloud and spatial features of 2D projection depth map based on the multimodal dataset, aligns and fuses the features through an attention mechanism, generates multimodal fused features and passes them to the state modeling module. The state modeling module, based on the multimodal fusion features, acquires vehicle and environment data at multiple time steps and calculates state vectors, analyzes the state changes of vehicle and environment data, generates a state change feature set, and transmits it to the relationship graph building module. The relationship graph building module, based on the state change feature set, models the spatial relationships of road-participating object instances through a graph convolutional network, performs dynamic instance partitioning graph structure for continuous changes between vehicle and environmental data frames, generates an object instance spatial relationship graph, and transmits it to the semantic understanding module. The semantic understanding module, based on the spatial relationship graph of the object instances, performs semantic label inference on the road-related object instances through the Transformer model, and generates a 3D scene semantic understanding result by comparing the multimodal fusion features with the semantic labels.

Citation Information

Cited By

  • Scene label generation method and device based on dynamic graph reasoning

    CN121564718A

  • Heterogeneous sensor-oriented multi-modal perception fusion and three-dimensional semantic scene reconstruction method

    CN121600195A