Scene perception method, training method, program product, medium, and electronic device
By combining Gaussian sphere features and Gaussian decoders, the problem of high computational resource consumption of dense feature maps is solved, and sparse feature map generation is realized, which improves the perception efficiency and accuracy in scenarios such as assisted driving.
Patent Information
- Application Number
- CN202511383648.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing scene perception methods based on dense feature maps consume a lot of computational resources and are difficult to meet the real-time requirements of scenarios such as assisted driving. In addition, the prediction results contain redundant information.
Gaussian sphere features are used to replace dense feature maps. Scene perception is achieved through a Gaussian decoder. Feature extraction and prediction are performed using multi-view images and the initial query vector of the Gaussian sphere. Combined with an adaptive attention mechanism and residual links, sparse feature maps are generated.
It reduces the demand for computing resources, improves the real-time performance and accuracy of scene perception, reduces redundant information, and improves computing efficiency.
Smart Images

Figure CN120877253A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic device technology, and in particular to a scene perception method, training method, program product, medium and electronic device. Background Technology
[0002] With the continuous development of the field of driver assistance, environmental perception methods based on BEV feature maps are playing an increasingly important role in obstacle perception.
[0003] Currently, BEV feature maps are typically dense feature maps, which results in a huge computational cost for perceiving dense feature maps.
[0004] In addition, in scenarios such as assisted driving, where computing resources are limited and the real-time requirements for scene perception are very high, perception methods based on dense feature maps cannot effectively meet the needs of such scenarios.
[0005] In practical applications, obstacles in the environment to be perceived are usually sparse, and using dense feature maps for obstacle prediction would waste a lot of computing resources. Summary of the Invention
[0006] This application provides a scene perception method, training method, program product, medium, and electronic device to address the problem of high computational resource consumption in scene perception in the prior art. The following describes this application from multiple aspects, and the embodiments and beneficial effects described below can be referenced interchangeably.
[0007] Firstly, this application provides a scene-aware method applied to electronic devices. The method includes: Acquire multi-view images of the target scene; extract features from the multi-view images to obtain a first feature map corresponding to the multi-view images; initialize multiple Gaussian spheres and determine initial query vectors for the multiple Gaussian spheres based on their initial attributes; input the first feature map and the multiple initial query vectors into a target Gaussian decoder to predict the first target attributes of the multiple Gaussian spheres corresponding to the first feature map, the first target attributes including the first predicted position and the first predicted semantic of the Gaussian spheres; obtain the perception result of the target scene based on the first target attributes of the multiple Gaussian spheres, the perception result including the semantic information of the spatial location of the target scene determined based on the first predicted position and the first predicted semantic.
[0008] According to embodiments of this application, a target scene perception method based on a Gaussian sphere is provided. Through a single path of multi-view features - Gaussian decoder - spatial perception results, the reconstruction and semantic annotation of the target scene perception results are completed in one step. It eliminates the need for dense feature maps, simplifying the perception process and reducing the demand for computational resources.
[0009] In some embodiments, a first feature map and multiple initial query vectors are input into a target Gaussian decoder to predict a first target attribute of multiple Gaussian spheres corresponding to the first feature map. This includes: generating multiple sampling points in each Gaussian sphere based on the initial query vectors, and sampling the first feature map based on the multiple sampling points to obtain a second feature map; performing feature fusion on the second feature map and the initial query vectors of the multiple Gaussian spheres to obtain a third feature map; updating the initial query vectors according to an adaptive attention mechanism to obtain a first query vector; performing a first processing on the third feature map and the first query vector to obtain a second query vector, the first processing including at least linear processing; superimposing the second query vector and the third query vector to obtain a fourth query vector, the third query vector being obtained by embedding the initial attribute; and decoding the fourth query vector to obtain the first target attribute of the multiple Gaussian spheres.
[0010] According to an embodiment of this application, a Gaussian decoder is provided. First, for the initial query vector of each Gaussian sphere, sampling points are dynamically generated within the sphere; local features are extracted from the first feature map to obtain a second feature map; then, feature fusion is performed to obtain a third feature map; then, the query vector is updated through an adaptive attention mechanism; the updated query vector is linearly transformed and superimposed with the attribute embedding vector, and finally decoded into the position and semantic attributes of the Gaussian sphere in one step.
[0011] The Gaussian decoder achieves hierarchical coupling from local to global features. It uses sampling points in the Gaussian sphere to collect features in dense feature maps and obtain sparse feature maps, thus reducing computational overhead.
[0012] In some embodiments, the number of sampling points is determined based on the computing resources of the electronic device.
[0013] In some embodiments, feature fusion is performed on the second feature map and the initial query vectors of multiple Gaussian spheres to obtain a third feature map, including: adaptive feature fusion of the second feature map to obtain a fourth feature map; and residual linking of the fourth feature map with the initial query vectors of multiple Gaussian spheres to obtain the third feature map.
[0014] In some embodiments, adaptive feature fusion of the second feature map includes: obtaining a first weight for adaptive fusion based on the second feature map, and adaptively fusion of the second feature map based on the feature channels and / or spatial locations of the second feature map.
[0015] In some embodiments, the first feature map is a multi-layer feature map; the third feature map is obtained by residual linking the fourth feature map with the initial query vectors of multiple Gaussian spheres, including: residual linking the fourth feature map with the initial query vectors of multiple Gaussian spheres to obtain a fifth feature map, the fifth feature map being a multi-layer feature map with the same number of layers as the first feature map; and layer normalization of the fifth feature map to obtain the third feature map.
[0016] According to embodiments of this application, a feature fusion method is provided. This method integrates feature information from feature maps at different scales, feature channels, or locations through dynamic weight fusion, suppressing redundancy and highlighting key features. Residual linking is used to preserve the original query semantics, alleviate gradient vanishing, and accelerate training convergence.
[0017] In some embodiments, updating the initial query vector according to the adaptive attention mechanism to obtain the first query vector includes: inputting multiple initial query vectors into a linear layer to obtain receptive field weights; obtaining attention weights based on the pairwise distance between the receptive field weights and the initial query vectors, wherein the attention weights decrease as the receptive field weights increase; and updating the initial query vectors according to the attention weights to obtain the first query vector.
[0018] According to an embodiment of this application, an adaptive attention calculation method is provided. Each initial query vector is mapped to a receptive field weight through a linear layer, and attention weights are calculated based on pairwise distances. The attention weights are then made to decrease monotonically as the receptive field weights increase, achieving both spatial and semantic sparsity.
[0019] This self-regulating mechanism adaptively suppresses queries that are far away or have a large receptive field, while strengthening queries that are close away and have a small receptive field. This mechanism automatically prunes redundant interactions while maintaining local precision, significantly reducing computational complexity and memory usage, and improving the accuracy and convergence stability of subsequent Gaussian attribute decoding.
[0020] In some embodiments, the first process further includes an activation process following the linear processing.
[0021] In some embodiments, obtaining a perception result of a target scene based on a first target attribute of a plurality of Gaussian spheres includes: determining a fifth query vector of a plurality of Gaussian spheres based on the first target attribute of the plurality of Gaussian spheres; inputting the fifth query vector and a first feature map into a target Gaussian decoder to obtain a second target attribute of the plurality of Gaussian spheres; and obtaining a perception result based on the second target attribute of the plurality of Gaussian spheres.
[0022] According to embodiments of this application, by repeatedly performing cyclic predictions on the Gaussian decoder, the target attributes of the Gaussian sphere are made to better match the target scene.
[0023] Secondly, this application provides a training method applied to an electronic device. The method includes: acquiring training sample data, which includes a first feature map corresponding to a multi-view image of a target scene and initial query vectors of multiple Gaussian spheres; inputting the training sample data into a Gaussian decoder to be trained to obtain the predicted attributes of the multiple Gaussian spheres, which include the predicted positions and predicted semantics of the multiple Gaussian spheres; and training the Gaussian decoder to be trained based on the predicted attributes of the multiple Gaussian spheres to obtain a target Gaussian decoder.
[0024] In some embodiments, training sample data is input into a Gaussian decoder to be trained to obtain predicted attributes of multiple Gaussian spheres, including: generating multiple sampling points in each Gaussian sphere according to an initial query vector, and sampling a first feature map according to the multiple sampling points to obtain a second feature map; performing feature fusion on the second feature map and the initial query vector of the multiple Gaussian spheres to obtain a third feature map; updating the initial query vector according to an adaptive attention mechanism to obtain a first query vector; performing a first processing on the third feature map and the first query vector to obtain a second query vector, the first processing including at least linear processing; superimposing the second query vector and the third query vector to obtain a fourth query vector, the third query vector being obtained by embedding the initial attributes; and decoding the fourth query vector to obtain predicted attributes of the multiple Gaussian spheres.
[0025] In some embodiments, training a Gaussian decoder to be trained based on the predicted attributes of multiple Gaussian spheres to obtain a target Gaussian decoder includes: acquiring label data, the label data including target location information and target semantic information at multiple locations in the target scene; calculating a first loss between the predicted locations of the multiple Gaussian spheres and the target location information; calculating a second loss between the predicted semantic information of the multiple Gaussian spheres and the target semantic information; and adjusting the Gaussian decoder to be trained based on the first loss and the second loss to obtain the target Gaussian decoder.
[0026] According to an embodiment of this application, a training method for a Gaussian decoder is provided, comprising four steps: input, output, loss calculation, and backpropagation. The target location information and target semantic information from the target attributes of the Gaussian sphere output by the Gaussian decoder are used as label values. The loss value between the label value and the predicted value is calculated, and the Gaussian decoder is trained through supervised learning.
[0027] This training method outputs the predicted values of the Gaussian sphere target attributes in an end-to-end manner by using the prior first feature map and the learnable Gaussian sphere target attributes, thus avoiding the error accumulation of traditional multi-stage training.
[0028] In addition, this training method improves the accuracy of Gaussian decoder prediction by constraining the prediction of dual label values of location and semantic information.
[0029] Thirdly, this application provides a computer program product, which includes instructions that, when executed by an electronic device, cause the electronic device to perform the scene perception method or the Gaussian decoder training method of any embodiment of the first or second aspect.
[0030] Fourthly, this application provides a computer medium storing instructions that, when executed on a computer, enable the computer to perform the scene perception method or the Gaussian decoder training method of any embodiment of the first or second aspect.
[0031] Fifthly, this application provides an electronic device, including: a memory for storing instructions executed by one or more processors of the electronic device; and a processor that, when executing the instructions in the memory, causes the electronic device to execute the scene perception method or the Gaussian decoder training method of any embodiment of the first or second aspect of this application.
[0032] It is understood that the beneficial effects that can be achieved by the third to fifth aspects described above can be referred to the beneficial effects of any embodiment of the first or second aspect, and will not be repeated here. Attached Figure Description
[0033] Figure 1 A schematic diagram illustrating an exemplary application scenario of the sensing method provided in the embodiments of this application; Figure 2 An exemplary flowchart of the scene perception method provided in the embodiments of this application; Figure 3 An exemplary flowchart of a Gaussian decoder provided in an embodiment of this application; Figure 4 A schematic diagram of the Gaussian decoder training method provided in this embodiment; Figure 5 A schematic diagram of the grid used in this application embodiment; Figure 6 Modular flowcharts provided for embodiments of this application; Figure 7 A block diagram of an electronic device provided in an embodiment of this application; Figure 8 This is a block diagram of a system-on-a-chip provided in an embodiment of this application. Detailed Implementation
[0034] This application provides a scene perception method, which can improve the problem of high computational resource consumption during the target scene perception process.
[0035] Figure 1An exemplary application scenario of this application is shown, specifically a vehicle-assisted driving scenario. (Reference) Figure 1 The road ahead may contain obstacles such as roadblocks, bicycles, and pedestrians. During assisted driving, the vehicle needs to detect these objects in order to avoid them.
[0036] In other words, Figure 1 In the example shown, the road scene is the target scene that the vehicle needs to perceive during the assisted driving process, and the road barrier 110, bicycle 120, pedestrian 130, etc. are the objects to be perceived.
[0037] In some technical solutions, target scene perception is based on bird's-eye-view (BEV) scene perception technology. In BEV scene perception technology, data collected by multiple sensors mounted on the vehicle (such as cameras, LiDAR, and millimeter-wave radar) is transformed into a unified, top-down two-dimensional coordinate system. Feature extraction, fusion, and subsequent perception tasks (such as object detection, map segmentation, and prediction) are then performed within this unified BEV space.
[0038] Currently, BEV feature maps are typically dense feature maps, which results in a huge computational cost for perceiving dense feature maps.
[0039] In addition, in scenarios such as assisted driving, where computing resources are limited and the real-time requirements for scene perception are very high, perception methods based on dense feature maps cannot effectively meet the needs of such scenarios.
[0040] Furthermore, predictions based on dense feature maps often include many components that are not of interest to downstream applications. For example, in... Figure 1 In the field of assisted driving, path planning based on the currently perceived target scene only requires information such as obstacles or curbs in the target scene. However, when using dense feature maps for prediction, the prediction results include redundant information such as the sky, road texture, or background buildings. This irrelevant information not only wastes valuable computational resources but can also cause problems such as false positives.
[0041] To address this, this application proposes a scene perception method. In an embodiment of this application, an electronic device acquires multi-view images of a target scene (e.g., images captured by multiple cameras on a vehicle). The electronic device performs feature extraction on the multi-view images (e.g., using a multi-scale feature extraction network) to obtain a first feature map (e.g., a multi-scale feature map) corresponding to the multi-view images. The electronic device initializes multiple Gaussian spheres (e.g., initializes the Gaussian spheres at random locations using random data) and determines initial query vectors for the multiple Gaussian spheres based on their initial attributes (e.g., center coordinates, rotation matrix, scaling matrix, opacity, or semantics) (e.g., encoding the initial attributes as high-dimensional vectors).
[0042] The electronic device inputs a first feature map and multiple initial query vectors into a target Gaussian decoder to predict the first target attributes of multiple Gaussian spheres corresponding to the first feature map (e.g., using the Gaussian decoder to predict the semantics or location of the object perceived by each Gaussian sphere). The first target attributes include the first predicted location and the first predicted semantics of the Gaussian spheres. Based on the first target attributes of the multiple Gaussian spheres, the electronic device obtains the perception result of the target scene (e.g., converting the predicted locations of the Gaussian spheres and the prediction results into point clouds or voxel networks). The perception result includes semantic information about the spatial location of the target scene determined based on the first predicted location and the first predicted semantics.
[0043] In this embodiment, a sparse feature map can be obtained by replacing the dense multi-view feature map with Gaussian sphere features. The corresponding query vector is then obtained using the Gaussian sphere, and a sparse occupancy grid is constructed.
[0044] Through the above-described solution, electronic devices can reduce computational load, improve perception performance, and enhance general obstacle perception capabilities. The following provides a detailed description of the scene perception method provided in the embodiments of this application.
[0045] It should be noted that, Figure 1 While the example uses an assisted driving scenario, this application is not limited to this. In other embodiments, this application can be applied to scenarios such as drone inspection, smart city management, or industrial equipment monitoring. For example, in a drone inspection scenario, the scene perception method proposed in this application can be used to quickly perceive target objects in the drone's working environment, including key components of equipment such as power transmission towers or photovoltaic panels.
[0046] in addition, Figure 1 While an in-vehicle infotainment system is used as an example of an electronic device, this application is not limited thereto. In other embodiments, the electronic device may be a smartphone, tablet, or controller for a logistics robot, etc., and this application is not limited thereto.
[0047] The following description uses the assisted driving scenario as an example of the target scenario and the vehicle system as an example of an electronic device to illustrate the scene perception method provided in this application.
[0048] Figure 2 An exemplary flowchart of the scene perception method provided in the embodiments of this application. Figure 2 The entity executing each step can be an electronic device, such as an in-vehicle infotainment system.
[0049] refer to Figure 2 The scene perception method provided in this application embodiment may include: Step S110: Acquire multi-view images of the target scene.
[0050] For example, the target scenario can be the environment in which the vehicle is currently located, such as urban roads, highways, tunnels, underground parking lots, port container yards, bridges or viaducts, etc. This application does not limit it.
[0051] For example, a vehicle may include multiple cameras, which can have different shooting angles to capture images from different perspectives. That is, by capturing images of the target scene using multiple cameras, multi-view images of the target scene can be obtained. In this embodiment, using cameras with different shooting angles can increase the vehicle's perception range of the target scene, enabling it to perceive objects in the target scene as much as possible and reduce blind spots. After obtaining multi-view images of the target scene, the multiple cameras can transmit the captured multi-view images to the vehicle's infotainment system through a data transmission interface or data transmission bus, so that the vehicle's infotainment system can perceive the multi-view images of the target scene.
[0052] For example, the vehicle system can acquire multi-view images captured by multiple cameras in step S110, which can be expressed by formula (1): (1) Here, I can represent a set of multi-view images captured by multiple cameras in step S110. Each element in the set... , which can represent the image of the target scene captured by the i-th camera.
[0053] In some embodiments of this application, the capture by multiple cameras is sequential. For example, when the vehicle's driver assistance function is activated, multiple cameras can continuously collect data on the target scene surrounding the vehicle. Therefore, I can also represent the set of multi-view images captured by multiple cameras at the current time t.
[0054] In some embodiments of this application, N can represent the total number of cameras. Since different cameras have different shooting angles, N different cameras can also represent N different shooting angles. It can be understood that in some embodiments of this application, N can be 4, capturing the front, rear, left, and right views of the vehicle, thus forming a 360-degree circular field of view centered on the vehicle. In other embodiments of this application, N can also be a value such as 4, 5, 6, 7, or 8, meaning that the vehicle is equipped with 4, 5, 6, 7, or 8 cameras.
[0055] W and H represent the width and height of the image captured by each of the multiple cameras, respectively. In some embodiments of this application, the number of pixels corresponding to the width or height of the image can be used directly.
[0056] Step S120: Extract features from the multi-view images to obtain the first feature map corresponding to the multi-view images.
[0057] In some embodiments of this application, the first feature map obtained through feature extraction includes multiple features that serve subsequent scene perception tasks. For example, the geometric and semantic features of the road surface include the feature texture and category information of ground elements such as lane lines, speed bumps, or parking spaces.
[0058] In some embodiments of this application, the first feature map may include multiple layers. Each layer may represent features extracted by the vehicle-mounted system at a certain scale in the multi-view image. For example, large-scale features can represent coarse semantic information within a relatively far range around the vehicle (e.g., 20m × 20m). At this scale, the features extracted by this layer are more representative of large structures such as lane lines, curbs, or buildings.
[0059] The vehicle-mounted system can acquire multiple features of a multi-view image at a certain scale through a feature extraction network (e.g., using multiple convolutional kernels for downsampling). In some embodiments of this application, pixels containing multiple feature channels can be used to represent multiple features of the multi-view feature map at that location. These multiple feature channels are used to store multiple feature values at that location.
[0060] For example, before extracting features from multi-view images, the intrinsic and extrinsic parameters of the camera must first be obtained. These parameters ensure that each pixel in the multi-view image is accurately projected onto the same world coordinate system, facilitating subsequent feature learning from the multi-view image.
[0061] The intrinsic parameters can be expressed using formula (2): (2) Here, K can represent a set of intrinsic parameters from multiple cameras, where the intrinsic parameters are a 3×3 matrix. For each element in the set... , which represents the intrinsic parameters corresponding to the i-th camera. N represents the total number of cameras.
[0062] The role of intrinsic parameters is to define the mapping relationship between the camera coordinate system and the image coordinate system, and to realize the mutual transformation of coordinates between the camera coordinate system and the image coordinate system.
[0063] The external parameters can be expressed using formula (3): (3) Here, T represents the set of extrinsic parameters from multiple cameras, which can be a 4×4 matrix. For each element in the set... , which represents the extrinsic parameter corresponding to the i-th camera. N represents the total number of cameras.
[0064] The role of extrinsic parameters is to determine the mapping relationship between the camera coordinate system and the world coordinate system, and to realize the mutual transformation of coordinates between the camera coordinate system and the world coordinate system.
[0065] For example, after obtaining a multi-view image in the world coordinate system using intrinsic and extrinsic parameters, a feature extraction network can be used to extract features from this multi-view image. The feature extraction network can extract various features from the multi-view image at different scales. These different features can be stored in multiple feature channels at each location in the multi-view image.
[0066] In some embodiments of this application, a backbone network and a feature pyramid network (FPN) can be used as multi-scale feature extraction networks to extract features from multi-view images and obtain first feature maps at different scales.
[0067] The first feature map extracted by the feature extraction network can be a multi-layered feature map, where each layer represents a feature map of a multi-view image at a certain scale. For example, the resolution of the first-scale feature map can be... The resolution of the feature map at the second scale can be In the matrix representation, H can represent the width of the multi-view image, and W can represent the length of the multi-view image.
[0068] Furthermore, each feature map layer includes a feature vector at each location point. For example, the feature map at the first scale could be... That is, each location point corresponds to a C-dimensional feature vector. Similarly, the feature map at the second scale can be derived from... .
[0069] In some embodiments of this application, C may be referred to as the number of channels in the feature map. This application does not limit the value of C. In some embodiments, C can be 128. In other embodiments of this application, C can also be 64, 96, 192, 256 or 512. This application does not impose any restrictions on this.
[0070] The reason why the feature map at the second scale has a smaller resolution than that at the first scale is that when performing feature sampling on the multi-view image according to the second scale, a convolutional kernel with a larger receptive field is used for downsampling, thus resulting in a feature map with a smaller resolution. For example, at the second scale, a convolution with a receptive field of 5*5 is used to downsample the multi-view image. Since the computational cost of a convolutional kernel with a larger receptive field is greater, in some embodiments of this application, the convolutional kernel with a larger receptive field has a larger stride during convolution (for example, the stride is 1 when using a 3*3 convolutional kernel and 2 when using a 5*5 convolutional kernel).
[0071] In some embodiments of this application, the multi-scale feature extraction network can be based on a Transformer-based multi-scale attention structure, in addition to using a backbone network and feature pyramids. This application does not impose any limitations on this method.
[0072] Step S130: Initialize multiple Gaussian spheres and determine the initial query vector of the multiple Gaussian spheres based on their initial properties.
[0073] The vehicle's infotainment system can randomly initialize multiple Gaussian spheres based on multi-view images corresponding to the target scene. Each Gaussian sphere can be uniquely identified by a set of initial attributes including multiple parameters. Based on the initial attributes of each Gaussian sphere, the vehicle's infotainment system can encode a query vector corresponding to that Gaussian sphere for interaction with the first feature map.
[0074] For ease of understanding, this application will first briefly describe the role of the Gaussian sphere. The specific methods for initializing the Gaussian sphere and generating the initial query vector for the Gaussian sphere will be described below.
[0075] For example, the Gaussian sphere can be determined by a set of initial properties containing multiple parameters, such as: the center position (mean) of the Gaussian sphere, the rotation matrix (quaternion), semantic logits, and / or the scaling matrix. In some embodiments of this application, the initial properties may also include properties such as the opacity of the Gaussian sphere.
[0076] The Gaussian sphere's center position (mean) represents its 3D position in the world coordinate system. The rotation matrix represents the angle between the Gaussian sphere and the coordinate axes. Semantic logits represent the type of object the Gaussian sphere corresponds to in the target scene, such as a curb or obstacle. The scaling matrix controls the scaling ratio of the Gaussian sphere along the coordinate axes.
[0077] For example, the initial properties can be represented by formula (4): (4) Here, g represents the set of initial properties of the Gaussian sphere. For each of the initial properties g... , which represents the initial properties of the i-th Gaussian sphere. The value of i ranges from 1 to P, where P represents the number of Gaussian spheres initialized. In some embodiments of this application, the value of P can be from 100 to 1000. In other embodiments of this application, P can be set according to the computing resources available to the electronic device on which the sensing method operates, and this application does not impose any limitations on this setting.
[0078] For example, g can be used to represent the initial properties of a Gaussian sphere. During the initialization of the Gaussian sphere, g can be set to a random value to complete the initialization.
[0079] Subsequently, based on g0 and formula (5), the vehicle-mounted system can encode the corresponding query vector for the Gaussian sphere: (5) Here, q represents the set of initial query vectors for all Gaussian spheres. For each of the initial query vectors q... , which represents the initial query vector for the i-th Gaussian sphere. q is an m-dimensional vector. When the vehicle system encodes the initial query vector, g can be implicitly encoded into q. The value of i ranges from 1 to P, where P represents the number of Gaussian spheres initialized.
[0080] For example, q can be used to represent the initial query vector of the Gaussian sphere. In some embodiments of this application, for each of the initial attributes g... The initial attributes g of a Gaussian sphere can be encoded into a Gaussian query q using methods such as dense embedding. In this process, the initial query vector of the Gaussian sphere is... The attribute values are filled into a one-dimensional vector of length L in a fixed order. When the dimension of an attribute is insufficient to make the length of the one-dimensional vector reach L, the one-dimensional vector is randomly padded to make its length meet the requirement of L. In some embodiments of this application, the length of L is 11. In other embodiments of this application, the length of L can also be any integer between 11 and 30, which is not limited here.
[0081] In addition to dense embedding, the embodiments provided in this application can also encode the initial attribute g through feature embedding or continuous value embedding to obtain the corresponding query vector q, which is not limited here.
[0082] Step S140: Input the first feature map and multiple initial query vectors into the target Gaussian decoder to predict the first target attributes of multiple Gaussian spheres corresponding to the first feature map. The first target attributes include the first predicted position and the first predicted semantic of the Gaussian sphere.
[0083] For example, the Gaussian decoder's function is to simultaneously receive a first feature map and multiple initial query vectors q, and predict the first target attribute corresponding to q based on the input data. Here, the first feature map carries multi-scale features of the target scene, and each query vector in q corresponds to a Gaussian sphere to be optimized.
[0084] Subsequently, the Gaussian decoder initiates the decoding process. Based on the center position of the Gaussian sphere corresponding to each query vector in q, it determines the location in the first feature map where the query should be perceived. This allows q to perform the relevant perception operation at the corresponding location. Using the results perceived by q at the query location (e.g., the semantics and position of the object at the query location), q is decoded to obtain the first target attribute that the Gaussian sphere corresponding to q should correct.
[0085] In some embodiments of this application, the first target attribute may include the predicted center positions of multiple Gaussian spheres (or "first predicted positions"), and the semantics of the perceived object that each Gaussian sphere should represent in the corresponding target scene (or "first predicted semantics"). For ease of understanding, the specific process of perceiving the first feature map through q is described below.
[0086] In addition, in some embodiments, the first target attribute may also include attributes such as rotation matrix (or "quaternion"), scaling matrix, and opacity.
[0087] Step S150: Based on the first target attributes of multiple Gaussian spheres, obtain the perception result of the target scene. The perception result includes the semantic information of the spatial location of the target scene determined based on the first predicted position and the first predicted semantic.
[0088] After decoding using a Gaussian decoder, the vehicle's infotainment system can combine the predicted first target attributes to obtain semantic information about objects at any location in the target scene space and the confidence level of that semantic information. By aggregating this information, the system can obtain a panoramic perception result of the target scene. To obtain more accurate perception results, the system can also decode the Gaussian sphere multiple times to improve the precision of the perception results.
[0089] For ease of understanding, this application will first briefly describe the perception results; how the perception results are obtained will be described in detail below.
[0090] In some embodiments of this application, the perception result of the target scene can be represented by a semantic point cloud. The semantic point cloud can consist of multiple discrete points in the three-dimensional space of the target scene. Each point can be represented by a set of first descriptive information. The first descriptive information may include the coordinates of the current point in the three-dimensional space of the target scene, and may also include the semantics of the target scene at the current point.
[0091] For example, based on the first predicted position in the first target attribute of each Gaussian sphere, the points covered by each Gaussian sphere in the target scene can be determined. Finally, the first descriptive information of the points covered by each Gaussian sphere is assigned using the first predicted semantics in the first target attribute of each Gaussian sphere. For example, the semantics of the points covered by each Gaussian sphere is assigned as the first predicted semantics of that Gaussian sphere.
[0092] In some embodiments of this application, the perception result of the target scene can also be represented using an occupancy grid. Specifically, the three-dimensional space where the target scene is located can be represented as multiple voxels in the occupancy grid, and each voxel can be represented by a set of second descriptive information. The second descriptive information may include the coordinates of the current voxel in the occupancy grid, and may also include the semantics of the current voxel. It is understood that different voxels correspond to different coordinates.
[0093] First, an occupancy grid corresponding to the target scene is established. The spatial range occupied by this occupancy grid can be greater than or equal to the three-dimensional space of the target scene. Then, based on the first predicted position in the first target attribute of each Gaussian sphere, the range of voxels covered by each Gaussian sphere in the target scene can be determined. Finally, the second descriptive information of the voxels covered by each Gaussian sphere is assigned using the first predicted semantics in the first target attribute of each Gaussian sphere. For example, the semantics of the voxels covered by each Gaussian sphere are assigned as the first predicted semantics of that Gaussian sphere.
[0094] Repeat the above steps to complete the assignment of the overlay voxel description information to all Gaussian spheres in order to obtain the perception results of the target scene.
[0095] For example, to ensure that the predicted first target attributes of the multiple Gaussian spheres are closer to the target scene, the target scene can be perceived multiple times: Based on the first target attributes of the multiple Gaussian spheres, a fifth query vector for the multiple Gaussian spheres is determined. The fifth query vector and the first feature map are input into the target Gaussian decoder to obtain the second target attributes of the multiple Gaussian spheres. Based on the second target attributes of the multiple Gaussian spheres, the perception result is obtained. The method for obtaining the perception result of the target scene based on the second target attributes is essentially the same as the method for obtaining the perception result of the target scene based on the first target attributes, and will not be elaborated upon here.
[0096] For example, the first target attribute of the predicted multiple Gaussian spheres output by the Gaussian decoder is re-encoded into a fifth query vector. The re-encoded fifth query vector is input into the Gaussian decoder to re-execute step S140 to obtain the re-predicted second target attribute. Based on the second target attribute, the perceptual result corresponding to the second target attribute is calculated.
[0097] In some embodiments of this application, after obtaining the re-predicted second target attribute, the second target attribute can be re-encoded into a sixth query vector. The re-encoded sixth query vector is input into a Gaussian decoder to re-execute step S140 to obtain the re-predicted third target attribute. Then, based on the third target attribute, a perception result of the target scene is obtained. The method for obtaining the perception result of the target scene based on the third target attribute is substantially the same as the method for obtaining the perception result of the target scene based on the first target attribute, and will not be described in detail here.
[0098] In some embodiments of this application, multiple iterations of prediction can be performed based on the target attributes predicted by the Gaussian decoder, and the vehicle's perception result of the target scene can be determined based on the target attributes of the last obtained Gaussian sphere. Since the number of iterations directly affects the computational resource requirements, those skilled in the art can determine the number of iterations according to the needs of the actual scenario, such as three, four, or five times, and this application does not limit this.
[0099] Figure 3 An exemplary flowchart of a Gaussian decoder provided in an embodiment of this application. Figure 3 The entity executing each step can be an electronic device, such as an in-vehicle infotainment system.
[0100] For example, refer to Figure 3 Step S140 may also include the following sub-steps: Step S141: Generate multiple sampling points in each Gaussian sphere according to the initial query vector, and sample the first feature map according to the multiple sampling points to obtain the second feature map.
[0101] The vehicle-mounted system uses the initial query vector of each Gaussian sphere as a basis to select several sampling points within the 3D region of the Gaussian sphere. The selection method can be uniform, ensuring that these sampling points are evenly distributed across each Gaussian sphere. Subsequently, the vehicle-mounted system performs feature sampling on the locations of these sampling points in the first feature map, converging them into a second feature map, thus making the originally dense feature map sparser.
[0102] For ease of understanding, this application will first briefly describe the function of sampling points, and the specific method of using sampling points for sampling will be described below.
[0103] For example, sampling points can be represented by the position coordinates of points used for feature sampling within a Gaussian sphere. After determining the center position and included range of each Gaussian sphere, the vehicle-mounted system can generate multiple evenly spaced sampling points within the Gaussian sphere based on the initial properties of the remaining Gaussian spheres. For example, X sampling points can be generated within the Gaussian sphere, where X can be 10. Since the computational resources required by the Gaussian decoder are positively correlated with the number of sampling points in the Gaussian sphere, those skilled in the art can determine the number of sampling points in the Gaussian sphere according to actual needs, and this application does not impose any limitation on this.
[0104] After generating multiple sampling points within each Gaussian sphere, in order to sample the first feature map, it is necessary to project the coordinates of the sampling points onto the first feature map based on the intrinsic and extrinsic parameters of the multi-view image, and then sample to obtain the second feature map.
[0105] In some embodiments of this application, the second feature map and the first feature map have the same storage structure except for resolution. Since the number of feature points in a Gaussian sphere is much smaller than the number of pixels in the first feature map, this sampling based on multiple Gaussian sphere sampling points is a form of downsampling. Downsampling makes the resolution of the second feature map much smaller than that of the first feature map. Because the Gaussian sphere samples the extraction positions of the first feature map, the second feature map is a sparse feature map.
[0106] Step S142: Perform feature fusion on the second feature map and the initial query vectors of multiple Gaussian spheres to obtain the third feature map.
[0107] For example, feature fusion refers to integrating feature vectors from different sources, scales, or modalities into a unified feature map, giving it richer semantics and more complete spatial details.
[0108] In some embodiments of this application, adaptive feature fusion can be used to perform feature fusion on the sampled second feature map.
[0109] Specifically, an adaptive feature fusion method can be used, which can be performed from either the dimension of sampling points or the two dimensions of feature channels.
[0110] In some embodiments of this application, adaptive feature fusion first requires stacking multiple temporal second feature maps. For example, first, T temporal second feature maps are collected, each feature map has S sampling points, and these second feature maps are stacked to obtain P = T * S sampling points.
[0111] When performing adaptive feature fusion from the feature channel dimension, for these stacked second feature maps, feature fusion is performed from the feature channel dimension according to formulas (6) and (7): (6) (7) in, This represents the dynamic weights used for feature fusion from the feature channel dimension, which are shared across all frames and sampling points. The Linear function represents a linear layer. q represents the initial query vector of the Gaussian sphere, LayerNorm is the layer normalization function, and ReLU represents the activation function. The output... This represents the third feature map after feature fusion along the feature channel dimension.
[0112] When performing adaptive feature fusion from the dimension of sampling points, for these stacked second feature maps, feature fusion is performed from the dimension of sampling points according to formulas (8) and (9): (8) (9) in, This represents the dynamic weights used for feature fusion from the sampling point dimension, which are shared across all frames and sampling points. The Linear function represents a linear layer. q represents the initial query vector of the Gaussian sphere, LayerNorm is the layer normalization function, and ReLU represents the activation function. The output... This represents the third feature map after feature fusion at the sampling point dimension.
[0113] For example, in some embodiments of this application, the feature map obtained after adaptive feature fusion is a fourth feature map. The output fourth feature map can also be residually linked with the initial query vectors of multiple Gaussian spheres to obtain a third feature map after feature fusion. Residual linking, by superimposing the fourth feature map with the initial query vectors of multiple Gaussian spheres, can reduce gradient explosion or gradient vanishing, making the training of the Gaussian decoder more stable.
[0114] Step S143: Update the initial query vector according to the adaptive attention mechanism to obtain the first query vector.
[0115] For ease of understanding, this application will first briefly describe the role of the attention mechanism. The specific method of how to use the attention mechanism to update the initial vector will be described below.
[0116] For example, the vehicle's infotainment system can utilize an adaptive attention mechanism to automatically adjust the size of the receptive field for each Gaussian query vector representing a potential obstacle, making it more focused on the most relevant visual region in the current frame. For instance, for a Gaussian sphere whose center is far from the vehicle, its receptive field can be automatically increased to focus more on distant obstacles.
[0117] Specifically, an adaptive attention mechanism can be implemented using adaptive weights. Updating the initial query vector based on attention weights can be achieved by performing the following steps: inputting multiple initial query vectors into a linear layer to obtain receptive field weights. Obtaining attention weights based on the pairwise distances between the receptive field weights and the initial query vectors; these attention weights decrease as the receptive field weights increase. Updating the initial query vectors based on these attention weights yields the first query vector.
[0118] Specifically, in the process of inputting multiple initial query vectors into a linear layer to obtain receptive field weights, the receptive field refers to the region where each initial query vector interacts with the third feature map. Its range and shape adaptively change with the third feature map.
[0119] Specifically, the receptive field weights of the Gaussian decoder can be obtained using the following formula (10).
[0120] (10) Where d represents the dimension of the initial query vector, H represents the number of attention heads used to generate the receptive field weights, and q represents the initial query vector. This represents the vector composed of the generated receptive field weights.
[0121] In some embodiments of this application, the Euclidean distance between the centers of two Gaussian spheres can be expressed using a formula as shown in formula (11): (11) in, This represents the Euclidean distance between the centers of two Gaussian spheres. and , and These represent the coordinates of the center position of the Gaussian sphere.
[0122] Based on the Euclidean distance between the centers of the two Gaussian spheres and the receptive field weights, the attention weights can be obtained according to formula (12): (12) in, The weights represent adaptive attention. Q is the activation function, K represents the initial query vector of the Gaussian sphere, V represents the set of features of all query vectors in all receptive fields, V represents the feature dimension of the actual vector to be aggregated or weighted and summed, and d represents the dimension of the initial query vector.
[0123] When the receptive field weight increases, the attention weight of the initial query vector of the Gaussian sphere that is farther away decreases, which will cause the receptive field to shrink accordingly.
[0124] In some embodiments of this application, when the receptive field weight is reduced to 0, the adaptive attention mechanism degenerates into an attention mechanism with a global receptive field.
[0125] Through an adaptive attention mechanism, the vehicle system can dynamically adjust the level of attention given to different Gaussian spheres, making the model more flexible and efficient in handling complex scenes.
[0126] In some embodiments of this application, the vehicle system determines the size of the receptive field of the Gaussian decoder based on the distance between the initial query vectors of the two Gaussian spheres, and then updates the initial query vectors to obtain the first query vector.
[0127] Step S144: Perform a first process on the third feature map and the first query vector to obtain the second query vector. The first process includes at least linear processing.
[0128] The vehicle-mounted system maps the third feature map to the first query vector, which has been associated with an attention mechanism, and then performs a fast mapping through a linear network to obtain the second query vector. In some embodiments of this application, the first process may further include activation processing following the linear processing.
[0129] For example, the linear processing serves to integrate the third feature map with the first query vector within a unified spatial dimension. In some embodiments of this application, linear processing and activation processing can be used together as a feedforward network to process the first query vector to obtain the second query vector.
[0130] In some embodiments of this application, linear processing can be performed using fully connected layers. Activation processing can be performed using activation functions, such as ReLU, GELU, or Swish, which increase the nonlinearity of the network and achieve abstraction of deep features.
[0131] Those skilled in the art can decide for themselves which linear networks or activation functions to use to form the feedforward network to perform the first processing, and this application does not impose any restrictions.
[0132] Step S145: Superimpose the second query vector and the third query vector to obtain the fourth query vector. The third query vector is obtained by embedding the initial attributes.
[0133] The vehicle's infotainment system takes the second and third query vectors, which have the same dimensions, and adds their elements along each dimension to obtain the fourth query vector. Since the third query vector is obtained by embedding the initial attributes of each Gaussian sphere, the system adjusts the third query vector through summation to obtain the fourth query vector corresponding to the predicted Gaussian attributes of the Gaussian spheres.
[0134] For example, embedding refers to transforming initial attributes, such as the center position of a Gaussian sphere, rotation matrix, scaling matrix, opacity, and semantic logits, into a third query vector in a high-dimensional space through a mapping process. This process enhances the expressive power of the attributes, enabling them to participate more effectively in subsequent feature fusion and information processing.
[0135] In some embodiments of this application, the embedding can be obtained by decoding the initial Gaussian properties of multiple Gaussian spheres using a multilayer perceptron.
[0136] By superimposing the second query vector and the third query vector, the fourth query vector corresponding to the predicted Gaussian sphere is obtained. The fourth query vector is a new vector that combines the information carried by the second query vector and the third query vector, containing the information of both.
[0137] Step S146: Decode the fourth query vector to obtain the first target attributes of multiple Gaussian spheres.
[0138] The vehicle system decodes the fourth query vector, converting the high-dimensional query vector into a first target attribute that can be used to describe the features of a Gaussian sphere.
[0139] In some embodiments of this application, a multilayer perceptron can be used to decode the fourth query vector into a first target attribute of multiple Gaussian spheres. This application does not limit this to specific embodiments.
[0140] Figure 4 This is a schematic diagram of the Gaussian decoder training method provided in the embodiments of this application. Figure 4 The execution entity for each step can be an electronic device, such as an in-vehicle infotainment system, or a server.
[0141] refer to Figure 4 The Gaussian decoder training method provided in this application embodiment may include: Step S210: Obtain training sample data, which includes the first feature map corresponding to the multi-view image of the target scene and the initial query vector of multiple Gaussian spheres.
[0142] For example, the initial query vector of the multiple Gaussian spheres included in the training samples can be obtained from the predicted values of the Gaussian sphere target attributes output by the target Gaussian decoder in the previous round.
[0143] Step S220: Input the training sample data into the Gaussian decoder to be trained to obtain the prediction attributes of multiple Gaussian spheres, including the predicted positions and predicted semantics of the multiple Gaussian spheres.
[0144] In some embodiments of this application, the parameters of each model in the Gaussian decoder can be initialized randomly or based on prior knowledge, and this application does not impose any limitations on this.
[0145] Step S230: Train the Gaussian decoder to be trained based on the prediction attributes of multiple Gaussian spheres to obtain the target Gaussian decoder.
[0146] In some embodiments of this application, the training method of the Gaussian decoder can be supervised learning. By obtaining the supervision value of the target scene and the loss value between the output of the Gaussian decoder in each round, the Gaussian decoder is backpropagated to adjust the weights of each network in the Gaussian decoder to obtain the target Gaussian decoder.
[0147] For example, step S230 may also include the following sub-steps: Step S231: Obtain tag data, which includes target location information and target semantic information for multiple locations in the target scene.
[0148] For example, the label data may include the semantic occupancy or point cloud of each pixel position in a multi-view image of the target scene. The label data can be generated manually or obtained through manual annotation, and this application does not limit this.
[0149] Step S232: Calculate the first loss between the predicted positions and target positions of multiple Gaussian spheres.
[0150] In some embodiments of this application, the chamfer distance between non-empty point clouds can be used as the first loss between the predicted position and the target position of multiple Gaussian spheres. The calculation method of the first loss can be expressed by formulas (13) to (15): (13) (14) (15) Wherein, P represents the set of center positions of the Gaussian spheres predicted by the Gaussian decoder in this application, that is, the point cloud composed of the center positions of the Gaussian spheres. A point cloud is a set of label values (Ground Truth) representing the target location. This represents the center position of each predicted Gaussian sphere. To Point Cloud The sum of the chamfer distances between them. This represents the label value from each target location. The sum of chamfer distances to point cloud P.
[0151] W(d) is a reweighting function implemented by a step function, the output of which increases with the independent variable. The purpose of this function is to penalize the prediction of the Gaussian sphere furthest from the ground truth, giving it a larger chamfer distance. For example, in some embodiments of this application, if d ≥ 0.2, the step function of W(d) can be 5; otherwise, it can be 1.
[0152] The symmetrical chamfer distance is represented quantitatively as the sum of the two chamfer distances mentioned above. It is used to measure the overall difference between the predicted position and the true position of the Gaussian sphere predicted by the Gaussian decoder, and therefore can be used as a representation of the first loss value in the embodiments of this application. The smaller the first loss, the smaller the difference between the predicted position and the true position, that is, the more accurate the center position of the Gaussian sphere predicted by the Gaussian decoder.
[0153] Step S233: Calculate the second loss between the predicted semantics and the target semantic information of multiple Gaussian spheres.
[0154] For example, Figure 5 A schematic diagram of a semantic occupancy grid provided in an embodiment of this application is shown. In some embodiments of this application, the semantic occupancy grid can be used to calculate a second loss between the Gaussian sphere predicted semantics and the target semantic information.
[0155] Since the scaling matrix of a Gaussian sphere can be used to control the scaling ratio of the Gaussian sphere along the coordinate axes, this application can utilize the scaling matrix to convert the Gaussian sphere into a semantic occupancy grid comprising multiple voxels. For each Gaussian sphere, the position of each voxel and the corresponding semantics of each voxel can be formed into a tuple. Since the position of each voxel can uniquely represent a tuple, a list of Gaussian sphere voxels can be constructed using the voxel positions as indices, based on the tuples corresponding to all voxels of the Gaussian sphere.
[0156] Therefore, in the occupied grid, each voxel can correspond to one or more lists, that is, each voxel may correspond to one or more Gaussian spheres.
[0157] Since the semantics of a voxel are more affected by the semantics of the Gaussian sphere that is closer to it, the weight of the influence of each Gaussian sphere on the voxel can be obtained based on the distance between the voxel and the center of the Gaussian sphere.
[0158] Specifically, the formula for calculating the weight can be expressed by formula (16): (16) Formula (16) utilizes the Mahalanobis distance of the voxel from the center of the Gaussian sphere. The closer the voxel is to the center of the Gaussian sphere, the closer the value of this weight is to 1; otherwise, it is closer to 0. By multiplying this weight by the semantics of all the Gaussian spheres corresponding to the voxel and summing them, the final semantic occupancy representation value of this voxel can be obtained.
[0159] Based on the semantic occupancy representation value of each voxel in the semantic occupancy grid, a second loss can be obtained between the predicted semantics and the target semantic information of multiple Gaussian spheres.
[0160] For example, a distance-weighted focal loss, Dice loss, scene-class affinity loss, or Lovász-Softmax loss between semantic occupancy grids can be used as a second loss between the predicted semantics and target semantic information of multiple Gaussian spheres.
[0161] The distance weight in the distance-weighted focal loss can be expressed by formula (17): (17) Formula (17) is a focal loss function with distance weights, where BEVCenterness represents the distance weights. and The coordinates represent the center position of the Gaussian sphere. Under this weight, samples located far from the center of the Gaussian sphere will receive a greater penalty for prediction error.
[0162] The Dice loss can be expressed by formula (18): (18) in and These represent the semantics of the predicted voxels and the actual semantics, respectively. The Dess loss can alleviate class imbalance.
[0163] Scene-Class Affinity Loss can be expressed by Equation (19): (19) in, This represents the second loss. Precision can represent the proportion of voxels in the grid whose predicted and labeled values are the same to the total number of voxels, i.e., accuracy. It can represent the proportion of all voxels that truly belong to category C that are correctly identified by the model, i.e., recall. It can represent the ability of a model to identify voxels that are not in that category, that is, the proportion of voxels that are not in that category that can be correctly classified into that category.
[0164] The scenario-category affinity loss is a second loss obtained by calculating the average of the above three indicators, and it has the characteristic of strong global optimization capability.
[0165] The Lovász-Softmax loss can be expressed by equations (20) and (21): (20) (twenty one) Lovász-Softmax loss is a loss function based on IoU. IoU represents the degree of overlap between the predicted value and the label value of class c, i.e., the Intersection over Union (IoU).
[0166] in, This represents the second loss. This represents the set of voxels of category c in the label values. The set of semantics representing category c among the predicted voxels.
[0167] when The closer to 0, the higher the overlap between the prediction and the label; conversely, the lower the overlap, the lower the overlap.
[0168] Step S234: Adjust the Gaussian decoder to be trained according to the first loss and the second loss to obtain the target Gaussian decoder.
[0169] In some embodiments of this application, the backpropagation method can be used to adjust the parameters of each network in the Gaussian decoder according to the first loss and the second loss. The specific backpropagation method is not limited here.
[0170] Figure 6 A modular flowchart provided in an embodiment of this application is shown, such as Figure 6 As shown, Figure 6 It includes a graph encoder training module 101, a Gaussian decoder training module 102, a graph encoder 103, and a Gaussian decoder 104.
[0171] The Gaussian decoder training module 102 corresponds to step S230 and its sub-steps S231 to S234 provided in the embodiments of this application. The Gaussian decoder training module 102 includes the calculation of the center position of the predicted Gaussian sphere based on the label value pair and the chamfer distance between the label values, as well as the calculation of the loss value between the voxel semantic prediction value and the label value of the semantic occupancy grid.
[0172] The image encoder 103 corresponds to steps S120 and S130, which are used to obtain the corresponding first feature map based on the input multi-view image, and initialize multiple Gaussian spheres to obtain the initial Gaussian query vector corresponding to the initial attributes of each Gaussian sphere.
[0173] The Gaussian decoder 104 corresponds to step S140 and its sub-steps S141 to S146 in the embodiments of this application.
[0174] For example, the graph encoder training module 101 first obtains the depth data of the encoder, and then calculates the label value of the depth information of the target scene based on the LiDAR data. Finally, the loss value of the depth information is obtained according to formulas (22) and (23): (twenty two) (twenty three) Equations (22) and (23) describe the depth supervision loss of an image. Specifically, Equations (22) and (23) use the logarithmic loss of a single image to represent the depth supervision loss. Wherein, the label value... The depth map is calculated from the projection of the lidar point cloud. λ is the logarithmic difference between the label value and the predicted value of the features extracted by the multi-scale extraction network. λ is a weighting parameter ranging from [0,1]. It is used to balance the influence of the two terms in the loss function. When λ=0.5, it indicates the use of absolute scale error. When λ=1, it indicates the use of scale-invariant error. n represents the total number of pixels in the image, used for normalization.
[0175] The graph encoder training module 101 can improve the accuracy of the multi-scale feature extraction network based on the loss values obtained by formulas (22) and (23), and perform sparse depth supervision on the label values calculated by the depth map after projection of lidar point cloud.
[0176] In some other embodiments of this application, the Gaussian decoder training module 102 may also use the cross-entropy loss function to evaluate the semantic loss of the Gaussian decoder output. Other loss functions that may be used are not limited here.
[0177] This application also provides a computer program product, which may be a software or program product including instructions, capable of running on a computing device or stored in any usable medium. When the computer program product runs on at least one computing device, it causes the at least one computing device to implement the scene perception method and Gaussian decoder training method provided in this application.
[0178] This application also provides a computer medium. The computer medium can be any storage medium (e.g., magnetic medium, optical medium, semiconductor medium, etc.) capable of storing and / or retrieving data by a computing device. The computer medium includes instructions that direct the computing device to implement the scene perception method and Gaussian decoder training method provided in this application.
[0179] This application also provides an electronic device, including one or more processors and one or more memories. The one or more processors can be used to execute instructions to implement the scene perception method and the Gaussian decoder training method provided in this application.
[0180] For example, Figure 7 A block diagram of an electronic device 400 according to an embodiment of this application is shown. The electronic device 400 may include one or more processors 401 coupled to a controller hub 403. In at least one embodiment, the controller hub 403 communicates with the processor 401 via a multi-branch bus such as a Front Side Bus (FSB), a point-to-point interface such as a QuickPath Interconnect (QPI), or a similar connection. The processor 401 executes instructions controlling general types of data processing operations. In one embodiment, the controller hub 403 may include, but is not limited to, a Graphics & Memory Controller Hub (GMCH) (not shown) and an Input / Output Hub (IOH) (which may be on separate chips) (not shown), wherein the GMCH may include memory and a graphics controller and is coupled to the IOH.
[0181] Electronic device 400 may also include a coprocessor 402 and a memory 404 coupled to a controller hub 403. Alternatively, one or both of the memory and GMCH may be integrated within the processor (as described in the embodiments of this application), with memory 404 and coprocessor 402 directly coupled to processor 401 and controller hub 403, which is located on a single chip with IOH.
[0182] Memory 404 may be, for example, Dynamic Random Access Memory (DRAM), Phase Change Memory (PCM), or a combination of both. Memory 404 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. The computer media stores instructions, specifically, temporary and permanent copies of those instructions. The instructions may include, when executed by at least one of the processors, causing the electronic device 400 to perform, as... Figures 1 to 6 The instructions for the method are shown. When the instructions are executed on a computer, the computer performs the scene perception method and the Gaussian decoder training method disclosed in the embodiments of this application.
[0183] In one embodiment, coprocessor 402 is a dedicated processor, such as, for example, a high-throughput many-integrated core (MIC) processor, a network or communication processor, a compression engine, a graphics processor, general-purpose computing on graphics processing units (GPGPU), or an embedded processor, etc. Optional properties of coprocessor 402 are indicated by dashed lines. Figure 7 middle.
[0184] In one embodiment, electronic device 400 may further include a Network Interface Controller (NIC) 406. The network interface 406 may include a transceiver for providing a radio interface for electronic device 400 to communicate with any other suitable device, such as a front-end module, antenna, etc. In various embodiments, the network interface 406 may be integrated with other components of electronic device 400. The network interface 406 can implement the functions of the communication unit in the above embodiments.
[0185] Electronic device 400 may further include input / output (I / O) devices 405. I / O 405 may include: a user interface designed to enable a user to interact with electronic device 400; a peripheral component interface designed to enable peripheral components to also interact with electronic device 400; and / or sensors designed to determine environmental conditions and / or location information related to electronic device 400.
[0186] It is worth noting that, Figure 7 This is merely an example. That is, although Figure 7The illustration shows an electronic device 400 including multiple devices such as a processor 401, a controller hub 403, and a memory 404. However, in actual applications, devices using the methods of the embodiments of this application may include only a portion of the devices in the electronic device 400. For example, it may include only the processor 401 and the network interface 406. Figure 7 The properties of the optional devices are shown by dashed lines.
[0187] Figure 8 A block diagram of a System-on-Chip (SoC) 500 applied to an electronic device according to an embodiment of this application is shown. Figure 8 In the diagram, similar components share the same reference numerals. Additionally, dashed boxes are an optional feature for more advanced SoCs. Figure 8 In this embodiment, SoC 500 includes: an interconnect unit 550 coupled to processor 510; a system proxy unit 580; a bus controller unit 590; an integrated memory controller unit 540; a group or one or more coprocessors 520, which may include integrated graphics logic, an image processor, an audio processor, and a video processor; a static random-access memory (SRAM) unit 530; and a direct memory access (DMA) unit 560. In one embodiment, coprocessor 520 includes a dedicated processor, such as, for example, a network or communication processor, a compression engine, general-purpose computing on graphics processing units (GPGPU), a high-throughput MIC processor, or an embedded processor.
[0188] Static Random Access Memory (SRAM) cell 530 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. The computer media stores instructions, specifically, temporary and permanent copies of those instructions. These instructions may include instructions that, when executed by at least one processor, cause the SoC to implement the methods shown in the embodiments of this application. When the instructions are executed on a computer, they cause the computer to perform the scene-aware method and the Gaussian decoder training method disclosed in the embodiments of this application.
[0189] It should be noted that the terminology used in the embodiment section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. In the description of the embodiments of this application, unless otherwise stated, " / " means "or", for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between associated obstacles, indicating that three relationships can exist, for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. In addition, in the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more, "at least one" or "one or more" means one, two or more.
[0190] It should be noted that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0191] References to "one embodiment" or "some embodiments" as used in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0192] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially as a computer program product. The computer program product may include one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer medium or transmitted through the computer medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer medium may be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium, etc.
[0193] Those skilled in the art will understand that implementing all or part of the processes in the above embodiments can be accomplished by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The aforementioned storage medium can include various media capable of storing program code, such as read-only memory, random access memory, magnetic disks, or optical disks.
[0194] The above description is merely a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be determined by the protection scope of the claims.
Claims
1. A scene perception method applied to electronic devices, characterized in that, The method includes: Acquire multi-view images of the target scene; Feature extraction is performed on the multi-view images to obtain a first feature map corresponding to the multi-view images; Initialize multiple Gaussian spheres and determine the initial query vector of the multiple Gaussian spheres based on their initial properties; The first feature map and the plurality of initial query vectors are input into the target Gaussian decoder to predict the first target attribute of the plurality of Gaussian spheres corresponding to the first feature map. The first target attribute includes the first predicted position and the first predicted semantic of the Gaussian sphere. Based on the first target attributes of the plurality of Gaussian spheres, a perception result of the target scene is obtained, the perception result including semantic information of the spatial location of the target scene determined based on the first predicted position and the first predicted semantics.
2. The method according to claim 1, characterized in that, The step of inputting the first feature map and the plurality of initial query vectors into a target Gaussian decoder to predict the first target attribute of the plurality of Gaussian spheres corresponding to the first feature map by the target Gaussian decoder includes: Multiple sampling points are generated in each Gaussian sphere according to the initial query vector, and the first feature map is sampled according to the multiple sampling points to obtain the second feature map; The second feature map and the initial query vectors of the plurality of Gaussian spheres are fused to obtain the third feature map; The initial query vector is updated according to the adaptive attention mechanism to obtain the first query vector; The third feature map and the first query vector are subjected to a first process to obtain a second query vector, wherein the first process includes at least linear processing; The second query vector and the third query vector are superimposed to obtain the fourth query vector, wherein the third query vector is obtained by embedding the initial attribute; Decoding the fourth query vector yields the first target attribute of the plurality of Gaussian spheres.
3. The method according to claim 2, characterized in that, The number of sampling points is determined based on the computing resources of the electronic device.
4. The method according to claim 2, characterized in that, The step of fusing features from the second feature map and the initial query vectors of the plurality of Gaussian spheres to obtain the third feature map includes: Adaptive feature fusion is performed on the second feature map to obtain the fourth feature map; The fourth feature map is residually linked with the initial query vectors of the multiple Gaussian spheres to obtain the third feature map.
5. The method according to claim 4, characterized in that, The adaptive feature fusion of the second feature map includes: Adaptive fusion of the second feature map is performed based on the feature channels and / or spatial location of the second feature map.
6. The method according to claim 4, characterized in that, The first feature map is a multi-layer feature map; The fourth feature map is residually linked with the initial query vectors of the multiple Gaussian spheres to obtain the third feature map, which includes: The fourth feature map is residually linked with the initial query vectors of the plurality of Gaussian spheres to obtain a fifth feature map. The fifth feature map is a multi-layer feature map, and the number of layers of the fifth feature map is the same as the number of layers of the first feature map. The fifth feature map is then subjected to layer normalization to obtain the third feature map.
7. The method according to claim 2, characterized in that, The step of updating the initial query vector according to the adaptive attention mechanism to obtain the first query vector includes: Multiple initial query vectors are input into a linear layer to obtain receptive field weights; The attention weight is obtained based on the pairwise distance between the receptive field weight and the initial query vector, and the attention weight decreases as the receptive field weight increases. The initial query vector is updated according to the attention weight to obtain the first query vector.
8. The method according to claim 2, characterized in that, The first process also includes an activation process following the linear process.
9. The method according to claim 1, characterized in that, The step of obtaining the perception result of the target scene based on the first target attributes of the plurality of Gaussian spheres includes: Based on the first target attribute of the plurality of Gaussian spheres, determine the fifth query vector of the plurality of Gaussian spheres; The fifth query vector and the first feature map are input into the target Gaussian decoder to obtain the second target attribute of the plurality of Gaussian spheres; The perception result is obtained based on the second target properties of the plurality of Gaussian spheres.
10. A training method for a Gaussian decoder, applied to electronic devices, characterized in that, The method includes: Acquire training sample data, which includes a first feature map corresponding to a multi-view image of the target scene, and initial query vectors of multiple Gaussian spheres; The training sample data is input into the Gaussian decoder to be trained to obtain the predicted attributes of the plurality of Gaussian spheres, the predicted attributes including the predicted positions and predicted semantics of the plurality of Gaussian spheres; The Gaussian decoder to be trained is trained based on the predicted properties of the multiple Gaussian spheres to obtain the target Gaussian decoder.
11. The method according to claim 10, characterized in that, The step of inputting the training sample data into the Gaussian decoder to be trained to obtain the predicted properties of the multiple Gaussian spheres includes: Multiple sampling points are generated in each Gaussian sphere according to the initial query vector, and the first feature map is sampled according to the multiple sampling points to obtain the second feature map; The second feature map and the initial query vectors of the plurality of Gaussian spheres are fused to obtain the third feature map; The initial query vector is updated according to the adaptive attention mechanism to obtain the first query vector; The third feature map and the first query vector are subjected to a first process to obtain a second query vector, wherein the first process includes at least linear processing; The second query vector and the third query vector are superimposed to obtain the fourth query vector, wherein the third query vector is obtained by embedding the initial properties of the Gaussian sphere; Decode the fourth query vector to obtain the predicted attributes of the multiple Gaussian spheres.
12. The method according to claim 10, characterized in that, The step of training the Gaussian decoder to be trained based on the predicted attributes of the plurality of Gaussian spheres to obtain the target Gaussian decoder includes: Acquire tag data, which includes target location information and target semantic information for multiple locations in the target scene; Calculate the first loss between the predicted positions of the plurality of Gaussian spheres and the target position information; Calculate a second loss between the predicted semantics of the plurality of Gaussian spheres and the target semantic information; The Gaussian decoder to be trained is adjusted based on the first loss and the second loss to obtain the target Gaussian decoder.
13. A computer medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the scene perception method as described in any one of claims 1 to 9 or the training method of the Gaussian decoder as described in any one of claims 10 to 12.
14. A computer program product, characterized in that, When the computer program product is run on an electronic device, it enables the electronic device to implement the scene perception method of any one of claims 1 to 9 or the training method of the Gaussian decoder of any one of claims 10 to 12.
15. An electronic device, characterized in that, include: Memory, used to store one or more programs; A processor for executing one or more programs to cause the electronic device to implement the scene perception method of any one of claims 1 to 9 or the training method of the Gaussian decoder of any one of claims 10 to 12.
Citation Information
Patent Citations
Law and regulation simulation test method, device and equipment for intelligent driving algorithm and storage medium
CN119783609A
Target detection method and device, computer equipment and vehicle
CN120014232A
3D target detection method for automatic driving sensor fault
CN120510419A
Scene perception model training method and device, robot control method and robot
CN120635678A
Gaussian coding model training method, indoor occupancy prediction method, equipment and medium
CN120635680A
Cited By
Digital asset insertion method, medium, program product and electronic device
CN121527371A