Target perception method and electronic device

By performing feature querying and correcting sampling positions on target point cloud data, the problem of reduced target perception accuracy in poor road conditions is solved, and high-precision target perception is achieved in harsh environments.

CN122134822APending Publication Date: 2026-06-02EACON TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
EACON TECHNOLOGY CO LTD
Filing Date
2026-02-09
Publication Date
2026-06-02

Smart Images

  • Figure CN122134822A_ABST
    Figure CN122134822A_ABST
Patent Text Reader

Abstract

This application provides a target perception method and electronic device, relating to the field of autonomous driving technology. The method includes: performing feature querying on target point cloud data to obtain point cloud query features; predicting 3D anchor point information and 2D sampling offset information based on the point cloud query features; correcting the offset of the sampling positions corresponding to the 3D anchor point information on the target image data based on the 2D sampling offset information to obtain the target sampling position, where the target image data is the image data corresponding to the target point cloud data; performing feature sampling on the target image data based on the target sampling position to obtain image sampling features; fusing the image sampling features and point cloud query features across modalities to obtain fused query features; and predicting the 3D target perception result based on the fused query features. For scenarios such as autonomous driving, automatic driving, and driverless vehicles, the embodiments of this application can improve the accuracy of target perception under poor road conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, specifically to a target perception method and electronic device. Background Technology

[0002] With the development of autonomous driving technology, more and more autonomous driving scenarios require the use of multi-sensor calibration technology to achieve automatic perception of targets.

[0003] In traditional multi-sensor calibration methods, either manual feature association is required, or the feature matching rules are not adapted to large-angle external parameter changes, which may lead to a decrease in the accuracy of target perception in scenarios with poor road conditions. Summary of the Invention

[0004] In view of this, the embodiments of this application aim to provide a target perception method and electronic device to solve the problem that the accuracy of target perception will be reduced in scenarios with poor road conditions in the prior art.

[0005] In a first aspect, one embodiment of this application provides a target perception method, comprising: performing feature query on target point cloud data to obtain point cloud query features; predicting three-dimensional anchor point information and two-dimensional sampling offset information based on the point cloud query features; performing offset correction on the sampling position corresponding to the three-dimensional anchor point information on target image data based on the two-dimensional sampling offset information to obtain a target sampling position, wherein the target image data is image data corresponding to the target point cloud data; performing feature sampling on the target image data based on the target sampling position to obtain image sampling features; performing cross-modal feature fusion on the image sampling features and the point cloud query features to obtain fused query features; and predicting a three-dimensional target perception result based on the fused query features.

[0006] In conjunction with the first aspect, in some implementations of the first aspect, after performing cross-modal feature fusion on the image sampling features and the point cloud query features to obtain fused query features, the method further includes: predicting target calibration parameters based on the fused query features, wherein the target calibration parameters include extrinsic parameters corresponding to the target image acquisition device, and the target image acquisition device is a device for acquiring the target image data.

[0007] In conjunction with the first aspect, in some implementations of the first aspect, the step of offsetting and correcting the sampling position of the three-dimensional anchor point information on the target image data according to the two-dimensional sampling offset information to obtain the target sampling position includes: acquiring image features corresponding to the target image data; converting the three-dimensional anchor point information to two dimensions according to a first calibration parameter to obtain two-dimensional reference point information; and offsetting and correcting the sampling position of the two-dimensional reference point information on the image features according to the two-dimensional sampling offset information to obtain the target sampling position. Correspondingly, the step of performing feature sampling on the target image data based on the target sampling position to obtain image sampling features includes: sampling on the image features based on the target sampling position to obtain image sampling features.

[0008] In conjunction with the first aspect, in some implementations of the first aspect, after predicting the target calibration parameters based on the fusion query features, the method further includes: updating the first calibration parameters based on the target calibration parameters.

[0009] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: obtaining the true value of the calibration parameter; determining the target loss value based on the difference between the target calibration parameter and the true value of the calibration parameter; and training the process of predicting the two-dimensional sampling offset information and the target calibration parameter based on the target loss value.

[0010] In conjunction with the first aspect, in some implementations of the first aspect, the step of converting the three-dimensional anchor point information to two dimensions according to the first calibration parameters to obtain two-dimensional reference point information includes: converting the three-dimensional anchor point information to two dimensions according to the first calibration parameters to obtain initial two-dimensional reference point information; determining target two-dimensional reference point information that meets the preset invisible condition from the initial two-dimensional reference point information; and performing masking processing on the target two-dimensional reference point information in the initial two-dimensional reference point information to obtain the two-dimensional reference point information.

[0011] In conjunction with the first aspect, in some implementations of the first aspect, the step of performing cross-modal feature fusion on the image sampling features and the point cloud query features to obtain fused query features includes: performing self-attention learning on a first target feature based on a multi-head self-attention mechanism to obtain a first feature, wherein the first target feature includes the point cloud query features; and performing cross-attention learning on the image sampling features and the first feature based on a multi-head cross-attention mechanism to obtain the fused query features.

[0012] In conjunction with the first aspect, in some implementations of the first aspect, the step of performing cross-attention learning on the image sampling features and the first feature based on a multi-head cross-attention mechanism to obtain the fused query feature includes: performing cross-attention learning on the image sampling features and the first feature based on a multi-head cross-attention mechanism to obtain a second feature; using the fourth feature as the first target feature, returning to perform self-attention learning on the first target feature based on a multi-head self-attention mechanism, until iterating a first preset number of times to obtain the fused query feature.

[0013] In conjunction with the first aspect, in some implementations of the first aspect, the step of performing feature query on the target point cloud data to obtain point cloud query features includes: acquiring point cloud features corresponding to the target point cloud data; determining multiple initial anchor points corresponding to the target point cloud data; and performing attention learning based on the point cloud features and the multiple initial anchor points to obtain the point cloud query features.

[0014] In conjunction with the first aspect, in some implementations of the first aspect, after performing feature query on the target point cloud data to obtain point cloud query features, the method further includes: predicting an initial three-dimensional target perception result based on the point cloud query features.

[0015] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: obtaining the ground truth value of a three-dimensional target; determining a first loss value based on a target loss function, according to the difference between the initial three-dimensional target perception result and the ground truth value of the three-dimensional target, and determining a second loss value based on the difference between the three-dimensional target perception result and the ground truth value of the three-dimensional target, wherein the target loss function includes at least one of a classification loss function, a regression loss function, and an intersection-union loss function; training the learning and prediction process of the initial three-dimensional target perception result based on the first loss value, and training the learning and prediction process of the three-dimensional target perception result based on the second loss value.

[0016] Secondly, one embodiment of this application provides a target perception device, comprising: a feature query module for performing feature query on target point cloud data to obtain point cloud query features; a first prediction module for predicting three-dimensional anchor point information and two-dimensional sampling offset information based on the point cloud query features; an offset correction module for correcting the offset of the sampling position corresponding to the three-dimensional anchor point information on target image data based on the two-dimensional sampling offset information to obtain a target sampling position, wherein the target image data is image data corresponding to the target point cloud data; a feature sampling module for performing feature sampling on the target image data based on the target sampling position to obtain image sampling features; a feature fusion module for performing cross-modal feature fusion on the image sampling features and the point cloud query features to obtain fused query features; and a second prediction module for predicting a three-dimensional target perception result based on the fused query features.

[0017] Thirdly, one embodiment of this application provides a computer-readable storage medium storing a computer program for performing the target perception method described in the first aspect.

[0018] Fourthly, one embodiment of this application provides an electronic device, the electronic device comprising: a processor; a memory for storing processor-executable instructions; the processor being configured to execute the target perception method described in the first aspect.

[0019] Fifthly, one embodiment of this application provides a computer program product including instructions that, when executed on an electronic device, cause the electronic device to implement the target perception method described in the first aspect.

[0020] In this application, by predicting 3D anchor point information and 2D sampling offset information based on target point cloud data, the 2D sampling offset information can be used in a timely manner to correct the sampling position when performing 2D projection sampling of target image data using 3D anchor point information. This compensates for calibration errors and projection distortion, resulting in an accurate target sampling position. Based on this target sampling position, feature sampling can be performed accurately, achieving accurate cross-modal feature fusion and ultimately obtaining accurate perception results. Since the sampling position offset correction process in this application can compensate for calibration errors and projection distortion, improving the robustness of target perception and effectively mitigating the problem of inaccurate projection association caused by calibration errors, this application can improve the accuracy of target perception in scenarios with poor road conditions. Attached Figure Description

[0021] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0022] Figure 1 The diagram shown is a flowchart of a target perception method provided in an embodiment of this application.

[0023] Figure 2 The diagram shown is a schematic flowchart of a point cloud data processing branch provided in an embodiment of this application.

[0024] Figure 3 The diagram shown is a schematic flowchart of a feature fusion branch provided in an embodiment of this application.

[0025] Figure 4 The image shown is a rendering of the initial three-dimensional target perception result in one embodiment of this application.

[0026] Figure 5 The diagram shown is a schematic diagram of the position offset correction process in one embodiment of this application.

[0027] Figure 6 The diagram shown is a schematic diagram of the target sensing device provided in an embodiment of this application.

[0028] Figure 7 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0030] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed embodiments. Those skilled in the art should understand that this application can be implemented even without certain specific details. In some instances, methods and means well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.

[0031] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0032] Furthermore, the terms “first,” “second,” “third,” and “fourth” are used only for distinguishing descriptions and should not be interpreted as indicating or implying relative importance.

[0033] In autonomous driving scenarios, multi-sensor calibration technology is frequently used to achieve automatic perception of target objects. Current multi-sensor calibration schemes mainly fall into two categories: one involves manually designing feature extraction methods for multi-sensor data and then associating features between different sensor data to achieve target calibration and perception. This method has poor portability, is highly dependent on the scene, and has poor adaptability in many scenarios. Furthermore, feature association requires constructing similarity calculation equations, which also fail in many scenarios. To improve scene adaptability, a second scheme emerged: model-based calibration. This scheme extracts features from different sensor data using a model and then performs matching through simple rule sampling to achieve target calibration and perception. However, the second scheme cannot adaptively select and associate features, leading to the need for brute-force calculation of all features during feature matching. In addition, in scenarios with poor road conditions, the significant impact on the vehicle can cause large-angle changes in extrinsic parameters, which rule-based matching is not suitable for, resulting in inaccurate target perception predictions.

[0034] To address the aforementioned issues, this application provides a target perception method and electronic device. The method includes: performing feature querying on target point cloud data to obtain point cloud query features; predicting 3D anchor point information and 2D sampling offset information based on the point cloud query features; correcting the offset of the sampling position corresponding to the 3D anchor point information on the target image data based on the 2D sampling offset information to obtain the target sampling position, where the target image data is the image data corresponding to the target point cloud data; performing feature sampling on the target image data based on the target sampling position to obtain image sampling features; performing cross-modal feature fusion on the image sampling features and the point cloud query features to obtain fused query features; and predicting a 3D target perception result based on the fused query features. Thus, by predicting 3D anchor point information and 2D sampling offset information based on the target point cloud data, the 2D sampling offset information can be used in a timely manner to correct the sampling position when performing 2D projection sampling on the target image data using the 3D anchor point information, compensating for calibration errors and projection distortion, obtaining an accurate target sampling position, and then accurately performing feature sampling based on this target sampling position, achieving accurate cross-modal feature fusion, and ultimately obtaining an accurate perception result. Since the sampling position offset correction process in this application can compensate for calibration errors and projection distortion, improve the robustness of target perception, and effectively alleviate the problem of inaccurate projection association caused by calibration errors, this application can improve the accuracy of target perception in scenarios with poor road conditions.

[0035] The following is combined with Figures 1 to 5 The target perception method provided in this application is described in detail.

[0036] Figure 1 The diagram shown is a schematic flowchart of a target perception method provided in an embodiment of this application. This method can be applied to electronic devices, such as computers, vehicle-mounted systems, and other devices with data processing capabilities. Figure 1 As shown, the method may include the following steps.

[0037] S110, perform feature query on the target point cloud data to obtain the point cloud query features.

[0038] In some examples, the target point cloud data can be any radar data, such as a set of point cloud data collected by a vehicle-mounted LiDAR. The point cloud query features can be the data obtained after adaptive feature querying of the target point cloud data. Specifically, a transformer-based codec can be used to implement feature querying of the target point cloud data.

[0039] For example, after acquiring the target point cloud data collected by radar, the target point cloud data can be input into an encoder based on the transformer architecture for feature encoding. Then, the encoded point cloud features and preset 3D anchor points can be input into a decoder based on the transformer architecture for attention learning, thereby completing the feature query process of the target point cloud data and obtaining the point cloud query features.

[0040] In addition, in some embodiments, step S110 may specifically include: acquiring point cloud features corresponding to the target point cloud data; determining multiple initial anchor points corresponding to the target point cloud data; and performing attention learning based on the point cloud features and the multiple initial anchor points to obtain point cloud query features.

[0041] In some examples, the specific process of obtaining point cloud features corresponding to the target point cloud data may include: converting the disordered target point cloud data into regular grid structure data (such as voxelization or columnarization); obtaining initial point cloud features based on the grid structure data through multi-layer feature extraction; performing position encoding on the initial point cloud features to obtain point cloud position encoding; and performing feature fusion between the initial point cloud features and the point cloud position encoding to obtain point cloud features.

[0042] For example, target point cloud data (C1,3) is obtained, where C1 typically represents more than 100,000 points, and 3 represents the initial feature dimension of each point in the point cloud (e.g., it may include three dimensions: x-coordinate, y-coordinate, and z-coordinate). The unordered target point cloud data (C1,3) is voxelized or columnarized to obtain regular data with shapes (W, H, C).

[0043] Multi-level feature extraction is performed on the rule data, and the final output is an initial point cloud feature with shape (W / M1, H / M2, C / M2, N2), which is expanded to (C2, ), where C2=W / M1 H / M2 C / M2, For feature dimensions.

[0044] For the C2 features in the initial point cloud The feature of the dimension is mapped one-to-one with the grid position in the physical world coordinate system. The grid position (u,v,k) of the feature is transformed to the radar coordinate system (x,y,z) through indexing to obtain the corresponding position information.

[0045] Sine coding is performed on the location information to generate The position vectors of the two dimensions are used to obtain the corresponding point cloud position codes. Specifically, each position coordinate (x, y, z) can be independently encoded according to the following formula (1) to generate vectors with alternating even and odd dimensions.

[0046] Where p represents the coordinate component value (e.g., x=1.2), and i represents the dimension index ( ), This represents the total dimension after encoding.

[0047] Additionally, if the dimension of the point cloud location encoding Dimensions of the initial point cloud features If they are not equal, in order to enhance the learning ability of subsequent anchor points, the feature dimension of the position encoding can be reduced from the following formula (2) through linear projection: dimensional transformation to The dimensions are calculated, and effective features are extracted and aligned with the point cloud features.

[0048] Where W represents the learnable weight matrix, and b represents the bias vector (e.g., ...). ), z represents the projected feature ( dimension).

[0049] The point cloud position code z aligned with the feature dimensions is compared with the initial point cloud features. By adding them together, we can obtain point cloud features with location information. Under the transformer architecture, these point cloud features can be used as the key vector K (Key) and value vector V (Value) corresponding to the target point cloud data.

[0050] In some examples, multiple initial anchor points corresponding to the target point cloud data can be determined according to a preset anchor point initialization method, and these initial anchor points are used as query vectors Q (Query) in the Transformer architecture. These initial anchor points are learnable and can guide the final selection of anchor points through model training. The preset anchor point initialization method can include at least one of the following: fixed-point initialization, heatmap initialization, and ground truth bounding box initialization.

[0051] The aforementioned fixed-point initialization can specifically involve selecting a fixed number of representative points from the target point cloud data using the farthest-point sampling method. For example, a fixed number of n (e.g., 200) representative points can be selected from the target point cloud data (e.g., 100,000 points) using farthest-point sampling as initial anchor points. These points can uniformly cover the point cloud space corresponding to the target point cloud data and carry geometric structure information (e.g., edges, corners).

[0052] The above heatmap initialization can specifically be as follows: select high-response location points (such as the top few points with the highest heat) from the heatmap corresponding to the bird's-eye view features of the target point cloud data as initial anchor points. The initial anchor points selected by this initialization method are easier to converge, which can reduce the number of layers of the subsequent decoder (i.e., the number of loops of attention learning).

[0053] The above-mentioned initialization of the ground truth (GT) box can be specifically achieved by adding noise to the ground truth (GT) box to generate initial anchor points.

[0054] It should be noted that during model training, any of the above initialization methods can be chosen as the method for obtaining the initial anchor point. During model inference, it is preferable to obtain the initial anchor point through heatmap initialization.

[0055] In addition, in some examples, the above-mentioned attention learning based on point cloud features and multiple initial anchor points to obtain point cloud query features may specifically include: performing position encoding on multiple initial anchor points to obtain anchor point position encoding; performing self-attention learning on anchor point position encoding based on a multi-head self-attention mechanism to obtain a third feature; and performing cross-attention learning on point cloud features and the third feature based on a multi-head cross-attention mechanism to obtain point cloud query features.

[0056] For example, the location information corresponding to multiple initial anchor points can be encoded into a preset artificial intelligence model. A dimensional vector is used as the anchor point position encoding. This pre-defined artificial intelligence model can be, for example, a multilayer perceptron (MLP). Specifically, the position information corresponding to multiple initial anchor points can be sinusoidally encoded (refer to formula (1)) to obtain... After encoding the anchor point positions of the dimension, a linear projection is then performed on the encoded anchor point positions (refer to formula (2)) to obtain Anchor point position encoding in dimension. If there are n initial anchor points, then the shape obtained is (n, Anchor point location encoding.

[0057] After obtaining the anchor position encoding, this anchor position encoding can be used as the input sequence during self-attention learning. Furthermore, based on a multi-head self-attention mechanism, self-attention learning is performed on the anchor point position encoding to obtain the third feature, Queries_self. Each head can represent a learning objective.

[0058] Specifically, the input sequence can be processed according to the following formula (3). Perform linear projection to generate Q, K, and V.

[0059] in, ,Should The number of heads.

[0060] Decompose Q, K, and V by their first digits. For example, decompose Q into... ,in Similarly, K is split into Split V into .

[0061] Based on the scaling dot product attention mechanism, the attention output features corresponding to each head are obtained according to the following formula (4).

[0062] in, Indicates the first The attention output features corresponding to the head. Indicates the first Attention score corresponding to the head, This indicates scaling of the attention score. Indicates the first Attention weights corresponding to the head.

[0063] The multi-head attention output features are concatenated according to the following formula (5) to obtain the third feature.

[0064] in, , Indicates the first Attention output features corresponding to the head, This represents the third feature, Queries_self.

[0065] When performing cross-attention learning on point cloud features and a third feature based on a multi-head cross-attention mechanism, the third feature can be used as the query vector Q in the input sequence, and the point cloud features can be used as the key vector K and value vector V in the input sequence. Based on this, for each head... The projection can be calculated according to the following formula (6).

[0066] Based on the above projection results, each head is processed according to the above formula (4). Intra-head attention calculation is performed. Finally, the multi-head attention calculation results are concatenated according to the above formula (5) to obtain the point cloud query features.

[0067] In addition, in some implementations, to enhance the feature representation capability of point cloud query features, a feed-forward network (FFN) can be used to add nonlinearity to the point cloud query features. The essence of FFN is a two-layer fully connected network (such as MLP), which can be expressed as the following formula (7).

[0068] Among them, input features (like =256), , Represents the learnable weight matrix. , This represents the bias vector. The process can be described as increasing the dimensionality of the input features through the first layer of an FFN. ,For example Add ReLU activation, then perform a second layer of dimensionality reduction. .

[0069] It should be noted that the above self-attention and cross-attention learning processes can undergo multiple iterations. That is, the result obtained after cross-attention learning is used as the input sequence for self-attention learning in the next iteration, and the calculation is performed in the next iteration. After multiple iterations (e.g., the number of iterations reaches a second preset number), the final output can be obtained. This output can be used as the final determined point cloud query feature lidar_queries, with a shape of (n, ).

[0070] S120 predicts 3D anchor point information and 2D sampling offset information based on point cloud query features.

[0071] In some examples, 3D anchor point information may include the location information of the most characteristic and representative 3D anchor points in the target point cloud data. 2D sampling offset information may include offset information that may occur when mapping 3D anchor points to 2D.

[0072] For example, after obtaining the point cloud query features, these features can be input into a trained target neural network. This target neural network performs dimensionality upscaling, dimensionality reduction, adds nonlinearity, and enhances the features, ultimately outputting 3D anchor point information and 2D sampling offset information. Different target neural networks can be set for different prediction results; for example, the target neural network could be an FFN.

[0073] For example, a dedicated FFN can be set up for 3D anchor point prediction. Specifically, the point cloud query features are input into this FFN, and the first layer of the FFN processes the point cloud query features from... Increased to multiple times Dimension, and then through the second layer of this FFN from multiple times Dimensional Descending to Dimensional Attainment .in, For example, it can be 3-dimensional, corresponding to x, y, and z three-dimensional coordinates. A shape of (n, ... The point cloud query features can be used to predict the position information of n three-dimensional anchor points.

[0074] For example, for two-dimensional sampling offset information, another FFN can be set up for two-dimensional offset prediction. Specifically, the point cloud query features are input into this FFN, and the first layer of the FFN processes the point cloud query features from... Increased to multiple times Dimension, and then through the second layer of this FFN from multiple times Dimensional Descending to Dimensional Attainment ,in, For example, it could be (M) 2) Dimension, M corresponds to the number of two-dimensional sampling points, and 2 corresponds to the xy two-dimensional coordinates. The number of two-dimensional sampling points M can be determined according to the following formula (8).

[0075] in, This indicates the number of heads attracting attention from the bulls. This represents the number of scales of the image features corresponding to the target image data. This represents the number of points sampled on the feature map corresponding to each scale. This target image data can be image data from the target point cloud data to be used for target perception and calibration.

[0076] In a specific example, if the number of attention heads is preset to 8, the target image data corresponds to image features with 4 scales, and each scale requires sampling 4 points on the image feature, then for a point cloud query feature... 3D features can predict the 2D coordinate offset corresponding to 128 2D sampling points. That is, for every 3D anchor point predicted, 128 2D coordinate offsets can be predicted.

[0077] In addition, in some embodiments, after step S110 above, the target perception method provided in this application embodiment may further include: predicting the initial three-dimensional target perception result based on point cloud query features.

[0078] In some examples, the initial 3D target perception result can be the result obtained by perceiving the target point cloud data based on its own point cloud features. Target perception can be the perception of target objects existing in the actual scene. These target objects can be living beings such as people and animals, or inanimate objects such as obstacles, vehicles, and road signs; there is no limitation on this. The initial 3D target perception result may include the 3D boundary positions, center point positions, and pose information of one or more target objects.

[0079] For example, based on the initial 3D target perception results, an FFN can also be set up for target perception prediction. Specifically, the point cloud query features are input into the FFN, and the point cloud query features are processed from the first layer of the FFN. Increased to multiple times Dimension, and then through the second layer of this FFN from multiple times Dimensional Descending to Dimensional Attainment .in, For example, it could be 7-dimensional, corresponding to the xyz three-dimensional coordinates at the center point, the whl length, width, and height at the three-dimensional boundary, and the yaw angle in the attitude information. A shape of (n, The number of target objects that can be predicted based on the point cloud query features is not limited here.

[0080] S130, based on the two-dimensional sampling offset information, the sampling position corresponding to the three-dimensional anchor point information on the target image data is offset and corrected to obtain the target sampling position. The target image data is the image data corresponding to the target point cloud data.

[0081] In some examples, two-dimensional sampling offset information can be used to correct the sampling positions of the three-dimensional anchor points on the target image data. The target image data can be image data that will be used for target perception and calibration using target point cloud data.

[0082] For example, in scenarios with poor road conditions, the large impact force on the vehicle body can lead to large-angle changes in extrinsic parameters, resulting in inaccurate calibration parameters. In such cases, projecting the 3D anchor point onto the 2D plane using inaccurate calibration parameters to determine the sampling position will lead to inaccurate estimation of the sampling position. Based on this, the embodiments of this application can compensate for the sampling position offset caused by inaccurate calibration parameters by utilizing 2D sampling offset information, thereby obtaining a more accurate target sampling position. Based on this target sampling position, image feature sampling can be performed to obtain accurate image sampling features, thereby improving the accuracy of subsequent target perception.

[0083] In some implementations, step S130 may specifically include: acquiring image features corresponding to the target image data; converting the three-dimensional anchor point information to two dimensions according to the first calibration parameters to obtain two-dimensional reference point information; and correcting the sampling position of the two-dimensional reference point information on the image features according to the two-dimensional sampling offset information to obtain the target sampling position.

[0084] In some examples, the above-mentioned acquisition of image features corresponding to the target image data may specifically include: extracting features from the target image data to obtain feature maps at multiple scales; performing positional encoding on the feature maps at multiple scales to obtain image positional encoding; and fusing the feature maps at multiple scales with the image positional encoding to obtain image features.

[0085] For example, target image data corresponding to the target point cloud data is acquired. Features are extracted from the target image data at different scales according to multiple preset scales to obtain feature maps at multiple scales. Each feature pixel in the feature map has a corresponding coordinate position (u, v). This coordinate position is sinusoidally encoded and its dimension is transformed to... The image location code is obtained. Adding the image location code, aligned to the feature dimensions, to feature maps of multiple scales yields image features with location information. In the transformer architecture, these image features can serve as the key vector K and value vector V corresponding to the target image data. The above image feature extraction process is similar to the point cloud feature extraction process and will not be elaborated upon here.

[0086] In some examples, the first calibration parameter can be a parameter used to transform point cloud data to the coordinate system corresponding to the target image acquisition device. Specifically, this first calibration parameter can include extrinsic parameters corresponding to the target image acquisition device, such as an extrinsic parameter matrix. The target image acquisition device can be a device used to acquire the target image data, such as a camera. Furthermore, the first calibration parameter can be a preset, fixed parameter, or it can be a trainable or updatable parameter.

[0087] For example, the three-dimensional anchor point information with shape (n,3), that is, the set of three-dimensional anchor points P, is transformed from three-dimensional to two-dimensional according to the following formula (9) to obtain the position information after projection as two-dimensional reference points, that is, two-dimensional reference point information.

[0088] in, It is an extrinsic parameter matrix. Here, is the normalized camera coordinates, is the distort function for camera distortion compensation, and K is the camera intrinsic parameter. The 3D anchor point pi in the target point cloud data can first be mapped to the camera coordinate system using the extrinsic parameter matrix, and then mapped to the image coordinate system using the camera intrinsic parameters to obtain the position information of the 2D reference point. ), which is the sampling location.

[0089] In some examples, this can be based on two-dimensional sampling offset information (such as two-dimensional coordinate offset). The sampling position of the two-dimensional reference point on the image features. After offset correction, the target sampling position can be obtained. As can be seen from the previous example, for each predicted 3D anchor point, M corresponding 2D coordinate offsets can be predicted. Therefore, in this example, for one 3D anchor point, after offset correction, M target sampling positions can be obtained.

[0090] In this way, by correcting the offset of the sampling position, the target perception capability can be effectively improved. Thus, the correlation between point cloud features and image features can be established without the need for accurate first calibration parameters. This can serve as the basis for accurately perceiving the target object and also provide an important support for the subsequent accurate prediction of target calibration parameters.

[0091] Furthermore, to improve the projection effectiveness when converting from 3D to 2D, visibility of the 2D reference points can be determined. Based on this, in some embodiments, the process of converting the 3D anchor point information to 2D according to the first calibration parameters to obtain 2D reference point information can specifically include: converting the 3D anchor point information to 2D according to the first calibration parameters to obtain initial 2D reference point information; determining target 2D reference point information that meets a preset invisible condition from the initial 2D reference point information; and performing masking processing on the target 2D reference point information in the initial 2D reference point information to obtain the 2D reference point information.

[0092] For example, for each initial two-dimensional reference point information (such as the position coordinates after projecting the three-dimensional anchor point into a two-dimensional reference point), If a preset invisibility condition is met, it can be identified as an invisible target two-dimensional reference point. The preset invisibility condition may include at least one of the following: or (i.e., the image width of the target image data); or (i.e., the image height of the target image data); depth value in the camera coordinate system. If satisfied This indicates that the two-dimensional reference point corresponding to the three-dimensional anchor point is located behind the camera, so theoretically it cannot be seen in the image and can therefore be considered invalid projection data.

[0093] The target 2D reference point information that meets the preset invisible condition is masked. This masking process can be, for example, by setting the score of the target 2D reference point information (i.e., invalid sampling positions) to a very small negative number (e.g., after calculating the attention score and before performing the softmax transform during subsequent cross-modal feature fusion). Thus, after the softmax transformation, its weights will approach 0.

[0094] In this way, after masking the target two-dimensional reference point information in the initial two-dimensional reference point information, the final projected valid two-dimensional reference point information can be obtained.

[0095] S140, Based on the target sampling location, feature sampling is performed on the target image data to obtain image sampling features.

[0096] In some examples, after obtaining the accurate target sampling location, feature sampling can be performed on the target image data based on that accurate target sampling location. For example, feature sampling can be performed at the target sampling location (or at a location within a preset range around it) on the target image data, thereby obtaining image sampling features.

[0097] In some implementations, step S140 may specifically include: sampling image features based on the target sampling location to obtain image sampling features.

[0098] For example, image features can contain feature maps of multiple scales, meaning sampling can be performed on feature maps of multiple different scales. When the target object to be perceived is a small object, it may only have one or two pixels on a high-level (low-resolution) feature map, making it difficult to locate, but it is rich in detail on a low-level (high-resolution) feature map. Therefore, the model can autonomously decide which scale of feature map to sample from. For example, it is more likely to sample from a high-resolution feature map for small objects.

[0099] In some examples, since the offset in the two-dimensional sampling offset information is generally a decimal, the calculated target sampling position is usually a continuous floating-point coordinate. However, the image feature map is a discrete two-dimensional grid (i.e., pixels or feature points), so the feature value cannot be directly obtained through integer indexes. To ensure the accuracy of sampling and the trainability of the entire model, this embodiment can use an interpolation algorithm to achieve feature sampling. This interpolation algorithm can be, for example, a bilinear sampling algorithm.

[0100] Taking the bilinear sampling algorithm as an example, for any target sampling position (i.e., a two-dimensional sampling point) P(x,y), we can first find the four surrounding neighboring points (e.g., ...). ,in, (This indicates rounding down). Calculate the relative distances between point P and its four neighboring points, and determine the bilinear weights (e.g., ...) based on these distances. Obtain the feature vectors (e.g., ...) of the four adjacent points on the corresponding feature map. , dimension And according to the following formula (10), the sampling features at point P can be obtained by weighted summation. .

[0101] The above operations can be performed on all The process is executed in parallel on each of the two-dimensional sampling points, ultimately yielding the image sampling features. Where n is the number of 3D anchor points, and M is the number of 2D sampling points corresponding to each 3D anchor point.

[0102] Thus, by introducing bilinear sampling, this embodiment achieves differentiable, sub-pixel-level feature extraction from continuous image coordinates to discrete feature maps. This enables the model to autonomously optimize the two-dimensional sampling offset through gradient descent, thereby dynamically compensating for projection deviations caused by calibration parameter errors or vibration. This mechanism not only improves the accuracy of cross-modal feature alignment but also ensures that the calibration parameter estimation branch can learn collaboratively with the target perception branch, jointly enhancing the system's perception robustness and online calibration capabilities under harsh conditions such as vibration and shock.

[0103] S150 performs cross-modal feature fusion on image sampling features and point cloud query features to obtain fused query features.

[0104] In some examples, fused query features can be cross-modal features that combine features from both image and point cloud modalities. Specifically, image sampling features and point cloud query features can be fused to obtain fused query features.

[0105] For example, a decoder based on a transformer architecture can be used to achieve cross-modal feature fusion of image sampling features and point cloud query features. Specifically, image sampling features and point cloud query features can be input into the decoder for attention learning, and the fused query features can be output.

[0106] In some implementations, step S150 may specifically include: performing self-attention learning on the first target feature based on a multi-head self-attention mechanism to obtain a first feature, wherein the first target feature includes a point cloud query feature; and performing cross-attention learning on the image sampling feature and the first feature based on a multi-head cross-attention mechanism to obtain a fused query feature.

[0107] For example, the first target feature (e.g., point cloud query features) can be used as the input sequence for self-attention learning. The system performs self-attention learning on the first target feature and the second feature, Queries_self, based on a multi-head self-attention mechanism. The specific process of self-attention learning can be found in step S110 above, where self-attention learning is performed on the anchor point position encoding; it will not be repeated here.

[0108] In other examples, when performing cross-attention learning on image sampling features and the first feature based on a multi-head cross-attention mechanism, the first feature can be used as the query vector Q in the input sequence, and the image sampling features can be used as the key vector K and value vector V in the input sequence. The specific process of cross-attention learning can be referred to in step S110 above, which describes the process of performing cross-attention learning on point cloud features and the third feature, and will not be repeated here.

[0109] Furthermore, the aforementioned self-attention and cross-attention learning processes can also undergo multiple iterations. Based on this, in some implementations, the aforementioned cross-attention learning of image sampling features and the first feature based on a multi-head cross-attention mechanism to obtain fused query features can specifically include: performing cross-attention learning of image sampling features and the first feature based on a multi-head cross-attention mechanism to obtain a second feature; using the fourth feature as the first target feature, returning to perform self-attention learning of the first target feature based on a multi-head self-attention mechanism, until iterating a first preset number of times to obtain the fused query features.

[0110] For example, the second feature obtained after cross-attention learning is used as the input sequence for self-attention learning in the next iteration, and the next iteration calculation is performed. After multiple iterations (e.g., the number of iterations reaches a first preset number), the final output can be obtained, which can be used as the final determined fusion query feature fusion_queries.

[0111] S160, predict the 3D target perception result based on the fusion query features.

[0112] In some examples, the 3D target perception result can be the result obtained by target perception based on point cloud features and image features. This 3D target perception result may include the 3D boundary position, center point position, and pose information of one or more target objects.

[0113] For example, after obtaining the fused query features, the fused query features can be input into a trained target neural network. The target neural network then performs dimensionality upscaling, dimensionality reduction, nonlinearity addition, and feature enhancement processing on the fused query features, thereby outputting a 3D target perception result. The target neural network can be, for example, an FFN.

[0114] For example, for 3D target perception results, an FFN can be set up for target perception prediction. Specifically, the fused query features are input into this FFN, and the first layer of the FFN processes the fused query features from... Increased to multiple times Dimension, and then through the second layer of this FFN from multiple times Dimensional Descending to Dimensional Attainment .in, For example, it can be 7-dimensional, corresponding to the xyz three-dimensional coordinates of the center point, the whl length, width, and height of the three-dimensional boundary position, and the yaw angle in the attitude information.

[0115] In some embodiments, after step S150 above, the target perception method may further include: predicting target calibration parameters based on fusion query features, wherein the target calibration parameters include extrinsic parameters corresponding to the target image acquisition device, and the target image acquisition device is a device for acquiring target image data.

[0116] In some examples, the target image acquisition device may be, for example, a camera that acquires target image data. Additionally, the target calibration parameters may be parameters obtained by correcting inaccurate calibration parameters (such as the first calibration parameter). These target calibration parameters enable accurate mapping between point cloud features and image features even in poor road conditions.

[0117] For example, another FFN can be set up for predicting calibration parameters (e.g., extrinsic parameters) for the target calibration parameters. Specifically, the fused query features are input into this FFN, and the fused query features are passed through the first layer of this FFN from... Increased to multiple times Dimension, and then through the second layer of this FFN from multiple times Dimensional Descending to Dimensional Attainment .in, For example, it can be 7-dimensional, corresponding to 3-dimensional translation and 4-dimensional rotation.

[0118] In addition, in order to continuously improve the accuracy of model prediction, in some embodiments, after predicting the target calibration parameters based on the fusion query features, the target perception method may further include: updating the first calibration parameters based on the target calibration parameters.

[0119] For example, the relatively accurate target calibration parameters predicted under the current road conditions (such as in mining areas where there is continuous vibration) can be used as the first calibration parameters for converting 3D anchor point information to 2D in the next target perception and calibration process, thus forming a continuously self-optimizing calibration closed loop. In this way, when performing a new round of 3D anchor point to 2D projection conversion, the system will no longer rely entirely on the initial offline calibration parameters, which may have become inaccurate due to vibration, but will instead perform the 2D projection conversion based on dynamically updated calibration parameters that are closer to the actual working conditions. This significantly improves the initial accuracy of the 2D reference point information, allowing subsequent 2D sampling offset information to compensate for a smaller range of residual errors, thereby enhancing the convergence and robustness of the entire perception and calibration process, and realizing the inter-frame transfer and progressive optimization of the calibration state in dynamic and harsh environments.

[0120] Furthermore, during model training, the accuracy of 2D sampling offset information prediction can be improved by bringing the target calibration parameters closer to the true values ​​of the calibration parameters. Based on this, in some implementations, the target perception method may further include: obtaining the true values ​​of the calibration parameters; determining the target loss value based on the difference between the target calibration parameters and the true values ​​of the calibration parameters; and training the process of predicting 2D sampling offset information and target calibration parameters based on the target loss value.

[0121] For example, during training, the true value of the calibration parameters can be determined based on actual measurement results or laboratory test results, and the loss value between the target calibration parameters and the true value of the calibration parameters can be calculated based on the preset loss function to obtain the target loss value.

[0122] In some specific examples, the target calibration parameters mentioned above may include translation and rotation. The translation may, for example, include a 3D vector. The rotation amount can include, for example, a 4-dimensional vector. .

[0123] For the translation amount, the translation loss value can be calculated according to the following formula (11). .

[0124] in, The translation amount in the predicted target calibration parameters, This is the translation amount in the true value of the calibration parameters.

[0125] For the amount of rotation, the angle difference loss value can be calculated according to the following formula (12). .

[0126] in, The rotation amount in the predicted target calibration parameters, This is the rotation amount in the true value of the calibration parameters.

[0127] Based on this translation loss value and angle difference loss value The target loss value can be obtained by weighted summation according to the following formula (13). .

[0128] in, Translation loss value The corresponding weights Angular difference loss value The corresponding weights. These two weights are adjustable and are used to balance the contributions of the two losses to ensure that they are of similar magnitude in the early stages of training and converge together in the later stages.

[0129] For example, after obtaining the target loss value, gradient descent can be used to update the relevant trainable parameters in the model to train the process of predicting the two-dimensional sampling offset information and the target calibration parameters. The relevant trainable parameters can be, for example, the network or computational process involved in predicting the two-dimensional sampling offset information and the target calibration parameters, such as FFN, linear projection, or attention learning.

[0130] Furthermore, to ensure that the anchor points fall more accurately on the target object being perceived and to improve the accuracy of target perception, the initial 3D target perception results and the learning and prediction process of the 3D target perception results can also be trained during model training. Based on this, in some embodiments, the target perception method may further include: obtaining the true value of the 3D target; determining a first loss value based on the target loss function and the difference between the initial 3D target perception results and the true value of the 3D target, and determining a second loss value based on the difference between the 3D target perception results and the true value of the 3D target; training the learning and prediction process of the initial 3D target perception results based on the first loss value, and training the learning and prediction process of the 3D target perception results based on the second loss value.

[0131] In some examples, the true value of the 3D target can be determined based on manual calibration results, and the loss values ​​between the initial 3D target perception result and the true value of the calibration parameters can be calculated according to a preset target loss function. This yields a first loss value corresponding to the initial 3D target perception result and a second loss value corresponding to the 3D target perception result. For example, during training, the model can associate each predicted target in the 3D target perception result (or the initial 3D target perception result) with the real target in the 3D target true value using bipartite graph matching to further calculate the loss value. Furthermore, the target loss function can include at least one of a classification loss function, a regression loss function, and an intersection-union (IUU) loss function.

[0132] In some specific examples, when calculating the first loss value and / or the second loss value, the loss values ​​obtained by calculating the three loss functions respectively can be calculated by weighted summation according to the following formula (14).

[0133] in, This represents the total loss value (e.g., the first loss value or the second loss value). This represents the loss value calculated based on the classification loss function. Represents the weights corresponding to the classification loss function. This represents the loss value calculated based on the regression loss function. This represents the weights corresponding to the regression loss function. This represents the loss value calculated based on the cross-union ratio loss function. These represent the weights corresponding to the Cross-Union Ratio (CUI) loss function. These three weights can also be adjustable to balance the contribution of each loss function to the loss value.

[0134] For example, after obtaining the first loss value, gradient descent can be used to update the relevant trainable parameters in the model to train the learning and prediction process of the initial 3D target perception result. The relevant trainable parameters can be, for example, the network or computational process involved in learning and predicting the initial 3D target perception result, such as FFN, linear projection, or attention learning.

[0135] Furthermore, after obtaining the second loss value, gradient descent can be used to update the relevant trainable parameters in the model to train the learning and prediction process of 3D target perception results. The relevant trainable parameters can be, for example, the network or computational processes involved in learning and predicting 3D target perception results, such as FFN, linear projection, and attention learning processes.

[0136] Regarding the target perception methods provided in the above embodiments, the following will be combined with... Figure 2 and Figure 3 To better illustrate the target-aware approach, let's take a concrete example.

[0137] Figure 2 The diagram shown is a schematic flowchart of a point cloud data processing branch provided in an embodiment of this application.

[0138] like Figure 2 As shown, the processing flow of the point cloud data processing branch can include: performing feature extraction and feature location encoding on the target point cloud data to obtain point cloud features with location information. These point cloud features are used as the inputs K and V for subsequent cross-attention learning. Simultaneously, the system generates a set of learnable initial 3D anchor points through methods such as fixed-point sampling, heatmap guidance, or real-world bounding box perturbation. These anchor points are then position-encoded and used as the input sequence (Q, K, V) for subsequent self-attention learning. Multi-head self-attention learning and multi-head feature concatenation are performed on the input sequence, and the result is used as the input Q for subsequent cross-attention learning. Based on the K and V generated from the point cloud features and the result Q obtained from self-attention learning, multi-head cross-attention learning is performed. The result is input into the FFN (Flexible Input Network) to add non-linearity, thus obtaining an output result. This output result is returned to the input of self-attention learning as a new input sequence for repeated multi-head self-attention and multi-head cross-attention learning. This process is iterated N times to obtain the point cloud query features. The point cloud query features are processed through three independent FFNs, which then output the initial 3D target perception results, 3D anchor point information, and 2D sampling offset information, respectively.

[0139] In some examples, the resulting image corresponding to the initial 3D target perception result can be, for example, as shown below. Figure 4 As shown, where, Figure 4 The multiple rectangular calibration boxes in the image represent the initial three-dimensional target perception results.

[0140] The point cloud data processing branch provides the entire system with preliminary perception results based on point cloud data and crucial geometric prior information.

[0141] Figure 3 The diagram shown is a schematic flowchart of a feature fusion branch provided in an embodiment of this application.

[0142] like Figure 3As shown, the processing flow of the feature fusion branch can include: extracting multi-scale features from the target image data through the backbone network and performing position encoding to form image features. Using the first calibration parameters, the 3D anchor point information predicted by the point cloud data processing branch is transformed from 3D to 2D to obtain 2D reference point information. This information is then combined with the 2D sampling offset information predicted by the point cloud data processing branch for dynamic position correction to obtain a precise target sampling position. Based on this target sampling position, corresponding image sampling features are extracted from the image features through bilinear sampling, and these image sampling features are used as the inputs K and V for subsequent cross-attention learning. Simultaneously, the point cloud query features obtained from the point cloud data processing branch are used as the input sequence (Q, K, V) for self-attention learning. Multi-head self-attention learning and multi-head feature concatenation are performed on the input sequence, and the result is used as the input Q for subsequent cross-attention learning. Based on the K and V generated from the image sampling features and the result Q obtained from self-attention learning, multi-head cross-attention learning is performed, and the result is input into the FFN to add nonlinearity, thus obtaining a single output result. The output is then fed back to the input of the self-attention learning process, serving as a new input sequence for repeated multi-head self-attention and multi-head cross-attention learning. This process is iterated N times to obtain the fused query features. These fused query features are then processed through two independent FFNs, resulting in the final 3D target perception result and target calibration parameters.

[0143] In some examples, the process of correcting the positional offset of two-dimensional reference point information can be, for example, as follows: Figure 5 As shown. Figure 5 As shown, when mapping the three-dimensional target object 51 perceived in the point cloud data to the two-dimensional target object 52 in the image data, the anchor point (i.e., black point) on the three-dimensional target object 51 may be mapped to the corresponding reference point on the two-dimensional offset object 53 due to calibration parameter errors. However, this application can correct the offset position of the reference point by using two-dimensional sampling offset information, thereby correcting the two-dimensional offset object 53 to the two-dimensional target object 52.

[0144] The core of the feature fusion branch lies in realizing the adaptive alignment and collaborative reasoning of visual and radar features, thereby improving perception accuracy while completing the self-calibration of calibration parameters (such as extrinsic parameters).

[0145] The above text combined Figures 1 to 5 The embodiments of the target perception method of this application are described in detail below, in conjunction with... Figure 6 This application describes in detail embodiments of the target sensing device. It should be understood that the descriptions of the target sensing method embodiments correspond to the descriptions of the target sensing device embodiments; therefore, any parts not described in detail can be found in the foregoing method embodiments.

[0146] Figure 6 The diagram shown is a structural schematic of a target sensing device according to an embodiment of this application. This device can be configured in an electronic device. Figure 6 As shown, the target sensing device 600 provided in this application embodiment includes: The feature query module 601 is used to perform feature query on the target point cloud data to obtain the point cloud query features. The first prediction module 602 is used to predict three-dimensional anchor point information and two-dimensional sampling offset information based on the point cloud query features. The offset correction module 603 is used to perform offset correction on the sampling position corresponding to the three-dimensional anchor point information on the target image data according to the two-dimensional sampling offset information, so as to obtain the target sampling position, wherein the target image data is the image data corresponding to the target point cloud data; Feature sampling module 604 is used to perform feature sampling on the target image data based on the target sampling position to obtain image sampling features; Feature fusion module 605 is used to perform cross-modal feature fusion on the image sampling features and the point cloud query features to obtain fused query features; The second prediction module 606 is used to predict the three-dimensional target perception result based on the fusion query features.

[0147] In one embodiment of this application, the target perception device 600 further includes: a third prediction module, used to predict target calibration parameters based on the fusion query features, wherein the target calibration parameters include extrinsic parameters corresponding to the target image acquisition device, and the target image acquisition device is a device for acquiring the target image data.

[0148] In one embodiment of this application, the offset correction module 603 is further configured to: acquire image features corresponding to the target image data; convert the three-dimensional anchor point information to two dimensions according to the first calibration parameters to obtain two-dimensional reference point information; and perform offset correction on the sampling position of the two-dimensional reference point information on the image features according to the two-dimensional sampling offset information to obtain the target sampling position. Correspondingly, the feature sampling module 604 is further configured to: sample on the image features based on the target sampling position to obtain image sampling features.

[0149] In one embodiment of this application, the target sensing device 600 further includes a parameter update module, used to update the first calibration parameter according to the target calibration parameter.

[0150] In one embodiment of this application, the target sensing device 600 further includes: a first acquisition module for acquiring the true value of calibration parameters; a first determination module for determining a target loss value based on the difference between the target calibration parameters and the true value of the calibration parameters; and a first training module for training the process of predicting the two-dimensional sampling offset information and the target calibration parameters based on the target loss value.

[0151] In one embodiment of this application, the offset correction module 603 is further configured to: convert the three-dimensional anchor point information to two dimensions according to the first calibration parameters to obtain initial two-dimensional reference point information; determine target two-dimensional reference point information that meets the preset invisible condition from the initial two-dimensional reference point information; and perform masking processing on the target two-dimensional reference point information in the initial two-dimensional reference point information to obtain the two-dimensional reference point information.

[0152] In one embodiment of this application, the feature fusion module 605 is further configured to: perform self-attention learning on the first target feature based on a multi-head self-attention mechanism to obtain a first feature, wherein the first target feature includes the point cloud query feature; and perform cross-attention learning on the image sampling feature and the first feature based on a multi-head cross-attention mechanism to obtain the fused query feature.

[0153] In one embodiment of this application, the feature fusion module 605 is further configured to: perform cross-attention learning on the image sampling features and the first feature based on a multi-head cross-attention mechanism to obtain a second feature; take the fourth feature as the first target feature, and return to perform self-attention learning on the first target feature based on a multi-head self-attention mechanism until iterates a first preset number of times to obtain the fused query feature.

[0154] In one embodiment of this application, the feature query module 601 is further configured to: acquire point cloud features corresponding to the target point cloud data; determine multiple initial anchor points corresponding to the target point cloud data; and perform attention learning based on the point cloud features and the multiple initial anchor points to obtain the point cloud query features.

[0155] In one embodiment of this application, the target perception device 600 further includes: a fourth prediction module, used to predict the initial three-dimensional target perception result based on the point cloud query features.

[0156] In one embodiment of this application, the target perception device 600 further includes: a second acquisition module for acquiring the true value of a three-dimensional target; a second determination module for determining a first loss value based on a target loss function and the difference between the initial three-dimensional target perception result and the true value of the three-dimensional target, and determining a second loss value based on the difference between the three-dimensional target perception result and the true value of the three-dimensional target, wherein the target loss function includes at least one of a classification loss function, a regression loss function, and an intersection-union loss function; and a second training module for training the learning and prediction process of the initial three-dimensional target perception result based on the first loss value, and training the learning and prediction process of the three-dimensional target perception result based on the second loss value.

[0157] Below, for reference Figure 7 This describes an electronic device according to embodiments of the present application. Figure 7 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application.

[0158] like Figure 7 As shown, the electronic device 700 includes one or more processors 701 and memory 702.

[0159] The processor 701 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 700 to perform desired functions.

[0160] The memory 702 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 701 may execute the program instructions to implement the target perception methods of the various embodiments of this application described above and / or other desired functions. The computer-readable storage medium may also store various contents such as three-dimensional anchor point information, two-dimensional sampling offset information, and three-dimensional target perception results.

[0161] In one example, the electronic device 700 may also include an input device 703 and an output device 704, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0162] The input device 703 may include, for example, a keyboard, a mouse, etc.

[0163] The output device 704 can output various information to the outside, including three-dimensional anchor point information, two-dimensional sampling offset information, and three-dimensional target perception results. The output device 704 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0164] Of course, for the sake of simplicity, Figure 7 Only some of the components of the electronic device 700 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 700 may include any other suitable components depending on the specific application.

[0165] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the target perception methods according to the various embodiments of this application described above.

[0166] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0167] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the target perception methods according to the various embodiments of this application described above.

[0168] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.

[0169] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.

[0170] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0171] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.

[0172] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0173] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A target perception method, characterized in that, include: Perform feature queries on the target point cloud data to obtain the point cloud query features; Predict 3D anchor point information and 2D sampling offset information based on the point cloud query features; Based on the two-dimensional sampling offset information, the sampling position corresponding to the three-dimensional anchor point information on the target image data is offset and corrected to obtain the target sampling position. The target image data is the image data corresponding to the target point cloud data. Based on the target sampling location, feature sampling is performed on the target image data to obtain image sampling features; Cross-modal feature fusion is performed on the image sampling features and the point cloud query features to obtain fused query features; The three-dimensional target perception result is predicted based on the fusion query features.

2. The method according to claim 1, characterized in that, After performing cross-modal feature fusion on the image sampling features and the point cloud query features to obtain fused query features, the method further includes: The target calibration parameters are predicted based on the fusion query features. The target calibration parameters include extrinsic parameters corresponding to the target image acquisition device, which is a device used to acquire the target image data.

3. The method according to claim 2, characterized in that, The step of offsetting and correcting the sampling position of the three-dimensional anchor point information on the target image data according to the two-dimensional sampling offset information to obtain the target sampling position includes: Obtain the image features corresponding to the target image data; The three-dimensional anchor point information is converted to two-dimensional according to the first calibration parameters to obtain two-dimensional reference point information; Based on the two-dimensional sampling offset information, the sampling position of the two-dimensional reference point information on the image feature is offset and corrected to obtain the target sampling position; The step of performing feature sampling on the target image data based on the target sampling location to obtain image sampling features includes: Based on the target sampling location, sampling is performed on the image features to obtain image sampling features.

4. The method according to claim 3, characterized in that, After predicting the target calibration parameters based on the fusion query features, the method further includes: Update the first calibration parameter according to the target calibration parameter.

5. The method according to any one of claims 2 to 4, characterized in that, Also includes: Obtain the true values ​​of the calibration parameters; The target loss value is determined based on the difference between the target calibration parameters and the true values ​​of the calibration parameters; The process of predicting the two-dimensional sampling offset information and the target calibration parameters is trained based on the target loss value.

6. The method according to claim 3, characterized in that, The step of converting the three-dimensional anchor point information to two dimensions according to the first calibration parameters to obtain two-dimensional reference point information includes: The three-dimensional anchor point information is converted to two-dimensional according to the first calibration parameters to obtain the initial two-dimensional reference point information; Target two-dimensional reference point information that meets the preset invisible condition is determined from the initial two-dimensional reference point information; The target two-dimensional reference point information in the initial two-dimensional reference point information is masked to obtain the two-dimensional reference point information.

7. The method according to claim 1, characterized in that, The cross-modal feature fusion of the image sampling features and the point cloud query features to obtain fused query features includes: The first target feature is obtained by performing self-attention learning on the first target feature based on the multi-head self-attention mechanism, and the first target feature includes the point cloud query feature. The image sampling features and the first feature are cross-attention learned based on a multi-head cross-attention mechanism to obtain the fused query features; The method of performing cross-attention learning on the image sampling features and the first feature based on a multi-head cross-attention mechanism to obtain the fused query features includes: The image sampling features and the first feature are cross-attention learned based on a multi-head cross-attention mechanism to obtain the second feature; The fourth feature is used as the first target feature, and the self-attention learning of the first target feature based on the multi-head self-attention mechanism is performed again until the first preset number of iterations are completed to obtain the fused query feature.

8. The method according to claim 1, characterized in that, The step of performing feature query on the target point cloud data to obtain point cloud query features includes: Obtain the point cloud features corresponding to the target point cloud data; Determine multiple initial anchor points corresponding to the target point cloud data; Attention learning is performed based on the point cloud features and the multiple initial anchor points to obtain the point cloud query features.

9. The method according to claim 1, characterized in that, After performing feature query on the target point cloud data to obtain the point cloud query features, the method further includes: Predict the initial 3D target perception result based on the point cloud query features; The method further includes: Obtain the true value of the three-dimensional target; Based on the target loss function, a first loss value is determined according to the difference between the initial 3D target perception result and the true value of the 3D target, and a second loss value is determined according to the difference between the 3D target perception result and the true value of the 3D target. The target loss function includes at least one of the classification loss function, regression loss function, and intersection-union loss function. The learning and prediction process of the initial three-dimensional target perception result is trained based on the first loss value, and the learning and prediction process of the three-dimensional target perception result is trained based on the second loss value.

10. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the target perception method according to any one of claims 1 to 9.