Image processing methods, apparatus, electronic devices, and storage media

The image processing method addresses the challenge of understanding environmental information using multi-view images by determining bird's-eye view features and processing static and dynamic elements, achieving accurate and cost-effective environmental understanding without high-precision maps.

JP7867634B2Active Publication Date: 2026-05-29BEIJING HORIZON INFORMATION TECH CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
BEIJING HORIZON INFORMATION TECH CO LTD
Filing Date
2023-09-11
Publication Date
2026-05-29

Smart Images

  • Figure 0007867634000001
    Figure 0007867634000001
  • Figure 0007867634000002
    Figure 0007867634000002
  • Figure 0007867634000003
    Figure 0007867634000003
Patent Text Reader

Abstract

[0013] The present disclosure provides an image processing method, an apparatus, an electronic device, and a storage medium, the method including: determining, based on a processing target image corresponding to each of at least one viewpoint, first image features corresponding to each viewpoint; determining, based on the first image features corresponding to each viewpoint; determining, based on the first bird's-eye view features, at least one task query feature selected from the group consisting of static element task query features, dynamic object task query features, and motion trajectory task query features; and determining, based on each task query feature selected from the group consisting of the at least one task query features, a task processing result corresponding to each task query feature. The present disclosure provides end-to-end single-task or multi-task processing by relying solely on multi-view environmental images, and can effectively obtain accurate surrounding environment information even in the absence of a high-precision map, thereby significantly improving versatility and effectively reducing costs.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-reference to related applications

[0001] This disclosure claims the priority of a Chinese patent application filed with the China National Intellectual Property Administration on November 11, 2022, with the application number CN202211417346.2 and the invention title "Image Processing Method, Apparatus, Electronic Device, and Storage Medium", and all of its contents are incorporated herein by reference.

Technical Field

[0002] This disclosure relates to computer vision technology, and particularly to an image processing method, apparatus, electronic device, and storage medium.

Background Art

[0003] In the field of autonomous driving, how to efficiently understand environmental information depending on multi-view environmental images is an extremely important technical problem.

Summary of the Invention

Problems to be Solved by the Invention

[0004] Embodiments of this disclosure provide an image processing method, apparatus, electronic device, and storage medium.

Means for Solving the Problems

[0005] An image processing method according to one aspect of the embodiments of this disclosure includes: determining, based on a processing target image respectively corresponding to each of at least one perspective, a first image feature respectively corresponding to each of the at least one perspective; determining a first bird's-eye view feature based on the first image feature respectively corresponding to each of the at least one perspective; determining at least one type of task query feature among static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view feature; and determining a task processing result respectively corresponding to each of the at least one type of task query feature based on each of the at least one type of task query feature.

[0006] An image processing device according to another embodiment of the present disclosure includes: a first processing module for determining a first image feature corresponding to each of the at least one viewpoints based on a processing image corresponding to each of the at least one viewpoints; a second processing module for determining a first bird's-eye view feature based on the first image feature corresponding to each of the viewpoints; a third processing module for determining at least one task query feature from static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view feature; and a fourth processing module for determining a task processing result corresponding to each of the at least one task query feature based on each of the task query features.

[0007] A computer-readable storage medium according to yet another embodiment of the embodiments of the present disclosure stores a computer program for performing the image processing method described in any one of the above embodiments of the present disclosure.

[0008] An electronic device according to yet another embodiment of the embodiments of the present disclosure includes a processor and a memory for storing instructions that the processor can execute, the processor being used to read and execute the instructions from the memory to implement the image processing method described in any one of the embodiments of the present disclosure. [Effects of the Invention]

[0009] The image processing method, apparatus, electronic device, and storage medium according to the above embodiment of this disclosure determine a bird's-eye view feature based on image features corresponding to each viewpoint determined by the processing target image for each viewpoint, determine at least one type of task query feature based on the bird's-eye view feature, and further acquire task processing results corresponding to each task based on each task query feature, thereby enabling end-to-end single-task or multi-task processing based on multi-view environmental images, thereby avoiding or reducing reliance on high-precision maps, and thereby enabling the effective acquisition of accurate surrounding environment information even when high-precision maps are unavailable, contributing to improved versatility and cost reduction.

[0010] The technical solutions of this disclosure will be described in more detail below with reference to drawings and embodiments. [Brief explanation of the drawing]

[0011] [Figure 1] This is one example application scenario of the image processing method related to this disclosure. [Figure 2] This is a flowchart of an image processing method relating to one exemplary embodiment of the present disclosure. [Figure 3] This is a flowchart of an image processing method relating to another exemplary embodiment of the present disclosure. [Figure 4] This is a flowchart of step 2031a relating to one exemplary embodiment of the present disclosure. [Figure 5] This is a schematic diagram of the network structure of a first decoding network according to one exemplary embodiment of the present disclosure. [Figure 6] This is a flowchart of step 2031b according to one exemplary embodiment of the present disclosure. [Figure 7] This is a flowchart of step 2031c according to one exemplary embodiment of the present disclosure. [Figure 8] This is a schematic diagram of a third decoding network structure according to one exemplary embodiment of the present disclosure. [Figure 9]It is a flowchart of a method for processing an image according to another exemplary embodiment of the present disclosure. [Figure 10] It is a flowchart of step 301 according to one exemplary embodiment of the present disclosure. [Figure 11] It is a schematic diagram of the principle for determining the initial motion trajectory query feature according to one exemplary embodiment of the present disclosure. [Figure 12] It is a flowchart of step 2021 according to one exemplary embodiment of the present disclosure. [Figure 13] It is a schematic diagram of the network structure of the encoder network according to one exemplary embodiment of the present disclosure. [Figure 14] It is a schematic diagram of the overall structure of the network model for image processing according to one exemplary embodiment of the present disclosure. [Figure 15] It is a schematic diagram of the structure of an image processing apparatus according to one exemplary embodiment of the present disclosure. <​​​​​​​​​​​​​​​​​​​​​​​​In the process of realizing the present disclosure, the inventor has found that in the field of autonomous driving, how to efficiently understand environmental information by relying on multi-view environmental images is a very important technical problem. Based on multi-view environmental images, if an understanding of the surrounding environmental information is realized in accordance with a high-precision map, it is likely to cause the accuracy of the environmental information obtained when there is no high-precision map to be low.

[0015] Exemplary overview FIG. 1 is one exemplary application scenario of the image processing method according to the present disclosure.

[0016] In an autonomous driving scenario, vehicle surrounding environment images can be collected as processing target images corresponding to each viewpoint based on an in-vehicle surround-view camera (a camera that can include multiple viewpoints). By using the image processing device of this disclosure to execute the image processing method of this disclosure, a first image feature corresponding to each viewpoint can be determined based on the processing target images corresponding to each viewpoint among at least one viewpoint. A first bird's-eye view feature can be determined based on the first image feature corresponding to each viewpoint, the first bird's-eye view feature being a feature in a grid coordinate system corresponding to a bird's-eye view (abbreviated as BEV). A static element task query feature, a dynamic object task query feature, and a motion trajectory task query feature can be determined based on the first bird's-eye view feature. Furthermore, a corresponding task processing result (including, but not limited to, task processing result 1, task processing result 2, and task processing result 3) can be determined based on the static element task query feature, the dynamic object task query feature, and the motion trajectory task query feature, respectively. For example, a static element task processing result (specifically, for example, a static element detection result of the vehicle surrounding environment in an autonomous driving scenario) can be determined based on the static element task query feature. The system determines the dynamic object task query features, determines the dynamic object task processing result (specifically, the 3D target detection result), determines the motion trajectory task processing result (specifically, the motion trajectory prediction result of the dynamic object) based on the motion trajectory task query features, realizes end-to-end single-task or multi-task processing based on multi-view environmental images, does not require matching to high-precision maps, contributes to avoiding or reducing dependence on high-precision maps, thereby enabling the effective acquisition of accurate surrounding environment information even when high-precision maps are unavailable, contributes to improved versatility, and reduces costs. Here, static elements may include static object elements such as lane markings, crosswalks, and road shoulders, dynamic objects may include objects with motion attributes such as surrounding vehicles and pedestrians, and motion trajectory is the motion trajectory of a dynamic object.

[0017] Furthermore, the image processing method disclosed herein is not limited to the above-mentioned autonomous driving scene, but can be applied to any other possible scene depending on actual needs. For example, in a security monitoring scene of a certain area, the bird's-eye view features of this area can be obtained from images collected by cameras at each viewpoint, and end-to-end task processing of static elements, dynamic objects and / or motion trajectories of dynamic objects within this area can be realized. The specific scene can be set according to actual needs.

[0018] Exemplary Method Figure 2 is a flowchart of an image processing method according to one exemplary embodiment of the present disclosure. This embodiment can be specifically applied to electronic devices such as automotive computing platforms and includes the following steps, as shown in Figure 2.

[0019] In step 201, a first image feature corresponding to each viewpoint is determined based on the processing image corresponding to each of the at least one viewpoints.

[0020] Here, the number of viewpoints can be set according to the actual needs. For example, in an autonomous driving scenario, the number of viewpoints is the number of surround-view cameras installed on the vehicle, with each camera corresponding to one viewpoint. For instance, a four-way surround-view system consisting of a left-front camera, a left-rear camera, a right-front camera, and a right-rear camera includes four viewpoints, but is not specifically limited. The first image features can be obtained using any feasible feature extraction method. For example, feature extraction is performed on each image to be processed based on a pre-trained feature extraction network, and the first image features corresponding to each viewpoint are obtained. Here, the feature extraction network can be set according to the actual needs. For example, a convolutional neural network can be used as the feature extraction network.

[0021] In one possible example, step 201 may be performed by the processor calling a corresponding instruction stored in memory, or by a first processing module executed by the processor.

[0022] In step 202, the first bird's-eye view feature is determined based on the first image feature corresponding to each viewpoint.

[0023] Here, the first bird's-eye view feature is the BEV feature in the grid coordinate system corresponding to the bird's-eye view. Based on a pre-trained encoder network, the first image features of each viewpoint can be encoded, and the first bird's-eye view feature can be obtained. The encoder network can be configured according to the actual needs.

[0024] In one possible example, step 202 may be performed by the processor calling a corresponding instruction stored in memory, or by a second processing module executed by the processor.

[0025] In step 203, based on the first bird's-eye view feature, at least one task query feature is determined from among the static element task query feature, the dynamic object task query feature, and the motion trajectory task query feature.

[0026] Here, static element task query features are task query features related to static elements extracted from the first bird's-eye view feature; similarly, dynamic object task query features are task query features related to dynamic objects extracted from the first bird's-eye view feature; and motion trajectory task query features are task query features related to the motion trajectory of dynamic objects extracted from the static element task query features. The specific types of task query features that need to be obtained can be set according to the actual needs. For example, one type may be obtained, two types may be obtained, or all three types may be obtained simultaneously. The first bird's-eye view feature can be decoded and obtained using a pre-trained decoding network corresponding to this task for any one type of task query feature. The specific decoding network can be set according to the actual needs.

[0027] In one possible example, step 203 may be performed by the processor calling a corresponding instruction stored in memory, or by a third processing module executed by the processor.

[0028] In step 204, the task processing result corresponding to each task query feature is determined based on each of the task query features among at least one type of task query feature.

[0029] Here, for any one type of task, a head network corresponding to that task can be set up and trained to obtain the trained head network. The trained head network is then used to perform output projection onto the task query features corresponding to that task and to obtain the task processing results corresponding to those task query features. The specific network structure of the head network can be set up according to the actual needs, and can be implemented, for example, using a multilayer perceptron (abbreviated as MLP).

[0030] In one possible example, step 204 may be performed by the processor calling a corresponding instruction stored in memory, or by a fourth processing module executed by the processor.

[0031] The image processing method according to this embodiment determines bird's-eye view features based on image features corresponding to each viewpoint determined by the processing target image for each viewpoint, determines at least one type of task query feature based on the bird's-eye view features, and further obtains task processing results corresponding to each task based on each task query feature. This enables end-to-end single-task or multi-task processing by relying only on multi-view environmental images, does not require matching with high-precision maps, avoids or reduces dependence on high-precision maps, and thereby enables the effective acquisition of accurate surrounding environment information even when high-precision maps are unavailable, contributing to improved versatility and reduced costs.

[0032] Figure 3 is a flowchart of an image processing method according to another exemplary embodiment of the present disclosure.

[0033] In one possible example, step 203 may specifically include the following steps:

[0034] In step 2031a, static element task query features are determined using a pre-trained first decoding network based on the first bird's-eye view feature and the initial static element query features, and the initial static element query features include an initial query feature corresponding to each static element of at least one static element.

[0035] Here, the initial static element query features can be set according to the actual needs. For example, initial static element query features obtained by initializing based on a first initialization rule can be set according to the actual needs. For example, N2 D-dimensional static queries (where a static query represents an initial query feature corresponding to a static element) can be randomly initialized to become the initial static element query features, where N2 is the number of static queries, i.e., the number of static elements, and N2 can be set according to the actual needs. D represents the dimension of the initial query feature corresponding to each static element. The first decoding network may include at least one decoder, which is used to query task query features related to static elements from the first overview feature based on the initial static element query features and to obtain static element task query features. The specific network structure of the first decoding network can be set according to the actual needs.

[0036] In one possible example, step 2031a may be performed by the processor calling a corresponding instruction stored in memory, or by a first processing unit executed by the processor.

[0037] The embodiments of this disclosure utilize a first decoding network that has been trained to obtain static element task query features related to static elements from a first bird's-eye view feature based on initial static element query features, thereby providing more accurate and effective feature data for subsequent static element detection tasks. This enables end-to-end task processing based on multi-view images, and when applied to map reconstruction scenes, it enables the generation of static map information online without relying on offline high-precision maps, further improving versatility.

[0038] In one possible example, step 203 may specifically include the following steps:

[0039] In step 2031b, dynamic object task query features are determined using a pre-trained second decoding network based on the first bird's-eye view features and initial dynamic object query features, and the initial dynamic object query features include initial query features corresponding to each dynamic object among at least one dynamic object.

[0040] Here, the initial dynamic object query features can be set according to the actual needs. For example, initial dynamic object query features obtained by initializing based on a second initialization rule can be set according to the actual needs. For example, N1 D-dimensional dynamic queries (where a dynamic query represents an initial query feature corresponding to a dynamic object) can be randomly initialized to become the initial dynamic object query features, where N1 is the number of dynamic queries, i.e., the number of dynamic objects, and N1 can be set according to the actual needs. D represents the dimension of the initial query feature corresponding to each dynamic object. The second decoding network may include at least one decoder and is used to query task query features related to dynamic objects from the first overview features based on the initial dynamic object query features and to obtain dynamic object task query features. The specific network structure of the second decoding network can be set according to the actual needs.

[0041] In one possible example, step 2031b may be performed by the processor calling a corresponding instruction stored in memory, or by a second processing unit executed by the processor.

[0042] The embodiments of this disclosure utilize a second decoding network that has been trained to obtain dynamic object task query features related to dynamic objects from the first bird's-eye view features based on initial dynamic object query features. By providing accurate and effective feature data for subsequent dynamic object detection tasks, end-to-end 3D target detection task processing based on multi-view images is achieved, avoiding the need to track dynamic objects. This contributes to reducing the computational complexity of the network model and avoiding the impact of target tracking errors on subsequent applications.

[0043] In one possible example, step 203 may specifically include the following steps:

[0044] In step 2031c, the motion trajectory task query features are determined using a pre-trained third decoding network based on the static element task query features and the initial motion trajectory query features, and the initial motion trajectory query features include an initial trajectory query feature corresponding to each dynamic object among at least one dynamic object.

[0045] Here, the initial motion trajectory query features need to be determined by combining dynamic object task query features and modality query features, where modality query features are used to represent the motion tendency of the dynamic object, with different modalities focusing on different types of future motion (e.g., high-speed straight, low-speed straight, left turn, right turn, etc.), and dynamic object task query features are used to represent features related to the dynamic object. By combining the modality query features and dynamic object task query features, the initial trajectory query features related to the dynamic object motion trajectory can be determined, and further, they interact with static element task query features in the third decoding network to decode the motion trajectory task query features. The third decoding network can be configured according to the actual needs.

[0046] In one possible example, step 2031c may be performed by the processor calling a corresponding instruction stored in memory, or by a third processing unit executed by the processor.

[0047] The embodiments of this disclosure utilize a trained third decoding network to query and obtain motion trajectory task query features related to the motion trajectory of a dynamic object from static element task query features based on initial motion trajectory query features, thereby providing more accurate and effective feature data for subsequent dynamic object motion trajectory prediction tasks, and enabling end-to-end motion trajectory prediction task processing based on multi-view images.

[0048] In one selectable example, step 203 may include at least two of the steps 2031a to 2031c described above, which can be specifically configured according to actual needs, thereby enabling end-to-end multitask processing solely by relying on multi-view images, obtaining task processing results corresponding to each of the multiple tasks, avoiding reliance on high-precision maps and laser radar, and contributing to further improved versatility.

[0049] Figure 4 is a flowchart of step 2031a according to one exemplary embodiment of the present invention.

[0050] In one selectable embodiment, step 2031a, which determines static element task query features using a pre-trained and acquired first decoding network based on a first bird's-eye view feature and an initial static element query feature, includes the following steps:

[0051] In step 20311a, the first query tensor, the first key tensor, and the first value tensor are determined based on the initial static element query features.

[0052] Here, initial static element query features can be mapped as a first query tensor based on the first query mapping rule, for example, based on the first query mapping matrix. Similarly, initial static element query features can be mapped as a first key tensor based on the first key mapping rule, and initial static element query features can be mapped as a first value tensor based on the first value mapping rule. The specific principles of the mapping will not be explained here.

[0053] In step 20312a, the first self-attention result is determined based on the first query tensor, the first key tensor, and the first value tensor, using the first self-attention network of the first decoder in the first decoding network.

[0054] Here, the first self-attention network is a network based on a self-attention mechanism, which can be configured according to actual needs, and is used to complete self-attention operations on the first query tensor, the first key tensor, and the first value tensor. Specifically, it performs self-attention operations on the first query tensor and the first key tensor to obtain the first weights, and then weights the first value tensor based on the first weights to obtain the first self-attention result. The specific principles of the self-attention mechanism will not be explained here.

[0055] In step 20313a, the first intermediate result is determined based on the first self-attention result and the initial static element query features, using the first additive normalization network of the first decoder in the first decoding network.

[0056] Here, the first Addition and Normalization Network (Add&Norm) has two functions: addition and normalization. Addition refers to adding the first self-attention result and the initial static element query features to obtain the first addition result, and then normalizing the first addition result to obtain the first intermediate result.

[0057] In step 20314a, the second query tensor is determined based on the first intermediate result.

[0058] Here, the first intermediate result can be used as the second query tensor, or the first intermediate result can be mapped to the second query tensor based on the corresponding mapping rule, and this can be specifically configured according to the actual needs.

[0059] In step 20315a, the second key tensor and the second value tensor are determined based on the first bird's-eye view feature.

[0060] Here, the determination principles for the second key tensor and the second value tensor can be found by referring to what was described above, so we will omit the explanation here.

[0061] In step 20316a, the first cross-attention result is determined using the first deformable cross-attention network of the first decoder in the first decoding network, based on the second query tensor, the second key tensor, and the second value tensor.

[0062] Here, the first deformable cross-attention network is a deformable convolution-based cross-attention network whose function is to extract features within local regions near the location corresponding to the initial static element query features from the first bird's-eye view features, thereby further improving the accuracy and effectiveness of feature extraction.

[0063] In step 20317a, the static element task query features are determined based on the first cross-attention result and the first intermediate result.

[0064] Here, in the first decoder of the first decoding network, other related networks such as an Add&Norm network or a FeedForward network may be included after the first deformable cross-attention network. Thus, after obtaining the first cross-attention result, the first cross-attention result and the first intermediate result are added together and normalized, and then the decoding result of the first decoder can be finally obtained by the other related networks. If the first decoding network includes multiple decoders, the decoding result of the first decoder can be used as input to the second decoder. Decoding is then performed again based on the decoding flow of the first decoder described above, and this process is repeated until decoding of all decoders is completed, thereby obtaining the final decoding result of the first decoding network, which is then used as a static element task query feature.

[0065] In one selectable embodiment, Figure 5 is a schematic diagram of the network structure of a first decoding network according to one exemplary embodiment of the present disclosure. As shown in Figure 5, the first decoding network includes six decoders. Here, ×6 indicates that the first decoding network includes six decoders within a dashed box, the decoders within this dashed box being, for example, the first decoder, Q1, K1, and V1 representing the first query tensor, the first key tensor, and the first value tensor, respectively, Self Attention representing the first self-attention network, Add&Norm representing the additive normalization network, Add&Norm connected to the first self-attention network representing the first additive normalization network, Q2, K2, and V2 representing the second query tensor, the second key tensor, and the second value tensor, respectively, Deformable Cross Attention representing the first deformable cross-attention network, and Feed Forward representing the feedforward network. After the initial static element query features are mapped as the first query tensor, first key tensor, and first value tensor, the first self-interaction is performed in the first self-attention network, and the first self-attention result is obtained. After the first self-attention result is added to the initial static element query features and normalized, the first intermediate result is obtained, and the first intermediate result is mapped as the second query tensor, while the first bird's-eye view features are mapped as the second key tensor and second value tensor, and the second query tensor and second key tensor are mapped. The sol and the second value tensor undergo cross-attention in the first deformable cross-attention network, enabling interaction between the first bird's-eye view features and the initial static element query features, obtaining the first cross-attention result. This first cross-attention result is then added to the first intermediate result and normalized to obtain the first normalized result. This first normalized result then passes through a feedforward network and another additive normalization network, after which the decoded result of the first decoder is obtained. This decoded result then undergoes further decoding by five decoders to obtain the static element task query features.

[0066] In one selectable example, the first self-attention network may be a multi-head self-attention network, and the first deformable cross-attention network may also be a multi-head deformable cross-attention network, which can be specifically configured according to actual needs.

[0067] In one selectable example, steps 20311a to 20317a above may be performed by the processor calling corresponding instructions stored in memory, or by a first processing unit executed by the processor.

[0068] The embodiments of this disclosure achieve self-interaction of static element query features by a self-attention network in the first decoding network, enabling the capture of internal correlations of static element query features. Furthermore, by interacting with bird's-eye view features in a deformable cross-attention network, sparse attention to the data is achieved, flexibly capturing features of relevant local regions and ensuring the acquisition of accurate and effective relevant features, thereby contributing to a reduction in computational complexity and improving network reasoning speed.

[0069] Figure 6 is a flowchart of step 2031b according to one exemplary embodiment of the present disclosure.

[0070] In one selectable example, step 2031b, which determines the dynamic object task query features using a pre-trained second decoding network based on the first bird's-eye view features and the initial dynamic object query features, includes the following steps:

[0071] In step 20311b, the third query tensor, third key tensor, and third value tensor are determined based on the initial dynamic object query features.

[0072] In step 20312b, the second self-attention result is determined based on the third query tensor, the third key tensor, and the third value tensor, using the second self-attention network of the first decoder in the second decoding network.

[0073] In step 20313b, the second intermediate result is determined based on the second self-attention result and the initial dynamic object query features, using the second additive normalization network of the first decoder in the second decoding network.

[0074] In step 20314b, the fourth query tensor is determined based on the second intermediate result.

[0075] In step 20315b, the fourth key tensor and the fourth value tensor are determined based on the first bird's-eye view feature.

[0076] In step 20316b, the second cross-attention result is determined using the second deformable cross-attention network of the first decoder in the second decoding network, based on the fourth query tensor, the fourth key tensor, and the fourth value tensor.

[0077] In step 20317b, the dynamic object task query features are determined based on the second cross-attention result and the second intermediate result.

[0078] The specific operating principles of steps 20311b to 20317b are the same as or similar to those of steps 20311a to 20317a described above. The difference is that step 20311b is based on the initial dynamic object query features rather than the initial static element query features in step 20311a, and this will not be explained here. Based on this, the network structure of the second decoding network is the same as or similar to that of the first decoding network, and this will not be explained here.

[0079] In one selectable example, steps 20311b to 20317b above may be performed by the processor calling corresponding instructions stored in memory, or by a second processing unit executed by the processor.

[0080] Figure 7 is a flowchart of step 2031c according to one exemplary embodiment of the present disclosure.

[0081] In one selectable embodiment, step 2031c, which determines the motion trajectory task query features using a pre-trained third decoding network based on static element task query features and initial motion trajectory query features, includes the following steps:

[0082] In step 20311c, the fifth query tensor, the fifth key tensor, and the fifth value tensor are determined based on the initial motion trajectory query features.

[0083] In step 20312c, the third self-attention result is determined using the third self-attention network of the first decoder in the third decoding network, based on the fifth query tensor, the fifth key tensor, and the fifth value tensor.

[0084] In step 20313c, the third intermediate result is determined based on the third self-attention result and the initial motion trajectory query features, using the third additive normalization network of the first decoder in the third decoding network.

[0085] In step 20314c, the sixth query tensor is determined based on the third intermediate result.

[0086] The specific operating principles of steps 20311c to 20314c are the same as or similar to those of steps 20311a to 20314a described above, and therefore, an explanation is omitted here.

[0087] In step 20315c, the sixth key tensor and the sixth value tensor are determined based on the static element task query features.

[0088] Step 6 key The tensor and the sixth value tensor are obtained by mapping them based on the static element task query features obtained in the example described above, and the mapping principle can be found in the previously mentioned content.

[0089] In step 20316c, the third cross-attention result is determined based on the sixth query tensor, the sixth key tensor, and the sixth value tensor, using the first cross-attention network of the first decoder in the third decoding network.

[0090] Here, the first cross-attention network can be any feasible cross-attention network, and specifically can be configured according to the actual needs. For example, the first cross-attention network may be a cross-attention network structure in a typical Vision Transformer, and is not specifically limited.

[0091] In step 20317c, the motion trajectory task query features are determined based on the third cross-attention result and the third intermediate result.

[0092] The specific operating principle of this step can be found in step 20317a described above, so we will omit the explanation here.

[0093] Exemplary, Figure 8 is a schematic diagram of the structure of a third decoding network according to one exemplary embodiment of the present disclosure. Here, the initial motion trajectory query feature includes a plurality of trajectory queries (each trajectory query represents an initial trajectory query feature corresponding to a dynamic object), Q5, K5, and V5 represent the fifth query tensor, fifth key tensor, and fifth value tensor, respectively, and Q6, K6, and V6 represent the sixth query tensor, sixth key tensor, and sixth value tensor, respectively. The meaning and inference process of the other codes can be found in the previously mentioned content, so an explanation is omitted here.

[0094] In one selectable example, steps 20311c to 20317c above may be performed by the processor calling corresponding instructions stored in memory, or by a third processing unit executed by the processor.

[0095] The embodiments of this disclosure achieve self-interaction of motion trajectory query features by a self-attention network in the third decoding network, enabling the capture of internal relationships between motion trajectory query features. Furthermore, they enable interaction with static element task query features in the cross-attention network to effectively capture motion trajectory-related features of dynamic objects. This implicitly grasps surrounding static information (e.g., surrounding road information) and contributes to providing accurate and effective feature data for accurately predicting more rational future motion trajectories of dynamic objects, thereby realizing end-to-end motion trajectory prediction based on multi-view image features.

[0096] Figure 9 is a flowchart of an image processing method according to yet another exemplary embodiment of the present disclosure.

[0097] In one selectable example, prior to step 2031c, which determines the motion trajectory task query features using a pre-trained third decoding network based on static element task query features and initial motion trajectory query features, the following steps are included:

[0098] In step 301, the initial motion trajectory query features are determined based on the dynamic object task query features and modality query features, and the modality query features include a first modality query feature corresponding to each of at least one modality, and the first modality query features corresponding to a modality are used to represent a type of motion tendency of the dynamic object.

[0099] Here, modality query features can be set according to actual needs. For example, they can be initialized and obtained based on a third initialization rule. Since modality query features can represent the motion tendencies of dynamic objects, initial motion trajectory query features can be determined in conjunction with dynamic object task query features that represent the dynamic object position.

[0100] For example, N3 D-dimensional modality queries (where each modality query represents a first modality query feature corresponding to a modality) can be randomly initialized and used as modality query features. N3 can be set according to the actual needs. The modality query features are merged with the dynamic object task query features to form N1 × N3 D-dimensional trajectory queries, which are used as initial motion trajectory query features. In other words, each dynamic object has N3 modalities representing its N3 motion tendencies.

[0101] In one possible example, step 301 may be performed by the processor calling a corresponding instruction stored in memory, or it may be performed by a fourth processing unit executed by the processor.

[0102] Embodiments of this disclosure obtain initial motion trajectory query features based on the fusion of newly acquired dynamic object task query features and modality query features, thereby allowing the initial motion trajectory query features to include task query features corresponding to each dynamic object and query features for multiple modalities of each dynamic object, where different modalities can focus on different types of future motion (e.g., high-speed straight, low-speed straight, left turn, right turn, etc.), contributing to providing effective data support for decoding the motion trajectory task query features by a subsequent third decoding network.

[0103] Figure 10 is a flowchart of step 301 according to one exemplary embodiment of the present disclosure.

[0104] In one selectable example, the dynamic object task query feature includes a task query feature for at least one dynamic object, and step 301, which determines the initial motion trajectory query feature based on the dynamic object task query feature and modality query feature, includes the following steps:

[0105] In step 3011, for each task query feature corresponding to a dynamic object, a first number of this task query feature is determined based on this task query feature, which is the number of first modality query features included in the modality query feature.

[0106] Here, for each dynamic object, a first modality query feature of the first number can be assigned to represent the first number of motion tendencies of this dynamic object. Therefore, the task query features and modality query features of this object can be merged, and the number of task query features of this dynamic object can be transformed to be the same as the number of modality query features. Thus, the first number of this task query feature can be determined based on the task query features of this dynamic object.

[0107] For example, a modality query feature may contain N3 D-dimensional modality queries. In this case, for each dynamic object's D-dimensional task query feature (which can be called a task query), this task query feature can be duplicated N3 times to obtain N3 identical D-dimensional task query features.

[0108] In one possible example, step 3011 may be performed by the processor calling a corresponding instruction stored in memory, or by a fourth processing unit executed by the processor.

[0109] In step 3012, the first task query feature is added to the first modality query feature corresponding to each modality in the modality query feature to obtain the initial trajectory query feature corresponding to this dynamic object.

[0110] For example, N3 identical task query features can be added together with N3 first modality query features (modality queries) to form N3 motion trajectory query features.

[0111] In one possible example, step 3012 may be performed by the processor calling a corresponding instruction stored in memory, or it may be performed by a fourth processing unit executed by the processor.

[0112] In step 3013, the initial motion trajectory query features are determined based on the initial trajectory query features corresponding to each dynamic object.

[0113] Exemplary, Figure 11 is a schematic diagram of the principle for determining initial motion trajectory query features according to one exemplary embodiment of the present disclosure. In this example, the number of dynamic objects is N1, task query represents the task query features corresponding to the dynamic objects, the number of first modality query features (modality queries) for each dynamic object is N3, and the final obtained initial motion trajectory query features include N1 × N3 D-dimensional initial trajectory query features.

[0114] In one possible example, step 3013 may be performed by the processor calling a corresponding instruction stored in memory, or by a fourth processing unit executed by the processor.

[0115] The embodiments of this disclosure obtain initial motion trajectory query features by fusing dynamic object task query features and modality query features, so that the initial motion trajectory query features include dynamic object task query-related features and dynamic object motion tendency-related information, and furthermore, accurate and effective motion trajectory task query features can be obtained by a third decoding network that has been trained and acquired.

[0116] Step 202, which determines the first bird's-eye view feature based on the first image feature corresponding to each viewpoint in one selectable example, includes the following steps:

[0117] In step 2021, the first bird's-eye view feature is determined based on the first image feature corresponding to each viewpoint, the initial bird's-eye view query feature, and the previous frame bird's-eye view feature obtained before the first bird's-eye view feature.

[0118] Here, the initial bird's-eye view query feature is an initialized feature in the bird's-eye view, its size can represent the size of the bird's-eye view, and the specific initialization rule can be set according to the actual needs. For example, H × W D-dimensional bird's-eye view queries can be initialized based on the required size of the bird's-eye view, and these will be the initial bird's-eye view query feature, where H and W are the height and width of the bird's-eye view, respectively, and D is the feature dimension of each bird's-eye view query. Each bird's-eye view query can correspond to a set of 3D coordinates in physical space, for example, 3D coordinates in a world coordinate system with the vehicle as the origin, and the specific details can be set according to the actual needs. The previous frame bird's-eye view feature is a bird's-eye view feature acquired in the current previous image processing flow, and its specific processing flow is the same as that of the first bird's-eye view feature.

[0119] In one possible example, step 2021 may be performed by the processor calling a corresponding instruction stored in memory, or by a first decision unit executed by the processor.

[0120] Since the previous frame bird's-eye view features include relevant historical information such as static elements and dynamic objects in the image collected in the previous frame, combining the previous frame bird's-eye view features, initial bird's-eye view query features, and first image features not only enables the determination of relevant features of static elements and dynamic objects in the currently processed image, but also enables the realization of changes in dynamic objects compared to the previous frame, thereby facilitating the tracking of dynamic objects.

[0121] Figure 12 is a flowchart of step 2021 according to one exemplary embodiment of the present disclosure.

[0122] Step 2021 for determining the first bird's-eye view feature in one selectable example, based on the first image feature corresponding to each viewpoint, the initial bird's-eye view query feature, and the previous frame bird's-eye view feature obtained before the first bird's-eye view feature, includes the following steps:

[0123] In step 20211, the time-series self-attention result is determined using the time-series self-attention network of the first encoder in the pre-trained encoder network, based on the previous frame bird's-eye view features and the initial bird's-eye view query features.

[0124] Here, the encoder network is used to obtain the first bird's-eye view feature in the bird's-eye view by encoding the first image feature corresponding to each extracted viewpoint based on the previous frame bird's-eye view feature and the initial bird's-eye view query feature. The encoder network may include one or more encoders, each encoder may include a time-series self-attention network, which uses the initial bird's-eye view query feature to find the corresponding position in the previous frame bird's-eye view of the vehicle's history as a reference position based on the vehicle's motion, and is subsequently used to extract the feature corresponding to this reference position region in the first image feature. The first bird's-eye view feature of the current frame can be obtained by the continuous encoding of multiple encoders.

[0125] In step 20212, the fourth intermediate result is determined using the fourth additive normalization network of the first encoder, based on the time-series self-attention results and the initial bird's-eye view query features.

[0126] The specific principles of the fourth-order normalized network can be found in the previously mentioned content, so we will omit the explanation here.

[0127] In step 20213, the spatial cross-attention result is determined using the spatial cross-attention network in the first encoder, based on the first image feature and the fourth intermediate result corresponding to each viewpoint.

[0128] Here, the spatial cross-attention network can uniformly sample the fourth intermediate result in height to obtain a set of 3D coordinates, then map the 3D coordinates to the corresponding positions in the first image feature corresponding to each viewpoint based on the intrinsic and extrinsic parameters of the camera (webcam), and further extract the features of the corresponding positions in the first image feature based on deformable convolution, which are used in subsequent processing to obtain the first bird's-eye view feature of the current frame.

[0129] In step 20214, the first bird's-eye view feature is determined based on the spatial cross-attention results and the fourth intermediate result.

[0130] Here, in each encoder, in addition to the time-series self-attention network, the fourth additive normalization network, and the spatial cross-attention network described above, several other related networks are further included. For example, an additive normalization network, a feedforward network, and another additive normalization network may be further included after the spatial cross-attention network, and these can be specifically configured according to the actual needs. Therefore, after obtaining the spatial cross-attention result, the spatial cross-attention result and the fourth intermediate result can be added and normalized, and then the encoded result of the first encoder can be obtained by other related networks. If the encoder network includes multiple encoders, the encoded result output from the first encoder can be further encoded by each subsequent encoder, and finally the first bird's-eye view feature can be obtained.

[0131] Exemplary, Figure 13 is a schematic diagram of the network structure of an encoder network according to one exemplary embodiment of the present disclosure. Here, BEV B(t-1) represents the previous frame bird's-eye view feature, BEV queries Q represents the initial bird's-eye view query feature, Temporal Self-Attention represents the time-series self-attention network, Spatial Cross-Attention represents the spatial cross-attention network, and the meanings of the other symbols can be found in the previously mentioned content. The previous frame bird's-eye view feature BEV B(t-1) and the initial bird's-eye view query feature BEV queries Q interact in a time-series self-attention network. Depending on the vehicle's motion, the network finds the corresponding position in the previous frame bird's-eye view of the initial bird's-eye view query feature as a reference position, obtains a time-series self-attention result, adds and normalizes the time-series self-attention result and the initial bird's-eye view query feature to obtain a fourth intermediate result. The fourth intermediate result and the first image features of each viewpoint undergo a spatial cross-attention operation in a spatial cross-attention network. The spatial cross-attention network extracts the corresponding position feature from the first image feature based on the reference position, based on deformable convolution, obtains a spatial cross-attention result, adds and normalizes the spatial cross-attention result and the fourth intermediate result to obtain a fifth intermediate result. The fifth intermediate result obtains the encoded result of the first encoder via a feed-forward network and an add-normalization network. This encoded result then undergoes encoding by subsequent encoders to obtain the final encoded result, the first bird's-eye view feature. It is understood that spatial position coding can be embedded for each first image feature in the network reasoning process, and this disclosure is not limited thereto.

[0132] In one possible example, steps 20211 to 20214 above may be performed by the processor calling corresponding instructions stored in memory, or by a first decision unit executed by the processor.

[0133] The embodiments of this disclosure achieve positional matching between the initial bird's-eye view query features of the current frame and the bird's-eye view of the previous frame using a time-series self-attention network in the encoder network, establish a time-series correlation between the current frame and the previous frame, and further extract local features near the reference position from the first image features based on a spatial cross-attention network. This contributes to reducing the computational load in extracting effective features, thereby improving image processing efficiency.

[0134] In one selectable example, step 204, which determines the task processing result corresponding to each task query feature based on each of at least one task query feature, includes the following steps:

[0135] In step 2041a, the static element detection result is determined using a pre-trained static element detection head network based on the static element task query features.

[0136] Here, the static element detection head network can be any feasible head network, such as a head network based on a multilayer perceptron.

[0137] In one possible example, step 2041a may be performed by the processor calling a corresponding instruction stored in memory, or by a second decision unit executed by the processor.

[0138] In step 2041b, the dynamic object detection result is determined using a pre-trained dynamic object detection head network based on the dynamic object task query features.

[0139] Here, the dynamic object detection head network can use any feasible head network, such as a head network based on a multilayer perceptron.

[0140] In one possible example, step 2041b may be performed by the processor calling a corresponding instruction stored in memory, or by a third decision unit executed by the processor.

[0141] In step 2041c, the motion trajectory prediction result is determined using a pre-trained motion trajectory prediction head network based on the motion trajectory task query features.

[0142] Here, the motion trajectory prediction head network can use any feasible head network, such as a head network based on a multilayer perceptron.

[0143] In one possible example, step 2041c may be performed by the processor calling a corresponding instruction stored in memory, or by a fourth decision unit executed by the processor.

[0144] The embodiments of this disclosure acquire a first bird's-eye view feature by encoding based on a first image feature corresponding to each viewpoint, decode task query features for different tasks based on the first bird's-eye view feature and decoding networks corresponding to different tasks, and further acquire task processing results corresponding to different tasks based on head networks corresponding to different tasks, thereby realizing end-to-end multitask processing based on surround view images of multiple viewpoints in multiple frames, contributing to improved task processing efficiency, avoiding or reducing reliance on offline-generated high-precision maps and laser radar, and enabling static element detection, 3D detection of dynamic objects, and prediction of motion trajectories, thereby contributing to cost reduction.

[0145] In one selectable example, Figure 14 is a schematic diagram of the overall structure of a network model for image processing according to one exemplary embodiment of the present disclosure. Based on a multi-viewpoint image to be processed, a feature extraction network is used to obtain a first image feature corresponding to each viewpoint, each first image feature passes through an encoder network to obtain a first bird's-eye view feature, the first bird's-eye view feature passes through a first decoding network to obtain a static element task query feature, and a static element detection head network is used to obtain static element detection results, the first bird's-eye view feature passes through a second decoding network to obtain a dynamic object task query feature, and a dynamic object detection head network is used to obtain dynamic object detection results, the dynamic object task query results and modality query features are fused to obtain an initial motion trajectory query feature, and based on the initial motion trajectory query feature and the static element task query feature obtained by the first decoding network, a third decoding network is used to obtain a motion trajectory task query feature, and a motion trajectory prediction head network is used to obtain a motion trajectory prediction result.

[0146] In one possible example, the network model can be obtained by pre-training. If the network model includes multiple tasks simultaneously, the tasks may be trained together, or they may be trained individually first and then comprehensively. Specifically, this can be configured according to the actual needs. For example, to ensure better motion trajectory prediction performance, the static element task and dynamic object task may be trained first to obtain a base model, and then the three tasks may be trained together based on the base model. The specific training principles will not be explained here.

[0147] The embodiments of this disclosure can perform task processing of static elements, dynamic objects, and motion trajectories using only multi-view images, and can acquire richer environmental information compared to laser radar, while also having lower hardware costs and being easier to deploy. Furthermore, the embodiments of this disclosure can achieve online generation of static map information through static element detection, eliminating the need to rely on offline high-precision maps, thus broadening the range of applicable scenarios. In addition, the embodiments of this disclosure do not require explicit dynamic object tracking, contributing to a reduction in the complexity of model calculations and avoiding the impact of tracking module errors on subsequent processing, thereby further improving the accuracy of task processing results.

[0148] In one selectable example, the viewpoint images and data collected by the laser radar can be further feature-fused and used for end-to-end single-task or multi-task processing, improving the richness of feature information and thereby contributing to further improvements in model performance.

[0149] Each of the embodiments or selectable examples described herein may be implemented individually or in any combination, provided they are not inconsistent, and can be specifically configured according to actual needs, and is not limited to these embodiments.

[0150] Any one of the image processing methods relating to the embodiments of this disclosure may be performed by any suitable device having data processing capabilities, including but not limited to terminal devices and servers. Alternatively, any one of the image processing methods relating to the embodiments of this disclosure may be performed by a processor, for example, by the processor calling a corresponding instruction stored in memory to perform any one of the image processing methods referred to in the embodiments of this disclosure. Further explanation is omitted below.

[0151] As those skilled in the art will understand, all or part of the steps of the embodiment of the above method can be completed by a program instructing the relevant hardware, the aforementioned program can be stored in a computer-readable storage medium, the program, when executed, performs the steps including the embodiment of the above method, the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks or optical disks.

[0152] Exemplary device Figure 15 is a schematic diagram of the structure of an image processing apparatus according to one exemplary embodiment of the present disclosure. The apparatus of this embodiment can be used to implement an embodiment of a corresponding method of the present disclosure, and the apparatus shown in Figure 15 includes a first processing module 501, a second processing module 502, a third processing module 503, and a fourth processing module 504.

[0153] The first processing module 501 is used to determine a first image feature corresponding to each viewpoint based on the image to be processed corresponding to each viewpoint among at least one viewpoint; the second processing module 502 is used to determine a first bird's-eye view feature based on the first image feature corresponding to each viewpoint; the third processing module 503 is used to determine at least one type of task query feature from static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view feature; and the fourth processing module 504 is used to determine a task processing result corresponding to each task query feature based on each of the at least one type of task query feature.

[0154] Figure 16 is a schematic diagram of the structure of an image processing apparatus according to another exemplary embodiment of the present disclosure.

[0155] In one selectable example, the third processing module 503 is: A first processing unit 5031 for determining static element task query features using a first decoding network that has been pre-trained and acquired based on a first bird's-eye view feature and an initial static element query feature, wherein the initial static element query features include an initial query feature corresponding to each static element of at least one static element.

[0156] In one selectable example, the third processing module 503 is: A second processing unit 5032 for determining dynamic object task query features using a second decoding network that has been pre-trained and acquired based on a first bird's-eye view feature and an initial dynamic object query feature, wherein the initial dynamic object query feature includes an initial query feature corresponding to each dynamic object among at least one dynamic object.

[0157] In one selectable example, the third processing module 503 is: A third processing unit 5033 for determining motion trajectory task query features using a pre-trained third decoding network based on static element task query features and initial motion trajectory query features, wherein the initial motion trajectory query features include initial trajectory query features corresponding to each dynamic object among at least one dynamic object.

[0158] In one selectable example, the third processing module 503 may include at least two of the first processing unit 5031, the second processing unit 5032, and the third processing unit 5033, and can be configured specifically according to the actual needs.

[0159] In one selectable example, the first processing unit 5031 specifically, Based on the initial static element query features, the first query tensor, first key tensor, and first value tensor are determined; based on the first query tensor, first key tensor, and first value tensor, the first self-attention result is determined using the first self-attention network of the first decoder in the first decoding network; based on the first self-attention result and the initial static element query features, the first intermediate result is determined using the first additive normalization network of the first decoder in the first decoding network; based on the first intermediate result, the second query tensor is determined; based on the first bird's-eye view features, the second key tensor and second value tensor are determined; based on the second query tensor, second key tensor, and second value tensor, the first cross-attention result is determined using the first deformable cross-attention network of the first decoder in the first decoding network; and based on the first cross-attention result and the first intermediate result, the static element task query features are determined.

[0160] In one selectable example, the second processing unit 5032 specifically, Based on the initial dynamic object query features, the third query tensor, third key tensor, and third value tensor are determined; based on the third query tensor, third key tensor, and third value tensor, the second self-attention result is determined using the second self-attention network of the first decoder in the second decoding network; based on the second self-attention result and the initial dynamic object query features, the second intermediate result is determined using the second additive normalization network of the first decoder in the second decoding network; based on the second intermediate result, the fourth query tensor is determined; based on the first bird's-eye view features, the fourth key tensor and fourth value tensor are determined; based on the fourth query tensor, fourth key tensor, and fourth value tensor, the second cross-attention result is determined using the second deformable cross-attention network of the first decoder in the second decoding network; and based on the second cross-attention result and the second intermediate result, the dynamic object task query features are determined.

[0161] In one selectable example, the third processing unit 5033 specifically, Based on the initial motion trajectory query features, the fifth query tensor, fifth key tensor, and fifth value tensor are determined. Based on the fifth query tensor, fifth key tensor, and fifth value tensor, the third self-attention result is determined using the third self-attention network of the first decoder in the third decoding network. Based on the third self-attention result and the initial motion trajectory query features, the third intermediate result is determined using the third additive normalization network of the first decoder in the third decoding network. Based on the third intermediate result, the sixth query tensor is determined. Based on the static element task query features, the sixth key tensor and sixth value tensor are determined. Based on the sixth query tensor, sixth key tensor, and sixth value tensor, the third cross-attention result is determined using the first cross-attention network of the first decoder in the third decoding network. Based on the third cross-attention result and the third intermediate result, motion trajectory task query features are determined.

[0162] Figure 17 is a schematic diagram of the structure of a third processing module 503 according to one exemplary embodiment of the present disclosure.

[0163] In one selectable example, the third processing module 503 is: A fourth processing unit 5034 for determining initial motion trajectory query features based on dynamic object task query features and modality query features, wherein the modality query features include a first modality query feature corresponding to each of at least one modality, and the first modality query features corresponding to a modality are used to represent a type of motion tendency of a dynamic object.

[0164] In one selectable example, the dynamic object task query feature includes at least one dynamic object task query feature, and the fourth processing unit 5034 specifically, For each dynamic object, a first number of this task query feature is determined based on this task query feature, which is the number of first modality query features included in the modality query feature. This first number of task query feature is then added to the first modality query feature corresponding to each modality in the modality query feature to obtain the initial trajectory query feature corresponding to this dynamic object. This is then used to determine the initial motion trajectory query feature based on the initial trajectory query feature corresponding to each dynamic object.

[0165] In one selectable example, the second processing module 502 is: The system includes a first decision unit 5021 for determining the first bird's-eye view feature based on the first image feature corresponding to each viewpoint, the initial bird's-eye view query feature, and the previous frame bird's-eye view feature acquired before the first bird's-eye view feature.

[0166] In one selectable example, the first decision unit 5021 specifically, Based on the previous frame bird's-eye view features and the initial bird's-eye view query features, the time-series self-attention results are determined using the time-series self-attention network of the first encoder in the pre-trained encoder network. Based on the time-series self-attention results and the initial bird's-eye view query features, the fourth intermediate result is determined using the fourth additive normalization network of the first encoder. Based on the first image features corresponding to each viewpoint and the fourth intermediate result, the spatial cross-attention results are determined using the spatial cross-attention network of the first encoder. Based on the spatial cross-attention results and the fourth intermediate result, the first bird's-eye view features are determined.

[0167] In one selectable example, the fourth processing module 504 is: The system includes a second decision unit 5041 for determining static element detection results using a pre-trained static element detection head network based on static element task query features, a third decision unit 5042 for determining dynamic object detection results using a pre-trained dynamic object detection head network based on dynamic object task query features, and a fourth decision unit 5043 for determining motion trajectory prediction results using a pre-trained motion trajectory prediction head network based on motion trajectory task query features.

[0168] In one selectable example, each of the above units of the Disclosure may be further subdivided according to actual needs, for example, by dividing the unit into multiple subunits, which may be configured specifically according to actual needs.

[0169] Each of the embodiments or selectable examples described herein may be implemented individually or in any combination, provided they are not inconsistent, and can be specifically configured according to actual needs, and is not limited to these embodiments.

[0170] The beneficial technical effects corresponding to the exemplary embodiment of this apparatus can be found by referring to the corresponding beneficial technical effects of the exemplary method described above, and are therefore omitted from this explanation.

[0171] Exemplary electronic device Figure 18 is a schematic diagram of the structure of one application embodiment of the electronic device of the present disclosure. In this embodiment, the electronic device 10 includes one or more processors 11 and memory 12.

[0172] The processor 11 may be a central processing unit (CPU) or another form of processing unit having data processing and / or instruction execution functions, and can control other components in the electronic device 10 to perform desired functions.

[0173] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. The computer-readable storage media can store one or more computer program instructions, and the processor 11 can execute the program instructions to implement the methods of each embodiment of the present disclosure and / or other desired functions. The computer-readable storage media can store various contents, such as input signals, signal components, and noise components.

[0174] In one example, the electronic device 10 may further include an input device 13 and an output device 14, and these components are connected to each other via a bus system and / or other forms of connection mechanisms (not shown).

[0175] In addition, this input device 13 may further include, for example, a keyboard, a mouse, and the like.

[0176] This output device 14 can output various types of information to the outside, and may include, for example, a display, speaker, printer, communication network, and remote output devices connected thereto.

[0177] Naturally, for the sake of simplification, Figure 18 shows only some of the components of the electronic device 10 relevant to this disclosure, omitting components such as buses and input / output interfaces. Beyond this, the electronic device 10 may further include any other appropriate components depending on the specific application.

[0178] Exemplary computer program products and computer-readable storage media The embodiments of this disclosure may also include, in addition to the above-described methods and apparatus, computer program products that, when executed by a processor, cause the processor to perform steps in the various embodiments of the disclosure described in the above-described “Exemplary Methods” portion of this specification.

[0179] The computer program product can produce program code for performing the operations of the embodiments of this disclosure in any combination of one or more programming languages, the programming languages ​​include object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as the C language or similar programming languages. The program code may run entirely on the user's computing device, partially on the user's device, run as separate software packages, run in part on the user's computing device and in part on a remote computing device, or run entirely on a remote computing device or server.

[0180] In addition, embodiments of the present disclosure may further include a computer-readable storage medium that stores computer program instructions, which, when executed by a processor, cause the processor to perform the steps in the various embodiments of the present disclosure described in the “Exemplary Methods” section above.

[0181] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any combination thereof. More specific examples (non-exclusive list) of readable storage media include electrical connections with one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0182] While the basic principles of this disclosure have been explained above with reference to specific examples, the advantages, advantages, and effects mentioned herein are not limited to those mentioned above, but are merely illustrative, and these advantages, advantages, and effects are not necessarily present in every example of this disclosure. Furthermore, the specific details disclosed above are not limited to those mentioned above, but are merely illustrative and intended to facilitate understanding, and these details do not necessarily limit this disclosure to being realized by those specific details.

[0183] Those skilled in the art can make various modifications and alterations to this disclosure without departing from the spirit and scope of this disclosure. Thus, if such modifications and alterations of this disclosure fall within the scope of the claims of this disclosure and the equivalent art, this disclosure is intended to include such modifications and alterations.

Claims

1. A step of determining a first image feature of a viewpoint based on an image to be processed from at least one viewpoint, A step of determining a first bird's-eye view feature based on the first image feature of the aforementioned viewpoint, The steps include determining at least one task query feature from among static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view feature, The process includes the step of determining the task processing result of the task query feature based on at least one of the task query features, A method of image processing performed by an image processing device.

2. The step of determining at least one task query feature from among static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view feature is: A step of determining the static element task query features using a first decoding network that has been pre-trained and acquired based on the first bird's-eye view features and the initial static element query features, wherein the initial static element query features include an initial query feature of at least one static element, and / or A step of determining the dynamic object task query features using a second decoding network that has been pre-trained and acquired based on the first bird's-eye view features and the initial dynamic object query features, wherein the initial dynamic object query features include the initial query features of at least one dynamic object, and / or The method according to claim 1, comprising the step of determining the motion trajectory task query features using a third decoding network that has been pre-trained and acquired based on the static element task query features and the initial motion trajectory query features, wherein the initial motion trajectory query features include the initial trajectory query features of at least one dynamic object.

3. The step of determining the static element task query features using a first decoding network that has been pre-trained and acquired based on the first bird's-eye view features and initial static element query features is as follows: The steps include determining a first query tensor, a first key tensor, and a first value tensor based on the initial static element query features, The steps include determining a first self-attention result based on the first query tensor, the first key tensor, and the first value tensor, using the first self-attention network of the first decoder in the first decoding network, The steps include determining a first intermediate result based on the first self-attention result and the initial static element query features, using the first additive normalization network of the first decoder in the first decoding network, The steps include determining the second query tensor based on the first intermediate result, The steps include determining the second key tensor and the second value tensor based on the first bird's-eye view features, The steps include determining a first cross-attention result using a first deformable cross-attention network of the first decoder in the first decoding network, based on the second query tensor, the second key tensor, and the second value tensor, The steps include determining the static element task query features based on the first cross-attention result and the first intermediate result, and / or The step of determining the dynamic object task query features using a second decoding network that has been pre-trained and acquired based on the first bird's-eye view features and the initial dynamic object query features is as follows: The steps include determining a third query tensor, a third key tensor, and a third value tensor based on the initial dynamic object query characteristics, The steps include determining the second self-attention result based on the third query tensor, the third key tensor, and the third value tensor, using the second self-attention network of the first decoder in the second decoding network, The steps include determining a second intermediate result based on the second self-attention result and the initial dynamic object query features, using the second additive normalization network of the first decoder in the second decoding network, Based on the second intermediate result, the fourth query tensor is determined, The steps include determining the fourth key tensor and the fourth value tensor based on the first bird's-eye view features, The steps include determining the second cross-attention result using the second deformable cross-attention network of the first decoder in the second decoding network, based on the fourth query tensor, the fourth key tensor, and the fourth value tensor, The method according to claim 2, comprising the step of determining the dynamic object task query features based on the second cross-attention result and the second intermediate result.

4. The step of determining the motion trajectory task query features using a third decoding network that has been pre-trained and acquired based on the static element task query features and initial motion trajectory query features is as follows: The steps include determining the fifth query tensor, the fifth key tensor, and the fifth value tensor based on the initial motion trajectory query features, The steps include determining the third self-attention result using the third self-attention network of the first decoder in the third decoding network, based on the fifth query tensor, the fifth key tensor, and the fifth value tensor, Based on the third self-attention result and the initial motion trajectory query features, the third intermediate result is determined using the third additive normalization network of the first decoder in the third decoding network. Based on the third intermediate result, the steps include determining the sixth query tensor, The steps include determining the sixth key tensor and the sixth value tensor based on the static element task query features, The steps include determining the third cross-attention result using the first cross-attention network of the first decoder in the third decoding network, based on the sixth query tensor, the sixth key tensor, and the sixth value tensor, The method according to claim 2, comprising the step of determining the motion trajectory task query features based on the third cross-attention result and the third intermediate result.

5. Before the step of determining the motion trajectory task query features using a third decoding network that has been pre-trained and acquired based on the static element task query features and initial motion trajectory query features, The method according to claim 2, further comprising the step of determining the initial motion trajectory query feature based on the dynamic object task query feature and modality query feature, wherein the modality query feature includes a first modality query feature of at least one modality, and the first modality query feature of the modality is used to represent a type of motion tendency of a dynamic object.

6. The dynamic object task query feature includes a task query feature for at least one dynamic object, and the step of determining the initial motion trajectory query feature based on the dynamic object task query feature and modality query feature is: The steps include determining a first number of task query features, which is the number of first modality query features included in the modality query features, based on the task query features of the dynamic object, The steps include adding the first number of task query features to the first modality query features of the modality in the modality query features to obtain the initial trajectory query features of the dynamic object, The method according to claim 5, comprising the step of determining the initial motion trajectory query features based on the initial trajectory query features of the dynamic object.

7. The step of determining a first bird's-eye view feature based on the first image feature of the viewpoint is: The method according to claim 1, comprising the step of determining the first bird's-eye view feature based on the first image feature of the viewpoint, the initial bird's-eye view query feature, and the previous frame bird's-eye view feature acquired before the first bird's-eye view feature.

8. The step of determining the first bird's-eye view feature based on the first image feature of the viewpoint, the initial bird's-eye view query feature, and the previous frame bird's-eye view feature acquired before the first bird's-eye view feature is: Based on the aforementioned previous frame bird's-eye view features and the aforementioned initial bird's-eye view query features, the time-series self-attention result is determined using the time-series self-attention network of the first encoder in the pre-trained encoder network. The steps include determining a fourth intermediate result using the fourth additive normalization network of the first encoder, based on the time-series self-attention results and the initial bird's-eye view query features, The steps include determining the spatial cross-attention result using the spatial cross-attention network in the first encoder, based on the first image features and the fourth intermediate result of the viewpoint, The method according to claim 7, comprising the step of determining the first bird's-eye view feature based on the spatial cross-attention result and the fourth intermediate result.

9. The step of determining the task processing result of the task query feature based on the at least one task query feature is: The steps include determining the static element detection result using a static element detection head network that has been pre-trained and acquired based on the static element task query features, The steps include determining the dynamic object detection result using a pre-trained dynamic object detection head network based on the dynamic object task query features, The method according to claim 1, comprising the step of determining a motion trajectory prediction result using a motion trajectory prediction head network that has been pre-trained and acquired based on the motion trajectory task query features.

10. A first processing module for determining a first image feature of a viewpoint based on a processing target image of at least one viewpoint, A second processing module for determining a first bird's-eye view feature based on the first image feature of the aforementioned viewpoint, A third processing module for determining at least one type of task query feature from among static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view feature, A fourth processing module for determining the task processing result of a task query feature based on at least one of the aforementioned task query features, is included. Image processing device.

11. A computer-readable storage medium storing a computer program for performing the image processing method described in any one of claims 1 to 9.

12. Processor and The processor includes a memory for storing executable instructions, The processor is an electronic device used to read and execute the executable instructions from the memory to realize the image processing method described in any one of claims 1 to 9.