Image processing method and device, electronic equipment and storage medium

By generating bird's-eye view features and extracting task query features through multi-view image processing, the problem of inaccurate environmental information caused by the lack of high-precision maps in autonomous driving is solved, and accurate environmental information acquisition and multi-task processing are achieved without high-precision maps.

CN115719476BActive Publication Date: 2026-07-21BEIJING HORIZON INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING HORIZON INFORMATION TECH CO LTD
Filing Date
2022-11-11
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In the field of autonomous driving, the accuracy of understanding environmental information from multi-view environmental images is poor in the absence of high-precision maps.

Method used

By using a multi-view image processing method, the image features of each viewpoint are determined, bird's-eye view features are generated, and a decoding network is used to extract static elements, dynamic objects, and motion trajectory task query features, thereby achieving end-to-end single-task or multi-task processing.

Benefits of technology

Even without high-precision maps, it can effectively obtain accurate information about the surrounding environment, improve versatility, and reduce costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115719476B_ABST
    Figure CN115719476B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a kind of processing method of image, apparatus, electronic equipment and storage medium, wherein, method includes: based on each view angle corresponding to the image to be processed of each view angle in at least one view angle, determine the first image feature corresponding to each view angle;Based on the first image feature corresponding to each view angle, determine the first bird's eye view feature;Based on the first bird's eye view feature, determine at least one task query feature in static element task query feature, dynamic object task query feature and motion trajectory task query feature;Based on each task query feature in at least one task query feature, determine the task processing result corresponding to each task query feature respectively.This embodiment of the present disclosure can realize end-to-end single task or multi-task processing only by relying on multi-view environmental image, even in the absence of high-precision map, accurate surrounding environment information can be effectively obtained, greatly improve versatility, and effectively reduce cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to computer vision technology, and in particular to an image processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] In the field of autonomous driving, how to efficiently understand environmental information by relying on multi-view environmental images is an extremely important technical problem. In related technologies, the understanding of the surrounding environment information is usually achieved by combining multi-view environmental images with high-precision maps. However, without high-precision maps, the accuracy of the environmental information obtained is easily poor. Summary of the Invention

[0003] To address the aforementioned technical problems, such as the poor accuracy of environmental information obtained without high-precision maps, this disclosure is proposed. Embodiments of this disclosure provide an image processing method, apparatus, electronic device, and storage medium.

[0004] According to one aspect of the present disclosure, an image processing method is provided, comprising: determining first image features corresponding to each of the at least one viewpoints based on an image to be processed corresponding to each of the viewpoints; determining first bird's-eye view features based on the first image features corresponding to each of the viewpoints; determining at least one task query feature among static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view features; and determining a task processing result corresponding to each of the at least one task query feature based on each of the task query features.

[0005] According to another aspect of the present disclosure, an image processing apparatus is provided, comprising: a first processing module, configured to determine a first image feature corresponding to each of the at least one viewpoints based on an image to be processed corresponding to each of the viewpoints; a second processing module, configured to determine a first bird's-eye view feature based on the first image feature corresponding to each of the viewpoints; a third processing module, configured to determine at least one task query feature among static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view feature; and a fourth processing module, configured to determine a task processing result corresponding to each of the at least one task query feature based on each of the task query features.

[0006] According to another aspect of the present disclosure, a computer-readable storage medium is provided, the storage medium storing a computer program for performing the image processing method described in any of the above embodiments of the present disclosure.

[0007] According to another aspect of the present disclosure, an electronic device is provided, the electronic device comprising: a processor; a memory for storing executable instructions of the processor; the processor being configured to read the executable instructions from the memory and execute the instructions to implement the image processing method described in any of the above embodiments of the present disclosure.

[0008] Based on the image processing method, apparatus, electronic device, and storage medium provided in the above embodiments of this disclosure, bird's-eye view features can be determined based on the image features corresponding to each viewpoint of the image to be processed from each viewpoint. At least one task query feature can be determined based on the bird's-eye view features. Then, based on each task query feature, the task processing result corresponding to each task can be obtained. End-to-end single-task or multi-task processing can be achieved by relying only on multi-view environmental images without the need for high-precision maps. This enables accurate surrounding environmental information to be effectively obtained even without high-precision maps, greatly improving versatility and effectively reducing costs.

[0009] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0010] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0011] Figure 1 This is an exemplary application scenario of the image processing method provided in this disclosure;

[0012] Figure 2 This is a schematic flowchart of an image processing method provided in an exemplary embodiment of this disclosure;

[0013] Figure 3 This is a schematic flowchart of an image processing method provided in another exemplary embodiment of this disclosure;

[0014] Figure 4 This is a flowchart illustrating step 2031a provided in an exemplary embodiment of this disclosure;

[0015] Figure 5 This is a schematic diagram of the network structure of a first decoding network provided in an exemplary embodiment disclosed in this paper;

[0016] Figure 6 This is a flowchart illustrating step 2031b provided in an exemplary embodiment of this disclosure;

[0017] Figure 7 This is a flowchart illustrating step 2031c provided in an exemplary embodiment of this disclosure;

[0018] Figure 8 This is a schematic diagram of the structure of a third decoding network provided in an exemplary embodiment of this disclosure;

[0019] Figure 9 This is a schematic flowchart of an image processing method provided in yet another exemplary embodiment of this disclosure;

[0020] Figure 10 This is a flowchart illustrating step 301 provided in an exemplary embodiment of this disclosure;

[0021] Figure 11 This is a schematic diagram illustrating the principle of determining the initial motion trajectory query features provided in an exemplary embodiment of this disclosure;

[0022] Figure 12 This is a flowchart illustrating step 2021 provided in an exemplary embodiment of this disclosure;

[0023] Figure 13 This is a schematic diagram of the network structure of an encoder network provided in an exemplary embodiment of this disclosure;

[0024] Figure 14 This is a schematic diagram of the overall structure of a network model for image processing provided in an exemplary embodiment of this disclosure;

[0025] Figure 15 This is a schematic diagram of the structure of an image processing apparatus provided in an exemplary embodiment of the present disclosure;

[0026] Figure 16 This is a schematic diagram of the structure of an image processing apparatus provided in another exemplary embodiment of the present disclosure;

[0027] Figure 17 This is a schematic diagram of the structure of the third processing module 503 provided in an exemplary embodiment of this disclosure;

[0028] Figure 18 This is a schematic diagram of the structure of an application embodiment of the electronic device disclosed herein. Detailed Implementation

[0029] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.

[0030] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of this disclosure.

[0031] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.

[0032] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.

[0033] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.

[0034] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.

[0035] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.

[0036] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.

[0037] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.

[0038] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0039] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0040] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.

[0041] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.

[0042] This disclosure outlines

[0043] In the process of realizing this disclosure, the inventors discovered that in the field of autonomous driving, how to efficiently understand environmental information by relying on multi-view environmental images is an extremely important technical problem. In related technologies, the understanding of surrounding environmental information is usually achieved by combining multi-view environmental images with high-precision maps. However, without high-precision maps, the accuracy of the obtained environmental information is easily poor.

[0044] Exemplary Overview

[0045] Figure 1 This is an exemplary application scenario of the image processing method provided in this disclosure.

[0046] In autonomous driving scenarios, images of the vehicle's surrounding environment can be acquired using onboard surround-view cameras (which may include cameras with multiple viewpoints) as images to be processed for each viewpoint. The image processing device of this disclosure executes the image processing method of this disclosure. Based on the images to be processed corresponding to each of the at least one viewpoint, first image features corresponding to each viewpoint can be determined. Based on the first image features corresponding to each viewpoint, a first bird's-eye view feature is determined. The first bird's-eye view feature is a bird's-eye view. Features in the grid coordinate system corresponding to the view (BEV) can be used to determine static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view features. Then, based on the static element task query features, dynamic object task query features, and motion trajectory task query features, the corresponding task processing results (such as, but not limited to, task processing result 1, task processing result 2, and task processing result 3) can be determined. For example, the static element task processing result can be determined based on the static element task query features (specifically, the static element detection result of the vehicle's surrounding environment in an autonomous driving scenario), the dynamic object task processing result can be determined based on the dynamic object task query features (specifically, the 3D target detection result), and the motion trajectory task processing result can be determined based on the motion trajectory task query features (specifically, the motion trajectory prediction result of dynamic objects). This enables end-to-end single-task or multi-task processing that relies solely on multi-view environmental images without the need for high-precision maps. This allows for the effective acquisition of accurate surrounding environment information even without high-precision maps, greatly improving versatility and effectively reducing costs. Static elements can include static object elements such as lane lines, zebra crossings, and curbs, while dynamic objects can include objects with motion attributes such as surrounding vehicles and pedestrians. Motion trajectory refers to the motion trajectory of dynamic objects.

[0047] It should be noted that the image processing method disclosed herein is not limited to the above-mentioned autonomous driving scenario. It can be applied to any other possible scenario according to actual needs, such as a security monitoring scenario of a certain area. By acquiring images from cameras at various perspectives, the bird's-eye view features of the area can be obtained, and end-to-end task processing of static elements, dynamic objects and / or the motion trajectory of dynamic objects in the area can be realized. Specific scenarios can be set according to actual needs.

[0048] Exemplary methods

[0049] Figure 2 This is a schematic flowchart illustrating an image processing method provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, specifically, for example, in-vehicle computing platforms. Figure 2 As shown, it includes the following steps:

[0050] Step 201: Based on the images to be processed corresponding to each viewpoint in at least one viewpoint, determine the first image features corresponding to each viewpoint.

[0051] The number of viewpoints can be set according to actual needs. For example, in autonomous driving scenarios, the number of viewpoints corresponds to the number of surround-view cameras installed on the vehicle, with each camera corresponding to one viewpoint. For instance, a four-way surround-view system consisting of a left front camera, a left rear camera, a right front camera, and a right rear camera includes four viewpoints, but the specific number is not limited. The first image features can be obtained using any feasible feature extraction method. For example, features can be extracted from each image to be processed based on a pre-trained feature extraction network to obtain the first image features corresponding to each viewpoint. The feature extraction network can be set according to actual needs; for example, a convolutional neural network can be used as the feature extraction network.

[0052] Step 202: Determine the first bird's-eye view features based on the first image features corresponding to each viewpoint.

[0053] The first bird's-eye view feature is the BEV feature in the grid coordinate system corresponding to the bird's-eye view. It can be obtained by encoding the first image features from each viewpoint based on a pre-trained encoder network. The encoder network can be set according to actual needs.

[0054] Step 203: Based on the features of the first bird's-eye view, determine at least one of the following task query features: static element task query features, dynamic object task query features, and motion trajectory task query features.

[0055] Among them, static element task query features are task query features related to static elements extracted from the first bird's-eye view features. Similarly, dynamic object task query features are task query features related to dynamic objects extracted from the first bird's-eye view features, and motion trajectory task query features are task query features related to the motion trajectory of dynamic objects extracted from the static element task query features. The specific type or types of task query features to be obtained can be set according to actual needs. For example, any one type, any two types, or all three types can be obtained simultaneously. For any type of task query feature, a pre-trained decoding network corresponding to that task can be used to decode the first bird's-eye view features. The specific decoding network can be set according to actual needs.

[0056] Step 204: Based on each task query feature in at least one task query feature, determine the task processing result corresponding to each task query feature.

[0057] For any given task, a head network can be set up and trained to obtain the trained head network. This trained head network is then used to project the task query features corresponding to that task onto the output, thereby obtaining the task processing result corresponding to those query features. The specific network structure of the head network can be set according to actual needs; for example, it can be implemented using a multilayer perceptron (MLP).

[0058] The image processing method provided in this embodiment can determine bird's-eye view features based on the image features corresponding to each viewpoint of the image to be processed from each viewpoint. Based on the bird's-eye view features, at least one task query feature can be determined. Then, based on each task query feature, the task processing result corresponding to each task can be obtained. End-to-end single-task or multi-task processing can be achieved by relying only on multi-view environmental images without the need for high-precision maps. This enables accurate surrounding environmental information to be obtained even without high-precision maps, greatly improving versatility and effectively reducing costs.

[0059] Figure 3 This is a schematic flowchart of an image processing method provided in another exemplary embodiment of this disclosure.

[0060] In one optional example, step 203 may specifically include the following steps:

[0061] Step 2031a: Based on the first bird's-eye view features and the initial static element query features, the static element task query features are determined using the first decoding network obtained through pre-training. The initial static element query features include the initial query features corresponding to each static element in at least one static element.

[0062] The initial static element query features can be set according to actual needs. For example, they can be obtained based on the first initialization rule, which can be set according to actual needs. For instance, N² D-dimensional static queries (where each static query represents the initial query feature corresponding to a static element) can be randomly initialized as initial static element query features. N² is the number of static queries, i.e., the number of static elements, and can be set according to actual needs. D represents the dimension of the initial query feature corresponding to each static element. The first decoding network can include at least one decoder, used to query task query features related to static elements from the first bird's-eye view features based on the initial static element query features, to obtain static element task query features. The specific network structure of the first decoding network can be set according to actual needs.

[0063] This disclosure utilizes a first decoding network obtained through training to query static element task query features related to static elements from the first bird's-eye view features based on initial static element query features. This provides more accurate and effective feature data for the implementation of subsequent static element detection tasks, enabling end-to-end task processing based on multi-view images. When applied to map reconstruction scenarios, static map information can be generated online without relying on offline-generated high-precision maps, further improving versatility.

[0064] In one optional example, step 203 may specifically include the following steps:

[0065] Step 2031b: Based on the first bird's-eye view features and the initial dynamic object query features, the second decoding network obtained through pre-training is used to determine the dynamic object task query features. The initial dynamic object query features include the initial query features corresponding to each dynamic object in at least one dynamic object.

[0066] The initial dynamic object query features can be set according to actual needs. For example, they can be obtained based on the second initialization rule, which can be set according to actual needs. For instance, N1 D-dimensional dynamic queries (each dynamic query representing an initial query feature corresponding to a dynamic object) can be randomly initialized as initial dynamic object query features, where N1 is the number of dynamic queries, i.e., the number of dynamic objects. N1 can be set according to actual needs, and D represents the dimension of the initial query feature corresponding to each dynamic object. The second decoding network can include at least one decoder, used to query task query features related to dynamic objects from the first bird's-eye view features based on the initial dynamic object query features, to obtain dynamic object task query features. The specific network structure of the second decoding network can be set according to actual needs.

[0067] This disclosure utilizes a second decoding network obtained through training to query dynamic object task query features related to dynamic objects from the features of a first bird's-eye view based on initial dynamic object query features. This provides more accurate and effective feature data for the subsequent implementation of dynamic object detection tasks, enabling end-to-end 3D object detection task processing based on multi-view images. This eliminates the need for dynamic object tracking, reduces the computational complexity of the network model, and avoids the impact of target tracking errors on subsequent applications.

[0068] In one optional example, step 203 may specifically include the following steps:

[0069] Step 2031c: Based on the static element task query features and the initial motion trajectory query features, the motion trajectory task query features are determined using a pre-trained third decoding network. The initial motion trajectory query features include the initial trajectory query features corresponding to each dynamic object in at least one dynamic object.

[0070] The initial motion trajectory query features need to be determined by combining dynamic object task query features and modal query features. Modal query features are used to characterize the motion trend of dynamic objects, with different modalities focusing on different future motion types (such as fast straight-ahead movement, slow straight-ahead movement, left turn, right turn, etc.). Dynamic object task query features are used to characterize features related to dynamic objects. By combining modal query features and dynamic object task query features, the initial trajectory query features related to the motion trajectory of dynamic objects can be determined. These features then interact with the static element task query features in the third decoding network to decode the motion trajectory task query features. The third decoding network can be configured according to actual needs.

[0071] This disclosure utilizes a third decoding network obtained through training to query motion trajectory task query features related to the motion trajectory of dynamic objects from static element task query features based on initial motion trajectory query features. This provides more accurate and effective feature data for the subsequent motion trajectory prediction task of dynamic objects, thereby realizing end-to-end motion trajectory prediction task processing based on multi-view images.

[0072] In an optional example, step 203 may include at least two of the steps 2031a-2031c above. The specific steps can be set according to actual needs, so that end-to-end multi-task processing can be achieved by relying solely on multi-view images, and the task processing results corresponding to multiple tasks can be obtained at the same time, avoiding dependence on high-precision maps and LiDAR, and further improving versatility.

[0073] Figure 4 This is a flowchart illustrating step 2031a provided in an exemplary embodiment of this disclosure.

[0074] In an optional example, step 2031a, based on the first bird's-eye view features and the initial static element query features, utilizes a pre-trained first decoding network to determine the static element task query features, including:

[0075] Step 20311a: Based on the initial static element query characteristics, determine the first query tensor, the first key tensor, and the first value tensor.

[0076] In this process, the initial static element query features can be mapped to the first query tensor based on the first query mapping rule, such as mapping based on the first query mapping matrix. Similarly, the initial static element query features can be mapped to the first key tensor based on the first key mapping rule, and the initial static element query features can be mapped to the first value tensor based on the first value mapping rule. The specific mapping principle will not be elaborated here.

[0077] Step 20312a: Based on the first query tensor, the first key tensor, and the first value tensor, the first self-attention network of the first decoder in the first decoding network is used to determine the first self-attention result.

[0078] The first self-attention network is a network based on the self-attention mechanism. It can be configured according to actual needs to perform self-attention operations on the first query tensor, the first key tensor, and the first value tensor. Specifically, self-attention operations are performed on the first query tensor and the first key tensor to obtain the first weight. Then, the first value tensor is weighted and summed based on the first weight to obtain the first self-attention result. The specific principle of the self-attention mechanism will not be elaborated further.

[0079] Step 20313a: Based on the first self-attention result and the initial static element query features, the first intermediate result is determined using the first summation normalization network of the first decoder in the first decoding network.

[0080] The first Add&Norm network has two functions: addition and normalization. The addition function adds the first self-attention result to the initial static element query feature to obtain the first addition result. The first addition result is then normalized to obtain the first intermediate result.

[0081] Step 20314a: Based on the first intermediate result, determine the second query tensor.

[0082] The first intermediate result can be used as the second query tensor, or the first intermediate result can be mapped to the second query tensor based on the corresponding mapping rules. The specific settings can be configured according to actual needs.

[0083] Step 20315a: Based on the features of the first bird's-eye view, determine the second key tensor and the second value tensor.

[0084] The principles for determining the second bond tensor and the second value tensor are described above and will not be repeated here.

[0085] Step 20316a: Based on the second query tensor, the second key tensor, and the second value tensor, the first cross-attention result is determined using the first deformable cross-attention network of the first decoder in the first decoding network.

[0086] Among them, the first deformable cross-attention network is a cross-attention network based on deformable convolution. Its function is to extract features from the first bird's-eye view features in the local area near the position corresponding to the initial static element query features, thereby further improving the accuracy and effectiveness of feature extraction.

[0087] Step 20317a: Based on the first cross-attention result and the first intermediate result, determine the static element task query features.

[0088] In the first decoder of the first decoding network, other related networks may be included after the first deformable cross-attention network, such as an add-normalization network (Add&Norm) or a feed-forward network. Therefore, after obtaining the first cross-attention result, the first cross-attention result needs to be added to the first intermediate result and normalized before being passed through other related networks to finally obtain the decoding result of the first decoder. When the first decoding network includes multiple decoders, the decoding result of the first decoder also needs to be used as the input of the second decoder. Decoding is then performed according to the decoding process of the first decoder, and so on, until the decoding of all decoders is completed to obtain the final decoding result of the first decoding network. This final decoding result is used as the static element task query feature.

[0089] In one optional example, Figure 5This is a schematic diagram of the network structure of a first decoding network provided in an exemplary embodiment of this disclosure. In this example, the first decoding network includes 6 decoders. ×6 indicates that the first decoding network includes 6 decoders within dashed boxes. Taking the first decoder as an example, Q1, K1, and V1 represent the first query tensor, the first key tensor, and the first value tensor, respectively. Self Attention represents a first self-attention network, Add&Norm represents an addition normalization network, and Add&Norm connected to the first self-attention network represents a first addition normalization network. Q2, K2, and V2 represent the second query tensor, the second key tensor, and the second value tensor, respectively. Deformable Cross Attention represents a first deformable cross attention network, and Feed Forward represents a feedforward network. The initial static element query features are mapped to a first query tensor, a first key tensor, and a first value tensor. These features then interact with each other in a first self-attention network to obtain a first self-attention result. This first self-attention result is added to the initial static element query features and normalized to obtain a first intermediate result. The first intermediate result is mapped to a second query tensor. Simultaneously, the first bird's-eye view features are mapped to a second key tensor and a second value tensor. The second query tensor, the second key tensor, and the second value tensor undergo cross-attention in a first deformable cross-attention network to achieve interaction between the first bird's-eye view features and the initial static element query features, obtaining a first cross-attention result. This first cross-attention result is added to the first intermediate result and normalized to obtain a first normalized result. This first normalized result is then passed through a feedforward network and another summation normalization network to obtain the decoding result of the first decoder. This decoding result is then processed by five decoders to obtain the static element task query features.

[0090] In one optional example, the first self-attention network can be a multi-head self-attention network, and the first deformable cross-attention network can also be a multi-head deformable cross-attention network, which can be set according to actual needs.

[0091] This disclosure achieves self-interaction of static element query features through a self-attention network in the first decoding network, captures the internal correlation of static element query features, and then interacts with bird's-eye view features in a deformable cross-attention network to achieve sparse attention to the data, flexibly captures features of relevant local areas, effectively reduces the amount of computation and improves the network inference speed while ensuring the acquisition of accurate and effective relevant features.

[0092] Figure 6 This is a flowchart illustrating step 2031b provided in an exemplary embodiment of this disclosure.

[0093] In an optional example, step 2031b, based on the first bird's-eye view features and the initial dynamic object query features, utilizes a pre-trained second decoding network to determine the dynamic object task query features, including:

[0094] Step 20311b: Based on the initial dynamic object query characteristics, determine the third query tensor, the third key tensor, and the third value tensor.

[0095] Step 20312b: Based on the third query tensor, the third key tensor, and the third value tensor, the second self-attention network of the first decoder in the second decoding network is used to determine the second self-attention result.

[0096] Step 20313b: Based on the second self-attention result and the initial dynamic object query features, the second intermediate result is determined using the second summation normalization network of the first decoder in the second decoding network.

[0097] Step 20314b: Based on the second intermediate result, determine the fourth query tensor.

[0098] Step 20315b: Based on the features of the first bird's-eye view, determine the fourth bond tensor and the fourth value tensor.

[0099] Step 20316b: Based on the fourth query tensor, the fourth key tensor, and the fourth value tensor, the second cross-attention result is determined using the second deformable cross-attention network of the first decoder in the second decoding network.

[0100] Step 20317b: Based on the second cross-attention result and the second intermediate result, determine the dynamic object task query features.

[0101] The specific operational principles of steps 20311b-20317b are the same as or similar to those of steps 20311a-20317a, except that step 20311b is based on initial dynamic object query features, which differs from the initial static element query features in step 20311a. These differences will not be elaborated upon here. Based on this, the network structure of the second decoding network is the same as or similar to that of the first decoding network, and will not be described further here.

[0102] Figure 7 This is a flowchart illustrating step 2031c provided in an exemplary embodiment of this disclosure.

[0103] In an optional example, step 2031c, based on the static element task query features and the initial motion trajectory query features, utilizes a pre-trained third decoding network to determine the motion trajectory task query features, including:

[0104] Step 20311c: Based on the initial motion trajectory query features, determine the fifth query tensor, the fifth key tensor, and the fifth value tensor.

[0105] Step 20312c: Based on the fifth query tensor, the fifth key tensor, and the fifth value tensor, the third self-attention network of the first decoder in the third decoding network is used to determine the third self-attention result.

[0106] Step 20313c: Based on the third self-attention result and the initial motion trajectory query features, the third intermediate result is determined using the third summation normalization network of the first decoder in the third decoding network.

[0107] Step 20314c: Based on the third intermediate result, determine the sixth query tensor.

[0108] The specific operating principles of steps 20311c-20314c are the same as or similar to those of the aforementioned steps 20311a-20314a, and will not be repeated here.

[0109] Step 20315c: Based on the static element task query features, determine the sixth key tensor and the sixth value tensor.

[0110] The sixth query tensor and the sixth value tensor in this step are obtained based on the static element task query feature mapping obtained in the previous example. The mapping principle is described in the previous content.

[0111] Step 20316c: Based on the sixth query tensor, the sixth key tensor, and the sixth value tensor, the third cross-attention result is determined using the first cross-attention network of the first decoder in the third decoding network.

[0112] The first cross-attention network can be any implementable cross-attention network, and can be set according to actual needs. For example, the first cross-attention network can adopt the cross-attention network structure in the conventional visual Transformer, without any specific limitation.

[0113] Step 20317c: Based on the third cross-attention result and the third intermediate result, determine the motion trajectory task query features.

[0114] The specific operating principle of this step is described in step 20317a above, and will not be repeated here.

[0115] For example, Figure 8This is a schematic diagram of the structure of the third decoding network provided in an exemplary embodiment of this disclosure. The initial motion trajectory query feature includes multiple trajectory queries (a trajectory query represents the initial trajectory query feature corresponding to a dynamic object). Q5, K5, and V5 represent the fifth query tensor, the fifth key tensor, and the fifth value tensor, respectively; Q6, K6, and V6 represent the sixth query tensor, the sixth key tensor, and the sixth value tensor, respectively. The meanings and reasoning processes of other symbols are described above and will not be repeated here.

[0116] This disclosure achieves self-interaction of motion trajectory query features through a self-attention network in a third decoding network, capturing the internal correlation of motion trajectory query features, and then interacting with static element task query features in a cross-attention network to effectively capture the motion trajectory-related features of dynamic objects. This implicitly reveals surrounding static information (such as surrounding road information), providing accurate and effective feature data for accurately predicting a more reasonable future motion trajectory of dynamic objects, and realizing end-to-end motion trajectory prediction based on multi-view image features.

[0117] Figure 9 This is a schematic flowchart of an image processing method provided in another exemplary embodiment of the present disclosure.

[0118] In an optional example, before determining the motion trajectory task query features using a pre-trained third decoding network based on the static element task query features and the initial motion trajectory query features in step 2031c, the method further includes:

[0119] Step 301: Based on the dynamic object task query features and modal query features, determine the initial motion trajectory query features. The modal query features include at least one first modal query feature corresponding to each modality. The first modal query feature corresponding to each modality is used to characterize a motion trend of the dynamic object.

[0120] Among them, modal query features can be set according to actual needs, such as being obtained through initialization by the third initialization rule. Since modal query features represent the motion trend of dynamic objects, combined with dynamic object task query features that represent the position of dynamic objects, the initial motion trajectory query features can be determined.

[0121] For example, N3 D-dimensional modal queries (where each modal query represents the first modal query feature corresponding to a modality) can be randomly initialized as modal query features. N3 can be set according to actual needs. The modal query features are fused with the dynamic object task query features to form N1×N3 D-dimensional trajectory queries, which serve as the initial motion trajectory query features. That is, each dynamic object has N3 modalities, representing its N3 motion trends.

[0122] This disclosure is based on the fusion of updated dynamic object task query features and modal query features to obtain initial motion trajectory query features. The initial motion trajectory query features include the task query features corresponding to each dynamic object and the query features of multiple modalities for each dynamic object. Different modalities focus on different future motion types (such as fast straight-ahead movement, slow straight-ahead movement, left turn, right turn, etc.), providing effective data support for subsequently decoding the motion trajectory task query features through a third decoding network.

[0123] Figure 10 This is a flowchart illustrating step 301 provided in an exemplary embodiment of this disclosure.

[0124] In an optional example, the dynamic object task query features include at least one dynamic object task query feature; step 301, based on the dynamic object task query features and modal query features, determines the initial motion trajectory query features, including:

[0125] Step 3011: For each task query feature corresponding to a dynamic object, determine a first number of task query features based on the task query feature. The first number is the number of first modal query features included in the modal query features.

[0126] Each dynamic object is assigned a first number of first modal query features to characterize the first number of motion trends of the dynamic object. Therefore, in order to fuse the task query features and modal query features of the object, it is necessary to transform the number of task query features of the dynamic object to be the same as the number of modal query features. Therefore, based on the task query features of the dynamic object, the first number of task query features is determined.

[0127] For example, if the modal query features include N3 D-dimensional modal queries, then for each dynamic object's D-dimensional task query feature (which can be called a task query), the task query feature is copied N3 times to obtain N3 identical D-dimensional task query features.

[0128] Step 3012: Add the first number of task query features to the first modal query features corresponding to each modality in the modal query features to obtain the initial trajectory query features corresponding to the dynamic object.

[0129] For example, N3 identical task query features are added to N3 first modal query features (modal queries) to form N3 motion trajectory query features.

[0130] Step 3013: Determine the initial motion trajectory query features based on the initial trajectory query features corresponding to each dynamic object.

[0131] For example, Figure 11 This is a schematic diagram illustrating the principle of determining the initial motion trajectory query features provided in an exemplary embodiment of this disclosure. In this example, the number of dynamic objects is N1, the task query represents the task query feature corresponding to the dynamic object, the number of first modal query features (modal queries) for each dynamic object is N3, and the final obtained initial motion trajectory query features include N1×N3 D-dimensional initial trajectory query features.

[0132] This disclosure obtains initial motion trajectory query features by fusing dynamic object task query features and modal query features, so that the initial motion trajectory query features include dynamic object task query related features and dynamic object motion trend related information. Furthermore, the third decoding network obtained through training can obtain accurate and effective motion trajectory task query features.

[0133] In an optional example, step 202, determining the first bird's-eye view features based on the first image features corresponding to each viewpoint, includes:

[0134] Step 2021: Determine the first bird's-eye view feature based on the first image features corresponding to each viewpoint, the initial bird's-eye view query features, and the bird's-eye view features obtained before the first bird's-eye view feature.

[0135] The initial bird's-eye view query features are features initialized based on the bird's-eye view. Their size represents the size of the bird's-eye view, and the specific initialization rules can be set according to actual needs. For example, H×W D-dimensional bird's-eye view queries can be initialized based on the required bird's-eye view size as the initial bird's-eye view query features, where H and W are the height and width of the bird's-eye view, respectively, and D is the feature dimension of each bird's-eye view query. Each bird's-eye view query corresponds to a set of three-dimensional coordinates in physical space, such as three-dimensional coordinates in a world coordinate system with the vehicle as the origin, which can be set according to actual needs. The previous frame bird's-eye view features are bird's-eye view features obtained in the previous image processing flow, and their specific processing flow is consistent with the first bird's-eye view features.

[0136] Since the bird's-eye view features of the previous frame contain relevant historical information such as static elements and dynamic objects in the image acquired in the previous frame, combining the bird's-eye view features of the previous frame, the initial bird's-eye view query features, and the first image features can not only determine the relevant features of static elements and dynamic objects in the current image to be processed, but also realize the changes of dynamic objects relative to the previous frame, thus facilitating the tracking of dynamic objects.

[0137] Figure 12 This is a flowchart illustrating step 2021 provided in an exemplary embodiment of this disclosure.

[0138] In an optional example, step 2021, determining the first bird's-eye view feature based on the first image features corresponding to each viewpoint, the initial bird's-eye view query features, and the bird's-eye view features obtained before the first bird's-eye view feature, includes:

[0139] Step 20211: Based on the bird's-eye view features in the previous frame and the initial bird's-eye view query features, the temporal self-attention network of the first encoder in the pre-trained encoder network is used to determine the temporal self-attention result.

[0140] The encoder network encodes the extracted first image features corresponding to each viewpoint based on the bird's-eye view features from the previous frame and the initial bird's-eye view query features, thus obtaining the first bird's-eye view features. The encoder network may include one or more encoders, each of which may include a temporal self-attention network. This temporal self-attention network is used to find the corresponding position of the initial bird's-eye view query features in the historical previous frame bird's-eye view, based on the vehicle's motion, as a reference position. This reference position is then used to extract features corresponding to the region in the first image features. Through continuous encoding by multiple encoders, the first bird's-eye view features of the current frame are obtained.

[0141] Step 20212: Based on the temporal self-attention results and the initial bird's-eye view query features, the fourth intermediate result is determined using the fourth phase addition normalization network of the first encoder.

[0142] The specific principles of the fourth phase addition normalization network are described above and will not be repeated here.

[0143] Step 20213: Based on the first image features and the fourth intermediate result corresponding to each viewpoint, determine the spatial cross-attention result using the spatial cross-attention network in the first encoder.

[0144] Among them, the spatial cross-attention network is used to uniformly sample the fourth intermediate result in height to obtain a set of three-dimensional coordinates. Then, according to the camera's intrinsic and extrinsic parameters, the three-dimensional coordinates are mapped to the corresponding positions in the first image features corresponding to each viewpoint. Then, based on deformable convolution, the features at the corresponding positions in the first image features are extracted. After subsequent processing, the first bird's-eye view features of the current frame are obtained.

[0145] Step 20214: Based on the spatial cross-attention results and the fourth intermediate results, determine the features of the first bird's-eye view.

[0146] Each encoder, in addition to the aforementioned temporal self-attention network, fourth summation normalization network, and spatial cross-attention network, also includes other related networks. For example, after the spatial cross-attention network, there may be a summation normalization network, a feedforward network, and another summation normalization network, etc., which can be configured according to actual needs. Therefore, after obtaining the spatial cross-attention result, it is necessary to add and normalize the spatial cross-attention result with the fourth intermediate result, and then pass it through other related networks to obtain the encoding result of the first encoder. If the encoder network includes multiple encoders, the encoding result output by the first encoder needs to be further encoded by subsequent encoders to finally obtain the first bird's-eye view features.

[0147] For example, Figure 13 This is a schematic diagram of the network structure of the encoder network provided in an exemplary embodiment of this disclosure. Wherein, BEV B(t-1) represents the bird's-eye view features in the previous frame, BEV queries Q represents the initial bird's-eye view query features, Temporal Self-Attention represents the temporal self-attention network, Spatial Cross-Attention represents the spatial cross-attention network, and the meanings of other symbols are as described above. The preceding bird's-eye view feature BEV B(t-1) and the initial bird's-eye view query feature BEV queries Q interact in a temporal self-attention network. Based on the vehicle's motion, the position corresponding to the initial bird's-eye view query feature in the previous frame's bird's-eye view is found as a reference position to obtain a temporal self-attention result. The temporal self-attention result is added to the initial bird's-eye view query feature and normalized to obtain a fourth intermediate result. The fourth intermediate result and the first image features from each viewpoint undergo a spatial cross-attention operation in a spatial cross-attention network. The spatial cross-attention network, based on deformable convolution, extracts features from the first image features based on the reference position to obtain a spatial cross-attention result. The spatial cross-attention result is added to the fourth intermediate result and normalized to obtain a fifth intermediate result. The fifth intermediate result is passed through a feedforward network and an add&norm network to obtain the encoding result of the first encoder. This encoding result is then passed through subsequent encoders to obtain the final encoding result, which is the first bird's-eye view feature. Understandably, spatial location encoding embedding can also be performed on each first image feature during network inference, and this disclosure does not limit this.

[0148] This disclosure achieves position matching between the initial bird's-eye view query features of the current frame and the bird's-eye view of the previous frame through a temporal self-attention network in the encoder network, establishes the temporal correlation between the current frame and the previous frame, and then extracts local features near the reference position from the first image features based on a spatial cross-attention network. While ensuring the extraction of effective features, it effectively reduces the amount of computation and improves the efficiency of image processing.

[0149] In an optional example, step 204, based on each task query feature among at least one task query feature, determines the task processing result corresponding to each task query feature, including:

[0150] Step 2041a: Based on the static element task query features, determine the static element detection results using the pre-trained static element detection head network.

[0151] The static element detection head network can be any implementable head network, such as a head network based on a multilayer perceptron.

[0152] Step 2041b: Based on the dynamic object task query features, determine the dynamic object detection result using the pre-trained dynamic object detection head network.

[0153] The dynamic object detection head network can be any implementable head network, such as a head network based on a multilayer perceptron.

[0154] Step 2041c: Based on the motion trajectory task query features, use the pre-trained motion trajectory prediction head network to determine the motion trajectory prediction result.

[0155] The motion trajectory prediction head network can be any implementable head network, such as a head network based on a multilayer perceptron.

[0156] This disclosure obtains first bird's-eye view features by encoding first image features corresponding to each viewpoint, decodes task query features for different tasks based on the first bird's-eye view features and decoding networks corresponding to different tasks, and then obtains task processing results for different tasks based on head networks corresponding to different tasks. This realizes end-to-end multi-task processing based on multi-frame, multi-view panoramic images, effectively improving task processing efficiency. It can achieve static element detection, dynamic object 3D detection, and motion trajectory prediction without relying on offline generated high-precision maps and LiDAR, effectively reducing costs.

[0157] In one optional example, Figure 14This is a schematic diagram of the overall structure of a network model for image processing provided in an exemplary embodiment of this disclosure. Based on the image to be processed from multiple perspectives, a feature extraction network is used to obtain first image features corresponding to each perspective. Each first image feature is passed through an encoder network to obtain first bird's-eye view features. The first bird's-eye view features are passed through a first decoding network to obtain static element task query features, and then a static element detection head network is used to obtain static element detection results. The first bird's-eye view features are passed through a second decoding network to obtain dynamic object task query features, and then a dynamic object detection head network is used to obtain dynamic object detection results. The dynamic object task query results are fused with modal query features to obtain initial motion trajectory query features. Based on the initial motion trajectory query features and the static element task query features obtained by the first decoding network, a third decoding network is used to obtain motion trajectory task query features, and then a motion trajectory prediction head network is used to obtain motion trajectory prediction results.

[0158] In one optional example, the network model needs to be pre-trained. When the network model includes multiple tasks, the tasks can be trained together, or they can be trained separately first and then combined. The specific settings can be configured according to actual needs. For example, to ensure better performance in motion trajectory prediction, the static element task and the dynamic object task can be trained first to obtain a base model, and then the three tasks can be trained together based on the base model. The specific training principle will not be elaborated here.

[0159] This disclosure uses only multi-view images to process static elements, dynamic objects, and motion trajectories. Compared with LiDAR, it can obtain richer environmental information, and has lower hardware costs and is easier to deploy. This disclosure can realize the online generation of static map information through static element detection, without relying on offline generated high-precision maps, and has a wider range of application scenarios. In addition, this disclosure does not require explicit dynamic target tracking, which effectively reduces the computational complexity of the model and can avoid the impact of tracking module errors on subsequent processing, further improving the accuracy of task processing results.

[0160] In an optional example, features can also be fused from images at different viewpoints with data acquired by LiDAR for end-to-end single-task or multi-task processing, thereby increasing the richness of feature information and further improving model performance.

[0161] The embodiments or optional examples disclosed above can be implemented individually or in any combination without conflict. The specific implementation can be set according to actual needs, and this disclosure does not limit it.

[0162] Any image processing method provided in the embodiments of this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any image processing method provided in the embodiments of this disclosure can be executed by a processor, such as a processor executing any image processing method mentioned in the embodiments of this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.

[0163] Exemplary device

[0164] Figure 15 This is a schematic diagram of the structure of an image processing apparatus provided in an exemplary embodiment of this disclosure. The apparatus of this embodiment can be used to implement the corresponding method embodiments of this disclosure, such as... Figure 15 The device shown includes: a first processing module 501, a second processing module 502, a third processing module 503, and a fourth processing module 504.

[0165] The first processing module 501 is used to determine a first image feature corresponding to each of the at least one viewpoints based on the images to be processed corresponding to each of the at least one viewpoints; the second processing module 502 is used to determine a first bird's-eye view feature based on the first image feature corresponding to each of the at least one viewpoints; the third processing module 503 is used to determine at least one task query feature among static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view feature; and the fourth processing module 504 is used to determine the task processing result corresponding to each of the at least one task query feature based on each of the task query features.

[0166] Figure 16 This is a schematic diagram of the structure of an image processing apparatus provided in another exemplary embodiment of the present disclosure.

[0167] In an optional example, the third processing module 503 includes:

[0168] The first processing unit 5031 is used to determine the static element task query features based on the first bird's-eye view features and the initial static element query features, using a pre-trained first decoding network. The initial static element query features include the initial query features corresponding to each of the static elements in at least one static element.

[0169] In an optional example, the third processing module 503 includes:

[0170] The second processing unit 5032 is used to determine the dynamic object task query features based on the first bird's-eye view features and the initial dynamic object query features, using a pre-trained second decoding network. The initial dynamic object query features include the initial query features corresponding to each of the dynamic objects in at least one dynamic object.

[0171] In an optional example, the third processing module 503 includes:

[0172] The third processing unit 5033 is used to determine the motion trajectory task query features based on the static element task query features and the initial motion trajectory query features, using a pre-trained third decoding network. The initial motion trajectory query features include the initial trajectory query features corresponding to each of the dynamic objects in at least one dynamic object.

[0173] In an optional example, the third processing module 503 may include at least two of the first processing unit 5031, the second processing unit 5032 and the third processing unit 5033 mentioned above, which can be set according to actual needs.

[0174] In an optional example, the first processing unit 5031 is specifically used for:

[0175] Based on the initial static element query features, a first query tensor, a first key tensor, and a first value tensor are determined. Based on the first query tensor, the first key tensor, and the first value tensor, a first self-attention result is determined using the first self-attention network of the first decoder in the first decoding network. Based on the first self-attention result and the initial static element query features, a first intermediate result is determined using the first summation normalization network of the first decoder in the first decoding network. Based on the first intermediate result, a second query tensor is determined. Based on the first bird's-eye view features, a second key tensor and a second value tensor are determined. Based on the second query tensor, the second key tensor, and the second value tensor, a first cross-attention result is determined using the first deformable cross-attention network of the first decoder in the first decoding network. Based on the first cross-attention result and the first intermediate result, the static element task query features are determined.

[0176] In an optional example, the second processing unit 5032 is specifically used for:

[0177] Based on the initial dynamic object query features, a third query tensor, a third key tensor, and a third value tensor are determined. Based on the third query tensor, the third key tensor, and the third value tensor, a second self-attention result is determined using the second self-attention network of the first decoder in the second decoding network. Based on the second self-attention result and the initial dynamic object query features, a second intermediate result is determined using the second summation normalization network of the first decoder in the second decoding network. Based on the second intermediate result, a fourth query tensor is determined. Based on the first bird's-eye view features, a fourth key tensor and a fourth value tensor are determined. Based on the fourth query tensor, the fourth key tensor, and the fourth value tensor, a second cross-attention result is determined using the second deformable cross-attention network of the first decoder in the second decoding network. Based on the second cross-attention result and the second intermediate result, the dynamic object task query features are determined.

[0178] In an optional example, the third processing unit 5033 is specifically used for:

[0179] Based on the initial motion trajectory query features, a fifth query tensor, a fifth key tensor, and a fifth value tensor are determined. Based on the fifth query tensor, the fifth key tensor, and the fifth value tensor, a third self-attention result is determined using the third self-attention network of the first decoder in the third decoding network. Based on the third self-attention result and the initial motion trajectory query features, a third intermediate result is determined using the third summation normalization network of the first decoder in the third decoding network. Based on the third intermediate result, a sixth query tensor is determined. Based on the static element task query features, a sixth key tensor and a sixth value tensor are determined. Based on the sixth query tensor, the sixth key tensor, and the sixth value tensor, a third cross-attention result is determined using the first cross-attention network of the first decoder in the third decoding network. Based on the third cross-attention result and the third intermediate result, the motion trajectory task query features are determined.

[0180] Figure 17 This is a schematic diagram of the structure of the third processing module 503 provided in an exemplary embodiment of this disclosure.

[0181] In an optional example, the third processing module 503 also includes:

[0182] The fourth processing unit 5034 is used to determine the initial motion trajectory query features based on the dynamic object task query features and modal query features. The modal query features include at least one first modal query feature corresponding to each of the modalities. The first modal query feature corresponding to the modal is used to characterize a motion trend of the dynamic object.

[0183] In an optional example, the dynamic object task query feature includes at least one task query feature of a dynamic object; the fourth processing unit 5034 is specifically used for:

[0184] For each dynamic object corresponding to the task query feature, based on the task query feature, a first number of task query features are determined, where the first number is the number of first modal query features included in the modal query features; the first number of task query features are added to the first modal query features corresponding to each modality in the modal query features to obtain the initial trajectory query features corresponding to the dynamic object; based on the initial trajectory query features corresponding to each dynamic object, the initial motion trajectory query features are determined.

[0185] In an optional example, the second processing module 502 includes:

[0186] The first determining unit 5021 is used to determine the first bird's-eye view feature based on the first image feature corresponding to each of the aforementioned viewpoints, the initial bird's-eye view query feature, and the bird's-eye view feature obtained before the first bird's-eye view feature in the previous frame.

[0187] In one optional example, the first determining unit 5021 is specifically used for:

[0188] Based on the bird's-eye view features in the previous frame and the initial bird's-eye view query features, a temporal self-attention result is determined using the temporal self-attention network of the first encoder in the pre-trained encoder network; based on the temporal self-attention result and the initial bird's-eye view query features, a fourth intermediate result is determined using the fourth summation normalization network of the first encoder; based on the first image features corresponding to each of the aforementioned viewpoints and the fourth intermediate result, a spatial cross-attention result is determined using the spatial cross-attention network in the first encoder; based on the spatial cross-attention result and the fourth intermediate result, the first bird's-eye view features are determined.

[0189] In an optional example, the fourth processing module 504 includes:

[0190] The second determining unit 5041 is used to determine the static element detection result based on the static element task query features and using a pre-trained static element detection head network; the third determining unit 5042 is used to determine the dynamic object detection result based on the dynamic object task query features and using a pre-trained dynamic object detection head network; the fourth determining unit 5043 is used to determine the motion trajectory prediction result based on the motion trajectory task query features and using a pre-trained motion trajectory prediction head network.

[0191] In an optional example, the units described above in this disclosure can be further divided into more granular units according to actual needs, such as dividing the unit into multiple sub-units, which can be set according to actual needs.

[0192] The embodiments or optional examples disclosed above can be implemented individually or in any combination without conflict. The specific implementation can be set according to actual needs, and this disclosure does not limit it.

[0193] The specific operation of each module and unit in the device disclosed herein is described in the foregoing method embodiments, and will not be repeated here.

[0194] Exemplary electronic devices

[0195] This disclosure also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program stored in the memory, wherein when the computer program is executed, it implements the image processing method described in any of the above embodiments of this disclosure.

[0196] Figure 18 This is a schematic diagram of an application embodiment of the electronic device disclosed herein. In this embodiment, the electronic device 10 includes one or more processors 11 and a memory 12.

[0197] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.

[0198] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute the program instructions to implement the methods of the various embodiments of this disclosure described above and / or other desired functions. Various contents such as input signals, signal components, and noise components may also be stored in the computer-readable storage medium.

[0199] In one example, the electronic device 10 may also include an input device 13 and an output device 14, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).

[0200] For example, the input device 13 may be the microphone or microphone array described above, used to capture the input signal of the sound source.

[0201] In addition, the input device 13 may also include, for example, a keyboard, a mouse, etc.

[0202] The output device 14 can output various information to the outside, including determined distance information, direction information, etc. The output device 14 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0203] Of course, for the sake of simplicity, Figure 18 Only some of the components of the electronic device 10 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 10 may include any other suitable components depending on the specific application.

[0204] Exemplary computer program products and computer-readable storage media

[0205] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products comprising computer program instructions that, when executed by a processor, cause the processor to perform the steps of the methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.

[0206] The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this disclosure. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0207] Furthermore, embodiments of this disclosure may also be computer-readable storage media having computer program instructions stored thereon, which, when executed by a processor, cause the processor to perform the steps in the methods according to various embodiments of this disclosure described in the "Exemplary Methods" section above.

[0208] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0209] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0210] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0211] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0212] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the order specifically described above unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the methods according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the methods according to this disclosure.

[0213] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.

[0214] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0215] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. An image processing method, comprising: Based on the images to be processed corresponding to each of the at least one viewpoint, determine the first image features corresponding to each of the at least one viewpoint; Based on the first image features corresponding to each of the aforementioned viewpoints, the first bird's-eye view features are determined. Based on the first bird's-eye view features, at least one of the following task query features is determined: static element task query features, dynamic object task query features, and motion trajectory task query features. Based on each of the at least one task query features, determine the task processing result corresponding to each of the task query features; The step of determining at least one task query feature among static element task query features, dynamic object task query features, and motion trajectory task query features based on the first bird's-eye view features includes: Based on the first bird's-eye view features and the initial static element query features, the static element task query features are determined using a pre-trained first decoding network. The initial static element query features include at least one initial query feature corresponding to each of the static elements; and / or, Based on the first bird's-eye view features and the initial dynamic object query features, the dynamic object task query features are determined using a pre-trained second decoding network. The initial dynamic object query features include initial query features corresponding to each of the dynamic objects in at least one dynamic object; and / or, Based on the static element task query features and the initial motion trajectory query features, the motion trajectory task query features are determined using a pre-trained third decoding network. The initial motion trajectory query features include the initial trajectory query features corresponding to each of the dynamic objects in at least one dynamic object.

2. The method according to claim 1, wherein, The step of determining the static element task query features based on the first bird's-eye view features and the initial static element query features, using a pre-trained first decoding network, includes: Based on the initial static element query characteristics, the first query tensor, the first key tensor, and the first value tensor are determined. Based on the first query tensor, the first key tensor, and the first value tensor, the first self-attention network of the first decoder in the first decoding network is used to determine the first self-attention result; Based on the first self-attention result and the initial static element query features, the first intermediate result is determined using the first summation normalization network of the first decoder in the first decoding network; Based on the first intermediate result, determine the second query tensor; Based on the features of the first bird's-eye view, determine the second key tensor and the second value tensor; Based on the second query tensor, the second key tensor, and the second value tensor, the first cross-attention result is determined using the first deformable cross-attention network of the first decoder in the first decoding network; Based on the first cross-attention result and the first intermediate result, determine the static element task query features; and / or, The step of determining the dynamic object task query features based on the first bird's-eye view features and the initial dynamic object query features, using a pre-trained second decoding network, includes: Based on the initial dynamic object query characteristics, the third query tensor, the third key tensor, and the third value tensor are determined. Based on the third query tensor, the third key tensor, and the third value tensor, the second self-attention result is determined using the second self-attention network of the first decoder in the second decoding network; Based on the second self-attention result and the initial dynamic object query features, the second intermediate result is determined using the second summation normalization network of the first decoder in the second decoding network; Based on the second intermediate result, determine the fourth query tensor; Based on the features of the first bird's-eye view, the fourth key tensor and the fourth value tensor are determined; Based on the fourth query tensor, the fourth key tensor, and the fourth value tensor, the second cross-attention result is determined using the second deformable cross-attention network of the first decoder in the second decoding network; Based on the second cross-attention result and the second intermediate result, the query features of the dynamic object task are determined.

3. The method according to claim 1, wherein, The process of determining the motion trajectory task query features based on the static element task query features and the initial motion trajectory query features, using a pre-trained third decoding network, includes: Based on the initial motion trajectory query features, the fifth query tensor, the fifth key tensor, and the fifth value tensor are determined. Based on the fifth query tensor, the fifth key tensor, and the fifth value tensor, the third self-attention result is determined using the third self-attention network of the first decoder in the third decoding network; Based on the third self-attention result and the initial motion trajectory query features, the third intermediate result is determined using the third summation normalization network of the first decoder in the third decoding network; Based on the third intermediate result, the sixth query tensor is determined; Based on the static element task query characteristics, the sixth key tensor and the sixth value tensor are determined; Based on the sixth query tensor, the sixth key tensor, and the sixth value tensor, the third cross-attention result is determined using the first cross-attention network of the first decoder in the third decoding network; Based on the third cross-attention result and the third intermediate result, the motion trajectory task query features are determined.

4. The method according to claim 1, wherein, Before determining the motion trajectory task query features using a pre-trained third decoding network based on the static element task query features and the initial motion trajectory query features, the method further includes: Based on the dynamic object task query features and modal query features, the initial motion trajectory query features are determined. The modal query features include at least one first modal query feature corresponding to each of the modalities. The first modal query feature corresponding to the modality is used to characterize a motion trend of the dynamic object.

5. The method according to claim 4, wherein, The dynamic object task query features include at least one dynamic object task query feature; determining the initial motion trajectory query features based on the dynamic object task query features and modal query features includes: For each task query feature corresponding to the dynamic object, a first number of task query features are determined based on the task query feature, where the first number is the number of first modal query features included in the modal query features; The first number of task query features are added to the first modal query features corresponding to each modality in the modal query features to obtain the initial trajectory query features corresponding to the dynamic object; The initial motion trajectory query features are determined based on the initial trajectory query features corresponding to each of the dynamic objects.

6. The method according to claim 1, wherein, The step of determining the first bird's-eye view features based on the first image features corresponding to each of the aforementioned viewpoints includes: The first bird's-eye view feature is determined based on the first image feature corresponding to each of the aforementioned viewpoints, the initial bird's-eye view query feature, and the bird's-eye view feature obtained before the first bird's-eye view feature in the previous frame.

7. The method according to claim 6, wherein, The process of determining the first bird's-eye view feature based on the first image feature corresponding to each of the aforementioned viewpoints, the initial bird's-eye view query feature, and the bird's-eye view feature obtained before the first bird's-eye view feature includes: Based on the bird's-eye view features in the previous frame and the initial bird's-eye view query features, the temporal self-attention result is determined by utilizing the temporal self-attention network of the first encoder in the pre-trained encoder network. Based on the temporal self-attention results and the initial bird's-eye view query features, the fourth intermediate result is determined using the fourth summation normalization network of the first encoder; Based on the first image features corresponding to each of the aforementioned viewpoints and the fourth intermediate result, the spatial cross-attention result is determined using the spatial cross-attention network in the first encoder; Based on the spatial cross-attention result and the fourth intermediate result, the features of the first bird's-eye view are determined.

8. The method according to claim 1, wherein, The step of determining the task processing result corresponding to each of the at least one task query features includes: Based on the static element task query features, the static element detection result is determined using a pre-trained static element detection head network. Based on the dynamic object task query features, the dynamic object detection result is determined using a pre-trained dynamic object detection head network. Based on the motion trajectory task query features, the motion trajectory prediction result is determined using a pre-trained motion trajectory prediction head network.

9. An image processing apparatus, comprising: The first processing module is used to determine the first image features corresponding to each of the at least one viewpoints based on the images to be processed corresponding to each of the at least one viewpoints. The second processing module is used to determine the first bird's-eye view features based on the first image features corresponding to each of the aforementioned viewpoints. The third processing module is used to determine at least one of the following task query features based on the first bird's-eye view features: static element task query features, dynamic object task query features, and motion trajectory task query features. The fourth processing module is used to determine the task processing result corresponding to each of the at least one task query features based on each of the task query features. The third processing module includes at least one of the first processing unit, the second processing unit, and the third processing unit. The first processing unit is configured to determine the static element task query features based on the first bird's-eye view features and the initial static element query features, using a pre-trained first decoding network. The initial static element query features include initial query features corresponding to each of the static elements in at least one static element. And / or, The second processing unit is configured to determine the dynamic object task query features based on the first bird's-eye view features and the initial dynamic object query features, using a pre-trained second decoding network. The initial dynamic object query features include initial query features corresponding to each of the dynamic objects in at least one dynamic object; and / or, The third processing unit is used to determine the motion trajectory task query features based on the static element task query features and the initial motion trajectory query features, using a pre-trained third decoding network. The initial motion trajectory query features include the initial trajectory query features corresponding to each of the dynamic objects in at least one dynamic object.

10. A computer-readable storage medium storing a computer program for performing the image processing method according to any one of claims 1-8.

11. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the image processing method according to any one of claims 1-8.