Target detection method and apparatus, electronic device, and storage medium

By extracting the target circumferential image feature and fusing the visual sensor with lidar information, and combining the transformer model for encoding and decoding, the problem of insufficient information sources and feature alignment in target detection is solved, and detection performance and efficiency are improved.

WO2025167487A1PCT designated stage Publication Date: 2025-08-14BEIJING JINGDONG YUANSHENG TECH CO LTD

Patent Information

Application Number
PCT/CN2025/072252
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-06
Filing Date
2025-01-14
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

In the prior art, the target detection performance is not perfect enough, the information source of a single sensor is insufficient, and the characteristics of multiple sensors are difficult to align, resulting in large errors in the detection result.

Method used

By extracting the target circumferential images collected at multiple consecutive moments, combining vision sensors and lidar information, perspective angle conversion and feature fusion are performed, and encoding and decoding are used for multimodal object detection.

Benefits of technology

It improves the performance and detection effect of target detection, alleviates the problems of single sensor information and multi-sensor feature alignment, and improves the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072252_14082025_PF_FP_ABST
    Figure CN2025072252_14082025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical fields of the Internet of Vehicles, the Internet of Things, artificial intelligence, smart cities, and the like, and provides a target detection method and apparatus, an electronic device, and a storage medium. The method comprises: respectively performing feature extraction on a plurality of target surround-view images acquired for a target area at a plurality of consecutive moments to obtain a temporal surround-view image feature sequence; for each target pixel point in each target surround-view image, performing viewpoint transformation on the target pixel point on the basis of pixel coordinate information and an initial depth feature of the target pixel point to obtain a target depth feature and first bird's-eye view coordinate information of the target pixel point from a bird's-eye view; determining a target temporal fusion result sequence on the basis of the surround-view image feature, the target depth feature and the first bird's-eye view coordinate information corresponding to a same target surround-view image; and coding / decoding the target temporal fusion result sequence to obtain a target detection result for the target area.
Need to check novelty before this filing date? Find Prior Art

Description

Target detection method, device, electronic device and storage medium

[0001] This application claims priority to Chinese patent application No. 202410171498.1 filed on February 6, 2024, the contents of which are incorporated herein by reference. Technical Field

[0002] The present disclosure relates to technical fields such as vehicle networks, the Internet of Things, artificial intelligence, and smart cities, and more specifically, to a target detection method, device, electronic device, and storage medium. Background Art

[0003] Robust and efficient target detection is a very important task in autonomous driving. It perceives objects in the vehicle's surrounding environment and synchronizes the perception results to the autonomous vehicle, guiding it to make correct decisions and control to ensure the safety of the vehicle during driving.

[0004] In the process of realizing the concept of the present disclosure, the inventors discovered that there are at least the following problems in the related art: the target detection performance is not yet perfect. Summary of the Invention

[0005] In view of this, the present disclosure provides a target detection method, device, electronic device, and storage medium.

[0006] One aspect of the present disclosure provides a target detection method, comprising: performing feature extraction on multiple target surround view images collected for a target area at multiple consecutive moments to obtain a time-series surround view image feature sequence, wherein the time-series surround view image feature sequence includes multiple surround view image features corresponding to the multiple target surround view images; performing perspective conversion on each target pixel point in each of the target surround view images based on pixel coordinate information and initial depth features of the target pixel point to obtain target depth features and first bird's-eye view coordinate information of the target pixel point under a bird's-eye view; determining a target time-series fusion result sequence based on the surround view image features, target depth features, and first bird's-eye view coordinate information corresponding to the same target surround view image; and performing encoding and decoding processing on the target time-series fusion result sequence to obtain a target detection result for the target area.

[0007] Another aspect of the present disclosure provides a target detection device, including: a first feature extraction module, configured to perform feature extraction on a plurality of target surround view images collected for a target area at a plurality of consecutive moments, respectively, to obtain a time-series surround view image feature sequence, wherein the time-series surround view image feature sequence includes a plurality of surround view image features corresponding to the plurality of target surround view images; a perspective conversion module, configured to perform perspective conversion on each target pixel point in each of the target surround view images based on the pixel coordinate information and the initial depth feature of the target pixel point, to obtain a target depth feature and first bird's-eye view coordinate information of the target pixel point under a bird's-eye view; a target fusion module, configured to determine a target time-series fusion result sequence based on the surround view image features, target depth features and first bird's-eye view coordinate information corresponding to the same target surround view image; and a target detection module, configured to perform encoding and decoding processing on the target time-series fusion result sequence to obtain a target detection result for the target area.

[0008] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the target detection method of the present disclosure.

[0009] Another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed, are used to implement the object detection method of the present disclosure.

[0010] Another aspect of the present disclosure provides a computer program product, which includes computer-executable instructions. When the instructions are executed, the computer program product is used to implement the target detection method of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0012] FIG1 schematically illustrates an exemplary system architecture to which a target detection method according to an embodiment of the present disclosure may be applied;

[0013] FIG2 schematically shows a flow chart of a target detection method according to an embodiment of the present disclosure;

[0014] FIG3 schematically illustrates a schematic diagram of a transformer-based cross-temporal multimodal target detection framework according to an embodiment of the present disclosure;

[0015] FIG4 schematically shows a block diagram of an object detection apparatus according to an embodiment of the present disclosure; and

[0016] FIG5 schematically shows a block diagram of an electronic device suitable for implementing a target detection method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0017] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0018] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0020] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0021] It should be noted that the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of the data involved in the technical solutions of this disclosure (including, but not limited to, user personal information) all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and maintain information and network security.

[0022] In the embodiments of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0023] The embodiments of the present disclosure provide a target detection method, device, electronic device, and storage medium. The method includes:

[0024] FIG1 schematically shows an exemplary system architecture 100 to which an object detection method according to an embodiment of the present disclosure may be applied.

[0025] It should be noted that FIG1 is merely an example of a system architecture to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.

[0026] As shown in FIG1 , a system architecture 100 according to this embodiment may include an autonomous driving vehicle 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the autonomous driving device 101 and the server 103. The network 102 may include various connection types, such as wired and / or wireless communication links.

[0027] The autonomous driving vehicle 101 can interact with the server 103 via the network 102 to receive or send data, etc.

[0028] The autonomous driving vehicle 101 may be equipped with a display screen for realizing a human-machine interface, and may also be equipped with various cameras, infrared scanning sensors and / or laser radars and other information collection devices for collecting information about the surrounding environment.

[0029] Server 103 can be a server that provides various services, such as a backend management server (for example only) that supports users in collecting information using the information collection device on driving device 101. The backend management server can analyze and process the received information and provide feedback to driving device 101. The server can be a cloud server, also known as a cloud computing server or cloud host. This server is a host product within the cloud computing service ecosystem, addressing the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or "VPS"). The server can also be a distributed system server or a server integrated with blockchain.

[0030] It should be noted that the target detection method provided in the embodiments of the present disclosure can generally be executed by the autonomous driving vehicle 101. Accordingly, the target detection device provided in the embodiments of the present disclosure can also be provided in the autonomous driving vehicle 101.

[0031] Alternatively, the target detection method provided by the embodiment of the present disclosure may also be generally performed by the server 103. Accordingly, the target detection device provided by the embodiment of the present disclosure may generally be set in the server 103. The target detection method provided by the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 103 and can communicate with the autonomous driving vehicle 101 and / or the server 103. Accordingly, the target detection device provided by the embodiment of the present disclosure may also be set in a server or server cluster that is different from the server 103 and can communicate with the autonomous driving vehicle 101 and / or the server 103.

[0032] For example, multiple target surround view images collected for a target area at multiple consecutive moments may be originally stored in any one of the vehicle-side storage device of the autonomous vehicle 101 and the server 103, or stored on an external storage device and imported into the vehicle-side storage device of the autonomous vehicle 101 and the server 103. The vehicle-side storage device or server 103 of the autonomous vehicle 101 may then locally execute the target detection method provided by the embodiments of the present disclosure, or send the multiple target surround view images to other terminal devices, servers, or server clusters, which then execute the target detection method provided by the embodiments of the present disclosure on the other terminal devices, servers, or server clusters that receive the multiple target surround view images.

[0033] It should be understood that the number of vehicles, networks, and servers in FIG1 is merely illustrative and any number of vehicles, networks, and servers may be provided as needed.

[0034] According to the embodiments of the present disclosure, the visual sensor is the earliest vehicle-mounted sensor. This sensor can obtain rich perception information from complex traffic environments, such as texture, color, etc., and is not limited to this. It has a strong ability to perceive the details of the target environment.

[0035] FIG2 schematically shows a flow chart of a target detection method according to an embodiment of the present disclosure.

[0036] As shown in FIG. 2 , the method includes operations S201 to S204 .

[0037] In operation S201 , feature extraction is performed on a plurality of target surround view images collected for a target area at a plurality of consecutive moments to obtain a time-series surround view image feature sequence, wherein the time-series surround view image feature sequence includes a plurality of surround view image features corresponding to the plurality of target surround view images.

[0038] According to embodiments of the present disclosure, a target area may represent the range of the environment that can be captured by an image capture device. At different times, the target capture device may be positioned differently within the environment, and the target area it can capture in the environment may also vary. A target surround image may represent an image captured by the image capture device during a surround view.

[0039] For example, six image acquisition devices can be positioned on an autonomous vehicle, facing six different directions. These devices can be used to capture images from six directions, namely, the left front, right front, left side, right side, left rear, and right rear of the autonomous vehicle, to produce six captured images. These six captured images can then be spliced ​​in their spatial distribution order to produce a surround view image of the target area captured by the six image acquisition devices at that moment. The image acquisition device can be, for example, at least one of a camera and a video camera, but is not limited thereto.

[0040] It should be noted that the number of image acquisition devices is not limited to six as described above, and may be other numbers as long as the target surround image can be acquired.

[0041] According to embodiments of the present disclosure, the target surround view image can be input into a backbone network for feature extraction to obtain surround view image features. The backbone network can be, for example, any of the following: resNet (a residual network), mobileNet (a lightweight network), VGG (Visual Geometry Group, a convolutional neural network), etc., but is not limited to these.

[0042] For example, the plurality of consecutive moments may include at least a first moment and a second moment, where the first moment may represent the moment immediately preceding the second moment. A first surround view image captured at the first moment for a first area is input into a backbone network for feature extraction to obtain first surround view image features. A similar method can be used to obtain second surround view image features for a second surround view image captured at the second moment for a second area. A temporal surround view image feature sequence can be determined based on the first and second surround view image features.

[0043] In operation S202, for each target pixel in each target surround image, a perspective conversion is performed on the target pixel based on the pixel coordinate information and the initial depth feature of the target pixel to obtain the target depth feature and first bird's-eye view coordinate information of the target pixel under the bird's-eye view.

[0044] According to embodiments of the present disclosure, pixel coordinate information can be represented as UV coordinates or XY coordinates, without limitation. Depth features can be extracted from the target surround image based on a depth feature extraction model to obtain initial depth features for each target pixel. The depth feature extraction model can be, for example, an LSS model, but is not limited thereto.

[0045] According to embodiments of the present disclosure, perspective conversion can represent the conversion of a target surround view captured for a target area based on a camera perspective into a bird's-eye view image of the target area from a BEV (bird's-eye view) perspective. Because the target surround view image has initial depth features, the bird's-eye view image obtained based on pixel coordinate information and initial depth features can have target depth features. The first bird's-eye view coordinate information can represent the coordinate information of pixel points in the bird's-eye view image. This information can be represented by UV coordinates or XY coordinates, without limitation herein.

[0046] In operation S203 , a target temporal fusion result sequence is determined based on the surround view image features, the target depth features, and the first bird's-eye view coordinate information corresponding to the same target surround view image.

[0047] According to embodiments of the present disclosure, the surround view image features of a target surround view image acquired at the same moment, along with the target depth features and first bird's-eye view coordinate information of the converted bird's-eye view image, can be spliced ​​together to obtain a fusion result at that moment. The temporal fusion result sequence can include multiple fusion results obtained by splicing the aforementioned information from multiple target surround view images acquired at multiple consecutive moments.

[0048] In operation S204 , encoding and decoding are performed on the target time series fusion result sequence to obtain a target detection result for the target area.

[0049] According to the embodiments of the present disclosure, encoder and decoder processing can be performed on the target time series fusion result sequence based on the transformei model to obtain the target detection result of the target area. The target detection result may include a three-dimensional detection result, but is not limited to this.

[0050] Through the above-mentioned embodiments of the present disclosure, target detection is performed based on multiple frames of target surround view images collected at multiple moments, which can provide more information and features for the detection process, and is conducive to improving detection performance and detection effects.

[0051] The method shown in FIG2 is further described below with reference to specific embodiments.

[0052] According to an embodiment of the present disclosure, before performing the above operation S202, an initial depth feature may be first determined. The method may include: determining the initial depth feature of the target pixel based on a set of view frustum points corresponding to the target pixel. The view frustum point set may represent multiple predicted depths determined for the target pixel.

[0053] For example, for a target pixel (u, v), the view frustum can be sampled based on the line connecting the camera position and the target pixel position. For example, d view frustum points can be obtained, where d is a positive integer. Each view frustum point can correspond to a sampling depth, for example, d k The depth of the kth frustum point can be represented. The depth of the d frustum points can be used as the d predicted depths determined for the target pixel point (u, v). Based on this, the frustum point set corresponding to the target pixel point (u, v) can be expressed as: , where k={1,2,3,...d}. Then, by transforming the cone point set p k (u,v) is input into the LSS model, and the initial depth feature predicted for the target pixel (u,v) can be obtained, for example, D K , the pixel coordinate information and initial depth feature of the target pixel point (u, v) with the initial depth feature can be represented as .

[0054] Through the above-mentioned embodiments of the present disclosure, a more accurate prediction of the depth of the target pixel point can be achieved, alleviating the problem of missing the actual size and depth information of the object when using only a camera, which is conducive to executing the subsequent perspective conversion process and improving the accuracy of the conversion results.

[0055] According to an embodiment of the present disclosure, the above-mentioned operation S202 may include: converting the target pixel point from the image coordinate system to the camera coordinate system according to the pixel coordinate information, the initial depth feature, and the camera intrinsic parameter corresponding to the target pixel point, to obtain the camera coordinate information and the first depth feature of the target pixel point in the camera coordinate system. Converting the target pixel point from the camera coordinate system to the radar coordinate system according to the camera coordinate information, the first depth feature, and the first conversion matrix from the camera coordinate system to the radar coordinate system, to obtain the radar coordinate information and the second depth feature of the target pixel point in the radar coordinate system. Converting the target pixel point from the radar coordinate system to the bird's-eye view coordinate system according to the radar coordinate information, the second depth feature, and the second conversion matrix from the radar coordinate system to the bird's-eye view coordinate system, to obtain the first bird's-eye view coordinate information and the target depth feature of the target pixel point in the bird's-eye view.

[0056] According to the embodiments of the present disclosure, since the target surround view image can be stitched together based on images captured by multiple cameras, target pixels in different areas of the target surround view image can have different camera parameters, including camera intrinsic parameters and camera extrinsic parameters. The pixel coordinate information of the target surround view image can represent the coordinates of the corresponding pixel in the image coordinate system. The process of obtaining the coordinates of the target surround view image under the BEV perspective can be decomposed into the process of calculating the coordinates of any target pixel point (u, v) in the target surround view image under the BEV perspective.

[0057] Corresponding to the above process, for example, for the target pixel point P containing the depth feature K (u,v) can first be multiplied by the camera intrinsic parameter matrix K to transform from the image coordinate system to the camera coordinate system. Then, the extrinsic parameter matrix T from the camera coordinate system to the radar coordinate system is multiplied by the camera intrinsic parameter matrix K. l c Multiply them together to transform from the camera coordinate system to the radar coordinate system. Then, add the transformation matrix B from the radar coordinate system to the BEV view coordinate system. l c Multiply and transform to the BEV coordinate system. This process can be achieved, for example, by the following formulas (1) to (2).

[0058] T im (u,v) = T l c K -1 P K (u,v) Formula (1)

[0059] In formula (1), T im (u,v) can represent P K (u,v) is the coordinate in the radar coordinate system.

[0060] C im (u,v) = B l c T l c K -1 P K (u,v) Formula (2)

[0061] In formula (2), C im (u,v) can represent P K (u,v) is the coordinate in the BEV coordinate system.

[0062] According to an embodiment of the present disclosure, the multiple moments may include a first moment and a second moment. For target surround images acquired at any moment, the coordinate information and depth features of the corresponding images under a bird's-eye view can be calculated using the above method.

[0063] According to an embodiment of the present disclosure, the perspective conversion process for the target surround view image collected at each moment may also depend on the radar coordinates at the moment before or after the moment. Accordingly, when performing a conversion of the target pixel from the radar coordinate system to the bird's-eye view coordinate system based on the radar coordinate information, the second depth feature, and the second conversion matrix from the radar coordinate system to the bird's-eye view coordinate system to obtain the first bird's-eye view coordinate information and target depth feature of the target pixel at the bird's-eye view, the process may include: obtaining the initial radar coordinate information and initial second depth feature of the target pixel collected at the first moment in the radar coordinate system at the first moment. Aligning the initial radar coordinate information and initial second depth feature to the radar coordinate system at the second moment based on the initial radar coordinate information, the initial second depth feature, and the third conversion matrix from the radar coordinate system at the first moment to the radar coordinate system at the second moment to obtain the target radar coordinate information and target second depth feature. According to the target radar coordinate information, the target second depth feature and the third transformation matrix from the radar coordinate system at the first moment to the bird's-eye view coordinate system, the target pixel point is transformed from the radar coordinate system at the first moment to the bird's-eye view coordinate system to obtain the first bird's-eye view coordinate information and target depth feature of the target pixel point collected at the first moment under the bird's-eye view.

[0064] According to an embodiment of the present disclosure, the radar coordinate system may change as the position of the radar changes. At different times, the position of the radar may be different, and thus the position of the radar coordinate system may be different.

[0065] For example, based on the above formulas (1) to (2), the target pixel point (u t ,v t ) The first bird's-eye view coordinate information containing the target depth feature under the BEV perspective, for example, C im t (u t ,v t ) = B l c T l c K -1 P K (u t ,v t ).

[0066] When calculating the target pixel point (u t-1 ,v t-1) in the BEV perspective, the initial radar coordinate information containing the initial second depth feature of the target pixel point in the target surround image collected at the first moment in the radar coordinate system can be calculated based on the above formula (1), for example, T im t-1 (u t-1 ,v t-1 ) = T l c K -1 P K t-1 (u t-1 ,v t-1 Then, the first bird's-eye view coordinate information containing the target depth feature of the target pixel point in the target surround image collected at the first moment under the BEV perspective can be calculated by combining formulas (3) to (4).

[0067] T im t-1 t (u t-1 ,v t-1 ) = T l(t) l(t-1) T l c K -1 P K t-1 (u t-1 ,v t-1 ) Formula (3)

[0068] In formula (3), T l(t) l(t-1) The third transformation matrix, T, can represent the radar coordinate system at time t-1 to the radar coordinate system at time t. im t-1 t (u t-1 ,v t-1 ) can be characterized by K After the information of (u,v) at time t-1 is aligned to time t, P K (u t-1 ,v t-1 ) The target radar coordinate information including the target second depth feature in the radar coordinate system at time t-1.

[0069] C im t-1 t (u t-1 ,v t-1 ) = B l c T l(t) l(t-1) Tl c K -1 P K t-1 (u t-1 ,v t-1 ) Formula (4)

[0070] In formula (4), C im t-1 t (u t-1 ,v t-1 ) can be characterized by K (u t-1 ,v t-1 ) After the information at time t-1 is aligned to time t, P K (u t-1 ,v t-1 ) The first bird's-eye view coordinate information containing the target depth feature from the BEV perspective at time t-1.

[0071] Through the above embodiments of the present disclosure, target detection can be performed in combination with time series features to further improve the detection effect.

[0072] LiDAR is another type of sensor that can accurately determine the location, shape, and size of objects in the surrounding environment. This sensor is less affected by real-time weather conditions, such as dim light and low visibility. Fusion of visual sensors and LiDAR generates multimodal data, which can improve the effectiveness of object detection.

[0073] While implementing the concepts of this disclosure, the inventors discovered that using a single sensor for object detection suffers from a single information source. For example, using a LiDAR lacks information about the object's color, texture, and so on. Combining information from both a visual sensor and LiDAR can lead to alignment issues, primarily due to inaccurate camera depth predictions. This can lead to errors in the final detection results.

[0074] According to embodiments of the present disclosure, a lidar can be deployed simultaneously with an image acquisition device. While the image acquisition device is collecting target surround images, the lidar is also collecting point cloud information. This includes: performing voxelization on the multiple point cloud information collected for the target area at multiple consecutive moments to obtain a time-series bird's-eye view image sequence, where the time-series bird's-eye view image sequence includes multiple bird's-eye view images represented by the multiple point cloud information from a bird's-eye view perspective. For each bird's-eye view image, feature extraction is performed to obtain bird's-eye view image features. Second bird's-eye view coordinate information and bird's-eye view image features corresponding to the same point cloud information are fused to obtain a first time-series fusion result sequence. In this case, during the execution of operation S202, the surround image features, target depth features, and first bird's-eye view coordinate information corresponding to the same target surround image can be fused to obtain a second time-series fusion result sequence. The first time-series fusion result sequence and the second time-series fusion result sequence corresponding to the same moment are fused to obtain a target time-series fusion result sequence.

[0075] According to an embodiment of the present disclosure, the point cloud information may be information of a three-dimensional point cloud. The voxelization operation may, for example, include: dividing the three-dimensional point cloud into units in the x-axis and y-axis directions, obtaining the information of the highest point of each unit obtained by the division on the xy plane in the z-axis direction, and combining the information of the highest point of each unit in the z-axis direction according to the position arrangement of each unit, so as to construct a two-dimensional image of the three-dimensional point cloud. When the xy plane is perpendicular to the bird's-eye view, the constructed two-dimensional image can represent the bird's-eye view image of the point cloud information. For example, a method such as PointPillar (a 3D target detection network) can be used to perform a voxelization operation on the point cloud information.

[0076] According to an embodiment of the present disclosure, the second bird's-eye view coordinate information of the bird's-eye view image obtained by voxelization may be coordinate information containing depth information. By performing a voxelization operation on the point cloud information collected at the second moment, the second bird's-eye view coordinate information of the point cloud information collected at the second moment can be obtained. By performing a voxelization operation on the point cloud information collected at the first moment, the initial second bird's-eye view coordinate information of the first moment can be obtained. The initial second bird's-eye view coordinate information can be aligned to the radar coordinate system of the second moment in combination with the fourth transformation matrix from the radar coordinate system of the first moment to the radar coordinate system of the second moment, and the second bird's-eye view coordinate information of the target can be obtained. The second bird's-eye view coordinate information of the target can be used as the second bird's-eye view coordinate information of the point cloud information collected at the first moment.

[0077] According to an embodiment of the present disclosure, the feature extraction process may also be implemented based on the aforementioned backbone network, and is not limited thereto.

[0078] According to the embodiments of the present disclosure, the surround view image features of the target surround view image at each moment, the first bird's-eye view coordinate information containing the target depth features at the same moment, the bird's-eye view image features of the point cloud information at the same moment, and the second bird's-eye view coordinate information at the same moment are fused. After all the target surround view image features and point cloud information features are aligned according to time, that is, the features correspond to the coordinates, a target temporal fusion result sequence can be formed. The target temporal fusion result sequence can be fed into the encoder of a transformer (a deep learning model based on the self-attention mechanism) for feature fusion and learning.

[0079] Through the above-mentioned embodiments of the present disclosure, the point cloud information features and the target surround image features at the same time are fused, which can alleviate the problem of difficulty in aligning multi-sensor features and is conducive to improving subsequent detection performance.

[0080] According to an embodiment of the present disclosure, operation S204 may include: performing feature extraction on the target time series fusion result sequence to obtain a fusion vector; decoding the fusion vector and a position encoding vector for detecting the position information of the object to be detected in the target area to obtain category information and location information of the object to be detected; and determining a target detection result based on the category information and location information.

[0081] According to an embodiment of the present disclosure, a 3D anchor coordinate can be initialized in the Position-guided Query module used to calculate the Q value in the transformer. The corresponding embedding features are then passed through a layer of MLP (Multi-Layer Perception) and input into the transformer's decoder module for training. Subsequently, during application, the transformer's Position-guided Query module can provide a position encoding vector for detecting the position information of the object to be detected in the target area. This position encoding vector can be a three-dimensional position encoding vector, but is not limited to this.

[0082] According to an embodiment of the present disclosure, the decoder module of the transformer can be connected to a detection head for detecting category information and position information to perform category and position prediction and regression. The target detection result can be determined based on the regression result.

[0083] Through the above-mentioned embodiments of the present disclosure, target detection is performed in combination with a position coding vector for detecting position information of an object to be detected in a target area, which can effectively improve detection efficiency and detection quality.

[0084] FIG3 schematically illustrates a schematic diagram of a transformer-based cross-temporal multimodal target detection framework according to an embodiment of the present disclosure.

[0085] As shown in FIG3 , the input of the transformer-based cross-temporal multimodal target detection framework 300 may include a camera part 310 and a radar part 320 .

[0086] The camera section 310 can input a first surround view image 311 at time t. After processing by a 2D first backbone 317, first surround view image features 313 can be obtained. In this process, LSS 318 can be combined to calculate the depth information of the first surround view image 311 and first BEV coordinate information 315 of the first surround view image 311, which contains target depth features. Simultaneously, the camera section 310 can also input a second surround view image 312 at time t-1. Following the same process as described above, second surround view image features 314 and second BEV coordinate information 316 of the second surround view image 312, which contains target depth features, can be obtained. By combining the aforementioned methods to stitch together camera-related information at the same time, a second temporal fusion result sequence 319 can be obtained, which can be used as input for image temporal sequence.

[0087] In the radar component, the first point cloud information 321 at time t is input. After processing by the voxelization module 327, a first bird's-eye view image 323 with depth information is obtained. The first bird's-eye view image 323 is processed by the second backbone 328 to obtain the first BEV image features 325. Simultaneously, the second point cloud information 322 at time t-1 is input. Following the same process as above, a second bird's-eye view image 324 and second BEV image features 326 are obtained. By combining the aforementioned methods to splice radar information at the same time, a first time series fusion result sequence 329 is obtained, which can be used as the time series input for the point cloud.

[0088] By combining the aforementioned methods to fuse the second time series fusion result sequence 319 and the first time series fusion result sequence 329 at the same moment, a target time series fusion result sequence 330 can be obtained.

[0089] The target time series fusion result sequence 330 can be input into the transformer 340. After feature extraction by the encoder 341 of the transformer 340, it is decoded by the decoder 342 under the guidance of the position encoding vector provided by the Position-guided Query 343 for detecting the position information of the object to be detected in the target area, and then enters the category detection head 344 and the position detection head 345 for processing to obtain the final target detection result 350. The Position-guided Query 343 can use the learnable anchor position parameters as the query input in the transformer 340.

[0090] It should be noted that both the first bird's-eye view image 323 and the second bird's-eye view image 324 may have the second bird's-eye view coordinate information at the corresponding moment. The first backbone 317 and the second backbone 318 may be the same backbone network, which is not limited here. The encoder 341 may be a 3D position encoding module, but is not limited thereto.

[0091] Through the above-mentioned embodiments of the present disclosure, an effective transformer-based cross-temporal multimodal target detection framework for autonomous driving is provided. This framework fuses visual and radar information, using a method based on transforming 3D position space into the same size of 2D features to fuse image and radar features. It then uses attention in the transformer to learn the fused new features, solving the problems of single-sensor information being single and the difficulty in aligning features from multiple sensors. In addition, the features of each modality at the previous moment are processed in the same way and fused with the features at the current moment to obtain temporal features. Using cross attention, global information sharing can be achieved, effectively improving the target detection effect.

[0092] FIG4 schematically shows a block diagram of an object detection apparatus according to an embodiment of the present disclosure.

[0093] As shown in FIG4 , the object detection apparatus 400 includes a first feature extraction module 410 , a perspective conversion module 420 , an object fusion module 430 , and an object detection module 440 .

[0094] The first feature extraction module 410 is used to extract features from multiple target surround view images collected for the target area at multiple consecutive moments to obtain a time-series surround view image feature sequence, wherein the time-series surround view image feature sequence includes multiple surround view image features corresponding to the multiple target surround view images.

[0095] The perspective conversion module 420 is used to perform perspective conversion on each target pixel point in each target surround image based on the pixel coordinate information and initial depth feature of the target pixel point, so as to obtain the target depth feature and first bird's-eye view coordinate information of the target pixel point under the bird's-eye view.

[0096] The target fusion module 430 is configured to determine a target temporal fusion result sequence based on the surround view image features, target depth features, and first bird's-eye view coordinate information corresponding to the surround view image of the same target.

[0097] The target detection module 440 is used to perform encoding and decoding processing on the target time series fusion result sequence to obtain the target detection result for the target area.

[0098] According to an embodiment of the present disclosure, the perspective conversion module includes an image-camera conversion unit, a camera-radar conversion unit, and a radar-bird's-eye perspective conversion unit.

[0099] The image-to-camera conversion unit is used to convert the target pixel point from the image coordinate system to the camera coordinate system based on the pixel coordinate information, the initial depth feature and the camera intrinsic parameter corresponding to the target pixel point, so as to obtain the camera coordinate information and the first depth feature of the target pixel point in the camera coordinate system.

[0100] The camera-to-radar conversion unit is used to convert the target pixel point from the camera coordinate system to the radar coordinate system based on the camera coordinate information, the first depth feature, and the first conversion matrix from the camera coordinate system to the radar coordinate system, so as to obtain the radar coordinate information and the second depth feature of the target pixel point in the radar coordinate system.

[0101] The radar-to-bird's-eye view conversion unit is used to convert the target pixel point from the radar coordinate system to the bird's-eye view coordinate system based on the radar coordinate information, the second depth feature, and the second conversion matrix from the radar coordinate system to the bird's-eye view coordinate system, so as to obtain the first bird's-eye view coordinate information and target depth feature of the target pixel point under the bird's-eye view.

[0102] According to an embodiment of the present disclosure, the plurality of moments include a first moment and a second moment. The radar-bird's-eye view conversion unit includes a radar information acquisition subunit, a radar-radar conversion subunit, and a radar-bird's-eye view conversion subunit.

[0103] The radar information acquisition subunit is used to obtain the initial radar coordinate information and the initial second depth feature of the target pixel point collected at the first moment in the radar coordinate system at the first moment.

[0104] The radar-to-radar conversion subunit is used to align the initial radar coordinate information and the initial second depth feature to the radar coordinate system at the second moment based on the initial radar coordinate information, the initial second depth feature, and a third conversion matrix from the radar coordinate system at the first moment to the radar coordinate system at the second moment, so as to obtain the target radar coordinate information and the target second depth feature.

[0105] The radar-to-bird's-eye view conversion subunit is used to convert the target pixel point from the radar coordinate system at the first moment to the bird's-eye view coordinate system based on the target radar coordinate information, the target second depth feature, and the third conversion matrix from the radar coordinate system at the first moment to the bird's-eye view coordinate system, so as to obtain the first bird's-eye view coordinate information and target depth feature of the target pixel point collected at the first moment under the bird's-eye view.

[0106] According to an embodiment of the present disclosure, the target detection device further includes a voxelization module, a second feature extraction module and a first fusion module.

[0107] The voxelization module is used to perform voxelization operations on multiple point cloud information collected for the target area at multiple consecutive moments to obtain a time-series bird's-eye view image sequence, wherein the time-series bird's-eye view image sequence includes multiple bird's-eye view images represented by the multiple point cloud information under a bird's-eye view perspective.

[0108] The second feature extraction module is used to extract features from each bird's-eye view image to obtain bird's-eye view image features.

[0109] The first fusion module is used to fuse the second bird's-eye view coordinate information and the bird's-eye view image features corresponding to the same point cloud information to obtain a first temporal fusion result sequence.

[0110] The target fusion module includes a second fusion unit and a third fusion unit.

[0111] The second fusion unit is used to fuse the surround view image features, the target depth features and the first bird's-eye view coordinate information corresponding to the same target surround view image to obtain a second temporal fusion result sequence.

[0112] The third fusion unit is used to fuse the first time series fusion result sequence and the second time series fusion result sequence corresponding to the same moment to obtain a target time series fusion result sequence.

[0113] According to an embodiment of the present disclosure, the object detection module includes a feature extraction unit, a decoding unit and an object detection unit.

[0114] The feature extraction unit is used to extract features from the target time series fusion result sequence to obtain a fusion vector.

[0115] The decoding unit is used to decode the fusion vector and the position coding vector used to detect the position information of the object to be detected in the target area to obtain the category information and position information of the object to be detected.

[0116] The target detection unit is used to determine the target detection result based on the category information and position information.

[0117] According to an embodiment of the present disclosure, the object detection apparatus further includes an initial depth feature determination module.

[0118] The initial depth feature determination module is used to determine the initial depth feature of the target pixel point based on the view cone point set corresponding to the target pixel point, wherein the view cone point set represents multiple predicted depths determined for the target pixel point.

[0119] According to the embodiments of the present disclosure, feature extraction is performed on multiple target surround view images collected for the target area at multiple consecutive moments to obtain a time-series surround view image feature sequence; for each target pixel point in each target surround view image, the target pixel point is converted from a perspective based on the pixel coordinate information and the initial depth feature of the target pixel point to obtain the target depth feature and the first bird's-eye view coordinate information of the target pixel point under a bird's-eye view; the target time-series fusion result sequence is determined based on the surround view image feature, the target depth feature and the first bird's-eye view coordinate information corresponding to the same target surround view image; and the target time-series fusion result sequence is encoded and decoded to obtain a target detection result for the target area. This technical means performs target detection based on multiple frames of target surround view images collected at multiple moments, which can provide more information and features for the detection process, and is conducive to improving detection performance and detection effect.

[0120] According to the embodiments of the present invention, any number of modules, units, and sub-units, or at least part of the functions of any number of them, can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, units, and sub-units can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, units, and sub-units can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in an appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, units, and sub-units can be at least partially implemented as a computer program module, which can perform the corresponding functions when the computer program module is executed.

[0121] For example, any number of the first feature extraction module 410, the perspective conversion module 420, the target fusion module 430, and the target detection module 440 can be combined into a single module / unit / sub-unit, or any one of these modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functionality of one or more of these modules / units / sub-units can be combined with at least part of the functionality of other modules / units / sub-units and implemented in a single module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the first feature extraction module 410, the perspective conversion module 420, the target fusion module 430, and the target detection module 440 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware by any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the first feature extraction module 410 , the perspective conversion module 420 , the target fusion module 430 and the target detection module 440 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0122] It should be noted that the target detection device part in the embodiment of the present disclosure corresponds to the target detection method part in the embodiment of the present disclosure. The description of the target detection device part specifically refers to the target detection method part and will not be repeated here.

[0123] Figure 5 schematically shows a block diagram of an electronic device suitable for implementing the target detection method according to an embodiment of the present disclosure. The electronic device shown in Figure 5 is only an example and should not bring any limitation to the functions and scope of use of the embodiment of the present disclosure.

[0124] As shown in FIG5 , an electronic device 500 according to an embodiment of the present disclosure includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage unit 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0125] Various programs and data required for the operation of the electronic device 500 are stored in the RAM 503. The processor 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The processor 501 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 502 and / or RAM 503. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and RAM 503. The processor 501 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0126] According to an embodiment of the present disclosure, electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to bus 504. System 500 may also include one or more of the following components connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN card or modem. Communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 510 as needed, so that computer programs read from the removable media can be installed into storage section 508 as needed.

[0127] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.

[0128] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.

[0129] According to embodiments of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0130] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 502 and / or the RAM 503 described above and / or one or more memories other than the ROM 502 and the RAM 503 .

[0131] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method provided by the embodiment of the present disclosure. When the computer program product runs on an electronic device, the program code is used to enable the electronic device to implement the method provided by the embodiment of the present disclosure.

[0132] When the computer program is executed by the processor 501, the above functions defined in the system / device of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0133] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 509, and / or installed from a removable medium 511. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0134] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0135] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or coupled in various ways, and all such combinations and / or couplings fall within the scope of the present disclosure.

[0136] The embodiments of the present disclosure are described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be used in combination to advantage. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A target detection method, comprising: Performing feature extraction on a plurality of target surround view images collected for a target area at a plurality of consecutive moments to obtain a time-series surround view image feature sequence, wherein the time-series surround view image feature sequence includes a plurality of surround view image features corresponding to the plurality of target surround view images; For each target pixel in each target surround image, based on the pixel coordinate information and the initial depth feature of the target pixel, perform perspective conversion on the target pixel to obtain the target depth feature and first bird's-eye view coordinate information of the target pixel under a bird's-eye view; Determining a target temporal fusion result sequence according to the surround view image features, target depth features, and first bird's-eye view coordinate information corresponding to the same surround view image of the target; and The target time series fusion result sequence is encoded and decoded to obtain a target detection result for the target area.

2. The method according to claim 1, wherein The performing perspective conversion on the target pixel point based on the pixel coordinate information and the initial depth feature of the target pixel point to obtain the target depth feature and first bird's-eye view coordinate information of the target pixel point under a bird's-eye view includes: Converting the target pixel from an image coordinate system to a camera coordinate system based on the pixel coordinate information, the initial depth feature, and a camera intrinsic parameter corresponding to the target pixel to obtain camera coordinate information and a first depth feature of the target pixel in the camera coordinate system; Convert the target pixel point from the camera coordinate system to the radar coordinate system according to the camera coordinate information, the first depth feature, and a first conversion matrix from the camera coordinate system to the radar coordinate system to obtain radar coordinate information and a second depth feature of the target pixel point in the radar coordinate system; and According to the radar coordinate information, the second depth feature and the second transformation matrix from the radar coordinate system to the bird's-eye view coordinate system, the target pixel point is transformed from the radar coordinate system to the bird's-eye view coordinate system to obtain the first bird's-eye view coordinate information and the target depth feature of the target pixel point under the bird's-eye view.

3. The method according to claim 2, wherein: The multiple moments include a first moment and a second moment; the transforming of the target pixel from the radar coordinate system to the bird's-eye view coordinate system according to the radar coordinate information, the second depth feature, and a second transformation matrix from the radar coordinate system to the bird's-eye view coordinate system to obtain the first bird's-eye view coordinate information and the target depth feature of the target pixel at the bird's-eye view includes: Obtaining initial radar coordinate information and an initial second depth feature of the target pixel point collected at the first moment in the radar coordinate system at the first moment; aligning the initial radar coordinate information and the initial second depth feature to the radar coordinate system at the second moment according to the initial radar coordinate information, the initial second depth feature, and a third transformation matrix from the radar coordinate system at the first moment to the radar coordinate system at the second moment, to obtain target radar coordinate information and target second depth feature; and According to the target radar coordinate information, the target second depth feature, and the third transformation matrix from the radar coordinate system at the first moment to the bird's-eye view coordinate system, the target pixel point is transformed from the radar coordinate system at the first moment to the bird's-eye view coordinate system to obtain the first bird's-eye view coordinate information and target depth feature of the target pixel point collected at the first moment under the bird's-eye view.

4. The method according to claim 1, further comprising: Performing a voxelization operation on each of the plurality of point cloud information collected for the target area at the plurality of consecutive moments to obtain a time-series bird's-eye view image sequence, wherein the time-series bird's-eye view image sequence includes a plurality of bird's-eye view images represented by the plurality of point cloud information at the bird's-eye view perspective; For each of the bird's-eye view images, extract features of the bird's-eye view image to obtain bird's-eye view image features; fusing the second bird's-eye view coordinate information and the bird's-eye view image features corresponding to the same point cloud information to obtain a first time series fusion result sequence; The determining of the target temporal fusion result sequence according to the surround view image features, the target depth features, and the first bird's-eye view coordinate information corresponding to the same surround view image of the target comprises: fusing the surround view image features, target depth features, and the first bird's-eye view coordinate information corresponding to the same surround view image of the target to obtain a second temporal fusion result sequence; The first time series fusion result sequence and the second time series fusion result sequence corresponding to the same moment are fused to obtain the target time series fusion result sequence.

5. The method according to any one of claims 1 to 4, wherein The encoding and decoding of the target time series fusion result sequence to obtain the target detection result for the target area includes: Extracting features from the target time series fusion result sequence to obtain a fusion vector; Decoding the fusion vector and the position coding vector for detecting the position information of the object to be detected in the target area to obtain category information and position information of the object to be detected; and The target detection result is determined according to the category information and the location information.

6. The method according to claim 1, further comprising: Before performing perspective conversion on the target pixel point based on the pixel coordinate information and the initial depth feature of the target pixel point, An initial depth feature of the target pixel is determined according to a view cone point set corresponding to the target pixel, wherein the view cone point set represents a plurality of predicted depths determined for the target pixel.

7. A target detection device comprising: a first feature extraction module, configured to extract features from a plurality of target surround view images collected for a target area at a plurality of consecutive moments, respectively, to obtain a time-series surround view image feature sequence, wherein the time-series surround view image feature sequence includes a plurality of surround view image features corresponding to the plurality of target surround view images; a perspective conversion module, configured to perform perspective conversion on each target pixel in each target surround image based on the pixel coordinate information and the initial depth feature of the target pixel, to obtain the target depth feature and first bird's-eye view coordinate information of the target pixel under a bird's-eye view; a target fusion module, configured to determine a target temporal fusion result sequence based on the surround view image features, target depth features, and first bird's-eye view coordinate information corresponding to the same surround view image of the target; and The target detection module is used to perform encoding and decoding processing on the target time series fusion result sequence to obtain the target detection result for the target area.

8. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and device for detecting and recognizing vehicle looking-around target based on deep learning

    CN110827197A

  • Unified space-time fusion look-around aerial view perception method

    CN115115713A

  • Three-dimensional target detection method, electronic equipment and storage medium

    CN115331025A

  • Multi-camera fusion sensing method and device under view angle of aerial view

    CN115797454A

  • Target detection method and device, electronic equipment and storage medium

    CN117911973A

Cited By

  • Multi-channel video image splicing and target identification method and device

    CN121095555A

  • Multi-task three-dimensional sensing method and device, electronic equipment and readable storage medium

    CN121121152A

  • Target detection model training method, parking information processing method, device and equipment

    CN121505586A

  • Target detection model training method, parking information processing method, device and equipment

    CN121505586B