Object detection method and device
By time-series fusion and strengthening processing of the bird's-eye view features of target objects, the problems of inter-frame jumps and object occlusion in single-frame detection are solved, and the detection accuracy in autonomous driving is improved.
Patent Information
- Application Number
- CN202510328737.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-18
AI Technical Summary
Single-frame detection has inter-frame jumps and object occlusion problems in autonomous driving, resulting in low accuracy of detection results.
By obtaining the first aerial view feature and the second aerial view feature of the target object, and processing the first aerial view feature based on the mobile information, obtaining the third aerial view feature, performing timing fusion, including channel fusion, reorganization and strengthening processing of convolution and jump connections, and using multi-scale feature extraction and spatial cross attention modules to improve feature accuracy.
It improves the accuracy of target object detection, solves the problems of inter-frame jumps and object occlusion, and enhances the safety of autonomous driving.
Smart Images

Figure CN120339938A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of computers, and more particularly, to a method and apparatus for detecting an object. Background Art
[0002] With the development of visual detection technology, detecting target objects through images is applied in various fields. For example, vehicle detection, personnel detection, etc.
[0003] Currently, single-frame detection is usually used to detect target objects in images. However, single-frame detection has problems such as frame-to-frame jumps and object occlusion. Taking vehicle detection as an example, with the development of autonomous driving, vehicles need to be detected during the autonomous driving process, for example, detecting lane lines, vehicle driving directions, vehicle positions, vehicle speeds, etc. Single-frame detection has frame-to-frame jumps and object occlusion, which result in inaccurate detection results and affect the safety of autonomous driving.
[0004] In view of the above problems, there is currently no effective solution. Summary of the Invention
[0005] Embodiments of the present application provide a method and apparatus for detecting an object, so as to at least solve the technical problem in the related art that the accuracy of object detection results is relatively low due to frame-to-frame jumps and object occlusion in single-frame detection.
[0006] According to one aspect of the embodiments of the present application, a method for detecting an object is provided, including: obtaining a first bird's-eye view feature and a second bird's-eye view feature of a target object, where the first bird's-eye view feature is a feature obtained based on a first image collected at a first moment, the second bird's-eye view feature is a feature obtained based on a second image collected at a second moment, and the first moment is before the second moment; obtaining movement information of the target object, and processing the first bird's-eye view feature based on the movement information to obtain a third bird's-eye view feature; performing temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature; and determining a detection result of the target object through the target bird's-eye view feature.
[0007] In an exemplary embodiment, performing temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature includes: performing convolution and skip connection on the second bird's-eye view feature and the third bird's-eye view feature for channel fusion to obtain a fused bird's-eye view feature; reorganizing the fused bird's-eye view feature to obtain a reorganized bird's-eye view sequence; and obtaining the target bird's-eye view feature through the reorganized bird's-eye view sequence.
[0008] In an exemplary embodiment, the fused bird's-eye view features are reorganized to obtain a reorganized bird's-eye view sequence, including: discretely serializing the fused bird's-eye view features to obtain a bird's-eye view sequence; reorganizing the bird's-eye view sequence in N directions respectively to obtain N reorganized bird's-eye view sequences, where N is an integer greater than 1.
[0009] In an exemplary embodiment, the target bird's-eye view features are obtained by reorganizing the bird's-eye view sequence, including: inputting the N reorganized bird's-eye view sequences into a strengthening module, and performing feature strengthening processing on the N reorganized bird's-eye view sequences through the strengthening module to obtain N strengthened bird's-eye view sequences; restoring the order of each of the N strengthened bird's-eye view sequences in the N directions respectively to obtain N strengthened bird's-eye view sequences; and performing mean integration processing on the N strengthened bird's-eye view sequences to obtain the target bird's-eye view features.
[0010] In an exemplary embodiment, processing the first bird's-eye view features based on the movement information to obtain third bird's-eye view features, including: determining the time difference between the first moment and the second moment as the target time difference; obtaining the target distance through the target time difference and the movement speed of the target object; moving the first bird's-eye view features by the target distance in the moving direction of the target object to obtain the third bird's-eye view features; where the movement information includes the movement speed and the moving direction.
[0011] In an exemplary embodiment, before obtaining the first bird's-eye view features and the second bird's-eye view features of the target object, the method further includes: collecting images of the target object through M cameras disposed on the target object at the second moment to obtain M second images; after performing feature extraction on each of the second images, obtaining multi-scale features of each of the second images through a feature pyramid; and performing spatial cross-attention processing on the multi-scale features to obtain the second bird's-eye view features.
[0012] In an exemplary embodiment, determining the detection result of the target object through the target bird's-eye view features, including: inputting the target bird's-eye view features into a first detection head to identify lane lines through the first detection head; and / or inputting the target bird's-eye view features into a second detection head to obtain at least one of the following through the second detection head: a detection frame of the target object, a central position of the target object, and an angle of the target object.
[0013] According to another aspect of the embodiments of the present application, there is also provided a detection device for an object, including: a first acquisition module, configured to acquire a first bird's-eye view feature and a second bird's-eye view feature of a target object, where the first bird's-eye view feature is a feature obtained based on a first image acquired at a first moment, the second bird's-eye view feature is a feature obtained based on a second image acquired at a second moment, and the first moment is before the second moment; a second acquisition module, configured to acquire movement information of the target object and process the first bird's-eye view feature based on the movement information to obtain a third bird's-eye view feature; a fusion module, configured to perform temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature; and a determination module, configured to determine a detection result of the target object through the target bird's-eye view feature.
[0014] According to still another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0015] According to still another aspect of the embodiments of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the above method embodiments.
[0016] According to still another aspect of the embodiments of the present application, there is also provided an electronic device including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the steps in any one of the above method embodiments through the computer program.
[0017] Through the present application, since the first bird's-eye view feature and the second bird's-eye view feature of the target object are acquired, the first bird's-eye view feature is a feature obtained based on the first image acquired at the first moment, the second bird's-eye view feature and the second bird's-eye view feature are features obtained based on the second image acquired at the second moment, and the first moment is before the second moment; the movement information of the target object is acquired, and the first bird's-eye view feature is processed based on the movement information to obtain a third bird's-eye view feature; the second bird's-eye view feature and the third bird's-eye view feature are temporally fused to obtain a target bird's-eye view feature; and the detection result of the target object is determined through the target bird's-eye view feature. Therefore, the technical problem of low accuracy of the object detection result caused by frame-to-frame jump and object occlusion in single-frame detection in the related art can be solved, and the effect of improving the accuracy of target object detection can be achieved. Description of the Drawings
[0018] Figure 1 It is a schematic diagram of an application scenario of an object detection method according to an embodiment of the present application;
[0019] Figure 2 It is a schematic flowchart of an optional object detection method according to an embodiment of the present application;
[0020] Figure 3 It is a schematic diagram of an optional Query Re-arrange module according to an embodiment of the present application;
[0021] Figure 4 It is a schematic diagram of an optional temporal fusion module according to an embodiment of the present application;
[0022] Figure 5 It is a schematic diagram of an optional training process of a model according to an embodiment of the present application;
[0023] Figure 6 It is a structural block diagram of an optional object detection device according to an embodiment of the present application;
[0024] Figure 7 It is a structural block diagram of a computer system of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0025] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0027] According to one aspect of the embodiments of the present application, a method for detecting an object is provided. Optionally, in this embodiment, the above-mentioned method for detecting an object may be, but is not limited to, applied to a hardware environment including a terminal device 102 and a server 104 as shown in Figure 1 The server 104 can be connected to the terminal device 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal device 102 or the client installed on the terminal device 102. A database can be set up on the server 104 or independently of the server 104 to provide data storage services for the server 104.
[0028] The above-mentioned network may include, but is not limited to, at least one of the following: a wired network, a wireless network. The above-mentioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, a local area network. The above-mentioned wireless network may include, but is not limited to, at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be, but is not limited to, a PC (Personal Computer), a mobile phone, a tablet computer, etc. The server 104 may be, but is not limited to, a cloud server, a server cluster or other server types.
[0029] The method for detecting an object in the embodiments of the present application may be executed by the server 104, or may be executed by the terminal device 102, or may also be jointly executed by the server 104 and the terminal device 102. Among them, when the terminal device 102 executes the method for detecting an object in the embodiments of the present application, it may also be executed by the client installed thereon.
[0030] Taking the execution of the method for detecting an object in this embodiment by the server 104 as an example, Figure 2 is a schematic flowchart of an optional method for detecting an object according to the embodiments of the present application, as shown in Figure 2 The process of this method may include the following steps:
[0031] Step S202, obtaining a first bird's-eye view feature and a second bird's-eye view feature of a target object, where the first bird's-eye view feature is a feature obtained based on a first image collected at a first moment, and the second bird's-eye view feature is a feature obtained based on a second image collected at a second moment, and the first moment is before the second moment;
[0032] Among them, the above-mentioned target object is the object to be detected, including but not limited to any object to be detected such as a vehicle, a person, etc. The first image and the second image may be two adjacent frames of images, or there may be multiple frames of images between the first image and the second image. The acquisition time of the first image is earlier than that of the second image. The second moment may be the current moment, and the first moment is a historical moment before the current moment.
[0033] Step S204: Obtain the movement information of the target object, and process the first bird's-eye view feature based on the movement information to obtain a third bird's-eye view feature;
[0034] Among them, the above movement information can be CAN information. The CAN information of a vehicle refers to the vehicle diagnostic and control data transmitted through the Controller Area Network (CAN) bus, including information such as engine speed, vehicle speed, throttle opening, brake state, and vehicle position. These information can be used by various control systems of the vehicle to monitor and control the performance and driving state of the vehicle. CAN information plays a crucial role in modern vehicles and can help the vehicle monitor and adjust various parameters in real time to ensure the safety and performance of the vehicle.
[0035] Step S206: Perform temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature;
[0036] Step S208: Determine the detection result of the target object through the target bird's-eye view feature.
[0037] The object detection method in this embodiment can be applied to the field of image recognition and applied to the scenario of autonomous driving. During the autonomous driving of a vehicle, it is necessary to detect the vehicle, for example, detect lane lines, the position, driving direction, and moving speed of the vehicle. Currently, single-frame detection is often used, but single-frame detection has problems such as frame-to-frame jumps and object occlusion, resulting in inaccurate detection results and affecting the safety of autonomous driving.
[0038] Through the embodiment provided in this application, since the first bird's-eye view feature and the second bird's-eye view feature of the target object are obtained, the first bird's-eye view feature is the feature obtained based on the first image collected at the first moment, and the second bird's-eye view feature and the second bird's-eye view feature are the features obtained based on the second image collected at the second moment, and the first moment is before the second moment; the movement information of the target object is obtained, and the first bird's-eye view feature is processed based on the movement information to obtain a third bird's-eye view feature; the second bird's-eye view feature and the third bird's-eye view feature are temporally fused to obtain a target bird's-eye view feature; the detection result of the target object is determined through the target bird's-eye view feature. Therefore, the technical problem of low accuracy of object detection results caused by frame-to-frame jumps and object occlusion in single-frame detection in the related art can be solved, and the effect of improving the detection accuracy of the target object can be achieved.
[0039] In an exemplary embodiment, performing temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature includes: performing convolution and skip connection on the second bird's-eye view feature and the third bird's-eye view feature for channel fusion to obtain a fused bird's-eye view feature; reorganizing the fused bird's-eye view feature to obtain a reorganized bird's-eye view sequence; and obtaining the target bird's-eye view feature through the reorganized bird's-eye view sequence. In this embodiment, the third bird's-eye view feature is obtained based on the first bird's-eye view feature, and the target bird's-eye view feature combines the bird's-eye view feature at the current moment (the second moment) and the bird's-eye view feature at the historical moment (the first moment). This avoids the problems of frame-to-frame jump and object occlusion in the bird's-eye view feature of a single-frame image.
[0040] Feature extraction and channel fusion are performed on the second bird's-eye view feature and the third bird's-eye view feature through multiple convolutional layers in a convolutional neural network, and skip connections are used to retain more low-level and high-level features. The purpose of channel fusion is to combine the feature information of the input image to improve the representational ability and accuracy of the image. Skip connections can help solve the problems of vanishing gradients and exploding gradients, and at the same time can accelerate the convergence speed of the model and improve the performance of the model. Channel fusion in this way can effectively improve the accuracy and efficiency of image processing tasks.
[0041] In an exemplary embodiment, reorganizing the fused bird's-eye view feature to obtain a reorganized bird's-eye view sequence includes: discretely serializing the fused bird's-eye view feature to obtain a bird's-eye view sequence; and reorganizing the bird's-eye view sequence in N directions respectively to obtain N reorganized bird's-eye view sequences, where N is an integer greater than 1.
[0042] The above N directions can be set according to the actual situation, for example: up, down, left, right. As Figure 3 shown, the input is a bird's-eye view sequence, and the bird's-eye view sequence is reorganized in the four directions of up, down, left, and right to obtain reorganized bird's-eye view sequence 1, reorganized bird's-eye view sequence 2, reorganized bird's-eye view sequence 3, and reorganized bird's-eye view sequence 4. In this embodiment, the fused bird's-eye view feature is discretely serialized, and the bird's-eye view sequence is reorganized in different directions, so as to improve the global perception ability of the BEV space, learn the temporality and relevance between the features in the bird's-eye view sequence, and improve the accuracy of the detection result.
[0043] In an exemplary embodiment, obtaining the target bird's-eye view feature by reorganizing the bird's-eye view sequence includes: inputting N of the reorganized bird's-eye view sequences into a strengthening module, and performing feature strengthening processing on the N reorganized bird's-eye view sequences through the strengthening module to obtain N strengthened bird's-eye view sequences; restoring the order of each of the N strengthened bird's-eye view sequences according to the N directions to obtain N strengthened bird's-eye view sequences; and performing mean integration processing on the N strengthened bird's-eye view sequences to obtain the target bird's-eye view feature.
[0044] The strengthening module can be a Mamba2 model, a Transformer model based on the attention mechanism, or a State-Space Model. After strengthening the sequence features through the strengthening module, the output is obtained. Then, they are recombined to restore the original order to obtain N strengthened bird's-eye view sequences, and mean integration processing is performed on the N strengthened bird's-eye view sequences to obtain the target bird's-eye view feature. In this embodiment, by strengthening the reorganized bird's-eye view sequence, BEV feature strengthening is achieved, and the accuracy of the detection result is improved.
[0045] In an exemplary embodiment, processing the first bird's-eye view feature based on the movement information to obtain a third bird's-eye view feature includes: determining the time difference between the first moment and the second moment as the target time difference; obtaining the target distance through the target time difference and the movement speed of the target object; and moving the first bird's-eye view feature by the target distance in the moving direction of the target object to obtain the third bird's-eye view feature; where the movement information includes the movement speed and the moving direction.
[0046] In this embodiment, taking the target object as a vehicle as an example, the moving direction of the vehicle is forward. The target distance is obtained through the moving speed and the target time difference, and the first bird's-eye view feature is moved forward by the target distance to obtain the third bird's-eye view feature. This makes the bird's-eye view feature at the first moment and the second bird's-eye view feature at the second moment spatially consistent.
[0047] In an exemplary embodiment, before obtaining the first bird's-eye view feature and the second bird's-eye view feature of the target object, the method further includes: collecting images of the target object through M cameras disposed on the target object at the second moment to obtain M second images; after performing feature extraction on each of the second images, obtaining multi-scale features of each of the second images through a feature pyramid; and performing spatial cross-attention processing on the multi-scale features to obtain the second bird's-eye view feature.
[0048] The value of M can be determined according to actual conditions. For example, six cameras can be installed in different directions of the vehicle, and the surround view perception model can use Resnet to perform preliminary feature extraction on the image information collected by multiple cameras.
[0049] ResNet (Residual Network) is a deep residual network structure that solves the problems of gradient vanishing and gradient exploding in the deep neural network training process by introducing the idea of residual learning. By introducing skip connections and residual blocks to implement skip connections and residual learning in the network layer, the optimization difficulty in the training process can be reduced and the convergence speed of the network can be accelerated.
[0050] The feature pyramid of FPN is used to obtain multi-scale feature information. The feature pyramid of FPN (Feature Pyramid Network) is a multi-scale feature image pyramid structure used to extract feature information of different scales in the image. FPN establishes pyramid connections between feature images at different levels to effectively extract information of different scales in the image.
[0051] The main features of the feature pyramid include:
[0052] 1. Multi-scale information: FPN extracts and fuses information of different scales of the image by establishing connections between feature images at different levels, allowing the model to focus on objects of different scales in the image at the same time.
[0053] 2. Upsampling and downsampling: FPN achieves the matching of feature images of different resolutions through upsampling and downsampling operations, so that feature images at different levels can be effectively fused and information transferred.
[0054] 3. Horizontal connection: FPN uses horizontal connection to connect feature images of different scales, so that high-level feature images can obtain low-level detail information, thereby improving the detection and classification performance of the model.
[0055] In general, the feature pyramid structure of FPN can effectively extract multi-scale information in images and improve the detection and classification performance of the model.
[0056] The multi-scale feature information is fed into the Spatial Cross Attention (SCA) module to transform the information obtained from multiple cameras into the Bird's Eye View (BEV) space. The Spatial Cross Attention (SCA) module is an attention mechanism for processing spatial information. In SCA, the input feature map is divided into two parts, called query and key respectively, and then the similarity between them is calculated to determine the correlation between different positions. This enables the network to better capture the spatial information between different positions when processing the feature map, thereby improving the performance of the model.
[0057] In an exemplary embodiment, determining the detection result of the target object based on the target bird's eye view feature includes: inputting the target bird's eye view feature into a first detection head to identify lane lines through the first detection head; and / or inputting the target bird's eye view feature into a second detection head to obtain at least one of the following through the second detection head: the detection box of the target object, the central position of the target object, and the angle of the target object.
[0058] The above-mentioned first detection head is a lane line detection head for detecting lane lines. The second detection head is a vehicle detection head for outputting the detection box of the vehicle, the driving direction of the vehicle, the central position, the angle of the vehicle in the horizontal or vertical direction, and the speed, etc.
[0059] Through the above embodiments, information such as the position and driving direction of the vehicle can be accurately detected, providing effective vehicle information for autonomous driving technology.
[0060] The following explains the vehicle detection method in the embodiments of the present application in combination with optional examples. The vehicle detection process may include the following steps:
[0061] Step S401, collect the image information of 6 cameras around the vehicle. The collected data includes data information at different time periods (including the first image at the above-mentioned first moment and the second image at the second moment), different lighting conditions (including backlighting conditions of some cameras), different weather conditions, and different road conditions.
[0062] Step S402, clean the data and delete duplicate and obviously incorrect data.
[0063] Step S403, preliminarily extract features from the images (including the first image and the second image) collected by multiple cameras through Resnet;
[0064] Step S404, use the feature pyramid of FPN to obtain multi-scale feature information, including multi-scale feature information of the first image and multi-scale feature information of the second image;
[0065] Step S405: Feed the multi-scale feature information into the Spatial Cross Attention module (SCA) to transform the information obtained from multiple cameras into Bird's Eye View (BEV) features (including the above-mentioned first BEV feature and second BEV feature).
[0066] Step S406: Encode the CAN information of the ego vehicle through MLP, and convert the first BEV feature into the third BEV feature based on the moving speed.
[0067] Step S407: Input the second BEV feature and the third BEV feature into the Figure 4 shown Temporal Fusion module. In the figure, history is the third BEV feature and current is the second BEV feature. As shown in the figure, merge history (the third BEV feature, with a size of 256) and current (the second BEV feature, with a size of 256) to obtain a BEV feature with a size of 512. Then perform convolution and skip connection for channel fusion to obtain a fused BEV feature of 256. Through the Figure 3 shown Query Re-arrange module, after discretely serializing the fused BEV feature map, reorganize it in four directions: positive left, positive up, negative left, and negative up to form a new sequence ( Figure 3 shown reorganized BEV map sequences 1, 2, 3, 4) as the input of the Mamba2 model. Mamba2 strengthens the sequence features and then outputs. Then recombine to restore the original order, and calculate the mean to obtain the target BEV feature.
[0068] Step S408: Feed the target BEV feature into the lane line detection head and 3D object detection head to obtain the detection results.
[0069] Through this optional example, a long-term temporal feature fusion module (TemporalMamba) based on the novel spatial state model Mamba2. The Mamba2 model is introduced into the temporal module of the autonomous driving perception system. Mamba2 is a sequence processing model based on the Structured State Space Dual (SSD), with characteristics such as tensor parallelism, sequence parallelism, and variable length.
[0070] Through the multi-directional feature sequence scanning mechanism, after discretely serializing the fused BEV feature map, reorganize it in four directions: positive left, positive up, negative left, and negative up to form a new sequence (reorganized BEV map sequence) as the input of the Mamba2 model. Mamba2 strengthens the sequence features and then outputs. Then recombine to restore the original order. It enhances the global perception ability of the BEV space and the cross-temporal feature interaction ability, and improves the inference speed and detection ability of the model.
[0071] Regarding the training process of the model, as Figure 5As shown below, the specific implementation includes the following steps:
[0072] Step S501: Use a collection vehicle to collect image information of six surround cameras. The collected data includes data information under different time periods, different lighting conditions (including backlighting of some cameras), different weather conditions, and different road conditions.
[0073] Step S502: Data annotation: Obtain lane line data based on the high-precision map and the high-precision GPS information of the collection vehicle, clean the lane line information (exclude oncoming lane lines that cannot be detected by the camera, and integrate short lines into long lines), and optimize the training ground truth to improve the model performance. Represent the 3D objects around the collection vehicle in the format of (cx, cy, cz, w, l, h, rot, vx, vy). The class is described as car, truck, construction_vehicle, bus, trailer, barrier, motorcycle, bicycle, pedestrian, traffic_cone.
[0074] Step S503: Dataset division: Divide the dataset into a training set, a test set, and a validation set (used to evaluate whether the model converges).
[0075] Step S504: Model training: The model initially extracts features from the image information collected by multiple cameras through Resnet, obtains multi-scale feature information using the feature pyramid of FPN, and sends the multi-scale feature information into the spatial cross-attention module (SCA) to transform the information obtained by multiple cameras into the bird's-eye view space (BEV). At the same time, perform mlp encoding on the CAN information of the vehicle itself. Through the temporal fusion module, fuse the historical BEV features and the current BEV features using multiple convolutional layers and skip connections in the channel dimension. Discretize the sequence of fused BEV feature maps and use multi-directional rearrangement (four directions) as the input to the Mamba2 module to enhance the global perception ability of the BEV space. Send the BEV features obtained by temporal fusion into the lane line detection head and the 3D object detection head to obtain the detection results.
[0076] The model training adopts the AdamW optimization method, trains for a preset number of epochs (for example, 48 epochs), and the value of the loss Loss gradually converges. The value of mAP gradually increases. When it reaches the expected level, stop the training.
[0077] For the prominent problems faced in the field of pure vision autonomous driving perception, such as inaccurate single-frame detection results, target occlusion, and inaccurate speed prediction, the present application is based on the long-term fusion module (TemporalMamba) of the new spatial state model Mamba2, which solves the disadvantages of traditional temporal fusion based on the global attention mechanism, such as large video memory consumption and long inference time, strengthens the global perception ability of the BEV space, overcomes problems such as target occlusion and frame-to-frame jumps in target detection results faced by single-frame detection, and is conducive to improving the performance of the model in downstream tasks such as target detection, online mapping, and end-to-end autonomous driving.
[0078] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0079] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0080] According to another aspect of the embodiments of the present application, there is also provided a detection device for an object, which can be used to implement the detection method for the object provided in the above embodiments, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0081] Figure 6 is a structural block diagram of an optional detection device for an object according to the embodiments of the present application. As Figure 6 shown in, the detection device for the object includes:
[0082] The first acquisition module 602 is configured to acquire the first bird's-eye view feature and the second bird's-eye view feature of the target object, where the first bird's-eye view feature is a feature obtained based on the first image acquired at the first moment, the second bird's-eye view feature is a feature obtained based on the second image acquired at the second moment, and the first moment is before the second moment;
[0083] The second acquisition module 604 is configured to acquire the movement information of the target object, and process the first bird's-eye view feature based on the movement information to obtain a third bird's-eye view feature;
[0084] The fusion module 606 is configured to perform temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature;
[0085] The determination module 608 is configured to determine the detection result of the target object through the target bird's-eye view feature.
[0086] In an exemplary embodiment, the above device is further configured to perform convolution and skip connection on the second bird's-eye view feature and the third bird's-eye view feature for channel fusion to obtain a fused bird's-eye view feature; reorganize the fused bird's-eye view feature to obtain a reorganized bird's-eye view sequence; and obtain the target bird's-eye view feature through the reorganized bird's-eye view sequence.
[0087] In an exemplary embodiment, the above device is further configured to discretize the fused bird's-eye view feature to obtain a bird's-eye view sequence; reorganize the bird's-eye view sequence in N directions respectively to obtain N reorganized bird's-eye view sequences, where N is an integer greater than 1.
[0088] In an exemplary embodiment, the above device is further configured to input the N reorganized bird's-eye view sequences into a strengthening module, perform feature strengthening processing on the N reorganized bird's-eye view sequences through the strengthening module to obtain N strengthened bird's-eye view sequences; restore the order of each of the N strengthened bird's-eye view sequences in the N directions respectively to obtain N strengthened bird's-eye view sequences; and perform mean integration processing on the N strengthened bird's-eye view sequences to obtain the target bird's-eye view feature.
[0089] In an exemplary embodiment, the above device is further configured to determine the time difference between the first moment and the second moment as a target time difference; obtain a target distance through the target time difference and the moving speed of the target object; and move the first bird's-eye view feature by the target distance in the moving direction of the target object to obtain the third bird's-eye view feature; where the movement information includes the moving speed and the moving direction.
[0090] In an exemplary embodiment, the above-mentioned device is further configured to, before obtaining the first bird's-eye view feature and the second bird's-eye view feature of the target object, collect images of the target object through M cameras arranged on the target object at the second moment to obtain M second images; after performing feature extraction on each of the second images, obtain multi-scale features of each of the second images through a feature pyramid; perform spatial cross-attention processing on the multi-scale features to obtain the second bird's-eye view feature.
[0091] In an exemplary embodiment, the above-mentioned device is further configured to input the target bird's-eye view feature into a first detection head to identify lane lines through the first detection head; and / or input the target bird's-eye view feature into a second detection head to obtain at least one of the following through the second detection head: a detection frame of the target object, a central position of the target object, and an angle of the target object.
[0092] It should be noted that the above-mentioned various modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are separately located in different processors in any combination form.
[0093] According to another aspect of the embodiments of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it executes the steps in any one of the above method embodiments.
[0094] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a ROM, a RAM, a mobile hard disk, a magnetic disk, or an optical disc that can store computer programs.
[0095] According to another aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is configured to execute the steps in any one of the above method embodiments through the computer program. In an exemplary embodiment, the above-mentioned electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above-mentioned processor, and the input / output device is connected to the above-mentioned processor.
[0096] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0097] According to another aspect of the embodiments of the present application, a computer program product is further provided. The computer program product includes computer programs / instructions, and the computer programs / instructions contain program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit 1201, various functions provided by the embodiments of the present application are executed. The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0098] Figure 7 Schematically shown is a block diagram of a computer system of an electronic device for implementing the embodiments of the present application. As Figure 7 shown, the computer system 1200 includes a CPU (Central Processing Unit) 1201, which can execute various appropriate actions and processes according to the program stored in the ROM 1202 or the program loaded from the storage part 1208 into the RAM 1203. In the random access memory 1203, various programs and data required for system operation are also stored. The central processing unit 1201, the read-only memory 1202, and the random access memory 1203 are connected to each other through a bus 1204. The I / O (Input / Output) interface 1205 is also connected to the bus 1204.
[0099] The following components are connected to the I / O interface 1205: an input part 1206 including a keyboard, a mouse, etc.; an output part 1207 including such as a CRT (Cathode Ray Tube), an LCD (Liquid Crystal Display), etc. and a speaker, etc.; a storage part 1208 including a hard disk, etc.; and a communication part 1209 including a network interface card such as a local area network card, a modem, etc. The communication part 1209 performs communication processing via a network such as the Internet. The drive 1210 is also connected to the input / output interface 1205 as required. A removable medium 1211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1210 as required, so that the computer program read from it can be installed into the storage part 1208 as required.
[0100] In particular, according to the embodiments of the present application, the processes described in each method flow chart can be implemented as computer software programs. For example, embodiments of the present application include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by the central processing unit 1201, various functions defined in the system of the present application are executed.
[0101] It should be noted that Figure 7 The computer system 1200 of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0102] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present application is not limited to any specific combination of hardware and software.
[0103] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present application shall be included within the protection scope of the present application.
Claims
1. A method for detecting an object, characterized in that, Including: Obtain a first bird's-eye view feature and a second bird's-eye view feature of a target object, where the first bird's-eye view feature is a feature obtained based on a first image collected at a first moment, and the second bird's-eye view feature is a feature obtained based on a second image collected at a second moment, and the first moment is before the second moment; Obtain movement information of the target object, and process the first bird's-eye view feature based on the movement information to obtain a third bird's-eye view feature; Perform temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature; Determine a detection result of the target object through the target bird's-eye view feature.
2. The method according to claim 1, characterized in that, Performing temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature, including: Perform convolution and skip connection on the second bird's-eye view feature and the third bird's-eye view feature for channel fusion to obtain a fused bird's-eye view feature; Recombine the fused bird's-eye view feature to obtain a recombined bird's-eye view sequence; Obtain the target bird's-eye view feature through the recombined bird's-eye view sequence.
3. The method according to claim 2, wherein Recombining the fused bird's-eye view feature to obtain a recombined bird's-eye view sequence, including: Perform discrete serialization on the fused bird's-eye view feature to obtain a bird's-eye view sequence; Recombine the bird's-eye view sequence in N directions respectively to obtain N recombined bird's-eye view sequences, where N is an integer greater than 1.
4. The method according to claim 3, wherein Obtain the target bird's-eye view feature through the recombined bird's-eye view sequence, including: Input the N recombined bird's-eye view sequences into a strengthening module, and perform feature strengthening processing on the N recombined bird's-eye view sequences through the strengthening module to obtain N strengthened bird's-eye view sequences; Restore the order of each of the N strengthened bird's-eye view sequences in the N directions respectively to obtain N strengthened bird's-eye view sequences; Perform mean integration processing on the N strengthened bird's-eye view sequences to obtain the target bird's-eye view feature.
5. The method according to claim 1, characterized in that, Processing the first bird's-eye view feature based on the movement information to obtain a third bird's-eye view feature, including: Determine the time difference between the first moment and the second moment as the target time difference; Obtain a target distance through the target time difference and the moving speed of the target object; Move the first bird's-eye view feature by the target distance in the moving direction of the target object to obtain the third bird's-eye view feature; Wherein, the movement information includes the moving speed and the moving direction.
6. The method according to claim 1, characterized in that, Before obtaining the first bird's-eye view feature and the second bird's-eye view feature of the target object, the method further includes: Collect images of the target object through M cameras arranged on the target object at the second moment to obtain M second images; After performing feature extraction on each of the second images, obtain multi-scale features of each of the second images through a feature pyramid; Perform spatial cross-attention processing on the multi-scale features to obtain the second bird's-eye view feature.
7. The method according to claim 1, wherein Determining the detection result of the target object through the target bird's-eye view feature, including: Input the target bird's-eye view feature into a first detection head, and identify lane lines through the first detection head; and / or Input the target bird's-eye view feature into the second detection head, and obtain at least one of the following through the second detection head: the detection frame of the target object, the central position of the target object, and the angle of the target object.
8. A detection device for an object, characterized in that, Comprising: A first acquisition module, configured to acquire a first bird's-eye view feature and a second bird's-eye view feature of a target object, wherein the first bird's-eye view feature is a feature obtained based on a first image acquired at a first moment, the second bird's-eye view feature is a feature obtained based on a second image acquired at a second moment, and the first moment is before the second moment; A second acquisition module, configured to acquire the movement information of the target object, and process the first bird's-eye view feature based on the movement information to obtain a third bird's-eye view feature; A fusion module, configured to perform temporal fusion on the second bird's-eye view feature and the third bird's-eye view feature to obtain a target bird's-eye view feature; A determination module, configured to determine the detection result of the target object through the target bird's-eye view feature.
9. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.