An object detection method, device, apparatus and storage medium

By using a preset three-dimensional detection model and user visual cue images in the autonomous driving system to correct detection errors, the problem of the 3D object detector being unable to update in real time is solved, ensuring the safety and reliability of autonomous driving.

CN118918565BActive Publication Date: 2025-10-10SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410999623.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-24
Publication Date
2025-10-10
Estimated Expiration
2044-07-24

AI Technical Summary

Technical Problem

Existing 3D object detectors cannot update parameters in real time in autonomous driving systems, resulting in the inability to perceive new objects or new scenes, posing a safety risk.

Method used

By acquiring the image frame of the current scene, the preset 3D detection model is used to determine the target information, and combined with the visual cue image and historical output provided by the user, the object box and confidence level are generated, and screening is performed to correct detection errors.

Benefits of technology

It achieves rapid response and correction to detection errors, identifies unidentified target objects, and improves the reliability and safety of autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118918565B_ABST
    Figure CN118918565B_ABST
Patent Text Reader

Abstract

The application discloses an object detection method, device and equipment and a storage medium. The method comprises the following steps: acquiring an image frame of a current scene, and determining target information of the image frame by using a preset three-dimensional detection model; generating an object frame of an object in the image frame and corresponding confidence by using the preset three-dimensional detection model according to a visual prompt image and the target information, wherein the visual prompt image comprises a prompt image associated with a missed detection object determined by a user, and the missed detection object is determined according to historical output of the preset three-dimensional detection model; and screening the object frame according to the confidence to obtain a first target object frame. The technical scheme of the embodiment of the application responds to model missed detection quickly by using the visual prompt image determined by the user, immediately corrects detection errors, avoids risks caused by detection errors, and realizes reliable and safe automatic driving.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic driving, and in particular to an object detection method, device, equipment and storage medium. BACKGROUND

[0002] With the continuous development of information technology, 3D target detection based on vision plays a crucial role in the automatic driving system. The automatic driving scheme relies on accurate 3D detection results to predict the future driving behavior of other vehicles and plan the driving trajectory of the ego vehicle.

[0003] However, existing 3D target detectors usually follow an offline training and deployment process. Once the detector is trained and deployed on an autonomous vehicle, these offline solutions cannot update their parameters in real time and immediately correct errors when the detector cannot perceive new objects or new scenes due to a change in the application field. Such defects pose a significant safety risk to reliable driving systems, such as incorrect lane changes, turns, and even crashes caused by abnormal 3D detector results. SUMMARY

[0004] The present application provides an object detection method, device, equipment and storage medium to solve the problem of poor error correction effect of object detection.

[0005] In a first aspect, the present application provides an object detection method, comprising:

[0006] obtaining an image frame of a current scene, and determining target information of the image frame by using a preset three-dimensional detection model, wherein the target information comprises image feature information and depth information;

[0007] generating an object box and a corresponding confidence of an object in the image frame according to a visual prompt image and the target information by using the preset three-dimensional detection model, wherein the visual prompt image comprises a prompt image associated with a missed object determined by a user, and the missed object is determined according to historical output of the preset three-dimensional detection model;

[0008] screening the object box according to the confidence to obtain a first target object box.

[0009] In a second aspect, the present application provides an object detection device, comprising:

[0010] a target information determination module configured to obtain an image frame of a current scene, and determine target information of the image frame by using a preset three-dimensional detection model, wherein the target information comprises image feature information and depth information;

[0011] an object frame determination module, configured to generate an object frame and a corresponding confidence score for an object in the image frame based on a visual cue image and the target information using the preset 3D detection model, wherein the visual cue image includes a received user-determined cue image associated with a missed object, and the missed object is determined based on a historical output of the preset 3D detection model;

[0012] The object frame screening module is used to screen the object frame according to the confidence level to obtain a first target object frame.

[0013] In a third aspect, the present invention provides an electronic device, comprising:

[0014] at least one processor;

[0015] and a memory communicatively coupled to the at least one processor;

[0016] The memory stores a computer program that can be executed by at least one processor, and the computer program is executed by at least one processor so that the at least one processor can perform the object detection method of the first aspect.

[0017] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer instructions, which are used to enable a processor to implement the object detection method of the first aspect when executed.

[0018] The object detection solution provided by the present invention obtains an image frame of the current scene and uses a preset three-dimensional detection model to determine the target information of the image frame, wherein the target information includes image feature information and depth information. The preset three-dimensional detection model is used to generate an object frame and a corresponding confidence level of the object in the image frame based on a visual cue image and the target information. The visual cue image includes a received cue image associated with a missed object determined by the user, and the missed object is determined based on the historical output of the preset three-dimensional detection model. The object frame is screened based on the confidence level to obtain a first target object frame. By adopting the above technical solution, the visual cue image determined by the user based on the historical output of the preset three-dimensional detection model, as well as the image feature information and depth information of the image frame, is input into the preset three-dimensional detection model, thereby achieving a rapid response to model detection failures, immediately correcting detection errors, identifying previously unrecognized target objects, avoiding risks caused by detection errors, and achieving reliable and safe autonomous driving.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on these drawings.

[0021] Figure 1 is a flow chart of a kind of object detection method provided according to the embodiment one of the present application;

[0022] Figure 2 is a kind of object detection system schematic diagram provided according to the embodiment one of the present application;

[0023] Figure 3 is a flow chart of a kind of object detection method provided according to the embodiment two of the present application;

[0024] Figure 4 is a flow chart of a kind of object detection method provided according to the embodiment three of the present application;

[0025] Figure 5 is a flow chart of a kind of object detection method provided according to the embodiment four of the present application;

[0026] Figure 6 is a kind of object detection device structural schematic diagram provided according to the embodiment five of the present application;

[0027] Figure 7 is a kind of electronic equipment structural schematic diagram provided according to the embodiment six of the present application. DETAILED DESCRIPTION

[0028] In order to make the person in the art better understand the present application scheme, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In the description of the present application, "a plurality of" means two or more, unless otherwise specified. The association relationship of the associated objects is described, and it means that there can be three relationships, for example, A and / or B can represent: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0030] Embodiment one

[0031] Figure 1 A flowchart of an object detection method is provided for the first embodiment of the present application. The present embodiment can be applicable to the case of detecting objects by an intelligent driving vehicle. The method can be executed by an object detection device, which can be realized in the form of hardware and / or software. The object detection device can be configured in an electronic device, which can be composed of two or more physical entities or one physical entity.

[0032] As shown in Figure 1 , the object detection method provided by the first embodiment of the present application specifically includes the following steps:

[0033] S101, acquiring an image frame of a current scene, and determining target information of the image frame by using a preset three-dimensional detection model, wherein the target information includes image feature information and depth information.

[0034] In the present embodiment, the image frame of the current scene can be acquired by using the camera configured in the vehicle. The image feature information of the image frame can be extracted by using the sub-modules in the preset three-dimensional detection model, and the depth information of the current scene can be predicted based on the feature information. The preset three-dimensional detection model can be understood as an online 3D detector (Online 3DDetector, preset three-dimensional detection model).

[0035] S102. Utilizing the preset three-dimensional detection model, generate an object frame and a corresponding confidence score of the object in the image frame according to a visual cue image and the target information, wherein the visual cue image includes a received user-determined cue image associated with a missed object, and the missed object is determined based on a historical output of the preset three-dimensional detection model.

[0036] In this embodiment, the preset three-dimensional detection model can be used to evenly set points in the image frame (the points are the set points), and the position information of the set points can be determined. The position information is the position information of the default mode. The position information, the visual cue image and the target information are then input into the preset three-dimensional detection model to obtain the object frame and the corresponding confidence of the object in the image frame. The image feature information and the depth information can be used as the key and value of the feature decoder in the preset three-dimensional detection model. The preset three-dimensional detection model can predict the position information of the missed object based on the visual cue image. After the position information is feature-encoded, it can be used as a query for the feature decoder in the preset three-dimensional detection model. Among them, the preset three-dimensional detection model can output a three-dimensional object frame, object frame position, object frame size, object frame confidence and object frame depth, etc.

[0037] S103: Filter the object frame according to the confidence level to obtain a first target object frame.

[0038] In this embodiment, the object frame with a larger confidence level may be retained to obtain a first target object frame, thereby achieving detection of the object in the image frame.

[0039] The object detection method provided by the embodiment of the present invention obtains an image frame of the current scene and determines the target information of the image frame using a preset three-dimensional detection model, wherein the target information includes image feature information and depth information. The preset three-dimensional detection model is used to generate an object frame and a corresponding confidence level of the object in the image frame based on a visual cue image and the target information, wherein the visual cue image includes a received cue image associated with a missed object determined by the user, and the missed object is determined based on the historical output of the preset three-dimensional detection model. The object frame is screened based on the confidence level to obtain a first target object frame. The technical solution of the embodiment of the present invention inputs the visual cue image determined by the user based on the historical output of the preset three-dimensional detection model, as well as the image feature information and depth information of the image frame, into the preset three-dimensional detection model, thereby achieving a rapid response to model detection failures, immediately correcting detection errors, identifying previously unrecognized target objects, avoiding risks caused by detection errors, and achieving reliable and safe autonomous driving.

[0040] Optionally, the above method also includes: if the visual prompt image is empty, using the preset three-dimensional detection model to determine the position information of the set point in the image frame; inputting the position information into the prompt encoder in the preset three-dimensional detection model to obtain a first position code; inputting the target information and the first position code into the feature decoder in the preset three-dimensional detection model to obtain the object frame and corresponding confidence of the object in the image frame, and filtering the object frame according to the confidence to obtain a second target object frame.

[0041] Specifically, Figure 2 is a schematic diagram of an object detection system. If there is no visual cue image, that is, the visual cue image is empty, such as Figure 2 As shown, the position information of the default mode can be input into the prompt encoder, that is, the preset three-dimensional detection model is first used to evenly set points in the image frame (this point is the set point), and the position information of the set point is determined, and the position information is the position information of the default mode. After the position information is input into the prompt encoder in the preset three-dimensional detection model, a first position code can be obtained. After the target information and the first position code are input into the feature decoder in the preset three-dimensional detection model, the preset three-dimensional detection model can determine the object frame and the corresponding confidence of the object in the image frame based on the output result of the feature decoder. The object frame with a larger confidence can be screened out to obtain the second target object frame.

[0042] Example 2

[0043] Figure 3 This is a flow chart of an object detection method provided in Example 2 of the present invention. The technical solution of this embodiment of the present invention is further optimized based on the above-mentioned optional technical solutions, and provides a specific method for intelligent driving vehicles to detect objects.

[0044] Optionally, before using the preset three-dimensional detection model to generate the object frame and the corresponding confidence of the object in the image frame according to the visual prompt image and the target information, it also includes: receiving the associated points and / or associated frames of the missed objects in the historical image frame input by the user, and determining the position information of the associated points and / or the position information of the associated frames; using the preset three-dimensional detection model, according to the position information of the associated points and / or the position information of the associated frames, and the target information of the historical image frame, determining the visual prompt frame, and determining the image corresponding to the visual prompt frame in the historical image frame as the first visual prompt image. The advantage of this setting is that whenever the model misses an object, the user can provide user feedback information to the model by clicking or framing the unrecognized object in the image, and the detection error is immediately corrected by using the feedback information.

[0045] Optionally, the generating the object box of the object in the image frame and the corresponding confidence by using the preset three-dimensional detection model according to the visual prompt image and the target information comprises: inputting the first visual prompt image into a prompt encoder in the preset three-dimensional detection model to obtain second position encoding; and inputting the second position encoding and the target information into a feature decoder in the preset three-dimensional detection model to obtain the object box of the object in the image frame and the corresponding confidence. In this way, the prompt encoder and the feature decoder in the preset three-dimensional detection model are used to realize accurate error correction of the missed object.

[0046] As shown in Figure 3 The object detection method provided by the second embodiment of the present application specifically comprises the following steps:

[0047] S201, acquiring an image frame of a current scene, and determining target information of the image frame by using a preset three-dimensional detection model.

[0048] S202, judging whether a visual prompt image is empty, if yes, executing step 203, and if no, executing step 205.

[0049] Specifically, as shown in Figure 2 A visual prompt buffer can be set in advance, which is used to store visual prompt images corresponding to the undetected object and the object of interest determined by the user in the online inference process of the preset three-dimensional detection model.

[0050] S203, determining position information of a set point in the image frame by using the preset three-dimensional detection model, and inputting the position information into a prompt encoder in the preset three-dimensional detection model to obtain first position encoding.

[0051] S204, inputting the target information and the first position encoding into a feature decoder in the preset three-dimensional detection model to obtain an object box of an object in the image frame and corresponding confidence, and screening the object box according to the confidence to obtain a second target object box.

[0052] S205, receiving an associated point and / or an associated box of a missed object in a historical image frame input by a user, and determining position information of the associated point and / or position information of the associated box.

[0053] Specifically, as shown in Figure 2 Whenever the preset three-dimensional detection model makes a mistake (i.e. misses an object), the user can provide human feedback information for the model by clicking or box selecting an unrecognized object in the image. The clicked point is the associated point of the missed object, and the box selected image is the associated box. The historical image frame can be the previous frame or the previous N frames, and N is greater than 1.

[0054] S206. Using the preset three-dimensional detection model, determine a visual prompt frame according to the position information of the associated point and / or the position information of the associated frame, and the target information of the historical image frame, and determine the image corresponding to the visual prompt frame in the historical image frame as the first visual prompt image.

[0055] Specifically, the preset three-dimensional detection model can be used to extract the position information of the associated points and / or the position information of the associated boxes, and the visual prompt box can be determined based on these position information and the target information of the historical image frame. The image corresponding to the prompt box is the first visual prompt image.

[0056] Optional, such as Figure 2 As shown in the figure, the visual cue buffer includes an "entry" and "de-entry" mechanism. The "entry" mechanism refers to receiving visual cue images. The "de-entry" mechanism refers to discarding unnecessary visual cues in the buffer using preset rules to prevent the buffer from growing indefinitely. For example, the following are deleted: (1) visual cue images with low confidence output by the preset 3D detection model, the objects referred to by the images may not exist in the current scene; (2) visual cue images corresponding to detection boxes with high intersection-over-union (IoU) output by the preset 3D detection model, which means that these visual cue images may appear multiple times in the buffer. The advantage of this setting is that by maintaining the visual cue buffer, the input records of the visual cue images can be tracked and the visual cue image queue can be dynamically adjusted, so that the model can use the visual cue images in the visual cue buffer to detect and track objects that were previously missed online, thereby improving the detection robustness of the model and achieving rapid error correction.

[0057] S207: Input the first visual cue image into the cue encoder in the preset three-dimensional detection model to obtain a second position code.

[0058] S208: Input the second position code and the target information into a feature decoder in the preset three-dimensional detection model to obtain an object frame and a corresponding confidence level of the object in the image frame.

[0059] Specifically, the second position code can be used as the initialization of the query of the feature decoder. Figure 2 As shown, when a visual cue image is present (i.e., the visual cue image is not empty), the position information of the set point in the image frame and the visual cue image can be input into the cue encoder to obtain a second position code. The second position code and target information can then be input into a feature decoder in a preset 3D detection model. The 3D bounding box prediction module in the preset 3D detection model outputs the results of the feature decoder to output a 3D object bounding box, object bounding box position, object bounding box size, object bounding box confidence, and object bounding box depth.

[0060] S209: Filter the object frame according to the confidence level to obtain a first target object frame.

[0061] The object detection method provided by an embodiment of the present invention is such that whenever the model misses an object, the user can provide user feedback information to the model by clicking or framing the unrecognized object in the image. By utilizing this feedback information, the detection error is immediately corrected, and accurate error correction of missed objects is achieved by presetting the prompt encoder and feature decoder in the three-dimensional detection model.

[0062] Example 3

[0063] Figure 4 This is a flowchart of an object detection method provided in Example 3 of the present invention. The technical solution of this embodiment of the present invention is further optimized based on the above-mentioned optional technical solutions, and provides another specific method for intelligent driving vehicles to detect objects.

[0064] Optionally, before utilizing the preset 3D detection model to generate the object frame and corresponding confidence level of the object in the image frame based on the visual cue image and the target information, the method further includes: receiving an image block associated with a missed object in the historical image frame input by a user; and determining the image block as a second visual cue image. This arrangement has the advantage of allowing the preset 3D detection model to perform cross-scene and cross-timestamp error correction on object detection by receiving image blocks associated with missed objects in the historical image frame directly uploaded by the user, thereby enabling online error correction of the detection results and enabling safer and more robust autonomous driving of the vehicle.

[0065] Optionally, the use of the preset three-dimensional detection model to generate the object frame and corresponding confidence of the object in the image frame according to the visual cue image and the target information includes: inputting the second visual cue image into the visual alignment module in the preset three-dimensional detection model to obtain the candidate position of the object corresponding to the second visual cue image in the image frame; inputting the candidate position into the cue encoder in the preset three-dimensional detection model to obtain a third position code; inputting the third position code and the target information into the feature decoder in the preset three-dimensional detection model to obtain the object frame and corresponding confidence of the object in the image frame. The advantage of this setting is that the visual alignment of the visual cue image and the image frame is achieved by using the visual alignment module, thereby accurately predicting the candidate position of the object corresponding to the visual cue image in the image frame, further ensuring the accuracy of error correction.

[0066] like Figure 4 As shown, the object detection method provided by the third embodiment of the present invention specifically includes the following steps:

[0067] S301: Acquire an image frame of a current scene, and determine target information of the image frame using a preset three-dimensional detection model.

[0068] S302 , determining whether the visual prompt image is empty, if so, executing step 303 , if not, executing step 305 .

[0069] S303: Determine position information of a set point in the image frame using the preset 3D detection model, and input the position information into a prompt encoder in the preset 3D detection model to obtain a first position code.

[0070] S304: Input the target information and the first position code into the feature decoder in the preset three-dimensional detection model to obtain the object frame and the corresponding confidence of the object in the image frame, and filter the object frame according to the confidence to obtain a second target object frame.

[0071] S305: Receive an image block associated with a missed-detection object in a historical image frame input by a user.

[0072] S306: Determine the image block as a second visual prompt image.

[0073] Specifically, the visual cue images can also be image blocks associated with objects that were missed in historical image frames, such as images from the Internet. Users can directly upload customized image blocks corresponding to frequently missed objects based on their experience and use them as visual cue images.

[0074] S307: Input the second visual cue image into a visual alignment module in the preset three-dimensional detection model to obtain a candidate position of the object corresponding to the second visual cue image in the image frame.

[0075] Specifically, such as Figure 2 As shown, the visual cue (i.e., the second visual cue image) can be input into a visual alignment module of a preset 3D detection model. The module can visually align the second visual cue image with the current input image to obtain a predicted position after alignment, i.e., a candidate position of the object corresponding to the second visual cue image in the image frame.

[0076] Optionally, the second visual cue image is input into the visual alignment module in the preset three-dimensional detection model to obtain the candidate position of the object corresponding to the second visual cue image in the image frame, including: using the visual alignment module to determine the likelihood of the image feature information in the target information and the image features of the second visual cue image, and determining the candidate position of the object corresponding to the second visual cue image in the image frame based on the likelihood.

[0077] Specifically, the visual alignment module can calculate the likelihood between the image features in the target information and the image features in the second visual cue image, and output a candidate position based on the likelihood. There can be multiple candidate positions. During the supervised training phase of the visual alignment module, the position with the minimum Dice loss value can be determined as the candidate position. Since there can be multiple second visual cue images, multiple candidate positions can be obtained.

[0078] S308: Input the candidate position into a hint encoder in the preset three-dimensional detection model to obtain a third position code.

[0079] Specifically, the position information of the set point in the image frame and the candidate positions may be input into a prompt encoder to obtain a third position code.

[0080] S309: Input the third position code and the target information into the feature decoder in the preset three-dimensional detection model to obtain the object frame and the corresponding confidence of the object in the image frame, and filter the object frame according to the confidence to obtain a first target object frame.

[0081] The object detection method provided by an embodiment of the present invention receives image blocks directly uploaded by users and associated with missed objects in historical image frames, so that the preset three-dimensional detection model can correct object detection across scenes and timestamps, thereby realizing online error correction of detection results, allowing vehicles to perform autonomous driving more safely and robustly, and utilizing a visual alignment module to realize visual alignment of visual cue images and image frames, thereby accurately predicting the candidate position of the object corresponding to the visual cue image in the image frame, further ensuring the accuracy of error correction.

[0082] Example 4

[0083] Figure 5 This is a flowchart of an object detection method provided in Example 4 of the present invention. The technical solution of this embodiment of the present invention is further optimized based on the above-mentioned optional technical solutions, and provides another specific method for an intelligent driving vehicle to detect objects.

[0084] like Figure 5 As shown, the fourth embodiment of the present invention provides an object detection method, which specifically includes the following steps:

[0085] S401: Acquire an image frame of the current scene, and determine target information of the image frame using a preset three-dimensional detection model.

[0086] S402 , determining whether the visual prompt image is empty, if so, executing step 403 , if not, executing step 405 .

[0087] S403, determine the position information of the set point in the image frame by using the preset three-dimensional detection model, and input the position information into the prompt encoder in the preset three-dimensional detection model to obtain a first position code.

[0088] S404, input the target information and the first position code into the feature decoder in the preset three-dimensional detection model to obtain the object frame of the object in the image frame and the corresponding confidence, and screen the object frame according to the confidence to obtain a second target object frame.

[0089] S405, receive the associated point and / or associated frame of the missed object in the historical image frame input by the user, and determine the position information of the associated point and / or the position information of the associated frame; receive the image block associated with the missed object in the historical image frame input by the user.

[0090] S406, use the preset three-dimensional detection model to determine a visual prompt frame according to the position information of the associated point and / or the position information of the associated frame and the target information of the historical image frame, and determine the corresponding image of the visual prompt frame in the historical image frame as a first visual prompt image; determine the image block as a second visual prompt image.

[0091] S407, input the first visual prompt image into the prompt encoder in the preset three-dimensional detection model to obtain a second position code.

[0092] S408, use the visual alignment module to determine the likelihood of the image feature information in the target information and the image feature of the second visual prompt image, and determine the candidate position of the object corresponding to the second visual prompt image in the image frame according to the likelihood; input the candidate position into the prompt encoder in the preset three-dimensional detection model to obtain a third position code.

[0093] S409, input the second position code, the third position code and the target information into the feature decoder in the preset three-dimensional detection model to obtain the object frame of the object in the image frame and the corresponding confidence, and screen the object frame according to the confidence to obtain a first target object frame.

[0094] Specifically, the object detection method can be evaluated from the following three aspects:

[0095] 1. When using points, frames or image blocks for error correction detection, how does the preset three-dimensional detection model perform compared with the traditional offline 3D detection mode?

[0096] 2. How does the preset three-dimensional detection model perform in terms of instant online error correction?

[0097] 3. When facing scenes beyond the training distribution, how does the preset three-dimensional detection model enhance the performance of the offline trained detector?

[0098] The experimental setup includes the following four aspects:

[0099] 1) Dataset:

[0100] Experiments can be conducted on the nuScenes dataset, which contains 1000 autonomous driving sequences and is one of the most commonly used large-scale datasets in autonomous driving research.

[0101] 2) Tasks:

[0102] The first experiment aims to test the performance of the preset three-dimensional detection model in online error correction, which is conducted by simulating user intervention. Specifically, user feedback is simulated by comparing the distance between the true value and the detected 3D box. Objects within a 2-meter range without matching detection boxes are considered missed and added to the prompt buffer.

[0103] Then, the ability of the preset three-dimensional detection model to correct detection errors in scenarios beyond the training distribution is verified. The following four tasks are set:

[0104] Drop 80% of distant objects beyond 30 meters during training, and test the error correction of detecting distant objects.

[0105] Drop 80% of objects labeled as vehicles (including "car", "truck", "commercial vehicle", "bus", and "trailer") during training, and test the error correction of detecting vehicle objects.

[0106] Drop all objects of "truck" and "bus" categories during training, and test the preset three-dimensional detection model's zero-shot detection ability on these discarded categories.

[0107] Drop all data in "night" and "rain" scenarios during training, and test the improvement of the preset three-dimensional detection model in scenarios with domain gaps.

[0108] Finally, experiments are designed to verify the effectiveness of each specific prompt, including point prompts, box prompts, image block prompts, and default modes. Compare the performance between the four prompts, and test the robustness of visual cues from different perspectives by comparing the recall rate of future 15-frame detection using the current frame's visual cues.

[0109] 3) Entity Detection Score:

[0110] The preset three-dimensional detection model's ability to locate and detect new objects in the test phase is prioritized over its classification ability. Therefore, all class annotations of 3D objects are removed during training, and all objects are simply treated as entities. Similarly, the Entity Detection Score (EDS), which is a class-free version of the nuScenes Detection Score (NDS), is used to evaluate the performance of different models during evaluation.

[0111] 4) Implementation details:

[0112] The above method was implemented based on the mmDet3D codebase, and all experiments were conducted on a server with eight A100 GPUs. In the pre-built 3D detection model, a ResNet101 with an FPN neck was used as the image encoder to extract the target image feature Z. The visual cue encoder was a ResNet18. All visual cues were resized to 224×224 before being sent to the visual cue encoder. For the depth branch, the depth map was directly predicted using the head structure of Zoedepth, with sparse point clouds used for supervision. The loss function was the same as that of MonoDETR, encoding the depth map into high-dimensional features. During training, the AdamW optimizer was used with a batch size of 16, distributed evenly across eight GPUs. The learning rate was initialized to 2e-4 and adjusted using a cosine annealing strategy. When training point and box cues, user input was simulated by adding perturbations to the ground truth. To ensure robustness of the visual cues, the selected visual cues were not only from the image patches of the current frame, but also from randomly selected corresponding patches within the five frames before and after. Flipping and resizing operations were used as image-level data augmentation.

[0113] Test results:

[0114] 1) Effectiveness of pre-set 3D detection models in real-time error correction:

[0115] Online error correction is a core capability of the pre-built 3D detection model. Without requiring any training, the pre-built 3D detection model achieved a significant improvement of 4.7% in mAP (Mean Average Precision) and 5.0% in EDS (Entity Detection Score). This demonstrates the effectiveness of the pre-built 3D detection model in instantly correcting test errors during online inference.

[0116] 2) The effectiveness of the pre-set 3D detection model in scenarios beyond its training:

[0117] The performance of the preset three-dimensional detection model on the vehicle objects improves the mAP by 20.8% and the EDS by 21.6%. The effectiveness of the preset three-dimensional detection model in successfully detecting new objects not labeled in the training set, the mAP and the EDS on these new objects reach 8.5% and 13.6% respectively, significantly exceeding the 0.0% mAP and EDS of the offline baseline. The effectiveness of the preset three-dimensional detection model in online error correction when driving into a scene with domain transition provides an improvement of 4.5% in mAP and 4.7% in EDS. These experiments show that the preset three-dimensional detection model is an effective and versatile instant error correction system that performs well in handling missing long-range objects, unseen object categories, and scene domain transitions. The outstanding performance of the preset three-dimensional detection model in these challenging real-world scenarios highlights its potential to make 3D object detection systems robust and adaptable.

[0118] 3) Effectiveness of the preset three-dimensional detection model:

[0119] Finally, the effectiveness of the preset three-dimensional detection model in utilizing different cues is verified. First, the ability of the preset three-dimensional detection model to handle flexible visual cues is studied, including arbitrary views of target objects across scenes and time. For each video clip of the nuScenes validation set, an image patch of the target object in the first frame is used as a visual cue, and then the target object is detected and tracked in the subsequent frames, and the recall rate is calculated to evaluate the performance of the model in handling arbitrary visual cues. The results show that although there is a large difference in the angle of view and object pose between the first frame and the subsequent frames, the recall rate does not decrease significantly as the ego vehicle moves. This highlights the robustness of the preset three-dimensional detection model in handling diverse visual cues.

[0120] Then, the performance of the preset three-dimensional detection model in handling other interactive cues of objects in the target frame is evaluated, including object query, box, point, and visual cues (corresponding image patches of objects in the target frame). The results show that the preset three-dimensional detection model can effectively handle various forms of cues. In addition, compared with object query cues, the preset three-dimensional detection model can effectively utilize human interactive cues (box, point, and visual) for better 3D detection, with continuous improvements in mAP of 4.3%, 7.2%, and 5.4% respectively. This highlights the potential of error correction through interactive feedback to improve the online performance of offline detectors by utilizing human feedback.

[0121] Embodiment five

[0122] Figure 6 A structural schematic diagram of an object detection device provided by the embodiment five of the present application is shown in FIG. 1. As shown in the figure, the device comprises a target information determination module 501, an object box determination module 502, and an object box screening module 503. Figure 6 The target information determination module 501 is configured to determine target information of a target object in a target frame of a video. The target information comprises a target object class and a target object pose.

[0123] a target information determination module, configured to acquire an image frame of a current scene, and determine target information of the image frame by using a preset three-dimensional detection model, wherein the target information comprises image feature information and depth information;

[0124] an object frame determination module, configured to generate an object frame of an object in the image frame and a corresponding confidence by using the preset three-dimensional detection model according to a visual prompt image and the target information, wherein the visual prompt image comprises a prompt image associated with a missed object determined by a user, and the missed object is determined according to historical output of the preset three-dimensional detection model;

[0125] an object frame screening module, configured to screen the object frame according to the confidence to obtain a first target object frame.

[0126] The object detection device provided by the embodiment of the application inputs the visual prompt image determined by the user according to the historical output of the preset three-dimensional detection model, and the image feature information and the depth information of the image frame into the preset three-dimensional detection model, so that the model detection failure can be responded quickly, the detection error can be corrected immediately, the target object that has not been recognized before can be recognized, the risk caused by the detection error can be avoided, and reliable and safe automatic driving is achieved.

[0127] Optionally, the device further comprises:

[0128] a first position information determination module, configured to determine position information of a set point in the image frame by using the preset three-dimensional detection model if the visual prompt image is empty;

[0129] a first position encoding determination module, configured to input the position information into a prompt encoder in the preset three-dimensional detection model to obtain first position encoding;

[0130] a second target object frame determination module, configured to input the target information and the first position encoding into a feature decoder in the preset three-dimensional detection model to obtain an object frame of an object in the image frame and a corresponding confidence, and screen the object frame according to the confidence to obtain a second target object frame.

[0131] Optionally, the device further comprises:

[0132] a second position information determination module, configured to receive an associated point and / or an associated frame of a missed object in a historical image frame input by a user before the object frame determination module generates the object frame of the object in the image frame and the corresponding confidence by using the preset three-dimensional detection model according to the visual prompt image and the target information, and determine position information of the associated point and / or position information of the associated frame;

[0133] The first prompt image determination module is used to use the preset three-dimensional detection model to determine the visual prompt frame according to the position information of the associated point and / or the position information of the associated frame, and the target information of the historical image frame, and determine the image corresponding to the visual prompt frame in the historical image frame as the first visual prompt image.

[0134] Optionally, the object frame determination module includes:

[0135] a second position code determining unit, configured to input the first visual cue image into a cue encoder in the preset 3D detection model to obtain a second position code;

[0136] The first object frame determination unit is used to input the second position code and the target information into the feature decoder in the preset three-dimensional detection model to obtain the object frame of the object in the image frame and the corresponding confidence.

[0137] Optionally, the device further includes:

[0138] a receiving module, configured to receive an image block associated with a missed object in a historical image frame input by a user before generating an object bounding box and a corresponding confidence score of the object in the image frame based on the visual cue image and the target information using the preset three-dimensional detection model;

[0139] The second prompt image determining module is configured to determine the image block as a second visual prompt image.

[0140] Optionally, the object frame determination module includes:

[0141] a candidate position determining unit, configured to input the second visual cue image into a visual alignment module in the preset three-dimensional detection model to obtain a candidate position of the object corresponding to the second visual cue image in the image frame;

[0142] a third position code determining unit, configured to input the candidate position into a hint encoder in the preset 3D detection model to obtain a third position code;

[0143] The second object frame determination unit is used to input the third position code and the target information into the feature decoder in the preset three-dimensional detection model to obtain the object frame of the object in the image frame and the corresponding confidence.

[0144] Optionally, the inputting the second visual prompt image into the visual alignment module in the preset three-dimensional detection model to obtain the candidate position of the object corresponding to the second visual prompt image in the image frame comprises: determining, by the visual alignment module, a likelihood of image feature information in the target information and image features of the second visual prompt image, and determining the candidate position of the object corresponding to the second visual prompt image in the image frame according to the likelihood.

[0145] The object detection device provided by the embodiments of the present application can execute the object detection method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0146] Embodiment six

[0147] Figure 7 A structural schematic diagram of an electronic device 70 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0148] As shown in Figure 7 The electronic device 60 includes at least one processor 61 and a memory, such as a read-only memory (ROM) 62, a random access memory (RAM) 63, etc., which is communicatively connected to the at least one processor 61, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 61 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 62 or loaded from the storage unit 68 into the random access memory (RAM) 63. In the RAM 63, various programs and data required for the operation of the electronic device 60 can also be stored. The processor 61, the ROM 62, and the RAM 63 are connected to each other through a bus 64. An input / output (I / O) interface 65 is also connected to the bus 64.

[0149] Multiple components in the electronic device 60 are connected to the I / O interface 65, including an input unit 66, such as a keyboard, a mouse, etc.; an output unit 67, such as various types of displays, speakers, etc.; a storage unit 78, such as a magnetic disk, an optical disk, etc.; and a communication unit 69, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 69 allows the electronic device 60 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0150] The processor 61 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 61 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 61 executes the various methods and processes described above, such as the object detection method.

[0151] In some embodiments, the object detection method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 68. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 60 via ROM 62 and / or communication unit 69. When the computer program is loaded into RAM 63 and executed by processor 61, one or more steps of the object detection method described above can be performed. Alternatively, in other embodiments, processor 61 can be configured to perform the object detection method in any other suitable manner (e.g., via firmware).

[0152] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] A computer program for implementing the method of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor, implements the functions / operations specified in the flow diagrams and / or the block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package and partially on a remote machine, or entirely on a remote machine or server.

[0154] The computer device provided above can be used to execute the object detection method provided by any of the embodiments above, and has the corresponding functions and advantages.

[0155] Embodiment Seven

[0156] In the context of the present application, the computer-readable storage medium can be a tangible medium, the computer-executable instructions of which, when executed by a computer processor, are used to perform an object detection method, which comprises:

[0157] An image frame of a current scene is acquired, and a target information of the image frame is determined using a preset three-dimensional detection model, wherein the target information comprises image feature information and depth information;

[0158] An object frame of an object in the image frame and a corresponding confidence are generated using the preset three-dimensional detection model according to a visual prompt image and the target information, wherein the visual prompt image comprises a prompt image associated with a missed detection object determined by a user, and the missed detection object is determined according to historical output of the preset three-dimensional detection model;

[0159] The object frame is screened according to the confidence to obtain a first target object frame.

[0160] In the context of the present application, a computer readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer readable storage medium can be a machine readable signal medium. More specific examples of a machine readable storage medium will include one or more lines of a program of instructions in a transitory signal form, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0161] The computer device provided above can be used to execute the object detection method provided by any of the embodiments above, and has the corresponding functions and advantages.

[0162] It is worth noting that the embodiments of the object detection device described above include various units and modules only according to the logical division of functions, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy mutual distinction, and do not limit the protection scope of the present application.

[0163] Note that the above are only preferred embodiments of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described above, and those skilled in the art can make various obvious changes, readjustments and substitutions without departing from the scope of the present application. Therefore, although the present application has been described in more detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the appended claims.

Claims

1. An object detection method, characterized in that: include: Acquire an image frame of the current scene, and determine target information of the image frame using a preset three-dimensional detection model, wherein the target information includes image feature information and depth information; generating, using the preset 3D detection model, an object frame and a corresponding confidence score for the object in the image frame based on a visual cue image and the target information, wherein the visual cue image includes a received user-determined cue image associated with a missed object, the missed object is determined based on a historical output of the preset 3D detection model, and the object frame is a 3D object frame; The object frame is screened according to the confidence level to obtain a first target object frame.

2. The method according to claim 1, characterized in that Also includes: If the visual prompt image is empty, determining the position information of the set point in the image frame using the preset three-dimensional detection model; Inputting the position information into a prompt encoder in the preset three-dimensional detection model to obtain a first position code; The target information and the first position code are input into the feature decoder in the preset three-dimensional detection model to obtain the object frame and the corresponding confidence of the object in the image frame, and the object frame is filtered according to the confidence to obtain a second target object frame.

3. The method according to claim 1, characterized in that Before generating the object frame and the corresponding confidence score of the object in the image frame according to the visual cue image and the target information using the preset three-dimensional detection model, the method further includes: Receiving associated points and / or associated frames of missed objects in historical image frames input by a user, and determining position information of the associated points and / or position information of the associated frames; Using the preset three-dimensional detection model, a visual prompt frame is determined according to the position information of the associated point and / or the position information of the associated frame, as well as the target information of the historical image frame, and the image corresponding to the visual prompt frame in the historical image frame is determined as the visual prompt image.

4. The method according to claim 3, characterized in that The method of using the preset three-dimensional detection model to generate an object frame and a corresponding confidence score of the object in the image frame according to the visual cue image and the target information includes: Inputting the visual cue image into a cue encoder in the preset three-dimensional detection model to obtain a second position code; The second position code and the target information are input into a feature decoder in the preset three-dimensional detection model to obtain an object frame and a corresponding confidence level of the object in the image frame.

5. The method according to claim 1, wherein Before generating the object frame and the corresponding confidence score of the object in the image frame according to the visual cue image and the target information using the preset three-dimensional detection model, the method further includes: receiving an image block associated with an object missed in a historical image frame input by a user; The image block is determined as a visual cue image.

6. The method according to claim 5, characterized in that The method of using the preset three-dimensional detection model to generate an object frame and a corresponding confidence score of the object in the image frame according to the visual cue image and the target information includes: Inputting the visual cue image into a visual alignment module in the preset three-dimensional detection model to obtain a candidate position of the object corresponding to the visual cue image in the image frame; Inputting the candidate position into a hint encoder in the preset three-dimensional detection model to obtain a third position code; The third position code and the target information are input into a feature decoder in the preset three-dimensional detection model to obtain an object frame and a corresponding confidence level of the object in the image frame.

7. The method according to claim 6, characterized in that Inputting the visual cue image into a visual alignment module in the preset three-dimensional detection model to obtain a candidate position of an object corresponding to the visual cue image in the image frame includes: The visual alignment module is used to determine the likelihood of the image feature information in the target information and the image feature of the visual cue image, and the candidate position of the object corresponding to the visual cue image in the image frame is determined according to the likelihood.

8. An object detection device, characterized in that: include: a target information determination module, configured to obtain an image frame of the current scene and determine target information of the image frame using a preset three-dimensional detection model, wherein the target information includes image feature information and depth information; an object frame determination module, configured to generate an object frame and a corresponding confidence score for an object in the image frame based on a visual cue image and the target information using the preset 3D detection model, wherein the visual cue image includes a received user-determined cue image associated with a missed object, the missed object is determined based on a historical output of the preset 3D detection model, and the object frame is a 3D object frame; The object frame screening module is used to screen the object frame according to the confidence level to obtain a first target object frame.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so as to enable the at least one processor to perform the object detection method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the object detection method according to any one of claims 1 to 7 when executed.

Citation Information

Patent Citations

  • Target detection method and device based on visual prompt, equipment and storage medium

    CN117710644A