Target contour detection method and device and electronic equipment
Through the ground point detection model and the pavement depth estimation model, the 3D edge profile information in the vehicle ranging of a monocular camera is restored, which solves the problem of low detection accuracy in the prior art, and achieves higher detection accuracy and better processing capabilities.
Patent Information
- Application Number
- CN202311623212.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the vehicle ranging method of a monocular camera has the problem of low detection accuracy, especially when small targets, occlusion conditions and road plane assumptions are affected by camera shake.
Through the grounding point detection model and pavement depth estimation model, the grounding condition and pavement depth information of the target object are detected, and the 3D edge profile information of the target is restored based on this information.
Compared with the prior art, the detection accuracy is higher, it can effectively handle small targets and occlusion conditions, and provide more accurate 3D information on sloped or undulating road surfaces.
Smart Images

Figure CN120070481A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image detection technology, and in particular to a method, device and electronic device for detecting the outline of an object. Background Art
[0002] For vehicle ranging with a monocular camera, the following three methods are usually used in the past: 1. Using the grounding wire of the pseudo 3D detection box (or the target grounding outline generated after semantic segmentation) and the road plane hypothesis to restore the target 3D information; 2. The Mono3D model directly predicts the 3D information of the target in the input image in one stage or two stages; 3. Restoring the vehicle 3D information according to the vehicle wheel grounding point and the prior information of the vehicle size.
[0003] However, the above three methods all have certain defects: in Scheme 1, there are errors in the grounding wire or grounding outline, especially for small targets, and it is impossible to judge and handle occlusion situations. The road plane hypothesis has certain errors due to camera shake, and this hypothesis cannot accurately estimate the road surface with slopes and undulations; Scheme 2 has poor prediction for truncated or occluded targets, large vehicles, special-shaped vehicles, etc., and requires a large amount of accurate target 3D ground truth; in Scheme 3, the accuracy of the vehicle size prior information is poor, and the size distributions of different vehicle models vary greatly. Summary of the Invention
[0004] The purpose of this application is to provide a method, device and electronic device for detecting the outline of an object. By using a grounding point detection model, the grounding situation of the target object is detected, and at the same time, the road surface depth information at the target grounding point is obtained through a road surface depth estimation model. Based on the above grounding situation and road surface depth information, the 3D edge contour information of the target is restored, and the detection accuracy is higher than that of the existing technology solutions.
[0005] In a first aspect, this application provides a method for detecting the outline of an object. The method includes: obtaining an image to be detected and a target image; the image to be detected is an image containing a target object captured by a camera; the target image is an image within the target detection box obtained by performing target object detection on the image to be detected; inputting the target image into a grounding point detection model, and outputting a grounding point detection result corresponding to the target image; the grounding point detection result includes: the position information of the target grounding point; inputting the image to be detected into a road surface depth estimation model, outputting a road surface depth detection result, and determining the depth information of the target grounding point based on the road surface depth detection result; determining the 3D edge contour information of the target object based on the position information and depth information of the target grounding point.
[0006] Furthermore, the above-mentioned ground contact point detection model includes: a first convolutional network encoder, a convolutional network decoder, a ground contact point detection network structure, and a ground contact point visibility classifier; wherein, the first convolutional network encoder is respectively connected to the convolutional network decoder and the ground contact point visibility classifier; the convolutional network decoder is further connected to the ground contact point detection network structure; the ground contact point visibility classifier is used to output the detection result of whether the ground contact point of the target object is visible; the ground contact point detection network structure is used to output the position information of the target ground contact point of the target object in the target image.
[0007] Furthermore, the training process of the above-mentioned ground contact point detection model is as follows: Obtain a target image training sample set; the samples in the sample set include: a spare target image obtained by detecting the target object, and the ground contact point category and ground contact point position information in the spare target image; Apply the samples in the target image training sample set to train the first convolutional network encoder, the convolutional network decoder, the ground contact point detection network structure, and the ground contact point visibility classifier until the model converges to obtain the ground contact point detection model.
[0008] Furthermore, during the model training process, the regression loss function is applied to calculate the ground contact point regression position error; the loss functions such as cross entropy are applied to calculate the classification error of whether the ground contact point is visible.
[0009] Furthermore, the above-mentioned road surface depth estimation model includes: a second convolutional network encoder and a depth encoder; the second convolutional network encoder is used to perform convolutional encoding processing on the image to be detected and output the encoding processing result; the depth encoder is used to perform depth encoding processing on the encoding processing result and output the road surface depth detection result.
[0010] Furthermore, the above-mentioned road surface depth detection result includes: the depth information of the road surface where the target object is located.
[0011] Furthermore, the step of determining the 3D edge contour information of the target object based on the position information and depth information of the target ground contact point includes: inputting the position information and depth information of the target ground contact point into the camera pinhole model to obtain the 3D edge contour information of the target object.
[0012] Second aspect, the present application also provides a target contour detection device, which includes: an image acquisition module, configured to acquire an image to be detected and a target image; the image to be detected is an image captured by a camera and containing a target object; the target image is an image within a target detection box obtained by performing target object detection on the image to be detected; a ground connection detection module, configured to input the target image into a ground connection point detection model and output a ground connection point detection result corresponding to the target image; the ground connection point detection result includes: position information of the target ground connection point; a depth detection module, configured to input the image to be detected into a road surface depth estimation model, output a road surface depth detection result, and determine depth information of the target ground connection point based on the road surface depth detection result; a contour restoration module, configured to determine 3D edge contour information of the target object based on the position information and the depth information of the target ground connection point.
[0013] Third aspect, the present application also provides an electronic device, including a processor and a memory, where the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method described in the first aspect above.
[0014] Fourth aspect, the present application also provides a computer-readable storage medium, which stores computer executable instructions. When the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the method described in the first aspect above.
[0015] In the target contour detection method, device and electronic device provided by the present application, an image to be detected and a target image are acquired; wherein, the image to be detected is an image captured by a camera and containing a target object; the target image is an image within a target detection box obtained by performing target object detection on the image to be detected; then the target image is input into a ground connection point detection model, and a ground connection point detection result corresponding to the target image is output; the ground connection point detection result includes: position information of the target ground connection point; the image to be detected is input into a road surface depth estimation model, a road surface depth detection result is output, and depth information of the target ground connection point is determined based on the road surface depth detection result; finally, 3D edge contour information of the target object is determined based on the position information and the depth information of the target ground connection point. In this way, the grounding condition of the target object is detected through the ground connection point detection model, and at the same time, the road surface depth information at the target ground connection point is obtained through the road surface depth estimation model. Based on the above grounding condition and road surface depth information, the 3D edge contour information of the target is restored, and the detection accuracy is higher than that of the existing technology solutions. Description of the Drawings
[0016] To more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the accompanying drawings required for the description of the specific embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0017] Figure 1 Flowchart of a target contour detection method provided by an embodiment of the present application;
[0018] Figure 2 Structural schematic diagram of a ground point detection model provided by an embodiment of the present application;
[0019] Figure 3 Schematic diagram of the output result of a ground point detection model provided by an embodiment of the present application;
[0020] Figure 4 Structural schematic diagram of a road surface depth estimation model provided by an embodiment of the present application;
[0021] Figure 5 Schematic diagram of the output result of a road surface depth estimation model provided by an embodiment of the present application;
[0022] Figure 6 Structural block diagram of a target contour detection device provided by an embodiment of the present application;
[0023] Figure 7 Structural schematic diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0024] The following will clearly and completely describe the technical solutions of the present application in combination with the embodiments. Obviously, the described embodiments are some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0025] In the existing target contour detection methods, there are problems of low detection accuracy. Based on this, the embodiments of the present application provide a target contour detection method, device and electronic device. Through the ground point detection model, the grounding situation of the target object is detected, and at the same time, the road surface depth information at the target ground point is obtained through the road surface depth estimation model. Based on the above grounding situation and road surface depth information, the 3D edge contour information of the target is restored, and the detection accuracy is higher than that of the existing technical solutions. For the convenience of understanding this embodiment, a target contour detection method disclosed in the embodiments of the present application will be introduced in detail first.
[0026] Figure 1 The flowchart of a target contour detection method provided by an embodiment of this application. The method specifically includes the following steps:
[0027] Step S102, obtain the image to be detected and the target image; the image to be detected is an image containing the target object captured by a camera; the target image is the image within the target detection frame obtained by performing target object detection on the image to be detected.
[0028] The above target object can be various vehicles, such as cars, tricycles, etc. The above target image can be an image cropped from the above image to be detected using a 2D target detection frame, that is, the target image determined by the detection frame during target object detection.
[0029] Step S104, input the target image into the ground contact point detection model, and output the ground contact point detection result corresponding to the target image; the ground contact point detection result includes: the position information of the target ground contact point.
[0030] In this embodiment, the above ground contact point detection model may include: a first convolutional network encoder, a convolutional network decoder, a ground contact point detection network structure, and a ground contact point visibility classifier; wherein, the first convolutional network encoder is respectively connected to the convolutional network decoder and the ground contact point visibility classifier; the convolutional network decoder is further connected to the ground contact point detection network structure; the ground contact point visibility classifier is used to output the detection result of whether the ground contact point of the target object is visible; the ground contact point detection network structure is used to output the position information of the target ground contact point of the target object in the target image.
[0031] The above ground contact point detection result includes: when the target object is grounded, the position information of the target ground contact point. For the detected ungrounded situation, the model does not output a result.
[0032] Step S106, input the image to be detected into the road surface depth estimation model, output the road surface depth detection result, and determine the depth information of the target ground contact point based on the road surface depth detection result;
[0033] The above road surface depth estimation model includes: a second convolutional network encoder and a depth encoder; the second convolutional network encoder is used to perform convolutional encoding processing on the image to be detected and output the encoding processing result; the depth encoder is used to perform depth encoding processing on the encoding processing result and output the road surface depth detection result.
[0034] The above road surface depth detection result includes: the depth information of the road surface where the target object is located. Further, restore the above target ground contact point to correspond to the road surface pixels in the image to be detected, so as to obtain the depth information of the target ground contact point.
[0035] Step S108: Determine the 3D edge contour information of the target object based on the position information and depth information of the target ground point.
[0036] In specific implementation, the 3D edge contour information of the target object can be restored through the camera pinhole model. In this embodiment, the occlusion situation of the target object can also be judged by whether the above-mentioned ground points are visible. For example, if all ground points are invisible, it must be occlusion. If it is determined which ground points are visible and which are invisible, and then combined with the mutual correlation information between the position of the target box and the box size, it can be inferred who occludes whom. That is, the method provided in the embodiment of the present application can not only effectively restore the 3D information of the grounded target object, but also further judge whether the target is occluded.
[0037] A target contour detection method provided by an embodiment of the present application detects the grounding situation of a target object through a ground point detection model, and at the same time obtains the road surface depth information at the target ground point through a road surface depth estimation model. Based on the above grounding situation and road surface depth information, the 3D edge contour information of the target is restored, and the detection accuracy is higher than that of the existing technology solutions.
[0038] The embodiment of the present application also provides another target contour detection method, which is implemented on the basis of the above embodiment; the model structure and training process are mainly described in this embodiment.
[0039] See Figure 2 As shown, the above-mentioned ground point detection model includes: a first convolutional network encoder, a convolutional network decoder, a ground point detection network structure, and a ground point visibility classifier; wherein, the first convolutional network encoder is respectively connected to the convolutional network decoder and the ground point visibility classifier; the convolutional network decoder is also connected to the ground point detection network structure; the ground point visibility classifier is used to output the detection result of whether the ground point of the target object is visible; if it is visible, output, if it is not visible, do not output. The ground point detection network structure is used to output the position information of the target ground point of the target object in the target image when the ground point is visible. The detection situation of the ground point of the target object is as Figure 3 shown, and multiple ground point positions are shown as circles in the figure.
[0040] The training process of the above-mentioned ground point detection model is as follows: obtain a target image training sample set; the samples in the sample set include: a spare target image obtained by detecting the target object, and the ground point category and ground point position information in the spare target image; apply the samples in the target image training sample set to train the first convolutional network encoder, the convolutional network decoder, the ground point detection network structure, and the ground point visibility classifier until the model converges to obtain the ground point detection model.
[0041] Further, during the model training process, regression loss functions such as L1 loss are applied to calculate the regression position error of the grounding point; loss functions such as cross-entropy are applied to calculate the classification error of whether the grounding point is visible.
[0042] See Figure 4 As shown, the above pavement depth estimation model includes: a second convolutional network encoder and a depth encoder; the second convolutional network encoder is used to perform convolutional encoding processing on the image to be detected and output the encoding processing result; the depth encoder is used to perform depth encoding processing on the encoding processing result and output the pavement depth detection result.
[0043] The input of the pavement depth estimation model is the image collected by the camera, that is, the aforementioned image to be detected, and the output is the depth obtained by projecting the fitted surface or plane of the pavement onto the image, as Figure 5 shown. That is, the above pavement depth detection result includes: the depth information of the road surface where the target object is located. Further, through pixel comparison, the depth information corresponding to the target grounding point can be determined.
[0044] Further, the step of determining the 3D edge contour information of the target object based on the position information and depth information of the target grounding point includes: inputting the position information and depth information of the target grounding point into the camera pinhole model to obtain the 3D edge contour information of the target object.
[0045] For example, if the pixel position (u, v) of the target grounding point is already known through the grounding point detection model, and the road surface depth z at the target grounding point is already known through the pavement depth estimation model, and the camera internal parameter is K, then the 3D position P(x, y, z) of the grounding point can be restored according to the camera pinhole model: P = K -1 [u, v, 1] T *z.
[0046] A target contour detection method provided by an embodiment of the present application can perform visibility and position prediction of the grounding point of the target object; then perform pavement depth estimation at the grounding point; and finally use the depth of the target grounding point to restore the 3D edge contour. The prediction results for the occluded and truncated parts of the target are more stable, and it is more convenient and stable to predict targets such as special-shaped vehicles or large vehicles.
[0047] Based on the above method embodiment, an embodiment of the present application further provides a target contour detection device. See Figure 6As shown in the figure, the device includes: an image acquisition module 52 for acquiring an image to be detected and a target image; the image to be detected is an image containing a target object captured by a camera; the target image is an image within a target detection box obtained by performing target object detection on the image to be detected; a ground connection detection module 54 for inputting the target image into a ground connection point detection model and outputting a ground connection point detection result corresponding to the target image; the ground connection point detection result includes: position information of the target ground connection point; a depth detection module 56 for inputting the image to be detected into a road surface depth estimation model, outputting a road surface depth detection result, and determining depth information of the target ground connection point based on the road surface depth detection result; a contour restoration module 58 for determining 3D edge contour information of the target object based on the position information and depth information of the target ground connection point.
[0048] Further, the above-mentioned ground connection point detection model includes: a first convolutional network encoder, a convolutional network decoder, a ground connection point detection network structure, and a ground connection point visibility classifier; wherein, the first convolutional network encoder is respectively connected to the convolutional network decoder and the ground connection point visibility classifier; the convolutional network decoder is further connected to the ground connection point detection network structure; the ground connection point visibility classifier is used to output a detection result on whether the ground connection point of the target object is visible; the ground connection point detection network structure is used to output position information of the target ground connection point of the target object in the target image.
[0049] Further, the above-mentioned device further includes: a model training module for performing the following training process of the ground connection point detection model: acquiring a target image training sample set; the samples in the sample set include: a standby target image obtained by target object detection, and the ground connection point category and ground connection point position information in the standby target image; applying the samples in the target image training sample set to train the first convolutional network encoder, the convolutional network decoder, the ground connection point detection network structure, and the ground connection point visibility classifier until the model converges to obtain the ground connection point detection model.
[0050] Further, during the above-mentioned model training process, a regression loss function is applied to calculate the ground connection point regression position error; a loss function such as cross-entropy is applied to calculate the classification error on whether the ground connection point is visible.
[0051] Further, the above-mentioned road surface depth estimation model includes: a second convolutional network encoder and a depth encoder; the second convolutional network encoder is used to perform convolutional encoding processing on the image to be detected and output an encoding processing result; the depth encoder is used to perform depth encoding processing on the encoding processing result and output a road surface depth detection result.
[0052] Further, the above-mentioned road surface depth detection result includes: depth information of the road surface where the target object is located.
[0053] Further, the above-mentioned contour restoration module 58 is configured to input the position information and depth information of the target ground point into the camera pinhole model to obtain the 3D edge contour information of the target object.
[0054] The device provided by the embodiment of the present application has the same implementation principle and the same technical effects as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiment.
[0055] The embodiment of the present application further provides an electronic device, as Figure 7 shown, which is a schematic structural diagram of the electronic device. Among them, the electronic device includes a processor 61 and a memory 60. The memory 60 stores computer-executable instructions that can be executed by the processor 61, and the processor 61 executes the computer-executable instructions to implement the above method.
[0056] In Figure 6 the illustrated embodiment, the electronic device further includes a bus 62 and a communication interface 63. Among them, the processor 61, the communication interface 63, and the memory 60 are connected through the bus 62.
[0057] Among them, the memory 60 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 63 (which may be wired or wireless), a communication connection is realized between the system network element and at least one other network element, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 62 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 62 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only a bidirectional arrow is used in
[0058] The processor 61 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 61 or the instructions in the form of software. The above-mentioned processor 61 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor 61 reads the information in the memory and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0059] The embodiments of the present application also provide a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above method. For the specific implementation, reference can be made to the foregoing method embodiments, and details will not be described herein again.
[0060] The computer program product of the method, device, and electronic device provided by the embodiments of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For the specific implementation, reference can be made to the method embodiments, and details will not be described herein again.
[0061] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present application.
[0062] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0063] In the description of this application, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0064] Finally, it should be noted that the above-mentioned embodiments are only specific implementation manners of this application, used to illustrate the technical solutions of this application, rather than limiting it. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed in this application can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application and should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A method for detecting a target contour, characterized in that, the method includes: Obtain an image to be detected and a target image; the image to be detected is an image containing a target object captured by a camera; the target image is an image within a target detection frame obtained by performing target object detection on the image to be detected; Input the target image into a ground point detection model, and output a ground point detection result corresponding to the target image; the ground point detection result includes: position information of the target ground point; Input the image to be detected into a road surface depth estimation model, output a road surface depth detection result, and determine depth information of the target ground point based on the road surface depth detection result; Determine 3D edge contour information of the target object based on the position information and depth information of the target ground point.
2. The method according to claim 1, characterized in that, the ground point detection model includes: a first convolutional network encoder, a convolutional network decoder, a ground point detection network structure, and a ground point visibility classifier; wherein, the first convolutional network encoder is respectively connected to the convolutional network decoder and the ground point visibility classifier; the convolutional network decoder is further connected to the ground point detection network structure; the ground point visibility classifier is used to output a detection result indicating whether the ground point of the target object is visible; the ground point detection network structure is used to output position information of the target ground point of the target object in the target image.
3. The method according to claim 2, characterized in that, the training process of the ground point detection model is as follows: Obtain a target image training sample set; the samples in the sample set include: a spare target image obtained by target object detection, and the ground point category and ground point position information in the spare target image; Use the samples in the target image training sample set to train the first convolutional network encoder, the convolutional network decoder, the ground point detection network structure, and the ground point visibility classifier until the model converges to obtain the ground point detection model.
4. The method according to claim 3, characterized in that, During the model training process, a regression loss function is used to calculate the ground point regression position error; a loss function such as cross entropy is used to calculate the classification error of whether the ground point is visible.
5. The method according to claim 1, characterized in that, the road surface depth estimation model includes: a second convolutional network encoder and a depth encoder; the second convolutional network encoder is used to perform convolutional encoding processing on the image to be detected and output an encoding processing result; the depth encoder is used to perform depth encoding processing on the encoding processing result and output a road surface depth detection result.
6. The method according to claim 1, characterized in that, the road surface depth detection result includes: depth information of the road surface where the target object is located.
7. The method according to claim 1, characterized in that, the step of determining 3D edge contour information of the target object based on the position information and depth information of the target ground point includes: Input the position information and depth information of the target grounding point into the camera pinhole model to obtain the 3D edge contour information of the target object.
8. A target contour detection device, characterized in that the device includes: an image acquisition module, configured to acquire an image to be detected and a target image; the image to be detected is an image containing a target object captured by a camera; the target image is an image within a target detection frame obtained by performing target object detection on the image to be detected; a grounding detection module, configured to input the target image into a grounding point detection model and output a grounding point detection result corresponding to the target image; the grounding point detection result includes: the position information of the target grounding point; a depth detection module, configured to input the image to be detected into a road surface depth estimation model, output a road surface depth detection result, and determine the depth information of the target grounding point based on the road surface depth detection result; a contour restoration module, configured to determine the 3D edge contour information of the target object based on the position information and depth information of the target grounding point.
9. An electronic device, characterized in that it includes a processor and a memory, the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer executable instructions, and when the computer executable instructions are called and executed by a processor, the computer executable instructions cause the processor to implement the method according to any one of claims 1 to 7.