Object detection method, intelligent device, and storage medium
Through a target detection method based on convolutional neural network and fisheye camera model, the problem that monocular camera distance measurement cannot accurately detect objects higher than the ground is solved, and high-precision object detection and distance measurement are achieved, reducing costs.
Patent Information
- Application Number
- PCT/CN2024/140815
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-26
AI Technical Summary
Existing monocular camera distance measurement methods cannot accurately estimate the distance between objects higher than the ground and the camera, and are susceptible to light disturbances and environmental impacts. Alternative methods such as high-cost lidar are costly, large in size, and complex in deployment.
A target detection method is proposed. By obtaining the image to be detected, a two-dimensional information of the key point of the target is obtained based on the convolutional neural network, a three-dimensional information of the key point is determined in combination with the fisheye camera model, and the distance between the key point and the second target is calculated.
The precise detection of the distance between objects higher than the ground and the camera is achieved, the target detection accuracy is improved, the cost is reduced, and the shortcomings of the monocular camera distance measurement method are solved.
Smart Images

Figure CN2024140815_26062025_PF_FP_ABST
Abstract
Description
Target detection method, intelligent device and storage medium
[0001] This application claims priority to Chinese patent application CN202311786192.9, filed on December 22, 2023, with the invention name “Target Detection Method, Intelligent Device and Storage Medium”. The entire contents of the above Chinese patent application are incorporated into this application by reference. Technical Field
[0002] The present application relates to the field of target detection technology, and specifically provides a target detection method, an intelligent device, and a storage medium. Background Art
[0003] Currently, monocular ranging involves estimating the distance from a pixel to the camera based on a single RGB camera image. Compared to other ranging methods, such as binocular cameras and lidar, monocular ranging offers the advantages of low cost and power consumption. However, because the resulting image completely loses three-dimensional depth information, this method can only determine the distance from a pixel on the ground plane in the image after performing camera calibration on the ground.
[0004] Distance measurement methods based on monocular cameras can only calculate the distance between points on the ground and the camera; they cannot estimate the distance between objects above the ground and the camera. While they can estimate the distance between other vehicles and the vehicle itself using contact points such as tires, they cannot accurately estimate the distance between objects above the ground and the camera when the car door is open. Depth estimation methods based on monocular cameras are also susceptible to light disturbances, weather conditions, imaging noise, and other factors, and the accuracy of distance estimation is far from meeting the requirements of practical applications. Other distance measurement methods for objects above the ground primarily use cameras with depth measurement capabilities, such as lidar, binocular cameras, and time-of-flight cameras. However, these cameras are typically more expensive, larger, and more complex to deploy than monocular cameras.
[0005] Accordingly, a new target detection solution is needed in this field to solve the above problems.
[0006] Application Contents
[0007] In order to overcome the above-mentioned defects, the present application is proposed to provide a solution or at least partially solve the above-mentioned technical problems. The present application provides a target detection method, an intelligent device and a storage medium.
[0008] In a first aspect, the present application provides a target detection method, the method comprising:
[0009] Acquire an image to be detected containing a first target;
[0010] Acquire two-dimensional information of key points of the first target based on the image to be detected;
[0011] Determining three-dimensional information of the key point based on the two-dimensional information of the key point;
[0012] The distance between the key point and the second object is determined based on the three-dimensional information of the key point.
[0013] In one embodiment, determining the three-dimensional information of the key point based on the two-dimensional information of the key point includes:
[0014] Determine a first horizontal coordinate and a first vertical coordinate based on the two-dimensional information of the key point;
[0015] Obtaining a preset value of the height of the key point from the ground;
[0016] Determining a scaling factor based on the first horizontal coordinate, the first vertical coordinate, and a preset value representing the height of the key point from the ground;
[0017] The three-dimensional information of the key point is determined based on the first horizontal coordinate, the first vertical coordinate, and the scale factor.
[0018] In one embodiment, determining the first horizontal coordinate and the first vertical coordinate based on the two-dimensional information of the key point includes: inputting the two-dimensional information of the key point into a fisheye camera model, and outputting the first horizontal coordinate and the first vertical coordinate, wherein the fisheye camera model contains a correspondence between the two-dimensional coordinate of the key point and the first horizontal coordinate and the first vertical coordinate.
[0019] In one embodiment, the fisheye camera model is constructed by the following steps:
[0020] Calibrate the fisheye camera parameters to obtain the internal and external parameters of the fisheye camera;
[0021] The fisheye camera model is constructed based on the intrinsic parameters and the extrinsic parameters.
[0022] In one embodiment, the proportional coefficient is determined according to the following formula: s=(Z w -t2) / (r 21 ×X c +r 22 ×Y c +r 23 ×1)
[0023] Among them, s is the proportional coefficient, Z w is the preset value representing the height of the key point from the ground, t2, r 21 、r 22 、r 23 is the calibration external parameter of the fisheye camera, X c is the first horizontal coordinate, Y cis the first vertical coordinate.
[0024] In one embodiment, the three-dimensional information of the key point includes a second horizontal coordinate, a second vertical coordinate and a first vertical coordinate;
[0025] The determining of the three-dimensional information of the key point based on the first horizontal coordinate, the first vertical coordinate, and the scale coefficient includes:
[0026] determining the second abscissa based on a product of the first abscissa and the proportional coefficient;
[0027] determining the second ordinate based on a product of the first ordinate and the proportional coefficient;
[0028] The first vertical coordinate is determined based on the scale factor.
[0029] In one embodiment, acquiring the two-dimensional information of the key points of the first target based on the image to be detected includes: inputting the image to be detected into a convolutional neural network, and outputting the two-dimensional information of the key points of the first target.
[0030] In one embodiment, when the first target is a smart vehicle, the key point is a pixel point in the image to be detected that is at a preset position of a door handle of the smart vehicle.
[0031] In a second aspect, the present application provides an intelligent device comprising at least one processor and at least one memory, wherein the memory is suitable for storing multiple program codes, and the program codes are suitable for being loaded and run by the processor to execute any of the target detection methods described above.
[0032] In a third aspect, a computer-readable storage medium is provided, wherein a plurality of program codes are stored in the computer-readable storage medium, wherein the program codes are suitable for being loaded and run by a processor to execute any one of the aforementioned target detection methods.
[0033] The above one or more technical solutions of this application have at least one or more of the following beneficial effects:
[0034] The target detection method disclosed in this application specifically includes: acquiring an image containing a first target to be detected; obtaining two-dimensional information about key points of the first target based on the image to be detected; determining three-dimensional information about the key points based on the two-dimensional information about the key points; and determining the distance between the key points and a second target based on the three-dimensional information about the key points. This method solves the technical problem of being unable to predict the distance between pixels located at different heights above the ground and the camera, enabling accurate detection of the distance between objects above the ground and the camera, improving target detection accuracy, and saving costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The disclosure of this application will be more easily understood with reference to the accompanying drawings. Those skilled in the art will readily appreciate that these drawings are for illustrative purposes only and are not intended to limit the scope of protection of this application. Furthermore, similar numbers in the figures represent similar components, where:
[0036] FIG1 is a schematic flow chart of the main steps of a target detection method according to an embodiment of the present application;
[0037] FIG2 is a schematic diagram of a process for determining three-dimensional information of a key point based on two-dimensional information of the key point according to an embodiment of the present application;
[0038] FIG3 is a schematic structural diagram of a smart device according to an embodiment of the present application. DETAILED DESCRIPTION
[0039] Some embodiments of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are only used to explain the technical principles of the present application and are not intended to limit the scope of protection of the present application.
[0040] In the description of this application, "module" and "processor" may include hardware, software, or a combination of both. A module may include hardware circuitry, various suitable sensors, communication ports, and memory. It may also include software components, such as program code, or a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. A processor has data and / or signal processing capabilities. A processor may be implemented in software, hardware, or a combination of both. Non-transitory computer-readable storage media include any suitable medium capable of storing program code, such as magnetic disks, hard disks, optical disks, flash memory, read-only memory, random access memory, etc. The term "A and / or B" refers to all possible combinations of A and B, such as only A, only B, or both A and B. The terms "at least one of A or B" or "at least one of A and B" have similar meanings to "A and / or B" and may include only A, only B, or both A and B. The singular forms "a" and "the" may also include the plural forms.
[0041] Currently, traditional monocular camera-based ranging methods can only calculate the distance between points on the ground and the camera, but cannot estimate the distance between objects above the ground and the camera. Although the distance between other vehicles and the vehicle itself can be estimated by ground contact points such as tires, it is impossible to accurately estimate the distance between objects above the ground and the camera when the car door is open. Monocular-based depth estimation methods are also susceptible to light disturbances, weather conditions, imaging noise, and other factors, and the accuracy of estimated distances is far from meeting the requirements of practical applications. Other ranging methods for objects above the ground are mainly implemented using cameras with depth measurement capabilities such as lidar, binocular cameras, and time-of-flight cameras. However, these cameras are generally costly, bulky, and more complex to deploy than monocular cameras.
[0042] To this end, this application proposes a target detection method, intelligent device, and storage medium. The method specifically includes: obtaining an image containing a first target to be detected; obtaining two-dimensional information about key points of the first target based on the image to be detected; determining three-dimensional information about the key points based on the two-dimensional information about the key points; and determining the distance between the key points and a second target based on the three-dimensional information about the key points. This solves the technical problem of being unable to predict the distance between pixels located at different heights above the ground and the camera, enabling accurate detection of the distance between objects above the ground and the camera, improving target detection accuracy, and saving costs.
[0043] Although the application on a vehicle is used as an example, the first target and second target in the target detection method proposed in this application can be extended to any object, wherein a fisheye camera is set on the second target, and the image around the second target captured by the fisheye camera can be used to identify whether a preset target object (first target) is present in the image. The first target can be any target object around the second target that needs to be detected, especially a target object whose distance from the second target needs to be detected.
[0044] Please refer to FIG1 , which is a flowchart illustrating the main steps of a target detection method according to an embodiment of the present application.
[0045] As shown in FIG1 , the target detection method in the embodiment of the present application mainly includes the following steps S100 to S400 .
[0046] Step S100: Acquire an image to be detected containing a first target.
[0047] Specifically, a method for first identifying an image around the second target captured by the fisheye camera may include:
[0048] Detection is achieved by using a deep learning-based target detection algorithm. For example, a Region Proposal Network (RPN) is first used to generate candidate boxes that may contain objects. Each candidate box is then classified and regressed into bounding boxes to determine whether the image contains a preset target object (the first target) and its precise location.
[0049] For another example, directly use a convolutional neural network, such as the YOLO (You Only Look Once) series network and the SSD (Single Shot Detector) network to predict the categories and bounding boxes of all target objects in the image.
[0050] Traditional computer vision methods, such as the sliding window method, can also be used. Specifically, windows of different sizes and proportions are slid across the image, and feature extraction and classification are performed on the image area within each window to determine whether it contains the preset target object.
[0051] The above method can be used to identify and acquire the image to be detected containing the first target.
[0052] Step S200: Acquire two-dimensional information of key points of the first target based on the image to be detected.
[0053] A key point is a point on a first object that represents the distance between the first and second objects, particularly a point that best represents the minimum distance between the first and second objects. For example, to monitor whether a first object will collide with a second object, a point on the first object most likely to hit the second object can be determined. This could be the most prominent point on the object's structure, or, in a vehicle door-opening scenario, the door handle that moves the most when the door is opened and is most likely to hit the second object. Alternatively, a point within a preset distance range representing the outermost edge of the door can be determined, using the door handle as a reference.
[0054] The two-dimensional information of the key point includes the pixel position of the key point on the image to be detected.
[0055] Step S300: Determine the three-dimensional information of the key point based on the two-dimensional information of the key point.
[0056] The three-dimensional information of the key point includes the corresponding position of the key point of the first target in the real world coordinate system.
[0057] Step S400: determining the distance between the key point and a second target based on the three-dimensional information of the key point.
[0058] The distance between the key point and the second target is specifically the distance between the point on the first target object that is most likely to hit the second target and the second target.
[0059] For example, when the second target is the ego vehicle and the first target is a target vehicle around the ego vehicle, the key point may be a point on the door of the target vehicle that is most likely to hit the ego vehicle.
[0060] Based on steps S100-S400 described above, an image containing a first target to be detected is first acquired; two-dimensional information of key points of the first target is obtained based on the image to be detected; three-dimensional information of the key points is determined based on the two-dimensional information of the key points; and the distance between the key points and a second target is determined based on the three-dimensional information of the key points. This solves the technical problem of being unable to predict the distance between pixels located at different heights above the ground and the camera, enabling detection of the distance between objects above the ground and the camera, improving target detection accuracy, and saving costs.
[0061] The above steps S200 to S400 are further explained below.
[0062] In a specific embodiment, when the first target is a smart vehicle, the key point is a pixel point in the image to be detected that is at a preset position of a door handle of the smart vehicle, so as to represent the outermost end of the door.
[0063] The preset position can be a predetermined value, or can be adaptively modified or adjusted according to actual conditions, and is not specifically limited thereto. For example, 10 cm, 15 cm, 20 cm, 30 cm, etc. can all be examples of the preset position.
[0064] For predicting the distance between the handle of the smart vehicle and the second target when the door is opened, the key point at this time can be the pixel point at a preset position of the door handle of the smart vehicle in the image to be detected. This pixel point is the point on the door of the smart vehicle that is most likely to hit the second target (such as the vehicle's door handle, or a door point at a preset distance outside the door handle).
[0065] In this way, by using the pixel points at the preset position of the door handle of the smart vehicle in the image to be detected, the distance between the door of the smart vehicle and the second target when the smart vehicle opens the door can be accurately detected, which is conducive to improving the detection accuracy.
[0066] Regarding the above-mentioned step S200, in a specific embodiment, obtaining the two-dimensional information of the key points of the first target based on the image to be detected includes: inputting the image to be detected into a convolutional neural network, and outputting the two-dimensional information of the key points of the first target.
[0067] Convolutional Neural Networks (CNNs) are a type of feedforward neural network with a deep structure that incorporates convolutional computations. They are a representative algorithm for deep learning. CNNs possess the ability to learn representations and perform translation-invariant classification of input information based on their hierarchical structure.
[0068] For example, the convolutional neural network in this embodiment can be composed of three layers of convolution and two layers of deconvolution connected in sequence. First, unsupervised training is performed using the encoder-decoder structure. After the training reaches model convergence, the encoder structure is used to splice the upsampling structure to obtain a small-size output result with 4 channels.
[0069] In this embodiment, the convolutional neural network is a pre-trained network, and the convolutional neural network is trained according to different needs for different first targets and key points of preset first targets.
[0070] Specifically, the original image to be detected containing the first target is input into the trained convolutional neural network, and the two-dimensional information of the key points of the first target can be output.
[0071] Exemplarily, the original image to be detected containing the first target is input into the trained convolutional neural network, and the two-dimensional coordinates (u, v) of the pixel point at the preset position in the image to be detected can be output.
[0072] In this way, by using a convolutional neural network to detect the two-dimensional information of the key points of the first target, the two-dimensional coordinates of the pixel points with higher accuracy can be obtained.
[0073] The above is a further description of step S200 , and the following further describes step S300 .
[0074] Specifically, as shown in FIG2 , the above step S300 can be implemented through the following steps 301 to S304 .
[0075] Step 301: Determine a first horizontal coordinate and a first vertical coordinate based on the two-dimensional information of the key point.
[0076] The first horizontal coordinate and the first vertical coordinate are rays of the camera coordinate system corresponding to the key point of the first object.
[0077] In a specific embodiment, determining the first horizontal coordinate and the first vertical coordinate based on the two-dimensional information of the key point includes: inputting the two-dimensional information of the key point into a fisheye camera model, and outputting the first horizontal coordinate and the first vertical coordinate, wherein the fisheye camera model contains a correspondence between the two-dimensional coordinate of the key point and the first horizontal coordinate and the first vertical coordinate.
[0078] The fisheye camera model is a wide-angle camera model that can map points in three-dimensional space to a two-dimensional image plane and convert two-dimensional coordinates to three-dimensional coordinates. In this process, the camera's intrinsic and extrinsic parameters are used to transform the coordinate system.
[0079] In the fisheye camera model, each pixel on the image plane corresponds to a ray in the camera coordinate system. The starting point of this ray is the optical center of the camera, and the direction is the direction from the optical center of the camera through the pixel. Therefore, for a given pixel, the corresponding ray in the camera coordinate system can be calculated by back projection. Next, the three-dimensional coordinates corresponding to this pixel can be calculated by solving the intersection of this ray and the object in three-dimensional space. It should be noted that when performing coordinate transformation, the camera's intrinsic and extrinsic parameters are needed to transform the coordinates in the camera coordinate system into the world coordinate system. Specifically, the camera's rotation matrix and translation vector can be used to transform the coordinates in the camera coordinate system into the world coordinate system.
[0080] Specifically, in this embodiment, the fisheye camera model is used to determine the ray of the camera coordinate system corresponding to the key point of the first target, that is, the first horizontal coordinate X c and the first ordinate Y c .
[0081] In a specific embodiment, the fisheye camera model is constructed by the following steps: calibrating the parameters of the fisheye camera to obtain the intrinsic parameters and extrinsic parameters of the fisheye camera; and constructing the fisheye camera model based on the intrinsic parameters and the extrinsic parameters.
[0082] Specifically, we first calibrate the camera using a checkerboard grid to obtain the camera's intrinsic parameters and distortion parameters. Checkerboard calibration can be implemented using Python and the OpenCV library.
[0083] Next, we need to obtain the camera's extrinsic parameters, that is, the camera's position and posture in the world coordinate system. We can use a camera pose estimation algorithm, such as the PnP algorithm, to obtain the camera's extrinsic parameters. The calibrated camera extrinsic parameters are:
[0084] Finally, we need to construct a fisheye camera model based on the camera's intrinsic and extrinsic parameters. We can use a unified projection model to construct a fisheye camera model. This model is applicable to conventional, wide-angle, fisheye, and refractive cameras, and can formulate a closed-form Jacobian matrix.
[0085] Step 302: Obtain a preset value representing the height of the key point from the ground.
[0086] Specifically, a preset value representing the height of the key point from the ground is obtained based on the type of the first target. For example, when the first target is an intelligent vehicle, and the corresponding preset values for the height of the key point from the ground are different for sedans, sports cars, and trucks, the preset values can be adjusted based on the vehicle type of the first target. When the first target is not a vehicle, the preset values for the height of the key point from the ground can be adjusted based on the different target objects and the different types of target objects.
[0087] Step 303: Determine a proportional coefficient based on the first horizontal coordinate, the first vertical coordinate, and a preset value representing the height of the key point from the ground.
[0088] In a specific embodiment, the proportional coefficient is determined according to the following formula: s=(Z w -t2) / (r 21 ×X c +r 22 ×Y c +r 23 ×1)
[0089] Among them, s is the proportional coefficient, Z w is the preset value representing the height of the key point from the ground, t2, r 21 、r 22 、r 23 is the calibration extrinsic parameter of the fisheye camera, where is the offset of the z axis, r 21 、r 22 、r 23 is the rotation parameter of the coordinate system; X c is the first horizontal coordinate, Y c is the first vertical coordinate.
[0090] Step 304: Determine the three-dimensional information of the key point based on the first horizontal coordinate, the first vertical coordinate, and the scale factor.
[0091] In a specific embodiment, the three-dimensional information of the key point includes a second horizontal coordinate, a second vertical coordinate and a first vertical coordinate; determining the three-dimensional information of the key point based on the first horizontal coordinate, the first vertical coordinate and the proportional coefficient includes: determining the second horizontal coordinate based on the product of the first horizontal coordinate and the proportional coefficient; determining the second vertical coordinate based on the product of the first vertical coordinate and the proportional coefficient; and determining the first vertical coordinate based on the proportional coefficient.
[0092] Specifically, the three-dimensional coordinates of the key point of the first target are further determined based on the first horizontal coordinate, the first vertical coordinate and the scale factor. c The product of the scale factor s is used as the second horizontal coordinate (i.e. the horizontal coordinate of the key point), and the first vertical coordinate Yc The product of the proportional coefficient s determines the second vertical coordinate (ie, the vertical coordinate of the key point), and the proportional coefficient s is used as the first vertical coordinate (ie, the vertical coordinate of the key point).
[0093] In this way, the first horizontal coordinate and the first vertical coordinate are first determined according to the two-dimensional coordinates of the key point of the first target, and the three-dimensional coordinates of the key point of the first target are further determined according to the first horizontal coordinate, the first vertical coordinate and the proportional coefficient, thereby obtaining the three-dimensional information of the key point of the first target with higher accuracy.
[0094] The above is a further description of step S300 , and the following further describes step S400 .
[0095] With respect to step S400 , after determining the three-dimensional coordinates of the key point of the first object in the real-world coordinate system, the distance between the key point of the first object and the second object is calculated in combination with the size of the second object.
[0096] For example, when the second target is the ego vehicle and the first target is a target vehicle at a preset distance from the ego vehicle, the key point can be the point on the target vehicle's door that is most likely to hit the ego vehicle. In this case, combined with the aforementioned embodiments, the distance between the key point of the first target and the ego vehicle can be predicted. This provides a highly accurate distance between the key point of the target vehicle and the ego vehicle. By monitoring the key point of the target vehicle (the point on the door most likely to hit the ego vehicle) in real time, a collision warning can be provided to detect the opening of an external vehicle's door, thereby improving vehicle safety.
[0097] It should be pointed out that although the various steps in the above embodiments are described in a specific order, those skilled in the art will understand that in order to achieve the effect of the present application, different steps do not have to be performed in such an order. They can be performed simultaneously (in parallel) or in other orders. These changes are within the scope of protection of the present application.
[0098] It will be understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment of the present application can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium can include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.
[0099] Furthermore, the present application also provides a smart device. In an embodiment of a smart device according to the present application, as shown in Figure 3, the smart device includes at least one processor 31 and at least one memory 32. The memory 32 can be configured to store a program for executing the target detection method of the above-mentioned method embodiment, and the processor 31 can be configured to execute the program in the memory, which includes but is not limited to a program for executing the target detection method of the above-mentioned method embodiment. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method section of the embodiment of the present application.
[0100] In the embodiment of the present application, the intelligent device may be a control device device formed by various devices. In some possible implementations, the intelligent device may include multiple memories and multiple processors. The program for executing the target detection method of the above method embodiment can be divided into multiple subroutines, and each subroutine can be loaded and run by the processor to execute different steps of the target detection method of the above method embodiment. Specifically, each subroutine can be stored in different memories respectively, and each processor can be configured to execute the program in one or more memories to jointly implement the target detection method of the above method embodiment, that is, each processor executes different steps of the target detection method of the above method embodiment respectively to jointly implement the target detection method of the above method embodiment.
[0101] The multiple processors may be processors deployed on the same device. For example, the smart device may be a high-performance device composed of multiple processors, and the multiple processors may be processors configured on the high-performance device. Furthermore, the multiple processors may be processors deployed on different devices. For example, the smart device may be a server cluster, and the multiple processors may be processors on different servers in the server cluster.
[0102] Furthermore, the present application also provides a computer-readable storage medium. In a computer-readable storage medium embodiment according to the present application, the computer-readable storage medium can be configured to store a program for executing the target detection method of the above-mentioned method embodiment, and the program can be loaded and run by the processor to implement the above-mentioned target detection method. For ease of explanation, only the parts related to the embodiment of the present application are shown. For specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The computer-readable storage medium can be a memory device formed by various smart devices. Optionally, the computer-readable storage medium in the embodiment of the present application is a non-temporary computer-readable storage medium.
[0103] Thus far, the technical solutions of the present application have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of the present application is obviously not limited to these specific embodiments. Without departing from the principles of the present application, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present application.
Claims
1. A target detection method, characterized in that: The method comprises: Acquire an image to be detected containing a first target; Acquire two-dimensional information of key points of the first target based on the image to be detected; Determine the three-dimensional information of the key point based on the two-dimensional information of the key point; The distance between the key point and the second target is determined based on the three-dimensional information of the key point.
2. The target detection method according to claim 1, characterized in that: The determining the three-dimensional information of the key point based on the two-dimensional information of the key point includes: Determine a first horizontal coordinate and a first vertical coordinate based on the two-dimensional information of the key point; Obtaining a preset value representing the height of the key point from the ground; Determine a proportionality coefficient based on the first horizontal coordinate, the first vertical coordinate, and a preset value representing the height of the key point from the ground; The three-dimensional information of the key point is determined based on the first horizontal coordinate, the first vertical coordinate and the scale factor.
3. The target detection method according to claim 2, characterized in that: The determining of the first horizontal coordinate and the first vertical coordinate based on the two-dimensional information of the key point includes: inputting the two-dimensional information of the key point into a fisheye camera model, and outputting the first horizontal coordinate and the first vertical coordinate, wherein the fisheye camera model contains a correspondence between the two-dimensional coordinate of the key point and the first horizontal coordinate and the first vertical coordinate.
4. The target detection method according to claim 3, characterized in that: The fisheye camera model is constructed by the following steps: Calibrate the parameters of the fisheye camera to obtain the internal and external parameters of the fisheye camera; The fisheye camera model is constructed based on the intrinsic parameters and extrinsic parameters of the fisheye camera.
5. The target detection method according to claim 2, characterized in that: The proportionality coefficient is determined according to the following formula: s = (Z w -t2) / (r 21 ×X c +r 22 ×Y c +r 23 ×1) Among them, s is the proportionality coefficient, Z w is the preset value representing the height of the key point from the ground, t2, r 21 、r 22 、r 23 is the calibration external parameter of the fisheye camera, X c is the first horizontal coordinate, Y c is the first vertical coordinate.
6. The target detection method according to claim 2, characterized in that: The three-dimensional information of the key point includes a second horizontal coordinate, a second vertical coordinate and a first vertical coordinate; The determining the three-dimensional information of the key point based on the first horizontal coordinate, the first vertical coordinate and the proportional coefficient includes: determining the second horizontal coordinate based on the product of the first horizontal coordinate and the proportionality coefficient; determining the second ordinate based on the product of the first ordinate and the proportionality coefficient; The first vertical coordinate is determined based on the scale factor.
7. The target detection method according to claim 1, characterized in that: The acquiring the two-dimensional information of the key points of the first target based on the image to be detected includes: inputting the image to be detected into a convolutional neural network, and outputting the two-dimensional information of the key points of the first target.
8. The target detection method according to any one of claims 1 to 7, characterized in that: When the first target is a smart vehicle, the key point is a pixel point in the image to be detected that is at a preset position of a door handle of the smart vehicle.
9. An intelligent device, comprising at least one processor and at least one memory, wherein the memory is suitable for storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by the processor to execute the target detection method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a plurality of program codes, characterized in that: The program code is suitable for being loaded and run by a processor to execute the target detection method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Target detection method, electronic equipment, roadside equipment and cloud control platform
CN112668460A
Target detection method, intelligent device and storage medium
CN117788844A
Target detection method, electronic device and medium
US20220114759A1