Distance determination method and device, vehicle and storage medium
By looking around the camera to acquire image coordinates and map them to three-dimensional space, the problem that traditional vibration detection sensors cannot accurately determine the distance between the vehicle and the surrounding objects is solved, and higher distance determination accuracy and safety are achieved.
Patent Information
- Application Number
- CN202410685210.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2025-09-02
AI Technical Summary
Traditional parking safety solutions rely on vibration detection sensors and cannot fully understand the external environment, resulting in inaccurate determination of the distance between the vehicle and the surrounding objects, especially affected by the object's attitude.
The camera is used to obtain the image to be detected around the vehicle, the image coordinates are determined through object recognition, the camera's internal and external parameters are mapped to the three-dimensional space, and the distance between the object and the vehicle is calculated based on the monocular ranging technology.
It improves the accuracy of the determination of the distance between objects around the vehicle, increases the detection range, avoids the influence of the object position, and improves parking and driving safety.
Smart Images

Figure CN120580286A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent sensing technology, and in particular to a distance determination method, device, vehicle, and storage medium. Background Art
[0002] Traditional parking safety solutions rely primarily on vibration detection sensors, which can only sense the vehicle's own vibrations and lack a comprehensive understanding of the external environment. Consequently, their application scenarios and accuracy are limited. With the increasing prevalence of in-vehicle cameras, more and more vehicles are using information collected by external cameras to identify pedestrian threats around the vehicle.
[0003] In related technologies, the size or length and width of the detection frame of the object in the image is identified, and the distance between the pedestrian and the vehicle is judged based on the size or aspect ratio of the detection frame. This method is easily affected by the posture of the object, resulting in inaccurate distance determination between the vehicle and surrounding objects. Summary of the Invention
[0004] The present application aims to solve one of the technical problems in the related art at least to a certain extent.
[0005] To this end, the present application proposes a distance determination method, device, vehicle, and storage medium to improve the accuracy of distance determination between a first object around the vehicle and the vehicle.
[0006] In one aspect, an embodiment of the present application provides a distance determination method, including:
[0007] Obtain the image to be detected captured by the vehicle's target surround view camera;
[0008] performing object recognition on the image to be detected to determine the image coordinates of a grounding point of a first object in the image to be detected;
[0009] Mapping the image coordinates of the grounding point of the first object to a three-dimensional space according to the internal and external parameters of the target surround view camera to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space;
[0010] The distance between the first object and the vehicle is determined based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle.
[0011] Another aspect of the present application provides a distance determination device, including:
[0012] An acquisition module is used to acquire the image to be detected captured by the vehicle's target surround view camera;
[0013] an identification module, configured to perform object identification on the image to be detected and determine the image coordinates of a grounding point of a first object in the image to be detected;
[0014] a mapping module, configured to map the image coordinates of the grounding point of the first object to a three-dimensional space according to the internal and external parameters of the target surround view camera, so as to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space;
[0015] The determining module is configured to determine the distance between the first object and the vehicle based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle.
[0016] Another embodiment of the present application provides a vehicle, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method described in the above aspect is implemented.
[0017] Another aspect of the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method described in the aforementioned aspect is implemented.
[0018] Another embodiment of the present application provides a computer program product having a computer program stored thereon, which implements the method described in the above aspect when the program is executed by a processor.
[0019] The distance determination method, device, vehicle and storage medium proposed in the present application obtain the image to be detected captured by the target surround-view camera of the vehicle, perform object recognition on the image to be detected, determine the image coordinates of the grounding point of the first object in the image to be detected, and map the image coordinates of the grounding point of the first object to three-dimensional space according to the internal and external parameters of the target surround-view camera to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space. According to the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle, the distance between the first object and the vehicle is determined. The present application captures the image to be detected by the surround-view camera, thereby increasing the detection range around the vehicle. For the captured image to be detected, the image coordinates of the grounding point of the first object in the image to be detected in the corresponding image coordinate system are determined, and the image coordinates are mapped to the three-dimensional space according to the internal and external parameters of the camera to obtain the three-dimensional coordinates. The three-dimensional coordinates are not affected by the posture of the first object, thereby improving the accuracy of the distance determination between the first object and the vehicle.
[0020] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0022] Figure 1 A flow chart of a distance determination method provided in an embodiment of the present application;
[0023] Figure 2 A schematic diagram of a camera coordinate system for determining a grounding point according to an embodiment of the present application;
[0024] Figure 3 A flow chart of another distance determination method provided in an embodiment of the present application;
[0025] Figure 4 A schematic diagram of the installation position of a surround-view camera in a vehicle provided in an embodiment of the present application;
[0026] Figure 5 A schematic diagram of the division of the area surrounding a vehicle provided in an embodiment of the present application;
[0027] Figure 6 A schematic structural diagram of a distance determination device provided in an embodiment of the present application;
[0028] Figure 7 The present application is a schematic structural diagram of a vehicle provided in an embodiment. DETAILED DESCRIPTION
[0029] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0030] The following describes the distance determination method, device, vehicle, and storage medium according to embodiments of the present application with reference to the accompanying drawings.
[0031] Figure 1 A flow chart of a distance determination method provided in an embodiment of the present application.
[0032] The embodiment of the present disclosure is illustrated by taking the distance determination method as being configured in a distance determination device. The distance determination device can be applied to any vehicle-mounted device so that the vehicle-mounted device can perform a distance determination function.
[0033] like Figure 1 As shown, the method may include the following steps:
[0034] Step 101: Acquire an image to be detected captured by a target surround view camera of a vehicle.
[0035] In related art, the field of view of cameras installed in vehicles is relatively small, for example, less than 120 degrees. This results in a limited field of view and an inability to fully cover the vehicle's surroundings. This application utilizes a surround-view camera, which has a larger field of view, typically greater than 180 degrees. Multiple surround-view cameras are evenly distributed around the vehicle, ensuring there are no blind spots around the vehicle. The target surround-view camera is one of the multiple surround-view cameras.
[0036] Step 102 : performing object recognition on the image to be detected, and determining the image coordinates of a grounding point of a first object in the image to be detected.
[0037] In one implementation of an embodiment of the present application, the trained object recognition model has learned the mapping relationship between the input detection image and the image coordinates of the grounding point of the first object in the detection image, so that the trained object recognition model is used to perform object recognition on the image to be detected, and the image coordinates of the grounding point of the first object in the image to be detected are determined. The first object, that is, the object in the image to be detected, refers to a living object around the vehicle, such as a person or an animal. "First" is used to identify the object in the image to be detected. The grounding point refers to the contact point between the body part of the first object that contacts the ground and the ground, for example, the feet of the first object, including the left foot and / or the right foot. The height of the grounding point in three-dimensional space is determined to be zero. The coordinate system corresponding to the three-dimensional space, such as the body coordinate system of the vehicle, or the world coordinate system, wherein the body coordinate system is also called the vehicle coordinate system.
[0038] Step 103 : Mapping the image coordinates of the grounding point to a three-dimensional space according to the internal and external parameters of the target surround view camera to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space.
[0039] In an embodiment of the present application, the three-dimensional coordinates of the ground point in three-dimensional space can be inferred based on the assumption that the ground point height is zero. Specifically, first, the image coordinates of the ground point of the first object are mapped to the camera coordinate system using the intrinsic parameters of the target surround view camera, thereby obtaining the direction information between the ground point and the optical center of the target surround view camera in the camera coordinate system. The intrinsic parameters include the intrinsic parameter matrix K, which is K = [[f_x, s, c_x], [0, f_y, c_y], [0, 0, 1]], and the mapping formula from the image coordinate system to the camera coordinate system is as follows:
[0040] [X_c, Y_c, Z_c] = K -1 *[u,v,1];
[0041] Where f_x and f_y are the focal lengths, u and v are the pixel coordinates of the ground point in the image coordinate system, s is the skew coefficient, and c_x and c_y are the coordinates of the image center point.
[0042] Through the above mapping, the image coordinates of the grounding point in the image coordinate system, also called pixel coordinates, can be mapped to the camera coordinate system. Since points in the same direction correspond to the same point on the imaging plane of the camera coordinate system, after mapping the grounding point to the camera coordinate system, the positions of multiple points in one direction are determined, that is, at least one candidate camera coordinate of the grounding point in the camera coordinate system is determined. Therefore, based on the optical center of the camera and the at least one candidate camera coordinate of the grounding point, the direction information between the optical center of the camera and the at least one candidate camera coordinate of the grounding point can be determined. Therefore, based on the setting that the height of the grounding point is zero, the three-dimensional coordinates of the grounding point in the camera coordinate system can be inferred. As an example, if Figure 2 As shown in , based on the direction information of the determined ground point and the camera optical center, a triangle can be constructed, and the length of each side can be determined by trigonometric functions. Figure 2 As shown, the vertical distance between the optical center of the target surround camera and the ground is obtained, that is, Figure 2 In h, the direction information can be used Figure 2 The angle θ in represents the target camera coordinates of the ground point of the first object in the camera coordinate system, determined based on the direction information and vertical distance. Furthermore, the extrinsic parameters of the target surround camera are used to map the target camera coordinates of the ground point of the first object in the camera coordinate system to the vehicle coordinate system, thereby determining the three-dimensional coordinates of the first object in the vehicle coordinate system. The vehicle coordinate system is a coordinate system with a single location on the vehicle as its origin, typically the center of mass of the vehicle.
[0043] The mapping relationship between the camera coordinate system and the vehicle coordinate system is as follows:
[0044] The external parameters include the external parameter matrix [R|t]: [R,t; 0,1];
[0045] Mapping formula: [X_w, Y_w, Z_w, 1] = [R, t; 0, 1] * [X_c, Y_c, Z_c, 1];
[0046] Where R is the rotation matrix and t is the translation vector. X_c, Y_c, Z_c are the target camera coordinates of the ground point in the camera coordinate system, and X_w, Y_w, Z_w are the coordinates of the ground point in the vehicle coordinate system.
[0047] It should be noted that this application describes the ranging principle using a pinhole imaging model. This application is also applicable to other camera models. It only requires changing the corresponding camera model formula in the internal parameter transformation step. This is not limited in the embodiments of this application. This embodiment of the application combines the ground point prediction technology of the first object with the monocular ranging technology to more accurately calculate the distance between the first object and the vehicle.
[0048] Step 104 : Determine the distance between the first object and the vehicle based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle.
[0049] In an embodiment of the present application, when the coordinates of the grounding point of the first object in the vehicle body coordinate system are determined, the three-dimensional coordinates of the reference point corresponding to the vehicle in the vehicle body coordinate system can be determined based on the positioning device on the vehicle. Thus, according to the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the reference point corresponding to the vehicle in the vehicle body coordinate system, the distance between the first object and the vehicle can be determined, thereby improving the accuracy of distance determination.
[0050] It should be noted that a vehicle is typically equipped with multiple surround-view cameras. Steps 101 through 104 of the present embodiment can be repeated multiple times to determine the distance between the first object in the images captured by each surround-view camera and the vehicle, thereby enabling comprehensive monitoring of the vehicle's surroundings using multiple surround-view cameras and improving the safety of vehicle parking and driving. Furthermore, for each image to be detected captured by a surround-view camera, the first object identified through object recognition can be one or more. If there are multiple first objects, the three-dimensional coordinates of the grounding point of each first object can be determined according to Steps 103 and 104 above. Furthermore, the distance between the grounding point of each first object and the vehicle can be determined, and the distance between the grounding point of each first object and the vehicle can be used as the distance between each first object and the vehicle.
[0051] In the distance determination method of the embodiment of the present application, the image to be detected captured by the target surround-view camera of the vehicle is obtained, object recognition is performed on the image to be detected, the image coordinates of the grounding point of the first object in the image to be detected are determined, and the image coordinates of the grounding point of the first object are mapped to the three-dimensional space according to the internal and external parameters of the target surround-view camera to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space. The distance between the first object and the vehicle is determined according to the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle. The present application captures the image to be detected by the surround-view camera, thereby increasing the detection range around the vehicle. For the captured image to be detected, the image coordinates of the grounding point of the first object in the image to be detected in the corresponding image coordinate system are determined, and the image coordinates are mapped to the three-dimensional space according to the internal and external parameters of the camera to obtain the three-dimensional coordinates, which are not affected by the posture of the first object, thereby improving the accuracy of the distance determination between the first object and the vehicle.
[0052] Based on the above embodiments, Figure 3 A flow chart of another distance determination method provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the method comprises the following steps:
[0053] Step 301: Acquire an image to be detected captured by a target surround view camera of a vehicle.
[0054] Among them, step 301 can refer to the explanation in the above embodiment, the principle is the same, and it will not be repeated here.
[0055] As an example, Figure 4 As shown, Figure 4 The positions of the four surround-view cameras installed in the vehicle are shown in FIG. 1 , where the cameras marked 10 , 20 , 30 , and 40 are the positions of the four surround-view cameras.
[0056] Step 302: Input the image to be detected into the trained object recognition model to obtain the image coordinates of the grounding point of the first object.
[0057] Among them, the object recognition model includes a processing module, a feature extraction module, a feature fusion module and a task module.
[0058] As an implementation method, first, the image to be detected is input into the processing module of the object recognition model for downsampling to obtain a downsampled image. The size of the image to be processed is reduced by downsampling to improve processing efficiency. As an implementation method, the image to be detected is scaled to a specified size while maintaining the aspect ratio, and the remaining part is filled with 114. If the channel arrangement order of the image to be detected is in the BGR format of blue, green, and red, the BGR image is converted into an RGB image.
[0059] Secondly, the downsampled image is input into the feature extraction module of the object recognition model for multi-scale feature extraction to obtain feature maps of multiple scales. The feature maps of different scales contain features with different granularity. Among them, the large-scale feature map contains finer features, such as fine-grained features such as a person's skin color, hair, and eye size. The smaller the scale, the coarser the granularity of the features included, such as coarse-grained features such as a person's height, weight, etc.
[0060] Furthermore, the feature maps of multiple scales are input into the feature fusion module of the object recognition model for feature fusion to obtain a fused feature map, wherein the fused feature map includes channel feature attention and spatial attention, that is, a three-dimensional matrix is weighted in depth and in plane space respectively. This is because the present application focuses on the grounding point of the object, that is, it is necessary to pay attention to the depth and space of the channel features where the foot is located. Thus, by setting weights for depth and space, attention to the features of the grounding location can be achieved. As an implementation method, an attention mechanism is used to perform feature fusion on feature maps of multiple scales to obtain a fused feature map, that is, the attention mechanism is used to make the fused fused feature map carry depth information and plane space information, so as to more accurately determine the features of the grounding point based on the depth information and plane space information.
[0061] Finally, the fused feature map is input into the task module of the object recognition model for detection, and the image coordinates of the grounding point of the first object in the image to be detected are obtained.
[0062] The object recognition model is trained using a supervised training method. The object recognition model of this application is based on a detection network such as a machine learning model YOLOX or YOLOv5, and adds a ground point detection head to output the three-dimensional coordinates of the ground point. The training method is as follows:
[0063] Obtain a sample image. The sample image can be obtained from driving images of multiple scenes, parking fisheye lens data, or some images in the training set coco2017. Input the sample image into the object recognition model to obtain predicted detection frame information of the second object and the predicted image coordinates of the second object's ground point. Determine a target loss function based on the predicted detection frame information of the second object, the predicted image coordinates of the second object's ground point, the true detection frame information of the second object in the annotation information of the sample image, and the true image coordinates of the second object's ground point. Then, adjust the parameters of the image recognition model based on the target loss function to obtain a trained object recognition model. The trained object recognition model can determine the mapping relationship between the input image to be detected and the image coordinates of the ground point of the first object included in the image to be detected. The second object is an object in the sample image that is used to distinguish it from the first object in the image to be detected, so that "second" is also used to identify different objects. The second object refers to a living object around the vehicle, such as a person or animal.
[0064] Among them, for the method of determining the target loss, the first loss function can be determined according to the difference between the predicted detection box information and the real detection box information in the annotation information, and the second loss function can be determined according to the difference between the predicted image coordinates and the real image coordinates of the grounding point of the second object. The first loss function and the second loss function are weightedly added according to the set weights to determine the target loss function, wherein the weights corresponding to the first loss function and the second loss function are different and can be set based on needs.
[0065] During the training of the object recognition model, the present application will output the predicted detection frame information of the second object. The predicted detection frame information includes the position of the predicted detection frame, a first confidence level of whether the predicted detection frame includes an object, and a second confidence level of whether the included object is a second object, where the second object refers to a person or an animal. The first confidence level and the second confidence level can exclude detection frames that are not the second object and only output the detection frame of the second object, that is, the person or animal, that the present application is concerned about, thereby improving the accuracy of detection. In addition, since the detection accuracy of the detection frame is high, the position of the second object in the detection frame is determined based on the position of the detection frame to identify the position of the grounding point of the second object, which can also improve the accuracy of the position determination of the grounding point.
[0066] Among them, the predicted detection box information includes the position of the predicted detection box, the first confidence of whether the predicted detection box includes the object, and the second confidence of whether the included object is the second object. That is to say, the task module includes 4 detection heads, each detection head performs a recognition task, and the outputs of the 4 detection heads are the position of the predicted detection box, the first confidence of whether the predicted detection box includes the object, the second confidence of whether the object in the predicted detection box is the second object, and the image coordinates of the ground point of the second object. Therefore, the first loss function includes a first sub-loss function, a second sub-loss function, and a third sub-loss function, so that the target loss function is determined by the following method:
[0067] Determine the first sub-loss function based on the difference between the predicted detection box and the true detection box;
[0068] Determining a second sub-loss function according to a difference between the first predicted confidence and the first true confidence;
[0069] Determining a third sub-loss function according to the difference between the second predicted confidence and the second true confidence;
[0070] determining a second loss function based on a difference between the predicted image coordinates and the true image coordinates of a ground point of the first object;
[0071] The target loss function is determined by performing weighted summation on the first sub-loss function, the second sub-loss function, the third sub-loss function and the second loss function.
[0072] Step 303 : Map the image coordinates of the grounding point of the first object to a three-dimensional space according to the internal and external parameters of the target surround view camera to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space.
[0073] Among them, step 303 can refer to the explanation in the above embodiment, the principle is the same and will not be repeated here.
[0074] Step 304 : Determine a set area around the vehicle where the grounding point of the first object is located based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle.
[0075] As an example, Figure 5 As shown, the set area around the vehicle is Figure 5 The two horizontal dotted lines and the two vertical dotted lines intersect to divide the set area around the vehicle into Figure 5 There are eight areas marked as 1-8, wherein the size of the area is not limited in this embodiment, wherein the eight areas are mainly divided into a first set area and a second set area according to the length direction or the width direction of the vehicle, wherein the first set area overlaps with the area in the length direction or the width direction of the vehicle, wherein, Figure 5 Any one of the regions 2, 4, 5 and 7 in the first setting region is the first setting region; the second setting region does not overlap with the region in the length direction and the region in the width direction of the vehicle, that is, Figure 5 Any one of area 1, area 3, area 6 and area 8 is the second set area.
[0076] In an embodiment of the present application, the three-dimensional coordinates of the vehicle may be the coordinates of the center point of the vehicle body. Based on the length, width and three-dimensional coordinates of the vehicle, the area around the vehicle body may be divided into first set areas and second set areas, and the coordinate range corresponding to each first set area under the vehicle body coordinates and the coordinate range corresponding to each second set area may be determined. Thus, based on the three-dimensional coordinates of the grounding point of the first object, it may be determined in which set area around the vehicle the grounding point of the first object is located.
[0077] Step 305 : In response to the grounding point of the first object being located in a first set area around the vehicle, determine the three-dimensional coordinates of a first reference point closest to the grounding point based on the three-dimensional coordinates of the grounding point of the first object.
[0078] The first reference point is a point in the projection of the vehicle body on the ground.
[0079] In one implementation of an embodiment of the present application, if the grounding point of the first object is located in a first set area around the vehicle, the three-dimensional coordinates of the first reference point closest to the grounding point are determined based on the three-dimensional coordinates of the grounding point of the first object. That is to say, a perpendicular line can be drawn from the grounding point of the first object to the projection of the vehicle body on the ground, and the intersection of the perpendicular line and the vehicle body is used as the first reference point.
[0080] Step 306 : Determine the distance between the first object and the vehicle based on the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the first reference point.
[0081] In an embodiment of the present application, the distance between the grounding point and the first reference point is determined based on the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the first reference point, and the distance between the grounding point and the first reference point is used as the distance between the first object and the vehicle.
[0082] As an example, Figure 5 As shown, if the first set area is the area indicated by area 4, the grounding point of the first object is Figure 5 If it is point A in , then the first reference point is Figure 5 Point B in the figure is the point closest to point A in the projection of the vehicle body on the ground, so the distance between A and B is the distance between the first object and the vehicle.
[0083] It should be noted that the first setting area where the grounding point is located is only an example. Figure 5 When the first object is in other first set areas, the method for determining the distance between the first object and the vehicle is the same and will not be repeated here.
[0084] Step 307 : In response to the grounding point of the first object being located in a second set area around the vehicle, acquiring the three-dimensional coordinates of a second reference point corresponding to the second set area.
[0085] In an embodiment of the present application, if the grounding point of the first object is located in a second set area around the vehicle, the three-dimensional coordinates of a second reference point corresponding to the second set area are obtained, where the second reference point is the point in the second set area with the shortest distance from the vehicle.
[0086] As an example, Figure 5 As shown, the grounding point C of the first object is located at Figure 5 In area 1, that is, the first object is in the left front area of the vehicle, then point D closest to the vehicle is determined from the second area, and point D is used as the second reference point. The correspondence between the coordinates of point D and the coordinates of the vehicle body reference point, such as the center point of the vehicle body, can be predetermined based on the shape and size of the vehicle. Therefore, in actual applications, after the coordinates of the reference point of the vehicle are determined based on positioning technology, the three-dimensional coordinates of the second reference point D can be determined based on the correspondence.
[0087] Step 308 : Determine the distance between the first object and the vehicle based on the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the second reference point.
[0088] In an embodiment of the present application, the distance between the grounding point and the second reference point is determined based on the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the second reference point, and the distance between the grounding point and the second reference point is used as the distance between the first object and the vehicle.
[0089] It should be noted that the second setting area where the grounding point is located is only an example. Figure 5 When the distance between the first object and the vehicle is within other second set areas, the method for determining the distance between the first object and the vehicle is the same and will not be repeated here.
[0090] It should be noted that the first object may have one or two grounding points. The above example is illustrated using one grounding point. When there are two grounding points, for example, the first object is a person, the two grounding points are usually the contact points of the left and right feet with the ground. Therefore, the first distance between the left foot and the vehicle and the second distance between the right foot and the vehicle can be calculated respectively according to the above method. As an implementation method, the smaller distance of the first distance and the second distance can be used as the distance between the person and the vehicle; as another implementation method, the first distance and the second distance can be averaged, and the average value can be used as the distance between the person and the vehicle.
[0091] Optionally, when determining the distance between the first object and the vehicle, it can be determined based on the distance whether the distance between the first object and the vehicle is greater than a safe distance. If it is not greater than the safe distance, an early warning prompt is issued to improve safety.
[0092] In the distance determination method of the embodiment of the present application, the image to be detected captured by the target surround-view camera of the vehicle is obtained, the object recognition is performed on the image to be detected, the image coordinates of the grounding point of the first object in the image to be detected are determined, and the image coordinates of the grounding point of the first object in the image to be detected are mapped to three-dimensional space according to the internal and external parameters of the target surround-view camera to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space. The distance between the first object and the vehicle is determined according to the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle. The present application increases the detection range around the vehicle by capturing the image to be detected by the surround-view camera. For the captured image to be detected, the image coordinates of the grounding point of the first object in the image to be detected in the corresponding image coordinate system are determined. The image coordinates are mapped to three-dimensional space according to the internal and external parameters of the camera to obtain the three-dimensional coordinates. The three-dimensional coordinates are not affected by the posture of the first object, thereby improving the accuracy of the distance determination between the first object and the vehicle. At the same time, when determining the distance between the vehicle and the first object in the surrounding area, the distance between the grounding point of the first object and the vehicle is calculated, which is simple and efficient.
[0093] In order to implement the above embodiment, the embodiment of the present application further proposes a distance determination device.
[0094] Figure 6 A schematic diagram of the structure of a distance determination device provided in an embodiment of the present application.
[0095] like Figure 6 As shown, the device may include:
[0096] The acquisition module 61 is used to acquire the image to be detected captured by the target surround view camera of the vehicle.
[0097] The recognition module 62 is configured to perform object recognition on the image to be detected, and determine the image coordinates of the grounding point of the first object in the image to be detected.
[0098] The mapping module 63 is configured to map the image coordinates of the grounding point of the first object to a three-dimensional space according to the internal and external parameters of the target surround view camera, so as to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space.
[0099] The determination module 64 is configured to determine the distance between the first object and the vehicle based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle.
[0100] Furthermore, in an implementation of the embodiment of the present application, the recognition module 62 is configured to execute: inputting the image to be detected into a trained object recognition model to obtain image coordinates of a grounding point of the first object.
[0101] In one implementation of the embodiment of the present application, the identification module 62 is configured to execute:
[0102] Inputting the image to be detected into the processing module of the object recognition model for downsampling to obtain a downsampled image;
[0103] Inputting the downsampled image into the feature extraction module of the object recognition model to perform multi-scale feature extraction to obtain feature maps of multiple scales;
[0104] Inputting the feature maps of the multiple scales into the feature fusion module of the object recognition model to perform feature fusion to obtain a fused feature map;
[0105] The fused feature map is input into the task module of the object recognition model for detection to obtain the image coordinates of the grounding point of the first object in the image to be detected.
[0106] In one implementation of the embodiment of the present application, the mapping module 63 is configured to execute:
[0107] Mapping the image coordinates of the ground point of the first object to a camera coordinate system using the intrinsic parameters of the target surround-view camera to obtain direction information between at least one candidate camera coordinate of the ground point and the optical center of the target surround-view camera in the camera coordinate system;
[0108] Obtaining the vertical distance between the optical center of the target surround view camera and the ground;
[0109] determining, according to the direction information and the vertical distance, target camera coordinates of the grounding point of the first object in the camera coordinate system;
[0110] The external parameters of the target surround view camera are used to map the target camera coordinates of the touchdown point in the camera coordinate system to the vehicle body coordinate system, and the three-dimensional coordinates of the touchdown point in the vehicle body coordinate system are determined.
[0111] In one implementation of the embodiment of the present application, the determination module 64 is configured to execute:
[0112] determining a set area around the vehicle where the grounding point of the first object is located based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle;
[0113] In response to the grounding point of the first object being located in a first set area around the vehicle, determining the three-dimensional coordinates of a first reference point that is closest to the grounding point based on the three-dimensional coordinates of the grounding point of the first object; wherein the first reference point is a point on the ground projected by the vehicle body;
[0114] The distance between the first object and the vehicle is determined based on the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the first reference point; wherein the first set area overlaps with an area in the length direction or an area in the width direction of the vehicle.
[0115] In one implementation of the embodiment of the present application, the determination module 64 is configured to execute:
[0116] In response to the grounding point of the first object being located in a second set area around the vehicle, obtaining three-dimensional coordinates of a second reference point corresponding to the second set area; wherein the second set area does not overlap with an area in the longitudinal direction or an area in the width direction of the vehicle; and the second reference point is a point in the second set area that is shortest from the vehicle;
[0117] The distance between the first object and the vehicle is determined based on the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the second reference point.
[0118] In one implementation of the embodiment of the present application, the apparatus further includes a training module configured to execute:
[0119] Get a sample image;
[0120] Inputting the sample image into the object recognition model to obtain predicted detection box information of the second object and predicted image coordinates of the ground point of the second object;
[0121] determining a target loss function based on the predicted detection box information of the second object, the predicted image coordinates of the ground point of the second object, the true detection box information of the second object in the annotation information of the sample image, and the true image coordinates of the ground point of the second object;
[0122] According to the target loss function, the parameters of the object recognition model are adjusted to obtain a trained object recognition model.
[0123] In one implementation of the embodiment of the present application, the apparatus further includes a training module configured to execute:
[0124] Determining a first loss function according to a difference between the predicted detection box information and the true detection box information;
[0125] determining a second loss function based on a difference between the predicted image coordinates and the true image coordinates of the ground point of the second object;
[0126] The first loss function and the second loss function are weightedly added according to the set weights to determine the target loss function.
[0127] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment and will not be repeated here.
[0128] In the distance determination device of the embodiment of the present application, the image to be detected captured by the target surround-view camera of the vehicle is obtained, object recognition is performed on the image to be detected, the image coordinates of the grounding point of the first object in the image to be detected are determined, and the image coordinates of the grounding point of the first object are mapped to the three-dimensional space according to the internal and external parameters of the target surround-view camera to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space. The distance between the first object and the vehicle is determined according to the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle. The present application captures the image to be detected by the surround-view camera, thereby increasing the detection range around the vehicle. For the captured image to be detected, the image coordinates of the grounding point of the first object in the image to be detected in the corresponding image coordinate system are determined, and the image coordinates are mapped to the three-dimensional space according to the internal and external parameters of the camera to obtain the three-dimensional coordinates, which are not affected by the posture of the first object, thereby improving the accuracy of the distance determination between the first object and the vehicle.
[0129] In order to implement the above embodiments, the present application also proposes a vehicle, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method described in the above method embodiments is implemented.
[0130] In order to implement the above embodiments, the present application also proposes a non-transitory computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the method described in the above method embodiments is implemented.
[0131] In order to implement the above embodiments, the present application further proposes a computer program product on which a computer program is stored. When the computer program is executed by a processor, the method described in the above method embodiments is implemented.
[0132] Figure 7 FIG6 is a block diagram illustrating a vehicle 600 according to an exemplary embodiment. For example, vehicle 600 may be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or another type of vehicle. Vehicle 600 may be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0133] Reference Figure 7 Vehicle 600 may include various subsystems, such as an infotainment system 610, a perception system 620, a decision control system 630, a drive system 640, and a computing platform 650. Vehicle 600 may also include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and each component of vehicle 600 may be interconnected via wired or wireless means.
[0134] In some embodiments, the infotainment system 610 may include a communication system, an entertainment system, a navigation system, and the like.
[0135] The perception system 620 may include several sensors for sensing information about the environment surrounding the vehicle 600. For example, the perception system 620 may include a global positioning system (which may be a GPS system, a BeiDou system, or other positioning systems), an inertial measurement unit (IMU), a laser radar, a millimeter-wave radar, an ultrasonic radar, and a camera.
[0136] The decision control system 630 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.
[0137] The drive system 640 may include components that provide power to the vehicle 600. In one embodiment, the drive system 640 may include an engine, an energy source, a transmission system, and wheels. The engine may be an internal combustion engine, an electric motor, an air compression engine, or a combination thereof. The engine is capable of converting energy provided by the energy source into mechanical energy.
[0138] Some or all functions of the vehicle 600 are controlled by a computing platform 650. The computing platform 650 may include at least one processor 651 and a memory 652. The processor 651 may execute instructions 653 stored in the memory 652.
[0139] The processor 651 can be any conventional processor, such as a commercially available CPU. The processor can also include a graphics processor (GPU), a field programmable gate array (FPGA), a system on chip (SOC), an application specific integrated circuit (ASIC), or a combination thereof.
[0140] The memory 652 may be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0141] In addition to instructions 653 , memory 652 may also store data, such as road maps, route information, and vehicle location, direction, speed, etc. The data stored in memory 652 may be used by computing platform 650 .
[0142] In the embodiment of the present disclosure, the processor 651 may execute the instruction 653 to complete all or part of the steps of the above method embodiment.
[0143] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0144] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0145] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0146] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0147] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0148] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0149] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0150] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A distance determination method, characterized in that: include: Obtain the image to be detected captured by the vehicle's target surround view camera; performing object recognition on the image to be detected to determine the image coordinates of a grounding point of a first object in the image to be detected; Mapping the image coordinates of the grounding point of the first object to a three-dimensional space according to the internal and external parameters of the target surround view camera to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space; The distance between the first object and the vehicle is determined based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle.
2. The method according to claim 1, wherein The performing object recognition on the image to be detected and determining the image coordinates of the grounding point of the first object in the image to be detected includes: The image to be detected is input into the trained object recognition model to obtain the image coordinates of the grounding point of the first object.
3. The method according to claim 2, wherein Inputting the image to be detected into a trained object recognition model to obtain the image coordinates of the grounding point of the first object includes: Inputting the image to be detected into the processing module of the object recognition model for downsampling to obtain a downsampled image; Inputting the downsampled image into the feature extraction module of the object recognition model to perform multi-scale feature extraction to obtain feature maps of multiple scales; Inputting the feature maps of the multiple scales into the feature fusion module of the object recognition model to perform feature fusion to obtain a fused feature map; The fused feature map is input into the task module of the object recognition model for detection to obtain the image coordinates of the grounding point of the first object in the image to be detected.
4. The method according to claim 1, wherein Mapping the image coordinates of the ground point of the first object to a three-dimensional space according to the internal and external parameters of the target surround view camera to obtain the three-dimensional coordinates of the ground point of the first object in the three-dimensional space includes: Mapping the image coordinates of the ground point of the first object to a camera coordinate system using the intrinsic parameters of the target surround-view camera to obtain direction information between at least one candidate camera coordinate of the ground point and the optical center of the target surround-view camera in the camera coordinate system; Obtaining the vertical distance between the optical center of the target surround view camera and the ground; determining, according to the direction information and the vertical distance, target camera coordinates of the grounding point of the first object in the camera coordinate system; The external parameters of the target surround view camera are used to map the target camera coordinates of the touchdown point in the camera coordinate system to the vehicle body coordinate system, and the three-dimensional coordinates of the touchdown point in the vehicle body coordinate system are determined.
5. The method according to claim 1, wherein The determining the distance between the first object and the vehicle according to the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle includes: determining a set area around the vehicle where the grounding point of the first object is located based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle; In response to the grounding point of the first object being located in a first set area around the vehicle, determining the three-dimensional coordinates of a first reference point that is closest to the grounding point based on the three-dimensional coordinates of the grounding point of the first object; wherein the first reference point is a point on the ground projected by the vehicle body; The distance between the first object and the vehicle is determined based on the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the first reference point; wherein the first set area overlaps with an area in the length direction or an area in the width direction of the vehicle.
6. The method according to claim 5, wherein The method further comprises: In response to the grounding point of the first object being located in a second set area around the vehicle, obtaining three-dimensional coordinates of a second reference point corresponding to the second set area; wherein the second set area does not overlap with an area in the longitudinal direction or an area in the width direction of the vehicle; and the second reference point is a point in the second set area that is shortest from the vehicle; The distance between the first object and the vehicle is determined based on the three-dimensional coordinates of the grounding point and the three-dimensional coordinates of the second reference point.
7. The method according to any one of claims 1 to 6, wherein: The object recognition model training method includes: Get a sample image; Inputting the sample image into the object recognition model to obtain predicted detection box information of the second object and predicted image coordinates of the ground point of the second object; determining a target loss function based on the predicted detection box information of the second object, the predicted image coordinates of the ground point of the second object, the true detection box information of the second object in the annotation information of the sample image, and the true image coordinates of the ground point of the second object; According to the target loss function, the parameters of the object recognition model are adjusted to obtain a trained object recognition model.
8. The method according to claim 7, wherein The determining of the target loss function according to the predicted detection box information of the second object, the predicted image coordinates of the ground point of the second object, the real detection box information of the second object in the annotation information of the sample image, and the real image coordinates of the ground point of the second object includes: Determining a first loss function according to a difference between the predicted detection box information and the true detection box information; determining a second loss function based on a difference between the predicted image coordinates and the true image coordinates of the ground point of the second object; The first loss function and the second loss function are weightedly added according to the set weights to determine the target loss function.
9. A distance determination device, characterized in that: include: An acquisition module is used to acquire the image to be detected captured by the vehicle's target surround view camera; an identification module, configured to perform object identification on the image to be detected and determine the image coordinates of a grounding point of a first object in the image to be detected; a mapping module, configured to map the image coordinates of the grounding point of the first object to a three-dimensional space according to the internal and external parameters of the target surround view camera, so as to obtain the three-dimensional coordinates of the grounding point of the first object in the three-dimensional space; The determining module is configured to determine the distance between the first object and the vehicle based on the three-dimensional coordinates of the grounding point of the first object and the three-dimensional coordinates of the vehicle.
10. A vehicle, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to: Implement the steps of the method according to any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform the steps of the method according to any one of claims 1 to 8.
12. A computer program product, characterized in that The invention comprises a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 8.