Distance measurement method and device and computer readable storage medium
By optimizing the target indicators of the target detection model, setting the preset height weight is greater than the preset width weight, and calculating the horizontal distance with the camera parameters, the problem of low distance detection accuracy in the existing ranging method is solved, and higher ranging accuracy is achieved.
Patent Information
- Application Number
- CN202510526865.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-26
AI Technical Summary
Among the existing ranging methods, the accuracy of distance detection is low, mainly because the target detection accuracy of the target detection model is low, and the accuracy of the existing evaluation methods is insufficient.
The target image is processed by the trained object detection model. By optimizing the target index, setting the preset height weight is greater than the preset width weight, improving the prediction accuracy of the detection frame height, and calculating the horizontal distance with camera parameters.
The prediction accuracy of the target detection box height of the target object is improved, thereby improving the accuracy of distance measurement.
Smart Images

Figure CN120544157A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and in particular relates to a distance measurement method, device, and computer-readable storage medium. Background Art
[0002] With the development of intelligent vehicles, advanced driver assistance technologies (ADAS) and autonomous driving technologies have gradually become important development directions in the automotive field. The core of these technologies lies in the vehicle's comprehensive perception of its surrounding environment. Accurately detecting the distance between the vehicle and other traffic participants (such as pedestrians) is crucial for achieving safe driving. In current vehicle control systems, distance detection typically relies on images captured by the vehicle's forward-facing camera and a built-in object detection model within the vehicle control system. In specific implementations, the object detection model is typically used to process the image captured by the vehicle's forward-facing camera and determine the distance between the vehicle and other traffic participants based on the object detection box information output by the object detection model. Therefore, the distance detection accuracy of these distance measurement methods is primarily affected by the object detection accuracy of the object detection model. However, existing technologies typically evaluate the object detection accuracy of the object detection model based on mean average precision (mAP). This evaluation method has low accuracy, resulting in low object detection accuracy of the object detection model, thereby reducing the accuracy of distance detection. Summary of the Invention
[0003] In view of this, embodiments of the present application provide a distance measurement method, device, and computer-readable storage medium to solve the technical problem of low distance detection accuracy in existing distance measurement methods.
[0004] In a first aspect, an embodiment of the present application provides a ranging method, including:
[0005] Processing a target image using a trained target detection model to obtain detection frame information of a target object in the target image; the detection frame information includes a height of the detection frame;
[0006] determining a horizontal distance between the camera and the target object in the actual scene according to the height of the detection frame and the camera parameters corresponding to the target image;
[0007] In which, the target detection model is optimized according to the target indicator; the target indicator is used to represent the target detection accuracy of the target detection model, and the target indicator is determined by a preset height weight, a preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image, and the predicted target detection frame information is obtained by inputting the sample image into the target detection model for processing; the preset height weight is greater than the preset width weight.
[0008] In an optional implementation of the first aspect, the process of determining the target indicator includes:
[0009] Determining a first indicator corresponding to each sample image according to a preset height weight, a preset width weight, actual target detection frame information corresponding to each sample image, and predicted target detection frame information corresponding to each sample image;
[0010] The target indicator is determined according to the first indicators corresponding to all the sample images.
[0011] In an optional implementation of the first aspect, the actual target detection frame information includes the height and width of the actual target detection frame, and the predicted target detection frame information includes the height and width of the predicted target detection frame. Correspondingly, determining the first indicator corresponding to each sample image based on a preset height weight, a preset width weight, the actual target detection frame information corresponding to each sample image, and the predicted target detection frame information corresponding to each sample image includes:
[0012] For each sample image, the first index corresponding to the sample image is calculated using the following formula based on the preset height weight, the preset width weight, the height and width of the actual target detection frame corresponding to the sample image, and the height and width of the predicted target detection frame corresponding to the sample image:
[0013]
[0014] Among them, Index is the first index, W gt is the width of the actual target detection frame, H gt is the height of the actual target detection frame, W pred is the width of the predicted target detection box, H pred is the height of the predicted target detection frame, Δw is the width difference between the actual target detection frame and the predicted target detection frame, Δh is the height difference between the actual target detection frame and the predicted target detection frame, α is the preset width weight, β is the preset height weight, and α is less than 1.
[0015] In an optional implementation of the first aspect, determining the target indicator according to the first indicators corresponding to all the sample images includes:
[0016] An average value of the first indicators corresponding to all the sample images is determined as the target indicator.
[0017] In an optional implementation of the first aspect, the target detection model includes a feature extraction network and a target detection network; correspondingly, processing the target image using the trained target detection model includes:
[0018] Extracting low-dimensional features of the target image using the feature extraction network;
[0019] The target detection network is used to determine detection frame information of the target object in the target image based on the low-dimensional features.
[0020] In an optional implementation of the first aspect, the feature extraction network includes a convolution layer and a pooling layer; correspondingly, extracting low-dimensional features of the target image using the feature extraction network includes:
[0021] Performing pixel scanning on the target image based on preset convolution kernel parameters through the convolution layer to extract local features of the target image;
[0022] The local features are subjected to dimensionality reduction processing by the pooling layer to obtain low-dimensional features of the target image.
[0023] In an optional implementation of the first aspect, the camera parameters include a longitudinal focal length, a pitch angle, and an installation height of the camera; correspondingly, determining the horizontal distance between the camera and the target object in the actual scene based on the height of the detection frame and the camera parameters corresponding to the target image includes:
[0024] The horizontal distance between the camera and the target object in the actual scene is calculated using the following formula based on the height of the detection frame, the longitudinal focal length of the camera, the pitch angle of the camera, and the installation height of the camera:
[0025]
[0026] Where D is the horizontal distance between the camera and the target object in the actual scene, H is the installation height of the camera, and f y is the longitudinal focal length of the camera, H real is the preset height of the target object, θ is the pitch angle of the camera, h pixel is the height of the detection frame.
[0027] In a second aspect, an embodiment of the present application provides a ranging device, comprising:
[0028] A target detection unit is configured to process a target image using a trained target detection model to obtain detection frame information of a target object in the target image; the detection frame information includes a height of the detection frame;
[0029] a first determining unit, configured to determine a horizontal distance between the camera and the target object in the actual scene according to a height of the detection frame and a camera parameter corresponding to the target image;
[0030] In which, the target detection model is optimized according to the target indicator; the target indicator is used to represent the target detection accuracy of the target detection model, and the target indicator is determined by a preset height weight, a preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image, and the predicted target detection frame information is obtained by inputting the sample image into the target detection model for processing; the preset height weight is greater than the preset width weight.
[0031] In a third aspect, an embodiment of the present application provides another ranging device, comprising a memory and a computer program stored in the memory and executable on a processor, wherein when the processor executes the computer program, the method described in any optional implementation of the first aspect is implemented.
[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the ranging method as described in any optional implementation of the first aspect above.
[0033] In a fifth aspect, an embodiment of the present application provides a computer program product. When the computer program product is run on a ranging device, the ranging device implements the ranging method described in any optional implementation manner of the first aspect.
[0034] The ranging method, device, computer-readable storage medium, and computer program product provided by the embodiments of the present application have the following beneficial effects:
[0035] The distance measurement method provided in the embodiment of the present application processes the target image using a trained target detection model to obtain the detection frame information of the target object in the target image, the detection frame information including the height of the detection frame; and the horizontal distance between the camera and the target object in the actual scene can be determined based on the height of the detection frame and the camera parameters corresponding to the target image. Since the target detection model is optimized based on the target index used to represent the target detection accuracy, and the target index is determined by the preset height weight, the preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image, by setting the preset height weight to be greater than the preset width weight, the target detection model's prediction accuracy for the target object's detection frame height can be improved, thereby improving the accuracy of distance measurement. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0037] Figure 1 A schematic flow chart of a distance measurement method provided in an embodiment of the present application;
[0038] Figure 2 A schematic diagram of a detection frame information of a target object provided in an embodiment of the present application;
[0039] Figure 3 A schematic diagram of the structure of a target detection model provided in an embodiment of the present application;
[0040] Figure 4 This is a schematic diagram of a specific implementation flow of S101 in a ranging method provided in an embodiment of the present application;
[0041] Figure 5 A schematic structural diagram of a distance measuring device provided in an embodiment of the present application;
[0042] Figure 6 A schematic structural diagram of a distance measuring device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0043] The following embodiments are only used to more clearly illustrate the technical solutions of the present application and are therefore only used as examples and are not intended to limit the scope of protection of the present application.
[0044] In the description of the embodiments of the present application, the technical terms "include", "comprise", "have" and any variations thereof mean "including but not limited to", unless otherwise specifically emphasized. In the description of the embodiments of the present application, unless otherwise specified, the technical term "multiple" refers to two or more than two, and the technical terms "at least one" and "one or more" refer to one, two or more. The technical terms "first" and "second" are only used to distinguish different objects, and cannot be understood as indicating or implying relative importance or implicitly indicating the number, specific order or primary and secondary relationship of the indicated technical features. The technical term "and / or" is merely a description of the association relationship of associated objects, indicating that there can be three relationships, such as A and / or B, which can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0045] The embodiments of the present application first provide a distance measurement method, the execution subject of which can be a distance measurement device. For example, the distance measurement method can be applied in application scenarios such as vehicle assisted driving, vehicle automatic driving, obstacle avoidance of a self-moving device (such as a robot), or navigation of a self-moving device. Based on this, the distance measurement device can be a built-in module of a vehicle control system, or a built-in module of a self-moving device, etc. The embodiments of the present application do not limit the application scenarios of the above-mentioned distance measurement method and the above-mentioned distance measurement device.
[0046] For example, see Figure 1 , is a schematic flow chart of a distance measurement method provided in an embodiment of the present application. Figure 1 As shown, the ranging method may include S101 to S102, which are described in detail as follows:
[0047] S101, using the trained target detection model to process the target image to obtain detection frame information of the target object in the target image, where the detection frame information includes the height of the detection frame.
[0048] The target image may be an image that does not contain depth information, for example, an image captured by a monocular camera.
[0049] In one specific implementation, the target image can be an image captured by a vehicle's forward-facing camera. Based on this, the target object can include pedestrians, cyclists, other vehicles, roadblocks, and other objects that may collide with the current vehicle. It is understood that by performing target detection on the image captured by the vehicle's forward-facing camera, it is convenient to subsequently calculate the horizontal distance between the current vehicle and objects such as pedestrians, cyclists, other vehicles, or roadblocks in the actual scene, thereby providing a reference for safe driving of the vehicle.
[0050] In another specific implementation, the target image can be an image captured by a forward-facing camera of a self-mobile device (e.g., a robot). Based on this, the target object can include objects such as obstacles or pedestrians that may affect the movement of the self-mobile device. It is understood that by performing target detection on the image captured by the forward-facing camera of the self-mobile device, it is convenient to subsequently calculate the distance between the self-mobile device and objects such as obstacles or pedestrians in the actual scene, thereby providing a reference for the safe movement of the self-mobile device.
[0051] It should be noted that the above is merely an exemplary description of the target image and the target object and does not constitute a limitation on the target image and the target object.
[0052] Optionally, the detection frame information of the target object may include the coordinates of the center point of the detection frame, the height of the detection frame, and the width of the detection frame. The coordinates of the center point of the detection frame may refer to the coordinates of the center point of the detection frame in the image coordinate system. The image coordinate system may be a plane rectangular coordinate system with the upper left corner vertex of the image as the origin, and the horizontal rightward and vertical downward lines intersecting the vertex as the horizontal and vertical axes respectively. For example, see Figure 2 , is a schematic diagram of a detection frame information of a target object provided in an embodiment of the present application. Figure 2 As shown, the image coordinate system may be a plane rectangular coordinate system established with the upper left corner vertex O of image 20 as the origin, the first side OA of image 20 as the x-axis, and the second side OB of image 20 as the y-axis. The detection frame information of the target object may include the coordinates (x1, y1) of the center point C of the detection frame 21, the height h1 of the detection frame 21, and the width w1 of the detection frame.
[0053] In practical applications, the target detection model can adopt architectures such as convolutional neural networks (CNN), feature pyramid networks (FPN), YOLO (you only look once) networks, or retina networks (RetinaNet). The embodiments of this application do not limit the specific architecture of the target detection model.
[0054] It is understandable that before using the trained target detection model to process the target image, the target detection model needs to be trained. Exemplarily, the target detection model can be trained based on a deep learning algorithm using a number of labeled sample images. The labeled sample image refers to an image after the actual target detection frame information of the target object in the image is marked, that is, the labeled sample image carries the actual target detection frame information of the target object in the sample image. The actual target detection frame information may include the coordinates of the center point of the actual target detection frame, the height of the actual target detection frame, and the width of the actual target detection frame. The coordinates of the center point of the actual target detection frame may refer to the coordinates of the center point of the actual target detection frame in the image coordinate system. It should be noted that the meaning of the image coordinate system can be referred to the above description and will not be elaborated here.
[0055] When training an object detection model, you can divide several labeled sample images into a training set and a validation set. The training set is used to train the object detection model, while the validation set is used to evaluate the object detection accuracy of the trained object detection model.
[0056] Based on this, the training process of the target detection model can include the following stages:
[0057] Model training stage: In this stage, each sample image in the training set can be used as the input of the target detection model, and the actual target detection box information carried by each sample image in the training set can be used as the output of the target detection model. The target detection model is trained to enable the target detection model to learn to identify the target object from the image and output the detection box information of the target object.
[0058] Model verification stage: In this stage, each sample image in the verification set can be input into the trained target detection model, and the trained target detection model is used to process each sample image in the verification set respectively to obtain the predicted target detection box information corresponding to each sample image in the verification set; and the target index is calculated based on the actual target detection box information and the predicted target detection box information corresponding to each sample image in the verification set; the target detection accuracy of the target detection model is evaluated based on the target index.
[0059] The predicted target detection frame information may include the coordinates of the center point of the predicted target detection frame, the height of the predicted target detection frame, and the width of the predicted target detection frame. The coordinates of the center point of the predicted target detection frame may refer to the coordinates of the center point of the predicted target detection frame in the image coordinate system. It should be noted that the meaning of the image coordinate system can be referred to the above description and is not further elaborated here.
[0060] The target indicator can be used to represent the target detection accuracy of the target detection model, that is, to evaluate the degree of error between the predicted results and the actual results of the target detection model.
[0061] Currently, IOU (intersection over union) is commonly used to represent the target detection accuracy (also known as positioning accuracy) of target detection models. IOU represents the intersection over union ratio of the predicted target detection box corresponding to the sample image to the actual target detection box, that is, the ratio of the intersection area of the predicted target detection box and the actual target detection box to the union area. In existing technologies, the IOU value range is [0,1]. The larger the IOU, the higher the overlap between the predicted detection box corresponding to the sample image and the actual detection box, that is, the higher the detection accuracy. In this calculation method, the height and width of the detection box contribute equally to the indicator calculation.
[0062] However, the ranging method provided by the embodiment of the present application mainly relies on the height of the detection frame. That is, the embodiment of the present application converts the height of the detection frame of the target object in the target image into the horizontal distance between the target object and the camera in the actual scene. It can be seen that the impact of the width of the detection frame of the target object on the ranging accuracy is far less than the impact of the detection frame height on the ranging accuracy. The current IOU metric ignores the different effects of the height and width of the detection frame on the ranging accuracy. Therefore, when the traditional target detection model using IOU as a metric is applied to the ranging scenario, the ranging accuracy is low.
[0063] In view of this, in order to improve the ranging accuracy, the embodiment of the present application improves the IOU by introducing a preset height weight and a preset width weight, and setting the preset height weight to be greater than the preset width weight, so as to distinguish the different effects of the height and width of the detection frame of the target object on the ranging accuracy. Based on this, the embodiment of the present application improves the IOU to obtain the target index, which can be determined based on the preset height weight, the preset width weight, the actual target detection frame information corresponding to each sample image, and the predicted target detection frame information corresponding to each sample image. Among them, the sample images used to determine the target index can refer to the sample images used in the above-mentioned model verification stage.
[0064] Based on this, the process of determining target indicators can include steps 1.1 to 1.2, as detailed below:
[0065] Step 1.1: For each sample image, determine the first indicator corresponding to the sample image based on the preset height weight, the preset width weight, the actual target detection frame information corresponding to each sample image, and the predicted target detection frame information corresponding to each sample image.
[0066] Step 1.2: Determine the target index based on the first index corresponding to all sample images.
[0067] Optionally, the average value of the first indicators corresponding to all sample images may be determined as the target indicator.
[0068] In an optional implementation, step 1.1 may specifically include:
[0069] For each sample image, the first index corresponding to the sample image is calculated using the following formula based on the preset height weight, the preset width weight, the height and width of the actual target detection frame corresponding to the sample image, and the height and width of the predicted target detection frame corresponding to the sample image:
[0070]
[0071] Among them, Index is the first index, W gt is the width of the actual target detection frame, H gt is the height of the actual target detection frame, W pred is the width of the predicted target detection box, H pred is the height of the predicted target detection frame, Δw is the width difference between the actual target detection frame and the predicted target detection frame, Δh is the height difference between the actual target detection frame and the predicted target detection frame, α is the preset width weight, β is the preset height weight, and α is less than 1.
[0072] According to the above formula, the embodiment of the present application can decouple the width and height of the detection frame by taking the inverse of the IOU; since the bases of α and β are both greater than or equal to 1, after decoupling the width and height of the detection frame, by configuring α to be less than 1 and β>α, the impact of the width of the detection frame on the target indicator can be reduced, and the target indicator's emphasis on the height of the detection frame can be increased.
[0073] In an optional implementation, β may be greater than or equal to 1, thereby further reducing the impact of the width of the detection box on the target index.
[0074] The embodiment of the present application optimizes the target detection model through the improved target index, which can improve the target detection model's prediction accuracy of the detection frame height, thereby improving the accuracy of ranging for traffic participants (such as pedestrians, etc.).
[0075] In some embodiments, taking the target detection model as CNN as an example, Figure 3As shown, the target detection model may include a feature extraction network and a target detection network sequentially connected between its input and output ends. The number of feature extraction networks may be one or more. Exemplarily, the feature extraction network may include a convolution layer and a pooling layer. The target detection network may include a fully connected layer. The convolution layer may be used to perform pixel scanning on the target object using preset convolution kernel parameters to extract local features of the target image. The pooling layer may be used to perform dimensionality reduction processing on the local features of the target image to obtain low-dimensional features of the target image. The fully connected layer may be used to determine the detection frame information of the target object in the target image based on the low-dimensional features of the target image. The detection frame information of the target object may include the height and width of the detection frame.
[0076] Optionally, when there are multiple feature extraction networks, the preset convolution kernel parameters used by the convolution layer in each feature extraction network can be different. In this way, the convolution layers in different feature extraction networks can extract different local features of the target image (such as edges, textures, shapes or objects, etc.), so that the entire feature extraction network can extract the global features of the target image, making it easier for the target detection model to cope with various complex target images. In addition, by further reducing the dimension of the local features extracted by the convolution layer through the pooling layer, the computational complexity of the fully connected layer can be reduced, the complexity of the target detection model can be reduced, thereby reducing the occurrence of model overfitting and improving the detection accuracy and efficiency of the target detection model.
[0077] Based on this, the step of using the trained target detection model to process the target image in S101 may include the following steps: Figure 4 S1011 to S1012 shown are described in detail as follows:
[0078] S1011, uses the feature extraction network in the target detection model to extract low-dimensional features of the target image.
[0079] Specifically, S1011 may include steps 2.1 and 2.2, which are described in detail as follows:
[0080] In step 2.1, the convolution layer in the feature extraction network performs pixel scanning on the target image based on preset convolution kernel parameters to extract local features of the target image.
[0081] In step 2.2, the local features of the target image are reduced in dimension through the pooling layer in the feature extraction network to obtain the low-dimensional features of the target image.
[0082] S1012: Using the target detection network in the target detection model to determine the detection frame information of the target object in the target image according to the low-dimensional features of the target image.
[0083] It is understood that the above is merely an exemplary description of the architecture of the target detection model and does not constitute a limitation on the architecture of the target detection model. In other embodiments, when the target detection model adopts other architectures, the corresponding specific image processing methods may be different.
[0084] S102 : Determine the horizontal distance between the camera and the target object in the actual scene according to the height of the detection frame and the camera parameters corresponding to the target image.
[0085] The camera parameters corresponding to the target image may include internal parameters and external parameters of the camera.
[0086] The internal parameters of a camera may include, for example, the camera's longitudinal focal length and principal point coordinates. The longitudinal focal length of a camera may refer to the camera's focal length in the vertical direction. The principal point coordinates of a camera may refer to the coordinates of the camera's principal point in the image coordinate system. The principal point of a camera may refer to the projection point of the camera's optical axis in the target image, i.e., the center point of the target image. It should be noted that the meaning of the image coordinate system can be found in the above description and will not be elaborated upon here.
[0087] Camera external parameters include the camera's pitch angle and installation height. The camera's pitch angle represents the angle between the camera's optical axis and the horizontal plane. The camera's installation height represents the horizontal distance between the camera's center and the ground.
[0088] Based on this, in a specific implementation, S102 may specifically include:
[0089] Based on the height of the target object's detection frame, the camera's vertical focal length, the camera's pitch angle, and the camera's installation height, the horizontal distance between the camera and the target object in the actual scene is calculated using the following formula:
[0090]
[0091] Where D is the horizontal distance between the camera and the target object in the actual scene, H is the installation height of the camera, and f y is the vertical focal length of the camera, H real is the preset height of the target object, θ is the pitch angle of the camera, h pixel is the height of the detection box.
[0092] In practical applications, the preset heights of different types of target objects may be different. For example, the preset height of a vehicle may be different from the preset height of a pedestrian.
[0093] For example, if the camera is a forward-facing camera of a vehicle, after determining the horizontal distance D between the camera and the target object in the actual scene, the distance between the target object and the vehicle can be obtained based on the vehicle's body parameters and the installation position of the forward-facing camera.
[0094] As can be seen from the above, the distance measurement method provided in the embodiment of the present application can obtain the detection frame information of the target object in the target image by processing the target image using the trained target detection model, and the detection frame information includes the height of the detection frame; by determining the horizontal distance between the camera and the target object in the actual scene based on the height of the detection frame and the camera parameters corresponding to the target image. Since the target detection model is optimized based on the target index used to represent the target detection accuracy, and the target index is determined by the preset height weight, the preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image, by setting the preset height weight to be greater than the preset width weight, the target detection model can improve the prediction accuracy of the detection frame height of the target object, thereby improving the accuracy of the distance measurement.
[0095] It can be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0096] Based on the distance measurement method provided in the above embodiment, the present application further provides an embodiment of a distance measurement device that implements the above method embodiment. Figure 5 , is a structural diagram of a distance measuring device provided in an embodiment of the present application. For ease of explanation, only the parts related to this embodiment are shown. Figure 5 As shown, the distance measuring device 50 may include: a target detection unit 501 and a first determination unit 502.
[0097] The target detection unit 501 is used to process the target image using the trained target detection model to obtain detection frame information of the target object in the target image; the detection frame information includes the height of the detection frame.
[0098] The first determining unit 502 is configured to determine a horizontal distance between the camera and the target object in the actual scene according to the height of the detection frame and the camera parameters corresponding to the target image.
[0099] In which, the target detection model is optimized according to the target indicator; the target indicator is used to represent the target detection accuracy of the target detection model, and the target indicator is determined by a preset height weight, a preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image, and the predicted target detection frame information is obtained by inputting the sample image into the target detection model for processing; the preset height weight is greater than the preset width weight.
[0100] Optionally, the process of determining the target indicator includes:
[0101] For each of the sample images, determining a first indicator corresponding to the sample image according to a preset height weight, a preset width weight, actual target detection frame information corresponding to the sample image, and predicted target detection frame information corresponding to the sample image;
[0102] The target indicator is determined according to the first indicators corresponding to all the sample images.
[0103] Optionally, the actual target detection frame information includes the height and width of the actual target detection frame, and the predicted target detection frame information includes the height and width of the predicted target detection frame; correspondingly, for each sample image, determining the first indicator corresponding to the sample image according to a preset height weight, a preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image includes:
[0104] For each sample image, the first index corresponding to the sample image is calculated using the following formula based on the preset height weight, the preset width weight, the height and width of the actual target detection frame corresponding to the sample image, and the height and width of the predicted target detection frame corresponding to the sample image:
[0105]
[0106] Among them, Index is the first index, W gt is the width of the actual target detection frame, H gt is the height of the actual target detection frame, W pred is the width of the predicted target detection box, H pred is the height of the predicted target detection frame, Δw is the width difference between the actual target detection frame and the predicted target detection frame, Δh is the height difference between the actual target detection frame and the predicted target detection frame, α is the preset width weight, β is the preset height weight, and α is less than 1.
[0107] Optionally, determining the target indicator according to the first indicators corresponding to all the sample images includes:
[0108] An average value of the first indicators corresponding to all the sample images is determined as the target indicator.
[0109] Optionally, the target detection model includes a feature extraction network and a target detection network. Correspondingly, the target detection unit 501 is specifically used to:
[0110] Extracting low-dimensional features of the target image using the feature extraction network;
[0111] The target detection network is used to determine detection frame information of the target object in the target image based on the low-dimensional features.
[0112] Optionally, the feature extraction network includes a convolution layer and a pooling layer; correspondingly, the target detection unit 501 is further specifically configured to:
[0113] Performing pixel scanning on the target image based on preset convolution kernel parameters through the convolution layer to extract local features of the target image;
[0114] The local features are subjected to dimensionality reduction processing by the pooling layer to obtain low-dimensional features of the target image.
[0115] Optionally, the camera parameters include the longitudinal focal length, pitch angle, and installation height of the camera; correspondingly, the first determining unit 502 is specifically configured to:
[0116] The horizontal distance between the camera and the target object in the actual scene is calculated using the following formula based on the height of the detection frame, the longitudinal focal length of the camera, the pitch angle of the camera, and the installation height of the camera:
[0117]
[0118] Where D is the horizontal distance between the camera and the target object in the actual scene, H is the installation height of the camera, and f y is the longitudinal focal length of the camera, H real is the preset height of the target object, θ is the pitch angle of the camera, h pixel is the height of the detection frame.
[0119] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units as needed, that is, the internal structure of the ranging device can be divided into different functional units to complete all or part of the functions described above. The functional units in the embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units are only for the convenience of distinguishing each other and are not used to limit the scope of protection of this application. The specific working process of each unit in the above-mentioned ranging device can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.
[0120] See also Figure 6 , Figure 6 This is a structural diagram of a distance measuring device provided in another embodiment of the present application. Figure 6 As shown, the distance measuring device 6 provided in this embodiment may include: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60, such as a program corresponding to the distance measuring method. When the processor 60 executes the computer program 62, the steps in the above-mentioned distance measuring method embodiment are implemented, such as Figure 1 Alternatively, the processor 60 executes the computer program 62 to implement the functions of each module / unit in the above-mentioned distance measuring device embodiment, for example Figure 5 The functions of the units 501 to 502 are shown.
[0121] For example, the computer program 62 can be divided into one or more modules / units, one or more modules / units are stored in the memory 61 and executed by the processor 60 to complete the present application. One or more modules / units can be a series of computer program instruction segments that can complete specific functions. The instruction segments are used to describe the execution process of the computer program 62 in the distance measuring device 6. For example, the computer program 62 can be divided into a target detection unit and a first determination unit. The specific functions of each unit can be seen in FIG. Figure 5 The relevant descriptions in the corresponding embodiments are not repeated here.
[0122] Those skilled in the art will understand that Figure 6 This is merely an example of the distance measuring device 6 and does not constitute a limitation on the distance measuring device 6 . The distance measuring device 6 may include more or fewer components than shown in the figure, or may combine certain components, or may include different components.
[0123] The processor 60 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0124] The memory 61 can be an internal storage unit of the distance measuring device 6, such as a hard disk or memory of the distance measuring device 6. The memory 61 can also be an external storage device of the distance measuring device 6, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, or a flash card equipped on the distance measuring device 6. Furthermore, the memory 61 can also include both the internal storage unit of the distance measuring device 6 and an external storage device. The memory 61 is used to store computer programs and other programs and data required by the distance measuring device. The memory 61 can also be used to temporarily store data that has been output or is about to be output.
[0125] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, each step of the ranging method in the above method embodiment is implemented.
[0126] An embodiment of the present application provides a computer program product. When the computer program product is run on a distance measuring device, the distance measuring device is enabled to implement the steps in the above-mentioned various method embodiments.
[0127] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0128] It should be noted that, unless otherwise specified, all technical terms used in the embodiments of this application have the same meanings as those commonly understood by those skilled in the art in the technical field of this application. The technical terms used in the embodiments of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.
[0129] The phrase "embodiment" mentioned in the description of the embodiments of the present application means that the specific features, structures, or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it refer to independent or alternative embodiments that are mutually exclusive with other embodiments. It is understood explicitly and implicitly by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0130] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0131] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A distance measurement method, characterized in that: include: Processing the target image using the trained target detection model to obtain detection frame information of the target object in the target image; The detection frame information includes the height of the detection frame; determining a horizontal distance between the camera and the target object in the actual scene according to the height of the detection frame and the camera parameters corresponding to the target image; In which, the target detection model is optimized according to the target indicator; the target indicator is used to represent the target detection accuracy of the target detection model, and the target indicator is determined by a preset height weight, a preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image, and the predicted target detection frame information is obtained by inputting the sample image into the target detection model for processing; the preset height weight is greater than the preset width weight.
2. The method according to claim 1, characterized in that The process of determining the target indicators includes: For each of the sample images, determining a first indicator corresponding to the sample image according to a preset height weight, a preset width weight, actual target detection frame information corresponding to the sample image, and predicted target detection frame information corresponding to the sample image; The target indicator is determined according to the first indicators corresponding to all the sample images.
3. The method according to claim 2, characterized in that The actual target detection frame information includes the height and width of the actual target detection frame, and the predicted target detection frame information includes the height and width of the predicted target detection frame. Correspondingly, for each of the sample images, according to a preset height weight, a preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image, determining a first indicator corresponding to the sample image includes: For each sample image, the first index corresponding to the sample image is calculated using the following formula based on the preset height weight, the preset width weight, the height and width of the actual target detection frame corresponding to the sample image, and the height and width of the predicted target detection frame corresponding to the sample image: Among them, Index is the first index, W gt is the width of the actual target detection frame, H gt is the height of the actual target detection frame, W pred is the width of the predicted target detection box, H pred is the height of the predicted target detection frame, Δw is the width difference between the actual target detection frame and the predicted target detection frame, Δh is the height difference between the actual target detection frame and the predicted target detection frame, α is the preset width weight, β is the preset height weight, and α is less than 1.
4. The method according to claim 2, characterized in that Determining the target indicator according to the first indicators corresponding to all the sample images includes: An average value of the first indicators corresponding to all the sample images is determined as the target indicator.
5. The method according to claim 1, characterized in that The target detection model includes a feature extraction network and a target detection network. Correspondingly, the trained target detection model is used to process the target image, including: Extracting low-dimensional features of the target image using the feature extraction network; The target detection network is used to determine detection frame information of the target object in the target image based on the low-dimensional features.
6. The method according to claim 5, characterized in that The feature extraction network includes a convolution layer and a pooling layer; correspondingly, the feature extraction network is used to extract low-dimensional features of the target image, including: Performing pixel scanning on the target image based on preset convolution kernel parameters through the convolution layer to extract local features of the target image; The local features are subjected to dimensionality reduction processing by the pooling layer to obtain low-dimensional features of the target image.
7. The method according to any one of claims 1 to 6, characterized in that The camera parameters include the longitudinal focal length, pitch angle, and installation height of the camera; correspondingly, determining the horizontal distance between the camera and the target object in the actual scene based on the height of the detection frame and the camera parameters corresponding to the target image includes: The horizontal distance between the camera and the target object in the actual scene is calculated using the following formula based on the height of the detection frame, the longitudinal focal length of the camera, the pitch angle of the camera, and the installation height of the camera: Where D is the horizontal distance between the camera and the target object in the actual scene, H is the installation height of the camera, and f y is the longitudinal focal length of the camera, H real is the preset height of the target object, θ is the pitch angle of the camera, h pixel is the height of the detection frame.
8. A distance measuring device, characterized in that: include: A target detection unit is used to process a target image using a trained target detection model to obtain detection frame information of a target object in the target image; The detection frame information includes the height of the detection frame; a first determining unit, configured to determine a horizontal distance between the camera and the target object in the actual scene according to a height of the detection frame and a camera parameter corresponding to the target image; In which, the target detection model is optimized according to the target indicator; the target indicator is used to represent the target detection accuracy of the target detection model, and the target indicator is determined by a preset height weight, a preset width weight, the actual target detection frame information corresponding to the sample image, and the predicted target detection frame information corresponding to the sample image, and the predicted target detection frame information is obtained by inputting the sample image into the target detection model for processing; the preset height weight is greater than the preset width weight.
9. A distance measuring device, characterized in that: The method comprises a memory and a computer program stored in the memory and executable on a processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the distance measurement method according to any one of claims 1 to 7 is implemented.