Monocular Camera-Based Target Ranging Method and System
By introducing a combination of feature fusion, effective information extraction and ranging modules into the monocular camera target ranging method, the problem of insufficient measurement accuracy and real-time performance in the prior art is solved, and higher ranging accuracy and real-time detection capabilities are achieved.
Patent Information
- Application Number
- CN202111479749.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-12-06
AI Technical Summary
The existing target ranging method based on monocular vision has the problem of low measurement accuracy and real-time performance, especially because the model generalization ability is affected by a variety of factors.
A target distance measurement method based on a monocular camera is adopted to measure the target depth through a combination of feature fusion module, effective information extraction module and distance measurement module. The feature fusion module enhances semantic information and spatial information through multi-dimensional feature fusion. The effective information extraction module filters and fuses local and global information. The ranging module reduces regression error through the depth range and depth residual modules.
The accuracy of target object distance measurement is improved, and real-time detection effect is achieved by accelerating the inference speed.
Smart Images

Figure CN114266919B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and in particular, to a method and system for target ranging based on a monocular camera. Background Art
[0002] With the development of autonomous driving technology, autonomous vehicles have gradually come into the public view. In the driving scenarios of autonomous vehicles, ranging of obstacles or the vehicle ahead is an essential link in the autonomous driving process. With the continuous development of computer vision technology, its role in intelligent vehicle systems has been continuously improved. Applying computer vision technology to vehicle detection, in order to achieve ranging of obstacles or the vehicle ahead, ranging methods based on monocular vision have emerged, which have played a significant role in improving the safety of vehicles. Currently, the ranging methods based on monocular vision are generally divided into two ways: traditional similar triangle ranging and direct regression of target depth by convolutional neural network (CNN). Since these two ways require a large amount of data for learning, and the generalization ability of the obtained model is restricted by many factors, there are problems of low measurement accuracy and real-time performance. Summary of the Invention
[0003] The purpose of the embodiments of the present invention is to provide a method and system for target ranging based on a monocular camera, which can improve the accuracy of target object ranging, and the inference speed of the adopted target ranging model is fast, and the effect of real-time detection can be achieved.
[0004] To achieve the above purpose, the embodiments of the present invention provide a method for target ranging based on a monocular camera, including:
[0005] Obtain a target image captured by a monocular camera;
[0006] Input the target image into a target ranging model, so that the target ranging model measures the target depth of the target image and outputs the distance between the vehicle itself and the target object;
[0007] Wherein, the target ranging model includes a feature fusion module, an effective information extraction module, and a ranging module; the feature fusion module performs several upsamplings on the target image to generate several feature maps of different sizes, performs upsampling on the feature maps to obtain a fused feature map, and fuses the fused feature map with the target position information output by the effective information extraction module to obtain a target feature map; the ranging module measures the target depth of the target feature map and outputs the measured distance between the vehicle itself and the target object.
[0008] As an improvement to the above solution, the feature fusion module includes N downsampling layers connected in series and M upsampling layers connected in series, where M ≤ N, and both N and M are integers greater than 0; each of the downsampling layers outputs a feature map of its corresponding size, and the upsampling layers perform upsampling on the feature maps in sequence according to a preset upsampling strategy to obtain a fused feature map; among them, after each current upsampling layer finishes upsampling, it fuses its upsampling output with the corresponding feature map of the same size and inputs it to the next upsampling layer for upsampling until the current upsampling layer is the last upsampling layer.
[0009] As an improvement to the above solution, the feature fusion module uses the backbone network of DLA34, including 5 downsampling layers connected in series and 3 upsampling layers connected in series.
[0010] As an improvement to the above solution, the input of the effective information extraction module is a standard feature map that meets a preset size among the several feature maps of different sizes.
[0011] As an improvement to the above solution, the effective information extraction module includes a global information extraction module, a local information extraction module, and a fusion module; among them,
[0012] The global information extraction module sequentially performs a convolution operation, a global average pooling operation, and a probability operation on the standard feature map to obtain a feature map of the first size;
[0013] The local information extraction module performs a convolution operation and a probability operation on the standard feature map to obtain a feature map of the second size;
[0014] The fusion module integrates the feature map of the first size and the feature map of the second size and restores the channels to output the target position information.
[0015] As an improvement to the above solution, the ranging module includes a category classification module, a size classification module, a depth range module, and a depth residual module, and the input of each module is the target feature map; among them,
[0016] The category classification module includes a first convolutional layer, a ReLU layer, a second convolutional layer, and a probability layer; the size classification module includes a third convolutional layer, a ReLU layer, and a fourth convolutional layer; the depth range module includes a fifth convolutional layer, a BN layer, a Mish layer, a sixth convolutional layer, a ReLU layer, a seventh convolutional layer, and a probability layer; the depth residual module includes an eighth convolutional layer, a ReLU layer, and a ninth convolutional layer.
[0017] As an improvement of the above solution, the category classification module is used to: obtain each pixel point of the target feature map, calculate the probability value of each pixel point on the channels of the category classification module, and calculate the position coordinates of the target pixel points whose probability values are greater than a preset threshold in the target feature map;
[0018] The size classification module is used to: take the position coordinates as the center point of the minimum circumscribed rectangle, obtain the width and height of the minimum circumscribed rectangle in the target feature map, and display the minimum circumscribed rectangle in the target feature map; wherein, the size classification module has two channels, and the values output by the two channels respectively correspond to the width and the height, and the target object in the target feature map is located in the minimum circumscribed rectangle.
[0019] As an improvement of the above solution, the ranging module is further used to:
[0020] The depth range module is used to: obtain the pixel point corresponding to the position coordinates in the target feature map through the position coordinates, obtain the channel with the largest probability value of the pixel point from several channels as the target channel, and obtain the depth range threshold corresponding to the target channel;
[0021] The depth residual module is used to: obtain the depth residual corresponding to the position coordinates in the target feature map through the position coordinates, calculate the sum of the depth range threshold and the depth residual as the depth of the target object, and use the depth as the measurement distance between the host vehicle and the target object.
[0022] As an improvement of the above solution, the depth range module includes several channels, and each channel is preset with its corresponding depth range threshold, and the depth range threshold is a preset distance value.
[0023] To achieve the above object, an embodiment of the present invention further provides a target ranging system based on a monocular camera, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the target ranging method based on a monocular camera as described in any one of the above embodiments.
[0024] Compared with the prior art, the object ranging method and system based on a monocular camera according to the embodiments of the present invention obtain an object image captured by the monocular camera and input the object image into an object ranging model, so that the object ranging model measures the object depth of the object image and outputs the distance between the vehicle and the object; wherein, the object ranging model includes a feature fusion module, an effective information extraction module, and a ranging module; the feature fusion module performs multiple upsamplings on the object image to generate multiple feature maps of different sizes, performs upsampling on the feature maps to obtain a fused feature map, and fuses the fused feature map with the object position information output by the effective information extraction module to obtain an object feature map, fusing the position information into the network to make the network pay more attention to the position information; the ranging module measures the object depth of the object feature map and outputs the measured distance between the vehicle and the object. By adopting the embodiments of the present invention, the accuracy of object ranging can be improved, and the inference speed of the adopted object ranging model is fast, and the effect of real-time detection can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a flowchart of an object ranging method based on a monocular camera provided by an embodiment of the present invention;
[0026] Figure 2 is a schematic structural diagram of a feature fusion module provided by an embodiment of the present invention;
[0027] Figure 3 is a schematic structural diagram of a downsampling layer in the feature fusion module provided by an embodiment of the present invention;
[0028] Figure 4 is a schematic structural diagram of an effective information extraction module provided by an embodiment of the present invention;
[0029] Figure 5 is a schematic structural diagram of a ranging module provided by an embodiment of the present invention;
[0030] Figure 6 is a structural block diagram of an object ranging system based on a monocular camera provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] See Figure 1 , Figure 1The following is a flowchart of a target ranging method based on a monocular camera provided by an embodiment of the present invention. The target ranging method based on a monocular camera includes:
[0033] S1. Obtain a target image captured by the monocular camera;
[0034] S2. Input the target image into a target ranging model so that the target ranging model measures the target depth of the target image and outputs the distance between the vehicle itself and the target object.
[0035] Specifically, the target ranging method based on a monocular camera described in the embodiment of the present invention can be implemented by a controller in the vehicle. The controller integrates multiple functions such as data processing and data communication, and has a powerful service scheduling function and data processing ability.
[0036] Specifically, in step S1, during the driving process of the vehicle, a target image captured in real time by a monocular camera set on the vehicle can be obtained. When the controller receives the target image, it can pre-process the target image, such as performing binary processing or grayscale processing on the target image, so as to reduce the amount of data to be processed and improve the detection efficiency. In addition, the target image can also be scaled and cropped to obtain an image with a fixed size for input into the target ranging model for processing.
[0037] Specifically, in step S2, the target ranging model includes a feature fusion module, an effective information extraction module, and a ranging module; the feature fusion module performs multiple upsamplings on the target image to generate multiple feature maps of different sizes, performs upsampling on the feature maps to obtain a fused feature map, and fuses the fused feature map with the target position information output by the effective information extraction module to obtain a target feature map; the ranging module measures the target depth of the target feature map and outputs the measured distance between the vehicle itself and the target object.
[0038] In the embodiment of the present invention, by performing feature fusion on the target image through the feature fusion module, semantic information and spatial information can be enhanced. The effective information extraction module can screen out important information of the target image and then perform fusion, considering both local information and global information, to obtain a more refined attention weight while reducing the amount of calculation. The ranging module first determines the depth range where the target is located, and then regresses the residual between the target depth and the range to reduce the regression error and improve the accuracy. By adopting the target ranging method based on a monocular camera described in the embodiment of the present invention, the distance between the vehicle itself and the vehicle in front can be measured in real time, and the driver can be reminded in time when the distance is too close to prevent collisions.
[0039] See Figure 2 ,Figure 2 It is a schematic structural diagram of the feature fusion module provided by an embodiment of the present invention. The feature fusion module includes N serially connected downsampling layers and M serially connected upsampling layers, where M ≤ N, and both N and M are integers greater than 0. Each downsampling layer outputs a feature map of its corresponding size, and the upsampling layer sequentially performs upsampling on the feature map according to a preset upsampling strategy to obtain a fused feature map. Among them, each current upsampling layer fuses its upsampling output with the corresponding feature map of the same size after upsampling is completed, and inputs it to the next upsampling layer for upsampling until the current upsampling layer is the last upsampling layer.
[0040] Specifically, the feature fusion module adopts the backbone network of DLA34, including 5 serially connected downsampling layers and 3 serially connected upsampling layers. Each downsampling layer uses Bottleneck (with higher accuracy than BasicBlock) to extract target features under different receptive fields, and then performs upsampling 3 times. Moreover, each upsampling will perform feature fusion with the feature map of the same size. Multi-Scale Feature Fusion enhances semantic information and spatial information.
[0041] Exemplarily, the structure of the downsampling layer can refer to Figure 3 . The downsampling layer sequentially includes a 1x1 convolutional layer, a ReLU layer, a 3x3 convolutional layer, a ReLU layer, and a 1x1 convolutional layer. The upsampling layer uses bilinear interpolation for upsampling and then is followed by a convolutional layer for output.
[0042] After passing the target image through 5 downsampling layers, 5 feature maps of different scales will be generated respectively, which are the features downsampled by 2 times Figure 1 , 4 times the features Figure 2 , the features 8 times Figure 3 , the features 16 times Figure 4 , and the features 32 times Figure 5 . Each downsampling layer corresponds to one downsampling operation (i.e., the output size is 1 / 2 of the input). The input of downsampling layer 2 is the output of downsampling layer 1, the input of downsampling layer 3 is the output of downsampling layer 2... After 5 times of downsampling, 5 feature maps of different sizes will be obtained. Because the features downsampled by 2 times Figure 1 have too large a size and a relatively large computational amount, the features downsampled by 2 times Figure 1 are not considered during feature fusion in upsampling, and only the feature maps of the remaining 4 sizes are fused.
[0043] The upsampling strategy includes: starting from the last downsampling layer, performing upsampling on the downsampling layers one by one from bottom to top until the target downsampling layer is reached. For example, in the present invention, the target downsampling layer is the second downsampling layer. First, start upsampling from the feature Figure 5 downsampled by 32 times to obtain a feature map downsampled by 16 times, and then fuse it with the feature Figure 4 output by the previous fourth downsampling layer with a scale of 16 times. After fusion, perform upsampling again to obtain a feature map downsampled by 8 times, and then fuse it with the feature Figure 3 output by the third downsampling layer with a scale of 8 times. Then perform upsampling again to obtain a feature map downsampled by 4 times, and fuse it with the feature Figure 2 output by the second downsampling layer with a scale of 4 times to obtain a fused feature map. After fusing the fused feature map and the target position information output by the effective information extraction module, a target feature map can be obtained, and then this target feature map is input into the ranging module.
[0044] In the embodiment of the present invention, setting 5 downsampling layers is to obtain richer feature information. There is relatively rich position information in the shallow layer of the network, and relatively rich semantic information in the deep layer of the network. If the number of downsampling layers is small, the network is relatively shallow and the effective information that can be learned is relatively small. However, if the number of downsampling layers is too large, the network is relatively deep, and at the same time, the computational amount will increase, and redundant layers may appear, making it impossible to achieve the effect of real-time detection.
[0045] See Figure 4 , Figure 4 which is a schematic structural diagram of the effective information extraction module provided by the embodiment of the present invention. The effective information extraction module includes a global information extraction module, a local information extraction module, and a fusion module. The input of the effective information extraction module is one of the several feature maps with different sizes that meets the preset size, such as a feature Figure 2 with a scale of 4 times.
[0046] Specifically, the global information extraction module includes a convolutional layer 1, a global average pooling layer, and a probability layer, and the probability layer is a Sigmoid layer. The global information extraction module sequentially performs a convolution operation, a global average pooling operation, and a probability operation on the standard feature map to obtain a feature map with the first size. The purpose of the convolutional layer 1 is to extract the global information of the standard feature map through the convolution operation, and then through the global average pooling layer, integrate the features of each channel, and use 1 pixel for each channel to represent the global information. The probability layer is used for probability calculation, and the network autonomously learns which channels are important information.
[0047] Exemplarily, the convolution kernel used in the sampling of convolutional layer 1 is 1*1, the stride is 1, and the number of output channels is 1 / 2 of the number of input channels. The purpose is to reduce the computational load. Finally, a feature map with the same size as the input standard feature map but with the number of channels being 1 / 2 of the previous one is obtained (it can be understood that the information of the previous C channels is compressed to C / 2 for representation), that is: assuming the input size is H*W*C (H and W represent the height and width of the feature map respectively, and C represents the number of channels), then the size after passing through the 1x1 convolutional layer is H*W*C / 2. Then, through the global average pooling operation (the pooling kernel size is H*W), the size of the feature map is changed to 1*1*C / 2. Each channel in the original standard feature map is represented by the information of H*W pixels. Through the feature integration of the global average pooling layer, the information represented by H*W pixels can be compressed to 1*1 pixel for representation. The purpose is to represent the importance of this channel through the compressed information and also to facilitate the calculation with the subsequent feature maps. Finally, after the sigmoid operation, the channels (C / 2) are probabilized and transformed to between 0 and 1. The sigmoid operation is a non-linear transformation operation on all elements in the entire feature map, and the value range is (0,1), similar to probability values. That is, assuming the input size of sigmoid is H*W*C, then there are HxWxC elements to be transformed, and the size of the final result is H*W*C. However, sigmoid performs a binary classification operation, and the output is only the probability of the importance.
[0048] Specifically, the local information extraction module includes convolutional layer 2 and a probabilization layer, and the probabilization layer is a Sigmoid layer. The local information extraction module performs a convolution operation and a probabilization operation on the standard feature map to obtain a feature map of the second size. The purpose of convolutional layer 2 is to extract local information through the convolution operation, and directly probabilize each pixel of the full feature map through Sigmoid to autonomously learn by the network which pixel information on the current feature map is important information.
[0049] Exemplarily, the convolution kernel used in the sampling of convolutional layer 2 is 1*1, and the stride is 1. Its working process is the same as that of convolutional layer 1 and will not be elaborated here.
[0050] Specifically, the fusion module includes a classifier and several fusion units X. The fusion module integrates the first-size feature map and the second-size feature map and restores the channels to output the target position information.
[0051] Exemplarily, the classifier is Softmax for multi-classification operation. Its operation result is to output the importance degrees of each class, such as the importance degree of class 1 and the importance degree of class 2, and the sum of the two importance degrees is 1. After the standard feature map passes through the operations of the global information extraction module (convolution + pooling + sigmoid), the size of the finally output first-size feature map is 1*1*C / 2. After the standard feature map passes through the operations of the local information extraction module (convolution + sigmoid), the size of the finally output second-size feature map is H*W*C / 2. Through the broadcast mechanism, the first-size feature map and the second-size feature map can be directly multiplied for integration. 1*1*C / 2 can be understood as a row vector, which multiplies each row and each column in H*W*C / 2. The size of the standard feature map input to the effective information extraction module is H*W*C. When passing through the global information extraction module and the local information extraction module, 1*1 convolution is used to reduce the number of channels. To reduce the calculation amount, the number of channels of the global information extraction module and the local information extraction module is C / 2. Channel restoration means restoring the number of channels from C / 2 to C through 1*1 convolution. Finally, the effective information extraction module outputs the target position information.
[0052] See Figure 5 , Figure 5 is a schematic structural diagram of the ranging module provided by an embodiment of the present invention. The input of the ranging module is the target feature map obtained by fusing the fused feature map and the target position information. The position information is fused into the network to make the network pay more attention to the position information. The ranging module includes 4 branches, namely: a class classification module, a size classification module, a depth range module, and a depth residual module. The input of each module is the target feature map; among them,
[0053] The class classification module includes a first convolutional layer, a ReLU layer, a second convolutional layer, and a probability layer; the size classification module includes a third convolutional layer, a ReLU layer, and a fourth convolutional layer; the depth range module includes a fifth convolutional layer, a BN layer, a Mish layer, a sixth convolutional layer, a ReLU layer, a seventh convolutional layer, and a probability layer; the depth residual module includes an eighth convolutional layer, a ReLU layer, and a ninth convolutional layer. The probability layer is a Sigmoid layer.
[0054] Exemplarily, the first convolutional layer is a 3x3 convolution, the second convolutional layer is a 1x1 convolution, the third convolutional layer is a 3x3 convolution, the fourth convolutional layer is a 1x1 convolution, the fifth convolutional layer is a 3x3 convolution, the sixth convolutional layer is a 3x3 convolution, the seventh convolutional layer is a 1x1 convolution, the eighth convolutional layer is a 3x3 convolution, and the ninth convolutional layer is a 1x1 convolution.
[0055] Specifically, the number of channels of the class classification module corresponds to the classes of the target objects. The class classification module is mainly used to identify the pixel points belonging to the target objects. For example, when the target object is a vehicle, the number of channels is 1; when the target objects are vehicles and people, the number of channels is 2. The number of channels of the size classification module is 2, and the two channels respectively correspond to the width and the height. The depth range module includes a plurality of channels, and each channel is preset with its corresponding depth range threshold, and the depth range threshold is a preset distance value. For example, the number of channels of the depth range module is 5, namely: channel 1, channel 2, channel 3, channel 4, and channel 5, and their corresponding index values are 0, 1, 2, 3, and 4, and the depth range thresholds corresponding to the index values are 10 meters, 20 meters, 30 meters, 40 meters, and 50 meters respectively. The number of channels of the depth residual module is 1, and it is used to output the residual value for compensating the depth range threshold.
[0056] Specifically, the class classification module is used to: obtain each pixel point of the target feature map, calculate the probability value of each pixel point on the channels of the class classification module, and calculate the position coordinates of the target pixel points with the probability value greater than the preset threshold in the target feature map.
[0057] Exemplarily, the preset threshold is 0.4. Find all the pixel points with the probability value greater than 0.4 in the target feature map in branch 1, and calculate the position coordinates (x, y) of this point in the target feature map. The position coordinates (x, y) are the position coordinates of the center point of the predicted target on this feature map. It should be noted that when there are multiple found position coordinates (x, y), it indicates that there are multiple target objects in the target feature map. The class classification module outputs the class of the target object (such as a vehicle), the probability value, and the center point position. The class can be displayed on the display screen, and the probability value and the center point position act on other branches.
[0058] Specifically, the size classification module is used to: take the position coordinates as the center point of the minimum bounding rectangle, obtain the width and height of the minimum bounding rectangle in the target feature map, and display the minimum bounding rectangle in the target feature map. Among them, the size classification module has two channels, and the values output by the two channels respectively correspond to the width and the height, and the target object in the target feature map is located in the minimum bounding rectangle.
[0059] Exemplarily, the width and height of the minimum bounding rectangle of the target object are obtained on the target feature map in branch 2 based on the position coordinates (x, y) output by the category classification module. The value corresponding to the position coordinates (x, y) in channel 1 is the height, and the value corresponding to the position coordinates (x, y) in channel 2 is the width. When the center point (x, y), width, and height of the minimum bounding rectangle are all available, the target object can be framed in the target feature map in the form of a rectangle and directly displayed on the display screen, facilitating the user to view the driving situation of the vehicle ahead. In case of an emergency, the driver can respond quickly.
[0060] Specifically, the depth range module is used to: obtain the pixel point corresponding to the position coordinates in the target feature map based on the position coordinates, and select the channel with the largest probability value of the pixel point from several channels as the target channel, and obtain the depth range threshold corresponding to the target channel.
[0061] Exemplarily, there are 5 channels at the pixel point of the position coordinates (x, y) on the target feature map in branch 3. For example, the probability values of the 5 channels corresponding to this pixel point are [0.3, 0.6, 0.4, 0.8, 0.64]. By comparing the sizes, it is found that the probability value 0.8 is the largest. Then the channel 4 corresponds to the 3rd index. Then, based on the previously preset depth range threshold of 40 meters corresponding to index 3, the predicted distance to the vehicle ahead is obtained as 40 meters.
[0062] Specifically, the depth residual module is used to: obtain the depth residual corresponding to the position coordinates in the target feature map based on the position coordinates, and calculate the sum of the depth range threshold and the depth residual as the depth of the target object, and use the depth as the measured distance between the host vehicle and the target object.
[0063] Exemplarily, there is 1 channel at the pixel point of the position coordinates (x, y) on the target feature map in branch 4, and the value corresponding to it is the depth residual. The depth residual can correct the depth range threshold, making the measured distance between the host vehicle and the vehicle ahead measured by the model closer to the actual distance, thereby improving the ranging accuracy.
[0064] It should be noted that the deep residual module described in the embodiments of the present invention can adopt the Deep Residual Network in the prior art, which is composed of multiple Bottleneck modules. Since the neural network needs to continuously propagate gradients during the backpropagation process, when the number of network layers increases, the gradients will gradually disappear during propagation (for example, when using the Sigmoid function, for a signal with an amplitude of 1, the gradient decays to 0.25 of the original value for each layer passed backward, and the more layers there are, the more severe the decay), resulting in ineffective adjustment of the weights of the previous network layers. Therefore, the error of the result directly output by the original neural network is relatively large. In order to increase the number of network layers, solve the problem of gradient disappearance, and improve the model accuracy, a residual network structure is introduced for correction. Through this residual network structure, the network layers can be made very deep, the output result can be corrected for errors, and the final classification effect is also very good.
[0065] In the embodiments of the present invention, the target pixel area calculated by the ranging module through the target size predicted by the CNN network can estimate the target depth, convert the problem of directly regressing the depth by the network into classification + regression. First, determine the depth range where the target is located, and then regress the residual between the target depth and the range, which can reduce the regression error and improve the accuracy at the same time.
[0066] Compared with the prior art, the target ranging method based on a monocular camera described in the embodiments of the present invention obtains a target image captured by the monocular camera and inputs the target image into a target ranging model, so that the target ranging model measures the target depth of the target image and outputs the distance between the vehicle itself and the target object; wherein, the target ranging model includes a feature fusion module, an effective information extraction module, and a ranging module; the feature fusion module performs multiple upsamplings on the target image to generate several feature maps of different sizes, performs upsampling on the feature maps to obtain a fused feature map, and fuses the fused feature map with the target position information output by the effective information extraction module to obtain a target feature map, fusing the position information into the network to make the network pay more attention to the position information; the ranging module measures the target depth of the target feature map and outputs the measured distance between the vehicle itself and the target object. By adopting the embodiments of the present invention, the accuracy of target object ranging can be improved, and the inference speed of the adopted target ranging model is fast, and the effect of real-time detection can be achieved.
[0067] See Figure 6 , Figure 6It is a structural block diagram of a target ranging system 10 based on a monocular camera provided by an embodiment of the present invention. The target ranging system 10 based on a monocular camera includes a processor 11, a memory 12, and a computer program stored in the memory 12 and executable on the processor 11. When the processor 11 executes the computer program, it implements the steps in each of the above embodiments of the target ranging method based on a monocular camera. Alternatively, when the processor 11 executes the computer program, it implements the functions of each module / unit in each of the above device embodiments.
[0068] Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory 12 and executed by the processor 11 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the target ranging system 10 based on a monocular camera.
[0069] The target ranging system 10 based on a monocular camera may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The target ranging system 10 based on a monocular camera may include, but is not limited to, a processor 11 and a memory 12. Those skilled in the art can understand that the schematic diagram is only an example of the target ranging system 10 based on a monocular camera, and does not constitute a limitation on the target ranging system 10 based on a monocular camera. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the target ranging system 10 based on a monocular camera may further include input / output devices, network access devices, a bus, etc.
[0070] The so-called processor 11 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 11 is the control center of the target ranging system 10 based on a monocular camera, and connects various parts of the entire target ranging system 10 based on a monocular camera through various interfaces and lines.
[0071] The memory 12 can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory 12, and invoking the data stored in the memory 12, the processor 11 realizes various functions of the monocular camera-based target ranging system 10. The memory 12 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 12 can include high-speed random access memory, and can also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0072] Among them, if the modules / units integrated in the monocular camera-based target ranging system 10 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 11, the steps of the above-mentioned various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0073] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative work.
[0074] Compared with the prior art, the target ranging system 10 based on a monocular camera according to the embodiment of the present invention obtains a target image captured by the monocular camera and inputs the target image into a target ranging model, so that the target ranging model performs target depth measurement on the target image and outputs the distance between the vehicle itself and the target object. The target ranging model includes a feature fusion module, an effective information extraction module, and a ranging module. The feature fusion module performs several upsamplings on the target image to generate several feature maps of different sizes, performs upsampling on the feature maps to obtain a fused feature map, and fuses the fused feature map with the target position information output by the effective information extraction module to obtain a target feature map, fusing the position information into the network to make the network pay more attention to the position information. The ranging module performs target depth measurement on the target feature map and outputs the measured distance between the vehicle itself and the target object. By adopting the embodiment of the present invention, the accuracy of target object ranging can be improved, and the inference speed of the adopted target ranging model is fast, and the effect of real-time detection can be achieved.
[0075] The above is the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A target ranging method based on a monocular camera, characterized in that, Including: Obtain a target image captured by a monocular camera; Input the target image into a target ranging model, so that the target ranging model performs target depth measurement on the target image and outputs the distance between the vehicle itself and the target object; wherein, the target ranging model includes a feature fusion module, an effective information extraction module and a ranging module; the feature fusion module performs multiple upsamplings on the target image to generate multiple feature maps of different sizes, performs upsampling on the feature maps to obtain a fused feature map, and fuses the fused feature map with the target position information output by the effective information extraction module to obtain a target feature map; the ranging module performs target depth measurement on the target feature map and outputs the measured distance between the vehicle itself and the target object, wherein the ranging module includes a depth range module and a depth residual module. After the ranging module obtains the position coordinates, the depth range module obtains the pixel point corresponding to the position coordinates in the target feature map through the position coordinates, obtains the channel with the largest probability value of the pixel point from several channels as the target channel, and obtains the depth range threshold corresponding to the target channel; the depth residual module obtains the depth residual corresponding to the position coordinates in the target feature map through the position coordinates, and calculates the sum of the depth range threshold and the depth residual as the depth of the target object, and uses the depth as the measured distance between the vehicle itself and the target object.
2. The target ranging method based on a monocular camera according to claim 1, characterized in that, The feature fusion module includes N serially connected downsampling layers and M serially connected upsampling layers, M ≤ N, and both N and M are integers greater than 0; each of the downsampling layers outputs a feature map of its corresponding size, and the upsampling layer sequentially performs upsampling on the feature maps according to a preset upsampling strategy to obtain a fused feature map; wherein, each current upsampling layer fuses its upsampling output with the corresponding feature map of the same size after upsampling and inputs it to the next upsampling layer for upsampling until the current upsampling layer is the last upsampling layer.
3. The target ranging method based on a monocular camera according to claim 2, characterized in that, The feature fusion module adopts the backbone network of DLA34, including 5 serially connected downsampling layers and 3 serially connected upsampling layers.
4. The target ranging method based on a monocular camera according to claim 1, characterized in that, The input of the effective information extraction module is one of the multiple feature maps of different sizes that meets a preset size, which is a standard feature map.
5. The target ranging method based on a monocular camera according to claim 4, characterized in that, The effective information extraction module includes a global information extraction module, a local information extraction module and a fusion module; wherein, the global information extraction module sequentially performs convolution operation, global average pooling operation and probability operation on the standard feature map to obtain a first-size feature map; the local information extraction module performs convolution operation and probability operation on the standard feature map to obtain a second-size feature map; the fusion module integrates and restores the channels of the first-size feature map and the second-size feature map to output target position information.
6. The target ranging method based on a monocular camera according to claim 1, characterized in that, The ranging module further includes a category classification module and a size classification module, and the input of each module is the target feature map; wherein, the category classification module includes a first convolutional layer, a ReLU layer, a second convolutional layer, and a probability layer; the size classification module includes a third convolutional layer, a ReLU layer, and a fourth convolutional layer; the depth range module includes a fifth convolutional layer, a BN layer, a Mish layer, a sixth convolutional layer, a ReLU layer, a seventh convolutional layer, and a probability layer; the depth residual module includes an eighth convolutional layer, a ReLU layer, and a ninth convolutional layer.
7. The target ranging method based on a monocular camera according to claim 6, characterized in that, The category classification module is configured to: obtain each pixel point of the target feature map, calculate the probability value of each pixel point on the channels of the category classification module, and calculate the position coordinates of the target pixel points whose probability values are greater than a preset threshold in the target feature map; the size classification module is configured to: take the position coordinates as the center point of the minimum bounding rectangle, obtain the width and height of the minimum bounding rectangle in the target feature map, and display the minimum bounding rectangle in the target feature map; wherein, the size classification module has two channels, and the values output by the two channels respectively correspond to the width and the height, and the target object in the target feature map is located in the minimum bounding rectangle.
8. The target ranging method based on a monocular camera according to claim 7, characterized in that, The depth range module includes a plurality of channels, and each channel is preset with its corresponding depth range threshold, and the depth range threshold is a preset distance value.
9. A target ranging system based on a monocular camera, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the monocular camera-based target ranging method according to any one of claims 1 to 8.
Citation Information
Patent Citations
A multi-scale target detection method fusing context information
CN109816012A
Target detection method and device, storage medium and terminal
CN112906794A