Target detection method and device, terminal equipment and storage medium
By combining the data of lidar and cameras in the roadside perception technology, and generating and analyzing bounding boxes to determine the detection target, the inaccurate target detection caused by multi-sensor coordination problems in traditional roadside perception technology is solved, and more efficient and accurate target detection is achieved.
Patent Information
- Application Number
- CN202311750890.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional roadside perception technology uses multiple sensors for fusion perception, and there is a problem of inaccurate object detection caused by the coordination problem between multiple sensors.
Lidar and cameras are used to collect relevant information on the current traffic section, and corresponding bounding boxes are generated according to the acquisition results, and then analyzed in combination with each bounding box to determine the detection target.
It improves the accuracy of object detection, reduces the complexity of analysis algorithms, and improves the efficiency of object detection.
Smart Images

Figure CN120219699A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of object detection, and particularly relates to an object detection method, device, terminal device, and storage medium. Background Art
[0002] With the continuous increase in traffic congestion and the number of traffic accidents, the development and application of intelligent transportation systems have gradually received attention. Roadside perception technology plays an important role in intelligent transportation systems, mainly used to accurately obtain real-time vehicle information and targets such as obstacles on the road. By collecting and analyzing these key information, roadside perception technology can provide valuable data and support for traffic planners and operators to achieve more efficient and safer traffic operations. However, traditional roadside perception technology uses multiple sensors for fusion perception, and there is a problem of inaccurate object detection due to the coordination problem between multiple sensors. Summary of the Invention
[0003] The purpose of this application is to provide an object detection method, device, terminal device, and storage medium, aiming to solve the problem of inaccurate object detection due to the coordination problem between multiple sensors.
[0004] The first aspect embodiment of this application provides an object detection method, and the object detection method includes:
[0005] Obtain the point cloud image collected by the lidar and the road section image collected by the camera in the current traffic section;
[0006] Generate at least one first object bounding box according to the point cloud image;
[0007] Generate at least one second object bounding box according to the road section image;
[0008] Determine the target image detected in the current traffic section according to the at least one first object bounding box and the at least one second object bounding box.
[0009] In an alternative embodiment, both the first object bounding box and the second object bounding box are two-dimensional images. The step of determining the target image detected in the current traffic section according to the at least one first object bounding box and the at least one second object bounding box includes:
[0010] Associate all the first object bounding boxes and all the second object bounding boxes to determine the planar matching relationship between each first object bounding box and the second object bounding box;
[0011] Determine the first object bounding box and the second object bounding box with a planar matching relationship as the same target image detected in the current traffic section.
[0012] In an alternative embodiment, both the first target bounding box and the second target bounding box are three-dimensional images. Determining the target image detected in the current traffic section based on the at least one first target bounding box and the at least one second target bounding box includes:
[0013] Associate all the first target bounding boxes and all the second target bounding boxes to determine the spatial matching relationship between each first target bounding box and the second target bounding box;
[0014] Determine the first target bounding box and the second target bounding box with a spatial matching relationship as the same target image detected within the current traffic section.
[0015] In an alternative embodiment, generating at least one first target bounding box based on the point cloud image includes:
[0016] Generate at least one three-dimensional target bounding box based on the point cloud image in combination with a three-dimensional object detection model;
[0017] Perform dimensionality reduction processing on each three-dimensional target bounding box to obtain the first target bounding box.
[0018] In an alternative embodiment, the associating all the first target bounding boxes and all the second target bounding boxes to determine the planar matching relationship between each first target bounding box and the second target bounding box includes:
[0019] Map all the first target bounding boxes and all the second target bounding boxes to the same two-dimensional spatial coordinate system;
[0020] If the ratio of the intersection area to the union area of a first target bounding box and a second target bounding box is greater than a set threshold, determine the planar matching relationship between the first target bounding box and the second target bounding box.
[0021] In an alternative embodiment, the associating all the first target bounding boxes and all the second target bounding boxes to determine the spatial matching relationship between each first target bounding box and the second target bounding box includes:
[0022] Map all the first target bounding boxes and all the second target bounding boxes to the same three-dimensional spatial coordinate system;
[0023] If the ratio of the intersection volume to the union volume of a first target bounding box and a second target bounding box is greater than a first set threshold, determine that the first target bounding box and the second target bounding box form a spatial matching relationship.
[0024] In an alternative embodiment, before generating at least one first target bounding box based on the point cloud image, the object detection method further includes:
[0025] Based on a map image with an absolute accuracy less than a set distance, depth information is added to the road segment image to generate a three-dimensional road segment image;
[0026] Based on the three-dimensional road segment image, by simulating the camera and lidar, the relative installation positions of the models of the camera and lidar are obtained in the simulation model;
[0027] According to the relative rotation vector and translation vector of the camera and the lidar in the simulation model, the point cloud image is mapped into the three-dimensional road segment image;
[0028] Adjust the relative rotation vector and translation vector to align the targets in the point cloud image and the three-dimensional road segment image; wherein, generating at least one second target bounding box according to the road segment image includes: generating at least one second target bounding box according to the road segment image and the map image.
[0029] In an alternative embodiment, generating at least one second target bounding box according to the road segment image includes:
[0030] Perform Mercator projection mapping on the map image to obtain the three-dimensional information of each map unit in the map image;
[0031] Select multiple groups of valid points from each map unit, and transform each map unit into a prior perception three-dimensional element through perspective transformation;
[0032] Generate the two-dimensional environmental integration image including depth information according to the prior perception three-dimensional element and the road segment image;
[0033] Input the two-dimensional environmental integration image into a preset deep learning model, and the deep learning model outputs the at least one second target bounding box; wherein the deep learning model is trained to convergence by using historical two-dimensional environmental integration images with all second target bounding boxes labeled.
[0034] In an alternative embodiment, the object detection method further includes:
[0035] Establish an initial model based on the YOLOV5 algorithm; the initial model includes multiple calculation layers, each calculation layer includes multiple nodes, and the input size of each node is a first scale and the output size is a second scale;
[0036] Add cross-layer weighted connections between all nodes with the same first scale and second scale to obtain an updated model; wherein the deep learning model is trained to convergence by using historical two-dimensional environmental integration images with all second target bounding boxes labeled to train the updated model.
[0037] In an alternative embodiment, after obtaining the updated model, the object detection method further includes:
[0038] Performing feature fusion on the outputs of each node in the top calculation layer and the calculation layer adjacent to the top layer of the updated model according to channels;
[0039] For non-top and non-bottom calculation layers, performing feature fusion on the outputs of two nodes on adjacent paths according to channels, and performing a feature map superposition operation on the outputs of two nodes on non-adjacent paths to obtain a fused output.
[0040] In an alternative embodiment, the performing a feature map superposition operation on the outputs of two nodes on non-adjacent paths includes:
[0041] Performing feature map fusion by combining the feature maps of the outputs of two nodes on adjacent paths according to the initial weight coefficient of each feature map;
[0042] Updating the initial weight coefficient within a preset weight range, and recombining the feature maps of the outputs of two nodes on adjacent paths to perform feature map fusion until no error occurs when the training speed is higher than a set speed threshold.
[0043] In an alternative embodiment, the object detection method further includes:
[0044] Setting a monitoring layer for sequentially performing object monitoring on the feature maps output by all calculation layers.
[0045] In an alternative embodiment, if the closest first object bounding box and second object bounding box do not intersect, when associating all first object bounding boxes and all second object bounding boxes to determine the planar matching relationship between each first object bounding box and the second object bounding box, it further includes:
[0046] Performing secondary association on the non-intersecting first object bounding box and second object bounding box using category information and ellipse-related thresholds until each first object bounding box has the spatial matching relationship.
[0047] An embodiment of the second aspect of the present application provides an object detection device, and the object detection device includes:
[0048] An acquisition module that acquires a point cloud image collected by a lidar and a road segment image collected by a camera in the current traffic road segment;
[0049] A first generation module that generates at least one first object bounding box according to the point cloud image;
[0050] A second generation module generates at least one second object bounding box according to the road segment image;
[0051] The target detection module determines the target image detected in the current traffic section according to the at least one first target bounding box and the at least one second target bounding box.
[0052] A third aspect of the embodiments of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in the first aspect above is implemented.
[0053] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0054] A fifth aspect of the embodiments of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device is enabled to execute the method described in the first aspect above.
[0055] Beneficial effects of the present application
[0056] A target detection method, device, terminal device, and storage medium provided by the present application first generate at least one first target bounding box according to the point cloud image, then generate at least one second target bounding box according to the section image, and finally determine the target image detected in the current traffic section according to the at least one first target bounding box and the at least one second target bounding box. The present application proposes a multi-sensor fusion perception method based on lidar and camera. By using the point cloud image and the section image and combining each bounding box for result-level fusion analysis to determine the detection target, the accuracy of target detection can be improved, and the complexity of the analysis algorithm is reduced compared with the prior art that performs analysis at the pixel level, thereby improving the efficiency of target detection. Description of the Drawings
[0057] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0058] Figure 1 It is one of the flowcharts of a joint calibration method provided by the embodiments of the present application;
[0059] Figure 2 It is the second of the flowcharts of a high-precision map application method provided by the embodiments of the present application;
[0060] Figure 3a One of the schematic diagrams of the combined calibration results provided by the embodiments of the present application;
[0061] Figure 3b Another one of the schematic diagrams of the combined calibration results provided by the embodiments of the present application;
[0062] Figure 4 The flow schematic diagram of a channel fusion feature method provided by the embodiments of the present application;
[0063] Figure 5 The flow schematic diagram of an intermediate layer feature fusion method provided by the embodiments of the present application;
[0064] Figure 6 The flow schematic diagram of a cross-layer connection method provided by the embodiments of the present application;
[0065] Figure 7 The flow schematic diagram of an object detection method provided by the embodiments of the present application;
[0066] Figure 8 The comparison diagram of the inference time before and after the model optimization provided by the embodiments of the present application;
[0067] Figure 9 The structural schematic diagram of an object detection device provided by the embodiments of the present application; Figure 10 The structural schematic diagram of a terminal device provided by the embodiments of the present application. Detailed implementation manners
[0068] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0069] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0070] It should also be understood that the term "and / or" used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.
[0071] As used in the specification and appended claims of the present application, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0072] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.
[0073] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present application means that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0074] It should be understood that the magnitudes of the sequence numbers of the steps in this embodiment do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0075] The development and application of intelligent transportation systems have gradually received attention. Roadside sensing technology plays an important role in intelligent transportation systems, mainly used to accurately obtain in real-time vehicle information and targets such as obstacles on the road. By collecting and analyzing these key information, roadside sensing technology can provide valuable data and support for traffic planners and operators to achieve more efficient and safer traffic operations. However, traditional roadside sensing technology uses multiple sensors for fusion sensing, and there is a problem of inaccurate target detection due to the coordination problem between multiple sensors.
[0076] To solve the above problems, the present application provides an object detection method, which uses a lidar and a camera to collect relevant information of the current traffic section, generates corresponding bounding boxes according to the collection results respectively, and then analyzes by combining each bounding box to determine the detection object, thereby improving the accuracy of object detection. In addition, compared with the prior art that analyzes at the pixel level to determine the detection object, the present application uses the bounding box for result-level fusion analysis, which can also reduce the complexity of the analysis algorithm and improve the object detection efficiency.
[0077] In the embodiment of the present application, three improvements are made. First, a result-level fusion method is combined with the point cloud image and the camera image, which reduces the algorithm complexity and improves the success rate of object fusion. The fusion perception scheme achieves an object fusion accuracy of 98.2% and a fusion recall rate of 95.1% on the aggregated dataset. The fusion perception distance is about 90 meters, the average positioning error within 50 meters is about 0.2 meters, and the average positioning error between 50-90 meters is about 0.35 meters.
[0078] Second, the present application proposes to use the lidar to build a model for calibration, and realizes joint calibration by building a model between the lidar and the camera, improving the joint calibration efficiency.
[0079] Finally, at the object detection algorithm level, on the basis of selecting the YOLO5s model, through cross-cascading, more features are fused using cross-layer connections, increasing the depth of the YOLOv5 algorithm and improving its ability to detect small objects.
[0080] The above three core invention points of the present application will be described separately below.
[0081] 1. Provide depth information to the camera pixel by pixel using a high-precision map to achieve joint calibration by building a model between the lidar and the camera.
[0082] In the embodiment of the present application, the lidar is a radar system that detects the position, speed and other characteristic quantities of the target by emitting laser beams. Its working principle is to emit a detection signal (laser beam) to the target, and then compare the received signal (target echo) reflected from the target with the emitted signal. After appropriate processing, relevant information of the target can be obtained, such as parameters of the target distance, azimuth, altitude, speed, attitude, and even shape, so as to realize the detection, tracking and recognition of targets such as vehicles.
[0083] Among them, the camera can be a monocular camera, which consists of an indicator, a lens, a photosensitive element, a processor, etc. The indicator is used to display the captured images and videos. The lens controls the focal length and clarity of the images. The photosensitive element is responsible for converting optical signals into electrical signals. The processor processes the electrical signals, converts them into digital information and stores them in a memory card or hard disk. In other preferred embodiments, the monocular camera also has wireless technologies such as WiFi and Bluetooth, and can perform data transmission and remote control with a mobile phone or a computer.
[0084] Among them, the current traffic section can be the section within the scanning range of the lidar and the camera at the current moment.
[0085] Among them, the point cloud image is three-dimensional point cloud image data.
[0086] Exemplarily, the lidar and the monocular camera are respectively installed at preset positions on the current traffic section. The lidar collects the three-dimensional point cloud image data of the traffic section, and the camera collects the section image of the traffic section, and stores and transmits the data respectively.
[0087] In current multi-sensor fusion, a 2D-3D conversion relationship is usually required. The lidar can directly output the perceived position of the target, while the positioning of the camera is generally obtained by calibrating the internal and external parameters of the camera.
[0088] In this embodiment, common means can be used for calibrating the lidar and the camera. This application is not limited thereto. Specifically, if common means are used for calibrating the lidar and the camera, the image data captured by the monocular camera is represented by (U, V), and the three-dimensional point cloud captured by the lidar is represented by (X, Y, Z). Create a transformation matrix M that maps the 3D point (x, y, z) to the 2D point (u, v), using Equation (1):
[0089]
[0090] The matrix (fu, fv, u0, v0) are camera parameters. fu and fv are the scale factors in the XY axis directions (the effective focal lengths in the horizontal and vertical directions), and u0, v0 are the center points of the image plane, also known as the principal point coordinates. R is the rotation matrix, and t is the translation vector. By substituting multiple sets of feature points, M can be solved.
[0091] In the embodiment of the second inventive point of the present application, coordinate calibration can be performed based on simulation. Before that, in this embodiment, first, the high-precision map is used to provide depth information for each pixel of the camera, so that the section image can be configured as "pseudo-three-dimensional" in the three-dimensional simulation model, that is, each pixel in the section image with the map image included contains three-dimensional information, and thus can be directly reflected in the three-dimensional simulation model.
[0092] Exemplarily, referring to Figure 1 , the steps of joint calibration include:
[0093] S01: Based on a map image with an absolute accuracy less than a set distance, depth information is added to the road section image to generate a three-dimensional road section image;
[0094] S02: Based on the three-dimensional road section image, by simulating the camera and lidar, the relative installation positions of the models of the camera and lidar are obtained in the simulation model;
[0095] S03: According to the relative rotation vector and translation vector of the camera and the lidar in the simulation model, the point cloud image is mapped into the three-dimensional road section image;
[0096] S04: Adjust the relative rotation vector and translation vector to align the targets in the point cloud image and the three-dimensional road section image.
[0097] Specifically, the absolute accuracy less than the set distance can generally be less than 5 meters. More precisely, the absolute accuracy can be set to less than 1 meter, or centimeter-level navigation can be adopted, and the absolute accuracy can be within 10 centimeters. This application does not limit this, that is, the absolute accuracy only needs to meet the model calibration requirements of this solution to be defined as high precision.
[0098] Specifically, in step S01, the following steps can be used to obtain the three-dimensional road section image, that is:
[0099] S011: Perform Mercator projection mapping on the map image to obtain the three-dimensional information of each map unit in the map image;
[0100] S012: Select multiple groups of valid points from each map unit, and transform each map unit into a priori perception three-dimensional elements through perspective transformation;
[0101] S013: According to the a priori perception three-dimensional elements and the road section image, generate a three-dimensional road section image including depth information.
[0102] Specifically, as Figure 2 shown, Figure 2 (a) is a map unit, Figure 2 (b) is a high-precision map, Figure 2 (c) is the road section image captured by the camera. The high-precision map forms a three-dimensional map unit (each pixel contains three-dimensional information) through Mercator projection mapping. Then, multiple groups of valid points are selected, and the map unit (a) is transformed into (b) through perspective transformation, thereby endowing it with depth information from the camera perspective (c) to obtain a three-dimensional road section image.
[0103] In steps S02 to S04 of this embodiment, based on the camera image, the relative installation positions of the camera and the lidar in the real world are simulated by simulation software. Using the relative rotation and translation vectors of the two simulated sensors, the point cloud data is mapped onto the image. By fine-tuning the rotation and translation vectors, the point cloud target is aligned with the image target, completing the joint calibration of the multi-sensors. This process is as Figure 3a and Figure 3b shown. It can be seen that through the above calibration, the positions of the targets in the point cloud image and the road segment image are calibrated to the same coordinate system, completing the alignment of the point cloud target and the image target.
[0104] As can be seen from the above embodiments, the second inventive point of this application first uses the high-precision map to provide the depth information for each pixel of the camera, which is a necessary condition for subsequent simulation. Compared with the common 3D-to-2D conversion method, this application directly fuses the data of the two sensors without going through the intermediate dimension conversion, so the obtained data is more accurate. At the same time, joint calibration is realized, improving the efficiency of the calibration process. 2. Target detection algorithm - YOLO5S
[0105] The current YOLO5 model algorithm focuses on capturing detailed information such as its shallow edge features. On the basis of obtaining simple features, it can help the network more accurately regress the target boundary. The focus of the deep layer of the network is to extract high-level semantic information, which can extract more complex features and help the network accurately detect the target.
[0106] In the embodiment of this application, for the model establishment steps, it specifically includes:
[0107] S051: Establish an initial model based on the YOLOV5 algorithm; the initial model includes multiple calculation layers, each calculation layer includes multiple nodes, and the input size of each node is the first scale and the output size is the second scale;
[0108] S052: Add cross-layer weighted connections between all nodes with the same first scale and second scale to obtain an updated model; wherein the deep learning model is obtained by training the updated model to convergence using the historical two-dimensional environment integration image annotated with all second target bounding boxes.
[0109] In the embodiments of the present application, in order to further enhance the model's attention to shallow semantics, fully integrate the semantic information extracted from each layer of the FPN (Feature Pyramid Network, object detection), and enhance the network's detection ability for small objects, the FPN of YOLOv5 is improved in the present application, which is called the Weighted Connections Across Layers - Path Aggregation network (WCAL - PAN), and cross - layer weighted connections are added between nodes with the same input and output sizes. The cross - layer cascading structure can effectively fuse information such as shallow details, edges, and contours into the deep network, fuse the shallow detail information of the target with almost no increase in computational complexity, make the network more accurate in the regression of small objects, and effectively improve the intersection - over - union IoU of the prediction box and the ground truth box. At the same time, considering that adding shallow features in the cross - layer cascade will have a certain impact on the deep semantic information, a learnable method is used for fusion.
[0110] In a specific embodiment, after obtaining the updated model, for model calculation, it specifically includes the following steps:
[0111] Perform feature fusion by channels on the outputs of each node in the top - layer calculation layer and the calculation layer adjacent to the top - layer of the updated model;
[0112] For non - top - layer and non - bottom - layer calculation layers, perform feature fusion by channels on the outputs of two nodes on adjacent paths, and perform a feature map superposition operation on the outputs of two nodes on non - adjacent paths to obtain a fused output.
[0113] Specifically, during the feature fusion process, the information flow speed of the top - layer and bottom - layer nodes is fast, and the number of convolutional operations experienced is small, so less detailed information is lost. To reduce the complexity of the model, a direct connection operation is used to fuse features by channels, as Figure 4 shown.
[0114] For nodes in other layers, a concat operation on adjacent paths and an add operation with learnable weights on non - adjacent paths are used for feature fusion. The add operation can not only reduce the computational amount but also reduce the fusion of invalid shallow features. The calculation is shown in Equation (2):
[0115]
[0116] In the formula, ix represents each feature map to be fused; ui is the weight coefficient of the feature map.
[0117] Furthermore, in the add operation of the intermediate nodes in the embodiments of the present application, the weight coefficient is learnable, that is: the operation of performing a feature map superposition on the outputs of two nodes on non - adjacent paths includes:
[0118] According to the initial weight coefficient of each feature map, combine the feature maps output by two nodes on adjacent paths for feature map fusion;
[0119] Update the initial weight coefficient within a preset weight range, and recombine the feature maps output by two nodes on adjacent paths for feature map fusion until no error occurs when the training speed is higher than the set speed threshold.
[0120] According to the above formula (2), through learning and updating, the initial weight coefficient is set to 1, indicating uniform fusion of two layers of feature maps; a very small value (3 - 10) can effectively prevent numerical instability. Normalize the weight value between 0 and 1 to improve the training speed and prevent the occurrence of unstable training. According to formula (2), the intermediate layer feature fusion method is as Figure 5 shown.
[0121] In Figure 4 and Figure 5 , given a certain input feature mapping layer 11wchFR ××∈ , the corresponding feature map layer of the top-down path 22whcFR ××∈ , and the functional map layer corresponding to the bottom-up path 33wchFR××∈, indicating the concat operation, and “+” indicating the addition operation. The weight values w18 and w28 of the feature maps respectively fuse the two paths.
[0122] It can be seen from the two figures that the feature fusion of the top layer and bottom layer output nodes is performed through the concat operation, while the feature fusion process of the intermediate layer nodes is to first perform the concat operation, and then perform weighted calculation by the input layer after channel alignment. The resulting feature map obtained at the output node contains composite features of details, edges, and high-level semantic information. For ease of understanding, taking the intermediate layer P4 as an example, the output calculations on each path are shown in formulas (3) and (4) as follows:
[0123] P4 td = Concat(P4 in , Resize(Conv(P5 in ))) (3)
[0124]
[0125] Where the formula Pk represents the k inputs of the first layer, td Pk represents the output of the intermediate node of the out Pk of the first layer in the top-down path, and k represents the output of the output node of the first layer in the top-down path.
[0126] In a further preferred embodiment, the YILOV5s model of the present application further adds a computing layer compared to the current YOLOV5 model, that is, a monitoring layer is set up. The monitoring layer is used to sequentially perform target monitoring on the feature maps output by all computing layers, which can further deepen the depth of the feature pyramid.
[0127] Specifically, the advanced receptive field of FPN contains higher-level semantic information, which can enhance the learning ability of the network and further improve the detection accuracy. The FPN of YOLOv5 consists of 3 layers. On the basis of improvement (1), the present application increases it to 4 layers, making full use of the proposed cross-level cascading structure. In addition, in order to match the depth of FPN, the present application adds a detection layer in the Detect part to sequentially perform target detection on the feature maps output by P3, P4, P5, and P6. After adding the detection layer, the anchor box spacing is more reasonable, and the training stability, as well as the convergence speed and accuracy of the model, can be effectively improved.
[0128] Figure 6 The dotted line in... indicates a cross-layer connection. As can be seen from the figure, only the two intermediate layers of P4 and P5 are subjected to cross-level weighted fusion. For the top layer P6 and the bottom layer P3, since the information flow loss is not large and considering the efficiency of the model at the same time, the present application directly splices the two parts of the feature maps along the channels.
[0129] It should be noted that the shallow network of deep learning focuses on detailed information, such as edge features. On the basis of obtaining simple features, it can help the network more accurately regress the target boundary; the deep network focuses on extracting high-level semantic information, which can extract more complex features and can help the network accurately detect the target. The FPN structure accordingly uses shallow features to distinguish simple targets and deep features to distinguish complex targets, aiming to obtain a more robust detection result. The FPN structure of YOLOv5 is based on PAN, creating a bottom-up path enhancement, accelerating the flow of bottom-layer information, and being able to well fuse the semantic information of each layer.
[0130] Therefore, the embodiment of the present application gives a solution to improve YOLOv5 to form YOLOv5s. However, it should be noted that the present application can also adopt other models, where the computing layer is the convolutional layer of other models, that is, the cross-layer weighted connection node involved in the core concept here of the present application is not only applicable to YOLOV5, but can also be further applicable to other models, such as CNN models and PNN models, etc. The present application does not limit this.
[0131] 3. Result-level fusion
[0132] According to different levels of image representation, fusion can be divided into three levels: pixel-level fusion, feature-level fusion, and result-level fusion. Considering the real-time requirement of road perception, this application selects to optimize the design at the result-level fusion.
[0133] Specifically, referring to Figure 7 , the object detection method provided by this application includes:
[0134] S1: Obtain the point cloud image collected by the lidar and the road section image collected by the camera in the current traffic section;
[0135] S2: Generate at least one first object bounding box according to the point cloud image;
[0136] S3: Generate at least one second object bounding box according to the road section image;
[0137] S4: Determine the object image detected in the current traffic section according to the at least one first object bounding box and the at least one second object bounding box.
[0138] Combined with the above embodiments, the optical image of each object in the current traffic section can be obtained according to the road section image. Currently, the object detection models all identify the images captured by the camera, and intercept the bounding boxes corresponding to the object images from the road section images captured by the camera, so as to achieve object detection.
[0139] Traditional roadside perception uses multi-sensors for fusion perception, and there are problems such as sensor joint calibration, small object recognition, and real-time requirements. This application proposes a multi-sensor fusion perception method based on lidar and camera. By using the point cloud image and the road section image, and combining each bounding box for result-level fusion analysis to determine the detection object, the accuracy of object detection can be improved, and the complexity of the analysis algorithm is reduced compared with the prior art in the pixel-level layer, and the object detection efficiency is improved.
[0140] In the embodiments of this application, for the three-dimensional point cloud image, the point cloud image is creatively detected by using an object detection model to obtain the first object bounding box.
[0141] In this application, the object bounding box can be obtained by an object detection model according to the point cloud image and the road section image. Specifically, for a three-dimensional image, the object bounding box can be obtained by a three-dimensional object detection model. For a two-dimensional object detection model, the three-dimensional image can be first reduced to two dimensions, and then the two-dimensional image is input into the two-dimensional object detection model to obtain the object bounding box.
[0142] In a preferred embodiment of the present application, the object detection model is further improved. Specifically, when the object detection model of the present application inputs a point cloud image, a three-dimensional object detection model can be used. In order to reduce the model complexity, the present application can reduce the three-dimensional point cloud image to a two-dimensional image, and then use a two-dimensional model for processing to obtain a two-dimensional object bounding box.
[0143] However, it should be noted that whether the present application uses three-dimensional object detection or two-dimensional object detection, the present application can use the Intersection over Union (IOU) algorithm to perform spatial matching on the first object bounding box and the second object bounding box. The difference is that for a three-dimensional object bounding box, the matching relationship is determined by the spatial distance and the volume intersection, while for a two-dimensional object bounding box, the matching relationship is determined by the in-plane distance and the area intersection.
[0144] Specifically, in some embodiments of the present application, both the first object bounding box and the second object bounding box are two-dimensional images. Determining the target image detected in the current traffic section according to the at least one first object bounding box and the at least one second object bounding box includes:
[0145] S41a: Associating all the first object bounding boxes and all the second object bounding boxes to determine the planar matching relationship between each first object bounding box and the second object bounding box;
[0146] S42a: Determining the first object bounding box and the second object bounding box with the planar matching relationship as the same target image detected in the current traffic section.
[0147] In other embodiments, both the first object bounding box and the second object bounding box are three-dimensional images. Determining the target image detected in the current traffic section according to the at least one first object bounding box and the at least one second object bounding box includes:
[0148] S41b: Associating all the first object bounding boxes and all the second object bounding boxes to determine the spatial matching relationship between each first object bounding box and the second object bounding box;
[0149] S42b: Determining the first object bounding box and the second object bounding box with the spatial matching relationship as the same target image detected in the current traffic section.
[0150] Specifically, the point cloud image can generate a three-dimensional object bounding box in combination with the above YOLOV5s model. That is, this step specifically includes:
[0151] Generating at least one three-dimensional object bounding box according to the point cloud image in combination with a three-dimensional object detection model;
[0152] Perform dimensionality reduction on each three-dimensional target bounding box to obtain the first target bounding box.
[0153] The three-dimensional object detection model can adopt the above YOLOV5s model, or a YOLOV5 model without cross-layer weighted connection, or a CNN and PNN model without cross-layer weighted connection, or a CNN and PNN model with cross-layer weighted connection. This application places no restrictions on this, that is, in the inventive point of result-level fusion of this application, the above preferred model can be adopted, or existing models can be adopted.
[0154] According to the above embodiments, for the planar matching relationship, the intersection area and the union area can be selected, that is, all the first target bounding boxes and all the second target bounding boxes are associated, and the planar matching relationship between each first target bounding box and the second target bounding box is determined, including:
[0155] Map all the first target bounding boxes and all the second target bounding boxes to the same two-dimensional space coordinate system;
[0156] If the ratio of the intersection area to the union area of a first target bounding box and a second target bounding box is greater than a set threshold, then determine the planar matching relationship between the first target bounding box and the second target bounding box.
[0157] Similarly, for the spatial matching relationship, the intersection volume and the union volume can be selected, that is, all the first target bounding boxes and all the second target bounding boxes are associated, and the spatial matching relationship between each first target bounding box and the second target bounding box is determined, including:
[0158] Map all the first target bounding boxes and all the second target bounding boxes to the same three-dimensional space coordinate system;
[0159] If the ratio of the intersection volume to the union volume of a first target bounding box and a second target bounding box is greater than a first set threshold, then determine that the first target bounding box and the second target bounding box form a spatial matching relationship.
[0160] Result-level fusion usually includes independent perception and detection of two sensors, and then fuses the perception results. In the result-level fusion of this application, if the first target bounding box and the second target bounding box are three-dimensional images, three-dimensional bounding boxes are obtained from two output ends, and the spatial matching of the two three-dimensional bounding boxes is performed through IOU. When a certain threshold is reached, it can be considered that the two bounding boxes overlap and the detected objects are the same.
[0161] Preferably, considering the real-time requirement of the fusion algorithm, the three-dimensional bounding box of the point cloud is reduced to a two-dimensional level, and the results are fused with the detection boxes in the two-dimensional image. That is, in the above solution of the present application, each three-dimensional target bounding box is dimensionally reduced to obtain the first target bounding box, and then two-dimensional plane matching is a preferred solution of the present application, which greatly reduces the complexity of the algorithm and is more conducive to data association.
[0162] In this preferred solution, the bounding box generated from the point cloud detection result is mapped to the two-dimensional coordinate system, and the point cloud bounding box and the image bounding box with the closest distance are matched in the two-dimensional space. During the matching process, the image box is given priority, and an IOU threshold is set to match with the point cloud bounding box.
[0163] Further, if the closest first target bounding box and second target bounding box do not intersect, associating all the first target bounding boxes and all the second target bounding boxes to determine the planar matching relationship between each first target bounding box and the second target bounding box further includes:
[0164] Performing secondary association on the non-intersecting first target bounding box and second target bounding box using the category information and the ellipse-related threshold until each first target bounding box has the spatial matching relationship.
[0165] Specifically, for non-intersecting targets, secondary association is performed using the category information and the ellipse-related threshold. If the boundary identity categories of the targets are inconsistent, the visual category is given priority. The output targets can be divided into three types: successfully fused targets, purely visually detected targets, and purely lidar detected targets.
[0166] Experimental data verification
[0167] The present application conducts experiments using a preset dataset, which includes multiple scenarios such as urban and tunnel environments, and a total of 15,302 annotated frames (including images and point clouds) for training visual and point cloud algorithms. The experiments include three parts: visual algorithm analysis, fusion accuracy analysis, and fusion positioning analysis.
[0168] The present application divides the 15,302 image data into a training set, a validation set, and a test set according to a ratio of 8:1:1. During the training process, basic data augmentation methods such as random brightness and contrast adjustment, flipping, etc. are used to enhance the robustness of the model. To verify the effectiveness of the proposed visual algorithm, it is optimized based on YOLOv5s.
[0169] Table 1 Performance comparison between WCAL-PAN and PAN based on YOLOv5s
[0170]
[0171] Table 1 shows the performance comparison before and after optimization. PAN is the original algorithm, WCAL-PAN_1 is the cross-layer weighted connection, and WCAL-PAN_2 is the PAN module after deepening the pyramid. It can be seen from Table 1 that although the FPS decreased slightly before and after optimization, it still meets the real-time requirements. Compared with PAN, the use of WCAL-PAN increased the mAP@0.5 metric by 4.1 percentage points, and under high IoU requirements, the mAP@0.5:0.95 metric increased by 4.9 percentage points. After deepening the pyramid, compared with WCAL-PAN_1, the mAP@0.5:0.95 metric increased by 3.5 percentage points, and compared with WCAL-PAN_2, it increased by 2.1 percentage points, proving that the cross-layer cascade structure can further improve the network's detection accuracy for targets.
[0172] Table 2 Performance Comparison of Objects of Different Scales
[0173]
[0174] According to the rules of dividing targets into small, medium, and large according to COCO, the targets were divided as shown in Table 2. The optimized model improved the detection of small targets by 4.5% in terms of mAP@0.5 and improved the detection of targets by 4.9% at a higher IoU threshold of mAP@0.5:0.95. The detection of medium and large targets also improved significantly, further indicating that the cross-layer fusion structure can improve the detection accuracy of small targets.
[0175] Figure 8 The comparison of the inference time before and after optimization is shown. On the left (a) is the inference time of the original YOLOv5, and on the right (b) is the inference time of the optimized model, with a batch size of 4. As shown in the figure, the basic inference time is stable at around 40 - 50 ms, and due to the cross-layer connection and the increase in network depth, the inference time did not increase significantly.
[0176] This experiment used the YOLOv5s-WCAL-PAN algorithm and the SECOND algorithm for inference, including six types of objects: people, cars, buses, trucks, tricycles, and bicycles. The detection results output by the two sensors were calibrated and unified to the two-dimensional image coordinates.
[0177] Table 3 Classification Accuracy of Different Target Classes in Different Ranges
[0178]
[0179] Table 4 Fusion Recall Rate of Different Classes of Targets in Different Perception Ranges
[0180]
[0181] Divided into 3 categories according to the perception range. Since there is a 5-meter blind area under the sensor, only data from 5 meters to 90 meters are included. They are the overall global fusion perception effect of the system, and the perception range is divided into 5 - 50 meters, 9 - 9 meters, and 50 - 90 meters. The accuracy and recall rate within different perception ranges are based on the field of view of the camera.
[0182] As can be seen from Table 3 and Table 4, the overall fusion accuracy is 98.2%, and the fusion recall rate is 95.1%. The fusion effect within 50 meters is better than that within 90 meters. The accuracy and recall rate of large targets such as cars, buses, and trucks are relatively close from near to far, with strong generalization ability. For small targets such as passengers and bicycles, the average precision at the near end is 98.1%, and the average precision at the far end is 94.9%. The reason for the decrease may be that small targets are more affected by the far-end positioning error.
[0183] This system compares the positioning output by the fusion perception system with the RTK navigation positioning of the target as the ground truth, and the comparison error is as Figure 8 shown. Through data comparison, the positioning error gradually increases with the increase of the perception distance. Due to the existence of a 5m blind area, only the positioning error within the range of 5 - 90m is considered. As can be seen from the figure, the average position error within 50 meters is about 0.2m, and the average position error between 50 - 90m only increases by 0.15m.
[0184] It can be seen that this application proposes a road perception system based on lidar and monocular cameras. The camera assigns depth information to each pixel through a high-precision map. By constructing a model and calibrating it with the lidar, the joint calibration efficiency is improved. The cross-level cascaded fusion method is adopted to fuse more features and increase the network depth, thereby improving the detection ability of the YOLOv5 algorithm for small targets and reducing the loss rate of fusion perception targets. The result-level fusion method is improved by combining the elliptical second-order correlation matching, reducing the algorithm complexity and improving the target fusion success rate. The overall implementation of the system promotes the development of vehicle-road collaboration and provides a reference for future similar scenarios.
[0185] Figure 9 It is a schematic structural diagram of a target detection device provided by an embodiment of this application. For the convenience of description, only parts related to the embodiment of this application are shown.
[0186] The target detection device may specifically include the following modules:
[0187] An acquisition module 901, which acquires the point cloud image collected by the lidar and the road segment image collected by the camera in the current traffic road segment;
[0188] A first generation module 902, which generates at least one first target bounding box according to the point cloud image;
[0189] The second generation module 903 generates at least one second target bounding box according to the road segment image;
[0190] The target detection module 904 determines the target image detected in the current traffic road segment according to the at least one first target bounding box and the at least one second target bounding box.
[0191] In specific implementation, the target detection device described in the embodiments of the present application includes, but is not limited to, other portable devices such as mobile phones, laptop computers or tablet computers having a touch-sensitive surface (for example, a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the above target detection device is not a portable communication device, but a computer having a touch-sensitive surface (for example, a touch screen display).
[0192] It should be understood that the target detection device may include one or more other physical user interface devices such as a physical keyboard, a mouse and / or a joystick. Various application programs that can be executed on the target detection device can use at least one common physical user interface device such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the terminal can be adjusted and / or changed between application programs and / or within the corresponding application programs. In this way, the common physical architecture of the terminal (for example, the touch-sensitive surface) can support various application programs with a user interface that is intuitive and transparent to the user.
[0193] Figure 10 This is a schematic structural diagram of a terminal device provided by an embodiment of the present application. The terminal device 1000 includes: at least one processor 1001 ( Figure 10 only one is shown in the figure), a processor, a memory 1002, and a computer program 1003 stored in the memory 1002 and executable on the at least one processor 1001. When the processor 1001 executes the computer program 1003, the steps in the above-mentioned target detection method embodiment are implemented.
[0194] The terminal device 1000 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor 1001 and a memory 1002. Those skilled in the art can understand that Figure 10 merely examples of the terminal device 1000, which do not constitute a limitation on the terminal device 1000. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0195] The so-called processor 1001 may be a Central Processing Unit (CPU), and the processor 1001 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0196] In some embodiments, the memory 1002 may be an internal storage unit of the terminal device 1000, such as the hard disk or memory of the terminal device 1000. In some other embodiments, the memory 1002 may also be an external storage device of the terminal device 1000, such as a plug-in hard disk equipped on the terminal device 1000, a SmartMedia Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 1002 may also include both the internal storage unit of the terminal device 1000 and the external storage device. The memory 1002 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 1002 may also be used to temporarily store data that has been output or is to be output.
[0197] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.
[0198] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not described in detail or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0199] An embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above various method embodiments can be implemented.
[0200] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above various method embodiments when executed.
[0201] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0202] In the embodiments provided by the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0203] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0204] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above units can be implemented in the form of hardware or in the form of software.
[0205] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0206] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A target detection method, characterized in that, The target detection method includes: Obtain the point cloud image collected by the lidar and the road segment image collected by the camera in the current traffic road segment; Generate at least one first target bounding box according to the point cloud image; Generate at least one second target bounding box according to the road segment image; Determine the target image detected in the current traffic road segment according to the at least one first target bounding box and the at least one second target bounding box.
2. The method according to claim 1, characterized in that Both the first target bounding box and the second target bounding box are two-dimensional images. The determining the target image detected in the current traffic road segment according to the at least one first target bounding box and the at least one second target bounding box includes: Associate all the first target bounding boxes and all the second target bounding boxes to determine the planar matching relationship between each first target bounding box and the second target bounding box; Determine the first target bounding box and the second target bounding box with the planar matching relationship as the same target image detected in the current traffic road segment.
3. The method according to claim 1, wherein Both the first target bounding box and the second target bounding box are three-dimensional images. The determining the target image detected in the current traffic road segment according to the at least one first target bounding box and the at least one second target bounding box includes: Associate all the first target bounding boxes and all the second target bounding boxes to determine the spatial matching relationship between each first target bounding box and the second target bounding box; Determine the first target bounding box and the second target bounding box with the spatial matching relationship as the same target image detected in the current traffic road segment.
4. The method according to claim 2, characterized in that, Generating at least one first target bounding box according to the point cloud image includes: Generate at least one three-dimensional target bounding box according to the point cloud image in combination with a three-dimensional target detection model; Perform dimensionality reduction processing on each three-dimensional target bounding box to obtain the first target bounding box.
5. The method according to claim 2, wherein The associating all the first target bounding boxes and all the second target bounding boxes to determine the planar matching relationship between each first target bounding box and the second target bounding box includes: Map all the first target bounding boxes and all the second target bounding boxes to the same two-dimensional space coordinate system; If the ratio of the intersection area to the union area of a first target bounding box and a second target bounding box is greater than a set threshold, determine the planar matching relationship between the first target bounding box and the second target bounding box.
6. The method according to claim 3, wherein The associating all the first target bounding boxes and all the second target bounding boxes to determine the spatial matching relationship between each first target bounding box and the second target bounding box includes: Map all the first target bounding boxes and all the second target bounding boxes to the same three-dimensional space coordinate system; If the ratio of the intersection volume to the union volume of a first target bounding box and a second target bounding box is greater than a first set threshold, determine that the first target bounding box and the second target bounding box form a spatial matching relationship.
7. The method according to claim 1, wherein Before generating at least one first target bounding box according to the point cloud image, the target detection method further includes: Based on a map image with an absolute accuracy less than a set distance, add depth information to the road segment image to generate a three-dimensional road segment image; Based on the three-dimensional road segment image, by simulating the camera and lidar, obtain the relative installation positions of the models of the camera and lidar in the simulation model; According to the relative rotation vector and translation vector of the camera and the lidar in the simulation model, map the point cloud image onto the three-dimensional road segment image; Adjust the relative rotation vector and translation vector to align the targets in the point cloud image and the three-dimensional road segment image; wherein, generating at least one second target bounding box according to the road segment image includes: generating at least one second target bounding box according to the road segment image and the map image.
8. The method according to claim 7, wherein The generating a three-dimensional road segment image by adding depth information to the road segment image based on a map image with an absolute accuracy less than a set distance includes: Perform Mercator projection mapping on the map image to obtain the three-dimensional information of each map unit in the map image; Select multiple groups of valid points from each map unit, and transform each map unit into a prior perception three-dimensional element through perspective transformation; Generate a three-dimensional road segment image including depth information according to the prior perception three-dimensional element and the road segment image.
9. The method according to claim 8, characterized in that, The target detection method further includes: Establish an initial model based on the YOLOV5 algorithm; the initial model includes multiple calculation layers, each calculation layer includes multiple nodes, and the input size of each node is a first scale and the output size is a second scale; Add cross-layer weighted connections between all nodes with the same first scale and second scale to obtain an updated model; wherein the deep learning model is obtained by training the updated model to convergence using the historical two-dimensional environment integration image with all second target bounding boxes annotated.
10. The method according to claim 9, characterized in that, After obtaining the updated model, the target detection method further includes: Perform feature fusion on the outputs of each node of the top calculation layer of the updated model and the calculation layer adjacent to the top layer according to channels; For non-top and non-bottom calculation layers, perform feature fusion on the outputs of two nodes on adjacent paths according to channels, and perform feature map superposition operation on the outputs of two nodes on non-adjacent paths to obtain a fusion output.
11. The method according to claim 10, wherein The performing feature map superposition operation on the outputs of two nodes on non-adjacent paths includes: According to an initial weight coefficient of each feature map, combine the feature maps of the outputs of two nodes on adjacent paths for feature map fusion; Update the initial weight coefficient within a preset weight range, and re-combine the feature maps of the outputs of two nodes on adjacent paths for feature map fusion until no error occurs when the training speed is higher than a set speed threshold.
12. The method according to claim 9, wherein The target detection method further includes: Set a monitoring layer for sequentially performing target monitoring on the feature maps output by all calculation layers.
13. The method according to claim 5, wherein If the closest first target bounding box and second target bounding box do not intersect, when associating all first target bounding boxes and all second target bounding boxes to determine the planar matching relationship between each first target bounding box and the second target bounding box, it further includes: Perform secondary association on the non-intersecting first target bounding box and second target bounding box using category information and ellipse-related thresholds until each first target bounding box has the spatial matching relationship.
14. A target detection device, characterized in that, The target detection device includes: An acquisition module that acquires a point cloud image collected by a lidar and a road segment image collected by a camera in the current traffic road segment; A first generation module that generates at least one first target bounding box according to the point cloud image; A second generation module that generates at least one second target bounding box according to the road segment image; A target detection module that determines a target image detected in the current traffic road segment according to the at least one first target bounding box and the at least one second target bounding box.
15. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method described in any one of claims 1 to 13 is implemented.
16. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method described in any one of claims 1 to 13 is implemented.