Target detection method, tracking method, device, vision sensor and medium

By establishing a mapping model between images and radar point clouds using visual sensors, the problem of missing depth information during radar malfunctions is solved, enabling three-dimensional detection and tracking of targets and ensuring traffic safety.

CN114170499BActive Publication Date: 2026-02-27VANJEE TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010837555.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-19
Publication Date
2026-02-27
Estimated Expiration
2040-08-19

AI Technical Summary

Technical Problem

When radar malfunctions, existing technologies struggle to provide depth information to traffic information centers, making it impossible to detect and track targets in road scenarios.

Method used

A mapping model between image pixels and radar sensor point cloud depth information is established using a visual sensor. Image data is processed using a target perception algorithm and a preset mapping model to generate two-dimensional and three-dimensional target detection results, including the target's pixel position, category, size, and three-dimensional position.

Benefits of technology

In the event of a radar malfunction, the vision sensor can automatically provide the target's depth information and 3D detection results, ensuring the traffic information center can detect and track the target, and guaranteeing safe driving in the target scenario.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170499B_ABST
    Figure CN114170499B_ABST
Patent Text Reader

Abstract

The application relates to a target detection method, a tracking method, a device, a visual sensor and a medium. The method comprises the following steps: acquiring an image of a target scene; processing the image by using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result comprises pixel positions and categories of targets; processing the two-dimensional target detection result by using a preset mapping model to obtain a three-dimensional target detection result; the mapping model comprises a mapping relationship between the positions of pixel points of the image and depth information of a point cloud of a radar sensor; and the three-dimensional target detection result comprises sizes and / or three-dimensional positions of the targets. The method can provide depth information, and the targets can be detected and tracked according to the depth information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a target detection method, tracking method, apparatus, visual sensor, and medium. Background Technology

[0002] Intelligent Transportation Systems (ITS) refer to the systems where traffic participants use sensors and transmission equipment installed on roads, vehicles, and other locations to provide real-time traffic information from various locations to a traffic information center. After receiving and processing this information, the traffic information center can provide traffic participants with other travel-related information, such as road traffic information. Based on this information, travelers can determine their travel mode and choose routes, thereby ensuring traffic safety.

[0003] In related technologies, when providing real-time traffic information from various locations to the traffic information center, image data of the road scene is usually collected by a monocular camera, and depth information such as distance in the road scene is collected by radar. Then, the camera and radar transmit the information they have collected to the traffic information center for processing.

[0004] However, when the radar malfunctions, the aforementioned technology cannot provide depth information to the traffic information center, which would prevent the traffic information center from detecting and tracking targets in the road scene. Summary of the Invention

[0005] Therefore, it is necessary to provide a target detection method, tracking method, device, visual sensor, and medium that can still provide depth information when the radar malfunctions, so as to detect and track targets in the scene.

[0006] An object detection method applied to a vision sensor, the method comprising:

[0007] Acquire images of the target scene;

[0008] The above image is processed using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0009] The two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results. The mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor. The three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0010] In one embodiment, the image is processed using a target perception algorithm to obtain a two-dimensional target detection result, including:

[0011] The above-mentioned target perception algorithm is used to process the above image to obtain the target bounding box of the target in the above image and the category of the target;

[0012] Based on the positions of each pixel in the target box, the pixel positions of the target are obtained.

[0013] The pixel positions and categories of the aforementioned targets are determined as the two-dimensional target detection results.

[0014] In one embodiment, obtaining the pixel position of the target based on the position of each pixel in the target box includes:

[0015] Based on the positions of each pixel in the target bounding box, the positions of the positioning pixels are determined; the positioning pixels are used to characterize the position of the target in the image.

[0016] The positions of the aforementioned location pixels are taken as the pixel positions of the aforementioned target.

[0017] In one embodiment, the location of the aforementioned positioning pixel is the center point of the bottom edge of the aforementioned target box.

[0018] In one embodiment, the above-mentioned processing of the two-dimensional target detection result using a preset mapping model to obtain the three-dimensional target detection result includes:

[0019] The pixel positions of the target are processed using the above mapping model to obtain the three-dimensional position of the target.

[0020] In one embodiment, obtaining the pixel position of the target based on the position of each pixel in the target box includes:

[0021] Based on the position of each pixel in the target box, the position of the bounding box corresponding to the target box is determined.

[0022] The position of the bounding box is taken as the pixel position of the target.

[0023] In one embodiment, the above-mentioned processing of the two-dimensional target detection result using a preset mapping model to obtain the three-dimensional target detection result includes:

[0024] The position of the bottom edge of the bounding box is obtained from the position of the bounding box, and the depth information corresponding to the position of the bottom edge is obtained by processing it using the mapping model.

[0025] Based on the depth information corresponding to the position of the bottom edge, construct the length and width of the target.

[0026] Obtain the height of the bounding box from its position, and construct the actual height of the target based on the height of the bounding box.

[0027] Based on the length, width, and actual height of the target, the dimensions of the target are obtained.

[0028] A target tracking method applied to a vision sensor, the method comprising:

[0029] Acquire video data of the target scene; this video data includes multiple frames of images.

[0030] The video data is processed using a target perception algorithm to obtain two-dimensional target detection results; the two-dimensional target detection results include the pixel position and category of the target;

[0031] The above two-dimensional target detection results are processed using an image tracking algorithm to obtain the two-dimensional target motion trajectory;

[0032] The above two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor; the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0033] Based on the above three-dimensional target detection results and the above two-dimensional target motion trajectory, the three-dimensional tracking results of the above target are obtained.

[0034] In one embodiment, the three-dimensional position of the target is the three-dimensional coordinates of the target in the world coordinate system, and the three-dimensional tracking result of the target includes the latitude and longitude information of the target; obtaining the three-dimensional tracking result of the target based on the three-dimensional target detection result and the two-dimensional target motion trajectory includes:

[0035] Based on the three-dimensional position of the target and the two-dimensional target motion trajectory, determine the three-dimensional coordinates corresponding to the pixel positions of the target in each frame image;

[0036] Obtain the latitude and longitude information of the origin of the radar coordinate system corresponding to the radar sensor, as well as the angle between any axis of the radar coordinate system and the candidate direction; the candidate direction is related to the latitude and longitude information of the origin.

[0037] Using the latitude and longitude information of the origin and the included angle, the three-dimensional coordinates of the target corresponding to each frame of the image are transformed to obtain the latitude and longitude information of the target corresponding to each frame of the image.

[0038] In one embodiment, the three-dimensional position of the target is the three-dimensional coordinates of the target in the world coordinate system, and the three-dimensional tracking result of the target includes the target's velocity and acceleration; obtaining the three-dimensional tracking result of the target based on the three-dimensional target detection result and the two-dimensional target motion trajectory includes:

[0039] Based on the three-dimensional position of the target and the two-dimensional target motion trajectory, determine the three-dimensional coordinates corresponding to the pixel positions of the target in each frame image;

[0040] Obtain the three-dimensional coordinates of the origin in the world coordinate system, and connect the three-dimensional coordinates of the origin with the three-dimensional coordinates of the target corresponding to any frame of the image to obtain the position vector of the target corresponding to the frame of the image.

[0041] Obtain the acquisition time corresponding to each of two adjacent frames, and obtain the time difference between the two adjacent frames based on their acquisition times.

[0042] Mathematical operations are performed on the position vectors of the aforementioned targets in two adjacent frames to obtain the position vector variables of the aforementioned targets;

[0043] By performing mathematical operations on the aforementioned position vector variables and time differences, the velocity and acceleration of the aforementioned target are obtained.

[0044] In one embodiment, the three-dimensional tracking result of the target includes the target's heading angle and angular velocity; the method further includes:

[0045] Obtain any axis in the world coordinate system described above, and use that axis as the reference axis;

[0046] Calculate the angle between the above position vector variable and the above reference axis, and determine the obtained angle as the heading angle of the above target;

[0047] By performing mathematical operations on the aforementioned heading angle and time difference, the angular velocity of the target is obtained.

[0048] A target detection device, applied to a vision sensor, the target detection device comprising:

[0049] The first acquisition module is used to acquire images of the target scene;

[0050] The first detection module is used to process the image using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0051] The first mapping processing module is used to process the above two-dimensional target detection results using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor, and the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0052] A target tracking device, applied to a vision sensor, the target tracking device comprising:

[0053] The second acquisition module is used to acquire video data of the target scene; the video data includes multiple frames of images.

[0054] The second detection module is used to process the video data using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0055] The two-dimensional tracking module is used to process the above two-dimensional target detection results using image tracking algorithms to obtain the two-dimensional target motion trajectory;

[0056] The second mapping processing module is used to process the above two-dimensional target detection results using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor; the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0057] The 3D tracking module is used to obtain the 3D tracking result of the target based on the 3D target detection result and the 2D target motion trajectory.

[0058] A vision sensor includes a camera, a memory, and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0059] Acquire images of the target scene;

[0060] The above image is processed using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0061] The two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results. The mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor. The three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0062] A vision sensor includes a camera, a memory, and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0063] Acquire video data of the target scene; this video data includes multiple frames of images.

[0064] The video data is processed using a target perception algorithm to obtain two-dimensional target detection results; the two-dimensional target detection results include the pixel position and category of the target;

[0065] The above two-dimensional target detection results are processed using an image tracking algorithm to obtain the two-dimensional target motion trajectory;

[0066] The above two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor; the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0067] Based on the above three-dimensional target detection results and the above two-dimensional target motion trajectory, the three-dimensional tracking results of the above target are obtained.

[0068] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0069] Acquire images of the target scene;

[0070] The above image is processed using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0071] The two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results. The mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor. The three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0072] A computer-readable storage medium having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0073] Acquire video data of the target scene; this video data includes multiple frames of images.

[0074] The video data is processed using a target perception algorithm to obtain two-dimensional target detection results; the two-dimensional target detection results include the pixel position and category of the target;

[0075] The above two-dimensional target detection results are processed using an image tracking algorithm to obtain the two-dimensional target motion trajectory;

[0076] The above two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor; the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0077] Based on the above three-dimensional target detection results and the above two-dimensional target motion trajectory, the three-dimensional tracking results of the above target are obtained.

[0078] The aforementioned target detection method, tracking method, device, visual sensor, and medium can acquire images of the target scene through a visual sensor, process the image using a target perception algorithm to obtain a two-dimensional target detection result, and then process the two-dimensional target detection result using a preset mapping model to obtain a three-dimensional target detection result. The two-dimensional target detection result includes the pixel position and category of the target, the mapping model includes the mapping relationship between the pixel position of the image and the depth information of the point cloud from the radar sensor, and the three-dimensional target detection result includes the size and / or three-dimensional position of the target. In this method, since the visual sensor can establish a mapping model including the pixel position of the image and the depth information of the point cloud, even when the radar cannot provide depth information, the visual sensor itself can obtain the depth information corresponding to the pixel position of the target through the pre-established mapping model, thus obtaining the three-dimensional target detection result. This allows the intelligent transportation information center to provide the target's depth information and three-dimensional target detection result, enabling the intelligent transportation information center to detect and track the target based on the depth information and the three-dimensional target detection result, thereby ensuring the safe driving of the target in the target scene. Attached Figure Description

[0079] Figure 1 This is a diagram showing the internal structure of a vision sensor in one embodiment;

[0080] Figure 2 This is a flowchart illustrating a target detection method in one embodiment;

[0081] Figure 3 This is a flowchart illustrating the target detection method in another embodiment;

[0082] Figure 3a This is a flowchart illustrating the process of obtaining depth information using an interpolation method in another embodiment;

[0083] Figure 3b This is an example diagram illustrating the use of interpolation methods to obtain depth information in another embodiment;

[0084] Figure 4 This is a flowchart illustrating the target tracking method in another embodiment;

[0085] Figure 5This is a structural block diagram of a target detection device in one embodiment;

[0086] Figure 6 This is a structural block diagram of a target tracking device in one embodiment. Detailed Implementation

[0087] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0088] The target detection method and target tracking method provided in the embodiments of this application can be applied to, for example... Figure 1 The visual sensor shown can be a monocular camera, such as a bullet camera, hemispherical camera, or spherical camera, etc., and possesses computing capabilities. The visual sensor may include a processor, memory, communication interface, display screen, and input device connected via a system bus. It may also include a camera, primarily used to acquire image data of targets in the scene, which can be connected to the processor to transmit the acquired image data to the processor for processing. The processor of the visual sensor provides computing and control capabilities. The memory of the visual sensor includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The communication interface of the visual sensor is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a target detection method and a target tracking method.

[0089] Those skilled in the art will understand that Figure 1 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the visual sensor applied thereto. A specific visual sensor may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0090] It should be noted that the execution subject in the embodiments of this application can be a vision sensor, or a target detection device and target tracking device inside the vision sensor. The following description will use a vision sensor as the execution subject.

[0091] In one embodiment, a target detection method is provided. This embodiment relates to the specific process of how to detect a target and obtain the three-dimensional target detection result. For example... Figure 2As shown, the method may include the following steps:

[0092] S202, acquire an image of the target scene.

[0093] The target scene can be a road scene, or any other scene. The road scene can be an outdoor road scene, an indoor playground road scene, etc. The image of the target scene may or may not include targets; this embodiment mainly describes the case where the target scene includes targets. Targets in the target scene can be vehicles, pedestrians, etc., in the road scene, and the number of targets can be one or more.

[0094] In addition, the vision sensor can be a monocular camera, which can usually be set on a side pole (such as a vertical pole or horizontal pole) in the road. In this way, the vision sensor can capture images of vehicles or pedestrians in the road scene.

[0095] Of course, a visual sensor can also be used to collect video data of the road scene. The images in the video data are generally images collected at different times in succession. For example, one image is collected every second from 1 to 10 seconds to obtain 10 images. Then, one frame is selected from the video data as the image of the target scene here.

[0096] S204, The above image is processed using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target.

[0097] The target perception algorithm here can be the YOLOv3 target detection algorithm, Kalman tracking algorithm, sort tracking algorithm, etc. Of course, semantic segmentation algorithms can also be used to process the images of the target scene mentioned above.

[0098] Specifically, target detection can be performed on the image of the target scene using a target perception algorithm to obtain a two-dimensional target detection result in the image. This two-dimensional target detection result includes the pixel position of the target, the category of the target, the identifier, the confidence level of the target, etc.

[0099] S206, the above two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor, and the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0100] In this step, before processing the two-dimensional target detection results using the mapping model, optionally, it is possible to first check whether the radar sensor is malfunctioning. If the radar is malfunctioning, this step can be executed, that is, the step of processing the two-dimensional target detection results using the preset mapping model to obtain the three-dimensional target detection results. Radar sensor malfunction here refers to situations such as radar sensor damage, partial data corruption in the data collected by the radar sensor, data distortion, data loss, or data unavailability caused by objective reasons such as weather conditions. The radar sensor here can be lidar, millimeter-wave radar, etc. Lidar can include 8-line, 16-line, 24-line, and 128-line lidar, and millimeter-wave radar can be 24G, 77G, etc.

[0101] Here, the mapping model can be a fitting mapping model or a deep learning model. So, before obtaining the corresponding depth information using the position of pixels in the image, we can first obtain a fitting mapping model or a deep learning model between the position of pixels in the image and the depth information of the point cloud.

[0102] To obtain the fitting mapping model and the deep learning model, historical image data and historical point cloud data at various times within the same scene can be collected. The historical image data includes the pixel positions of historical objects, which can be measured using a visual sensor. The historical point cloud data includes the depth information of the historical object at that pixel, representing the distance between the historical object and the acquisition device, as well as its x, y, z coordinates and related angles in the physical coordinate system. The acquisition device here refers to a radar sensor, and the historical object can be a vehicle, pedestrian, etc., within the scene. Then, the pixel positions in the historical image data at the same time and the depth information in the point cloud data are correlated to obtain the mapping relationship between the two, thus obtaining the fitting mapping model. Similarly, the pixel positions in the historical image data at the same time can be used as input to the initial deep learning model, and the depth information from the historical point cloud data at that time can be used as labels to train the initial deep learning model, resulting in the deep learning model.

[0103] It should be noted that, taking the roadside as an example, visual sensors and radar sensors are usually installed at the same location on the roadside. Therefore, the depth information in the historical point cloud data mentioned above can represent the distance between the historical object and the radar sensor, which is essentially the distance between the object and the visual sensor. It can also describe the x, y, z coordinates and related angles of the historical object in the physical coordinate system, which can also be called the three-dimensional position of the historical object.

[0104] After establishing the fitting mapping model or deep learning model, the pixel positions of the target in the image can be input into the fitting mapping model or deep learning model. The fitting mapping model or deep learning model can then be used to process the two-dimensional target detection results to obtain the depth information corresponding to the pixel positions of the target, which is to say, the three-dimensional position of the target can be obtained. Then, the size of the target can be obtained through the three-dimensional position of the target.

[0105] The visual sensor in this embodiment can obtain both the pixel position of the target in the image and the depth information corresponding to the pixel position. Therefore, this embodiment can achieve the mapping of pixel position and depth information with a single visual sensor. In practical applications, this eliminates the need to install numerous sensors on the roadside or vehicle; a single visual sensor can perform the functions of both a camera and radar, thereby reducing the overall system cost. Furthermore, maintenance is relatively simple, requiring only the visual sensor to be maintained, thus reducing maintenance costs. Additionally, it saves system space and reduces the number of sensors required in the system.

[0106] In the aforementioned target detection method, an image of the target scene can be acquired using a visual sensor. This image is then processed using a target perception algorithm to obtain a two-dimensional target detection result. A pre-defined mapping model is then used to further process the two-dimensional target detection result to obtain a three-dimensional target detection result. The two-dimensional target detection result includes the pixel position and category of the target. The mapping model includes the mapping relationship between the pixel position of the image and the depth information of the point cloud from the radar sensor. The three-dimensional target detection result includes the size and / or three-dimensional position of the target. In this method, because the visual sensor can establish a mapping model that includes the pixel position of the image and the depth information of the point cloud, even when the radar cannot provide depth information, the visual sensor itself can obtain the depth information corresponding to the pixel position of the target through the pre-established mapping model, thus obtaining the three-dimensional target detection result. This allows the intelligent transportation information center to provide the target's depth information and three-dimensional target detection result, enabling the intelligent transportation information center to detect and track targets based on the depth information and the three-dimensional target detection result, thereby ensuring safe driving of targets in the target scene.

[0107] In another embodiment, a different object detection method is provided. This embodiment relates to the specific process by which a visual sensor processes an image using an object perception algorithm to obtain a two-dimensional object detection result. Based on the above embodiments, as follows... Figure 3 As shown, the above S204 may include the following steps:

[0108] S302, The image is processed using a target perception algorithm to obtain the bounding box of the target in the image and the category of the target.

[0109] Specifically, target detection can be performed on the images of the aforementioned target scene using target perception algorithms to obtain the bounding boxes of each target in the image, as well as the category, confidence level, and identifier of each target.

[0110] The target category can be the category to which the target belongs, such as whether the target is a man or a woman, or what kind of animal the target is, etc.; the target confidence score represents the probability that the target in the target box belongs to each category of the target.

[0111] S304. Based on the positions of each pixel in the target box, obtain the pixel position of the target.

[0112] In this step, the target's pixel position can be the position of a single pixel or the position of multiple pixels. The following will explain these two cases separately.

[0113] In one possible implementation, taking the pixel position of the target as the position of a single pixel, optionally, the position of the positioning pixel can be determined based on the positions of each pixel in the target bounding box; the positioning pixel is used to characterize the position of the target in the image; the position of the positioning pixel is taken as the pixel position of the target. Optionally, the position of the positioning pixel is the center point of the bottom edge of the target bounding box.

[0114] Here, after obtaining the target bounding box, we can also obtain the position of each pixel within the bounding box on the image. Having obtained the positions of each pixel within the bounding box, we can then obtain the boundaries of the target bounding box. We can then select the bottom edge of the target bounding box and choose the center pixel from the pixels on the bottom edge, denoted as the bottom edge center point. Typically, the target bounding box is a rectangle, with the bottom edge being the boundary closest to the ground. Therefore, selecting the center point of the bottom edge of the target bounding box to represent the target ensures that the obtained target is relatively close to the ground, which is consistent with reality.

[0115] In addition, in the actual process of target detection and tracking, using all the pixels in the target box to represent the target to be tracked is computationally intensive. Therefore, using localized pixels to represent the target can save computation and improve the efficiency of target detection and tracking.

[0116] In another possible implementation, where the target pixel position is the pixel position of multiple pixels, optionally, the position of the bounding box corresponding to the target box can be determined from the position of each pixel in the target box; and the position of the bounding box is used as the pixel position of the target.

[0117] In other words, after obtaining the position of each pixel in the target box, we can also obtain the boundaries of the target box, and then obtain the position of the pixel on each boundary. Usually, the target box is a rectangle, so there will be four boundaries. Here, the position of the pixel on these four boundaries can be used as the pixel position of the target.

[0118] S306, The pixel position and category of the target are determined as the two-dimensional target detection result.

[0119] In this step, similar to S304 above, the pixel position of the target can be the pixel position of a single pixel or the pixel position of multiple pixels. The following will also explain these two cases.

[0120] In one possible implementation, taking the pixel position of the target as the pixel position of a single pixel, after obtaining the pixel position of the target, i.e. the position of the center point of the bottom edge of the target box, the pixel position of the target can optionally be processed using the above mapping model to obtain the three-dimensional position of the target.

[0121] In other words, after determining the pixel position of the bottom center point of the target bounding box, the pixel position of the bottom center point can be processed by the mapping model to obtain the depth information corresponding to the pixel position of the bottom center point, that is, the depth information of the target can be obtained. The depth information of the target can also be called the three-dimensional position of the target.

[0122] In another possible implementation, assuming the target's pixel position is the position of multiple pixels, after obtaining the target's pixel position, i.e., the positions of the pixels on the four boundaries of the target bounding box, the target's size can optionally be obtained using the following steps A1-A4:

[0123] A1. Obtain the position of the bottom edge of the bounding box from the position of the bounding box, and process it using the mapping model to obtain the depth information corresponding to the position of the bottom edge.

[0124] A2. Based on the depth information corresponding to the position of the bottom edge, construct the length and width of the target.

[0125] A3. Obtain the height of the bounding box from its position, and construct the actual height of the target based on the height of the bounding box.

[0126] A4. Based on the length, width, and actual height of the target, the dimensions of the target are obtained.

[0127] Specifically, after obtaining the depth information of each point on the bottom edge of the bounding box (where each point has three-dimensional coordinates of (x, y, z), the length and width of the target can be constructed using these coordinates. The height of the target can be calculated using the corresponding pixel heights on different fitted curves, for example, h1 = ∑α k h k2 h1 is the actual height of the target, h2 is the target pixel height (i.e., the height of the bounding box), k is the index of the point, and α is the height scaling factor, which can be deduced from the actual situation and is a known value here. Of course, other methods can also be used to obtain the actual height of the target, such as derivation using geometric relationships, etc. In short, this method can obtain the length, width, and height of the target, and the combination of these three dimensions gives the size of the target.

[0128] It should be noted that the deep learning model described above can map any pixel and obtain the depth information corresponding to any pixel. However, since the fitting mapping model cannot cover the depth information at every pixel's location—meaning some pixels have corresponding depth information while others do not—further methods are needed to obtain the depth information of a given pixel. The following method can be used to obtain this depth information:

[0129] Based on the pixel's position, determine if depth information corresponding to that pixel's position exists in the mapping model. If it does, obtain the depth information corresponding to that pixel's position. If not, determine multiple fitting curves formed by the depth information in the mapping model. Based on the two-dimensional coordinate axes of the image, determine an extension line with one of the pixel's position coordinates as the center point and along the direction of the other pixel's position coordinates. Obtain the intersection points of the extension line with the multiple fitting curves. Select the two intersection points closest to the pixel from the intersection points with the multiple fitting curves as multiple target pixels. Based on the positions of each target pixel and the aforementioned pixels, obtain the distance between each target pixel and the aforementioned pixels. Based on the distance between each target pixel and the aforementioned pixels, perform proportional interpolation on the depth information corresponding to the position of each target pixel to obtain the depth information corresponding to the aforementioned pixel's position.

[0130] For example, for the process of obtaining depth information for pixels that do not have depth information in the mapping model, please refer to [link to documentation]. Figure 3a and Figure 3bAs shown, the position of the pixel that does not have depth information is (x_0, y_0), which is... Figure 3b Taking point 1 in the above example, where the fitted curve obtained is a fitted circle, we can obtain two fitted circles c_1 and c_2 that closely surround this pixel. By drawing a straight line x = x_0, we can find the intersection of this line with the two fitted circles, thus obtaining the two target pixel points. Assuming the coordinates of the two target pixel points are (x_0, y_1) and (x_0, y_2), respectively... Figure 3b Points 2 and 3 in the graph are assumed to correspond to depth information (X_1,Y_1,Z_1) and (X_2,Y_2,Z_2) on the fitted curve, respectively. Then, by proportionally interpolating (X_1,Y_1,Z_1) and (X_2,Y_2,Z_2) from (x_0,y_0) to (x_0,y_1) and (x_0,y_2), the 3D coordinates of (x_0,y_0) can be obtained. Alternatively, weights can be assigned to (X_1,Y_1,Z_1) and (X_2,Y_2,Z_2), and then interpolated according to these weights to obtain the 3D coordinates of (x_0,y_0).

[0131] In this embodiment, the image is processed using a target perception algorithm to obtain the target bounding box and the target category. The pixel position of the target is then determined based on the position of each pixel within the target bounding box. The pixel position and target category are then used to define the two-dimensional target detection result. This method provides a relatively simple and accurate way to obtain the two-dimensional target detection result, which in turn makes the subsequent three-dimensional target detection result obtained based on this two-dimensional target detection result more accurate.

[0132] In one embodiment, a target tracking method is provided. This embodiment relates to the specific process of detecting and tracking a target to obtain the three-dimensional target detection results and the three-dimensional tracking results. For example... Figure 4 As shown, the method may include the following steps:

[0133] S402, acquire video data of the target scene; the video data includes multiple frames of images.

[0134] The target scene can be a road scene, or any other scene. The road scene can be an outdoor road scene, an indoor playground road scene, etc. The images in the target scene may or may not include targets; this embodiment mainly describes the case where the target scene includes targets. Targets in the target scene can be vehicles, pedestrians, etc., in the road scene, and the number of targets can be one or more. The frames in the video data are generally images captured continuously at different times, for example, one image per second from 1 to 10 seconds, resulting in 10 frames.

[0135] Specifically, the vision sensor can be a monocular camera, which can usually be set on a side pole (such as a vertical pole or horizontal pole) in the road. In this way, the vision sensor can continuously collect images of vehicles or pedestrians in the road scene, thus obtaining video data of the target scene.

[0136] S404, The above video data is processed using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target.

[0137] In this step, referring to the explanation of S204 above, each frame of the video data can be processed to obtain the pixel position and category of the target in each frame.

[0138] S406, The above two-dimensional target detection results are processed using an image tracking algorithm to obtain the two-dimensional target motion trajectory.

[0139] The image tracking algorithm here can be one of the target perception algorithms mentioned above, such as the Kalman tracking algorithm or the sort tracking algorithm.

[0140] Specifically, after obtaining the pixel position and category of the target in each frame of the image, the image tracking algorithm and the category of the target in each frame of the image can be used to find the target belonging to the same category in each frame of the image. Then, the same target is tracked, and the pixel position of the same target in each frame of the image is curve fitted (for example, by connecting the pixel positions), and a trajectory line can be obtained, which is called the two-dimensional target motion trajectory.

[0141] S408, the above two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor, and the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0142] For an explanation of this step, please refer to the explanation of S206 above. This embodiment will not repeat it here.

[0143] S410, Based on the above three-dimensional target detection results and the above two-dimensional target motion trajectory, the three-dimensional tracking results of the above target are obtained.

[0144] In this step, after obtaining the target's two-dimensional trajectory and its size and three-dimensional position in the actual scene, the target can be tracked to obtain its latitude and longitude information, velocity, acceleration, angular velocity, etc., which are recorded as the target's three-dimensional tracking result.

[0145] In the aforementioned target tracking algorithm, video data of the target scene can be acquired through a visual sensor. This video data is processed to obtain a two-dimensional target detection result and a two-dimensional target motion trajectory. A pre-defined mapping model is then used to process the two-dimensional target detection result to obtain a three-dimensional target detection result. Finally, based on the three-dimensional target detection result and the two-dimensional target motion trajectory, a three-dimensional target tracking result is obtained. The mapping model includes the mapping relationship between the pixel positions of the image and the depth information of the point cloud from the radar sensor. The three-dimensional target detection result includes the target's size and / or three-dimensional position. In this method, since the visual sensor can establish a mapping model that includes the pixel positions of the image and the depth information of the point cloud, even when the radar cannot provide depth information, the visual sensor itself can obtain the depth information corresponding to the target's pixel positions through the pre-established mapping model. This allows it to obtain the three-dimensional target detection result, which can then be provided to the intelligent transportation information center. The intelligent transportation information center can then detect and track the target based on the depth information and the three-dimensional target detection result, thereby ensuring the safe movement of the target in the scene.

[0146] In another embodiment, a different target tracking method is provided. Based on the above embodiments, after obtaining the three-dimensional position or size of the target in each frame image, the three-dimensional position refers to the target's three-dimensional coordinates in the world coordinate system. Therefore, the target's latitude and longitude information, velocity, acceleration, angular velocity, heading angle, and other three-dimensional tracking results can be obtained through the three-dimensional target detection results and the two-dimensional target motion trajectory. The following are possible implementation methods for obtaining these parameters:

[0147] One implementation method: Obtain the latitude and longitude information of the target.

[0148] Optionally, based on the three-dimensional position of the target and the two-dimensional target motion trajectory, the three-dimensional coordinates corresponding to the pixel positions of the target in each frame of the image can be determined; the latitude and longitude information of the origin of the radar coordinate system corresponding to the radar sensor and the angle between any axis of the radar coordinate system and the candidate direction can be obtained; the candidate direction is related to the latitude and longitude information of the origin; using the latitude and longitude information of the origin and the angle, the three-dimensional coordinates of the target corresponding to each frame of the image can be transformed to obtain the latitude and longitude information of the target corresponding to each frame of the image.

[0149] In other words, after obtaining the pixel position of the target, the target tracking here is mainly described using the positioning pixel point on the target box (i.e. the center point of the bottom edge) as the target. After obtaining the three-dimensional position of the target and the two-dimensional motion trajectory of the target, the three-dimensional coordinates of the positioning pixel point on each frame of the image can be obtained from them.

[0150] Furthermore, after the radar or vision sensor is fixed in place on the roadside, the radar coordinate system or vision sensor coordinate system can be obtained by using the radar or vision sensor as the origin. If the radar sensor malfunctions, the radar coordinate system can also be obtained by transforming the vision sensor coordinate system. Once the radar coordinate system is obtained, the coordinates of the radar coordinate system origin can be determined. Using relevant latitude and longitude calculation methods and the coordinates of the radar coordinate system origin, the latitude and longitude information of the origin in the radar coordinate system can be obtained.

[0151] After obtaining the latitude and longitude information of the origin in the radar coordinate system, the angle between any axis (x-axis or y-axis) of the radar coordinate system and the candidate direction (the candidate direction can be due north, due south, due east, due west, northwest, etc.) can be calculated. After obtaining the angle, the three-dimensional coordinates of each positioning pixel can be rotated or translated using the angle. This way, the latitude and longitude information of each positioning pixel can be obtained, that is, the latitude and longitude information of the target at each time. This allows for more accurate target tracking.

[0152] Another implementation method: obtain the target's velocity and acceleration.

[0153] Optionally, based on the target's three-dimensional position and the target's two-dimensional trajectory, the three-dimensional coordinates corresponding to the pixel positions of the target in each frame image can be determined; the three-dimensional coordinates of the origin in the world coordinate system can be obtained, and the three-dimensional coordinates of the origin and the target's three-dimensional coordinates corresponding to any frame image can be connected to obtain the target's position vector corresponding to the frame image; the acquisition time corresponding to each of two adjacent frames can be obtained, and the time difference between the two adjacent frames can be obtained based on the acquisition time of the two adjacent frames; mathematical operations can be performed on the target's position vectors in the two adjacent frames to obtain the target's position vector variables; mathematical operations can be performed on the position vector variables and the time difference to obtain the target's velocity and acceleration.

[0154] Optionally, both velocity and acceleration here are vectors. The world coordinate system can be the radar coordinate system. Connecting the origin of the world coordinate system to the target's three-dimensional coordinates at any given moment yields the target's position vector from the origin to its current position, denoted as s'. Then, based on the target's identifier, correlation tracking is performed, allowing us to determine the target's position vector at various moments. By subtracting the position vector at each moment from the previous moment's position vector, we obtain the target's position vector variable, denoted as Δs'. We can also obtain the two moments for calculating the position vector variable Δs'. Subtracting these two moments gives us the time taken for the time-position vector to travel from one moment to another, denoted as Δt. Quoting Δs' and Δt yields the target's velocity vector v', i.e., v' = Δs' / Δt. Similarly, multiple velocities can be calculated and subtracted to obtain the velocity vector variable Δv'. Finally, using the formula a' = Δv' / Δt, we can obtain the target's acceleration vector a'.

[0155] Another implementation method: Obtain the target's heading angle and angular velocity.

[0156] Optionally, any axis in the world coordinate system can be obtained and used as the reference axis; the angle between the above position vector variable and the above reference axis can be calculated and the obtained angle can be determined as the heading angle of the above target; mathematical operations can be performed on the above heading angle and the above time difference to obtain the angular velocity of the above target.

[0157] In this process, any coordinate axis of the world coordinate system can be used as the reference axis. After obtaining the target's position vector variable Δs', the position vector variable is generally a straight line. Therefore, the angle between this straight line and the reference axis can be calculated. This calculated angle is the heading angle of the target relative to the position vector variable, that is, the heading angle of the target at that moment, denoted as w. Similarly, the heading angle of the target at each moment can be calculated. Then, the difference between the heading angles at any two moments can be obtained to get the change in heading angle, denoted as Δw. At the same time, the time difference between these two moments can also be obtained, denoted as Δt, as above. Then, the quotient of Δw and Δt can be obtained to get the target's angular velocity γ, that is, γ = Δw / Δt.

[0158] The target tracking method provided in this embodiment can obtain the three-dimensional tracking result of the target based on the three-dimensional target detection result and the two-dimensional target motion trajectory, including the target's latitude and longitude information, velocity, acceleration, angular velocity, heading angle, etc. Through this method, the relevant operational parameters of the target can be obtained, thereby accurately grasping the target's real-time direction of travel and velocity, and thus achieving accurate target tracking.

[0159] It should be understood that, although Figures 2-4 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figures 2-4 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0160] In one embodiment, such as Figure 5 As shown, a target detection device is provided, applied to a vision sensor. The device may include: a first acquisition module 10, a first detection module 11, and a first mapping processing module 12, wherein:

[0161] The first acquisition module 10 is used to acquire images of the target scene;

[0162] The first detection module 11 is used to process the image using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0163] The first mapping processing module 12 is used to process the above two-dimensional target detection results using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixel points of the image and the depth information of the point cloud of the radar sensor, and the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0164] For specific limitations on target detection devices, please refer to the limitations on target tracking methods mentioned above, which will not be repeated here.

[0165] In another embodiment, another target detection device is provided. Based on the above embodiments, the first detection module 11 may include a first detection unit, a pixel position determination unit, and a two-dimensional detection result determination unit, wherein:

[0166] The first detection unit is used to process the image using the target perception algorithm described above, and to obtain the target bounding box of the target in the image and the category of the target.

[0167] The pixel position determination unit is used to obtain the pixel position of the target based on the position of each pixel in the target box.

[0168] The two-dimensional detection result determination unit is used to determine the pixel position of the target and the category of the target as the two-dimensional target detection result.

[0169] Optionally, the pixel position determination unit may include a positioning point determination subunit and a first pixel position determination subunit, wherein:

[0170] The positioning point determination subunit is used to determine the position of the positioning pixel based on the position of each pixel in the target box; the positioning pixel is used to characterize the position of the target in the image.

[0171] The first pixel position determination subunit is used to take the position of the above-mentioned positioning pixel as the pixel position of the above-mentioned target.

[0172] Optionally, the position of the aforementioned positioning pixel is the center point of the bottom edge of the aforementioned target box.

[0173] Optionally, the first mapping processing module 12 is specifically used to process the pixel position of the target using the above mapping model to obtain the three-dimensional position of the target.

[0174] Optionally, the pixel position determination unit may include a bounding box determination subunit and a second pixel position determination subunit, wherein:

[0175] The bounding box determination subunit is used to determine the position of the bounding box corresponding to the target box based on the position of each pixel in the target box.

[0176] The second pixel position determination subunit is used to take the position of the bounding box as the pixel position of the target.

[0177] Optionally, the first mapping processing module 12 is specifically used to: obtain the position of the bottom edge of the bounding box from the position of the bounding box; process the bounding box using the mapping model to obtain the depth information corresponding to the position of the bottom edge; construct the length and width of the target based on the depth information corresponding to the position of the bottom edge; obtain the height of the bounding box from the position of the bounding box; construct the actual height of the target based on the height of the bounding box; and obtain the size of the target based on the length, width, and actual height of the target.

[0178] For specific limitations on target detection devices, please refer to the limitations on target tracking methods mentioned above, which will not be repeated here.

[0179] In one embodiment, such as Figure 6 As shown, a target tracking device is provided, applied to a vision sensor. The device may include: a second acquisition module 20, a second detection module 21, a two-dimensional tracking module 22, a second mapping processing module 23, and a three-dimensional tracking module 24, wherein:

[0180] The second acquisition module 20 is used to acquire video data of the target scene; the video data includes multiple frames of images.

[0181] The second detection module 21 is used to process the video data using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0182] The two-dimensional tracking module 22 is used to process the above two-dimensional target detection results using an image tracking algorithm to obtain the two-dimensional target motion trajectory;

[0183] The second mapping processing module 23 is used to process the above two-dimensional target detection results using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixel points of the image and the depth information of the point cloud of the radar sensor; the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0184] The three-dimensional tracking module 24 is used to obtain the three-dimensional tracking result of the target based on the three-dimensional target detection result and the two-dimensional target motion trajectory.

[0185] For specific limitations on target tracking devices, please refer to the limitations on target tracking methods mentioned above, which will not be repeated here.

[0186] In another embodiment, another target tracking device is provided. Based on the above embodiments, the three-dimensional position of the target is the three-dimensional coordinates of the target in the world coordinate system, and the three-dimensional tracking result of the target includes the latitude and longitude information of the target; the three-dimensional tracking module 24 may include a three-dimensional coordinate determination unit, a reference information acquisition unit, and a latitude and longitude information determination unit, wherein:

[0187] The three-dimensional coordinate determination unit is used to determine the three-dimensional coordinates corresponding to the pixel positions of the target in each frame image based on the three-dimensional position of the target and the two-dimensional target motion trajectory.

[0188] The reference information acquisition unit is used to acquire the latitude and longitude information of the origin of the radar coordinate system corresponding to the radar sensor, and the angle between any axis of the radar coordinate system and the candidate direction; the candidate direction is related to the latitude and longitude information of the origin.

[0189] The latitude and longitude information determination unit is used to transform the three-dimensional coordinates of the target corresponding to each frame of the image by using the latitude and longitude information of the origin and the included angle, so as to obtain the latitude and longitude information of the target corresponding to each frame of the image.

[0190] In another embodiment, another target tracking device is provided. Based on the above embodiments, the three-dimensional position of the target is the three-dimensional coordinates of the target in the world coordinate system, and the three-dimensional tracking result of the target includes the target's velocity and acceleration. The three-dimensional tracking module 24 may include a three-dimensional coordinate determination unit, a position vector determination unit, a time difference determination unit, a position vector variable determination unit, and a velocity and acceleration determination unit, wherein:

[0191] The three-dimensional coordinate determination unit is used to determine the three-dimensional coordinates corresponding to the pixel positions of the target in each frame image based on the three-dimensional position of the target and the two-dimensional target motion trajectory.

[0192] The position vector determination unit is used to obtain the three-dimensional coordinates of the origin in the world coordinate system, and connect the three-dimensional coordinates of the origin with the three-dimensional coordinates of the target corresponding to any frame image to obtain the position vector of the target corresponding to the frame image.

[0193] The time difference determination unit is used to obtain the acquisition time corresponding to each of two adjacent frames, and to obtain the time difference between the two adjacent frames based on the acquisition time of the two adjacent frames.

[0194] The position vector variable determination unit is used to perform mathematical operations on the position vectors of the above-mentioned targets in two adjacent frames of images to obtain the position vector variables of the above-mentioned targets;

[0195] The velocity and acceleration determination unit is used to perform mathematical operations on the aforementioned position vector variables and time differences to obtain the velocity and acceleration of the aforementioned target.

[0196] In another embodiment, another target tracking device is provided. Based on the above embodiments, the three-dimensional tracking result of the target includes the target's heading angle and angular velocity. The device may further include a reference axis acquisition module, a heading angle determination module, and an angular velocity determination module, wherein:

[0197] The reference axis acquisition module is used to acquire any axis in the world coordinate system mentioned above and use the aforementioned axis as the reference axis.

[0198] The heading angle determination module is used to calculate the angle between the above position vector variable and the above reference axis, and determine the obtained angle as the heading angle of the above target;

[0199] The angular velocity determination module is used to perform mathematical operations on the aforementioned heading angle and time difference to obtain the angular velocity of the aforementioned target.

[0200] For specific limitations on target tracking devices, please refer to the limitations on target tracking methods mentioned above, which will not be repeated here.

[0201] The modules in the aforementioned target detection and target tracking devices can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device containing the vision sensor, or stored in the memory of the computer device containing the vision sensor, so that the processor can call and execute the corresponding operations of each module.

[0202] In one embodiment, a vision sensor is provided, including a camera, a memory, and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0203] Acquire images of the target scene;

[0204] The above image is processed using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0205] The two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results. The mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor. The three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0206] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0207] The image is processed using the target perception algorithm described above to obtain the bounding box containing the target and the category of the target. The pixel position of the target is obtained based on the position of each pixel in the bounding box. The pixel position and category of the target are then determined as the two-dimensional target detection result.

[0208] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0209] Based on the positions of each pixel in the target bounding box, the position of the positioning pixel is determined; the positioning pixel is used to represent the position of the target in the image; the position of the positioning pixel is taken as the pixel position of the target.

[0210] In one embodiment, the position of the aforementioned positioning pixel is the center point of the bottom edge of the target box.

[0211] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0212] The pixel positions of the target are processed using the above mapping model to obtain the three-dimensional position of the target.

[0213] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0214] Based on the positions of each pixel in the target box, the position of the bounding box corresponding to the target box is determined; the position of the bounding box is used as the pixel position of the target.

[0215] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0216] The position of the bottom edge of the bounding box is obtained from its position, and the depth information corresponding to the bottom edge position is obtained by processing it using the mapping model. The length and width of the target are constructed based on the depth information corresponding to the bottom edge position. The height of the bounding box is obtained from its position, and the actual height of the target is constructed based on the height of the bounding box. The dimensions of the target are obtained based on its length, width, and actual height.

[0217] In one embodiment, a vision sensor is provided, including a camera, a memory, and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:

[0218] Acquire video data of the target scene; this video data includes multiple frames of images.

[0219] The video data is processed using a target perception algorithm to obtain two-dimensional target detection results; the two-dimensional target detection results include the pixel position and category of the target;

[0220] The above two-dimensional target detection results are processed using an image tracking algorithm to obtain the two-dimensional target motion trajectory;

[0221] The above two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor; the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0222] Based on the above three-dimensional target detection results and the above two-dimensional target motion trajectory, the three-dimensional tracking results of the above target are obtained.

[0223] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0224] Based on the three-dimensional position of the target and the two-dimensional trajectory of the target, the three-dimensional coordinates corresponding to the pixel positions of the target in each frame of the image are determined; the latitude and longitude information of the origin of the radar coordinate system corresponding to the radar sensor is obtained, as well as the angle between any axis of the radar coordinate system and the candidate direction; the candidate direction is related to the latitude and longitude information of the origin; using the latitude and longitude information of the origin and the angle, the three-dimensional coordinates of the target corresponding to each frame of the image are transformed to obtain the latitude and longitude information of the target corresponding to each frame of the image.

[0225] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0226] Based on the three-dimensional position of the target and the two-dimensional trajectory of the target, determine the three-dimensional coordinates corresponding to the pixel positions of the target in each frame of the image; obtain the three-dimensional coordinates of the origin in the world coordinate system, and connect the three-dimensional coordinates of the origin with the three-dimensional coordinates of the target corresponding to any frame of the image to obtain the position vector of the target corresponding to the frame of the image; obtain the acquisition time corresponding to each of the two adjacent frames of the image, and obtain the time difference between the two adjacent frames of the image based on the acquisition time of the two adjacent frames of the image; perform mathematical operations on the position vectors of the target in the two adjacent frames of the image to obtain the position vector variables of the target; perform mathematical operations on the position vector variables and the time difference to obtain the velocity and acceleration of the target.

[0227] In one embodiment, the processor, when executing a computer program, also performs the following steps:

[0228] Obtain any axis in the world coordinate system and use it as the reference axis; calculate the angle between the position vector variable and the reference axis and determine the angle as the heading angle of the target; perform mathematical operations on the heading angle and the time difference to obtain the angular velocity of the target.

[0229] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0230] Acquire images of the target scene;

[0231] The above image is processed using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes the pixel position and category of the target;

[0232] The two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results. The mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor. The three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0233] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0234] The image is processed using the target perception algorithm described above to obtain the bounding box containing the target and the category of the target. The pixel position of the target is obtained based on the position of each pixel in the bounding box. The pixel position and category of the target are then determined as the two-dimensional target detection result.

[0235] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0236] Based on the positions of each pixel in the target bounding box, the position of the positioning pixel is determined; the positioning pixel is used to represent the position of the target in the image; the position of the positioning pixel is taken as the pixel position of the target.

[0237] In one embodiment, the position of the aforementioned positioning pixel is the center point of the bottom edge of the target box.

[0238] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0239] The pixel positions of the target are processed using the above mapping model to obtain the three-dimensional position of the target.

[0240] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0241] Based on the positions of each pixel in the target box, the position of the bounding box corresponding to the target box is determined; the position of the bounding box is used as the pixel position of the target.

[0242] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0243] The position of the bottom edge of the bounding box is obtained from its position, and the depth information corresponding to the bottom edge position is obtained by processing it using the mapping model. The length and width of the target are constructed based on the depth information corresponding to the bottom edge position. The height of the bounding box is obtained from its position, and the actual height of the target is constructed based on the height of the bounding box. The dimensions of the target are obtained based on its length, width, and actual height.

[0244] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0245] Acquire video data of the target scene; this video data includes multiple frames of images.

[0246] The video data is processed using a target perception algorithm to obtain two-dimensional target detection results; the two-dimensional target detection results include the pixel position and category of the target;

[0247] The above two-dimensional target detection results are processed using an image tracking algorithm to obtain the two-dimensional target motion trajectory;

[0248] The above two-dimensional target detection results are processed using a preset mapping model to obtain three-dimensional target detection results; the above mapping model includes the mapping relationship between the position of the pixels in the image and the depth information of the point cloud of the radar sensor; the above three-dimensional target detection results include the size and / or three-dimensional position of the target.

[0249] Based on the above three-dimensional target detection results and the above two-dimensional target motion trajectory, the three-dimensional tracking results of the above target are obtained.

[0250] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0251] Based on the three-dimensional position of the target and the two-dimensional trajectory of the target, the three-dimensional coordinates corresponding to the pixel positions of the target in each frame of the image are determined; the latitude and longitude information of the origin of the radar coordinate system corresponding to the radar sensor is obtained, as well as the angle between any axis of the radar coordinate system and the candidate direction; the candidate direction is related to the latitude and longitude information of the origin; using the latitude and longitude information of the origin and the angle, the three-dimensional coordinates of the target corresponding to each frame of the image are transformed to obtain the latitude and longitude information of the target corresponding to each frame of the image.

[0252] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0253] Based on the three-dimensional position of the target and the two-dimensional trajectory of the target, determine the three-dimensional coordinates corresponding to the pixel positions of the target in each frame of the image; obtain the three-dimensional coordinates of the origin in the world coordinate system, and connect the three-dimensional coordinates of the origin with the three-dimensional coordinates of the target corresponding to any frame of the image to obtain the position vector of the target corresponding to the frame of the image; obtain the acquisition time corresponding to each of the two adjacent frames of the image, and obtain the time difference between the two adjacent frames of the image based on the acquisition time of the two adjacent frames of the image; perform mathematical operations on the position vectors of the target in the two adjacent frames of the image to obtain the position vector variables of the target; perform mathematical operations on the position vector variables and the time difference to obtain the velocity and acceleration of the target.

[0254] In one embodiment, when the computer program is executed by a processor, it also performs the following steps:

[0255] Obtain any axis in the world coordinate system and use it as the reference axis; calculate the angle between the position vector variable and the reference axis and determine the angle as the heading angle of the target; perform mathematical operations on the heading angle and the time difference to obtain the angular velocity of the target.

[0256] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0257] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0258] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A target detection method characterized by, The method is applied to a visual sensor arranged on a side pole in a road, and comprises: acquiring an image of a target scene; processing the image by using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result comprises pixel positions and categories of targets; detecting whether a radar sensor is invalid, and if the radar sensor is invalid, processing the two-dimensional target detection result by using a preset mapping model to obtain a three-dimensional target detection result; the mapping model comprises a mapping relationship between positions of pixel points of the image and depth information of point clouds of the radar sensor, the mapping model is determined according to historical image data and historical point cloud data at each time in the same scene, the depth information comprises a distance between the target and the visual sensor, and the three-dimensional target detection result comprises a size and / or a three-dimensional position of the target; a method for acquiring the depth information of the point clouds, comprising: determining whether there is depth information corresponding to the position of the pixel point in the mapping model according to the position of the pixel point; if there is, acquiring the depth information corresponding to the position of the pixel point; if there is not, determining a plurality of fitting curves formed by the depth information in the mapping model on the image, determining an extension line based on two-dimensional coordinate axes on the image, and selecting two intersection points closest to the pixel point from intersection points of the extension line and the plurality of fitting curves as a plurality of target pixel points, and performing interpolation processing on the depth information corresponding to the positions of the target pixel points to obtain the depth information corresponding to the position of the pixel point.

2. The method of claim 1, wherein, processing the image by using the target perception algorithm to obtain a two-dimensional target detection result, comprising: processing the image by using the target perception algorithm to obtain a target frame in which a target is located and a category of the target; obtaining pixel positions of the target according to positions of each pixel point in the target frame; determining the pixel positions of the target and the category of the target as the two-dimensional target detection result.

3. The method of claim 2, wherein, obtaining pixel positions of the target according to positions of each pixel point in the target frame, comprising: determining positions of locating pixel points from the positions of each pixel point in the target frame; the locating pixel points are used to represent positions of the target in the image; taking the positions of the locating pixel points as the pixel positions of the target.

4. The method of claim 3, wherein, The positions of the locating pixel points are center points of bottom edges of the target frame.

5. The method of claim 4, wherein, processing the two-dimensional target detection result by using the preset mapping model to obtain a three-dimensional target detection result, comprising: processing the pixel positions of the target by using the mapping model to obtain a three-dimensional position of the target.

6. The method of claim 2, wherein, obtaining pixel positions of the target according to positions of each pixel point in the target frame, comprising: determining positions of a bounding box corresponding to the target frame from the positions of each pixel point in the target frame; taking the positions of the bounding box as the pixel positions of the target.

7. The method of claim 6, wherein, The method comprises the following steps: obtaining the position of the bottom edge of the bounding box from the position of the bounding box, and processing the position of the bottom edge by using the mapping model to obtain the depth information corresponding to the position of the bottom edge; constructing the length and width of the target according to the depth information corresponding to the position of the bottom edge; obtaining the height of the bounding box from the position of the bounding box, and constructing the actual height of the target according to the height of the bounding box; obtaining the size of the target according to the length, width and actual height of the target.

8. A target tracking method characterized by, The method is applied to a visual sensor arranged on a side pole in a road, and comprises the following steps: obtaining video data of a target scene; the video data comprises multiple images; processing the video data by using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result comprises the pixel position and category of a target; processing the two-dimensional target detection result by using an image tracking algorithm to obtain a two-dimensional target motion trajectory; detecting whether a radar sensor is invalid, and if the radar sensor is invalid, processing the two-dimensional target detection result by using a preset mapping model to obtain a three-dimensional target detection result; the mapping model comprises the mapping relationship between the position of a pixel of an image and the depth information of a point cloud of a radar sensor, the mapping model is determined according to historical image data and historical point cloud data of the same scene at each time, the depth information comprises the distance between the target and the visual sensor, and the three-dimensional target detection result comprises the size and / or three-dimensional position of the target; a method for obtaining the depth information of the point cloud, comprising: determining whether there is depth information corresponding to the position of the pixel in the mapping model according to the position of the pixel; if there is, obtaining the depth information corresponding to the position of the pixel; if there is not, determining multiple fitting curves formed by the depth information in the mapping model on the image, determining an extension line based on the two-dimensional coordinate axes on the image, selecting the two closest intersection points of the extension line and the multiple fitting curves as multiple target pixels, and performing interpolation processing on the depth information corresponding to the position of each target pixel according to the distance between each target pixel and the pixel to obtain the depth information corresponding to the position of the pixel. obtaining a three-dimensional tracking result of the target according to the three-dimensional target detection result and the two-dimensional target motion trajectory.

9. The method of claim 8, wherein, The three-dimensional position of the target is the three-dimensional coordinates of the target in a world coordinate system, and the three-dimensional tracking result of the target comprises the latitude and longitude information of the target. The method comprises the following steps: determining the three-dimensional coordinates corresponding to the pixel position of the target on each image according to the three-dimensional position of the target and the two-dimensional target motion trajectory. Obtain the longitude and latitude information of an origin in a radar coordinate corresponding to the radar sensor, and an included angle between any one axis of the radar coordinate system and a candidate direction; the candidate direction is related to the longitude and latitude information of the origin; Convert the three-dimensional coordinates of the target corresponding to each frame of image by using the longitude and latitude information of the origin and the included angle, to obtain the longitude and latitude information of the target corresponding to each frame of image.

10. The method of claim 8, wherein, The three-dimensional position of the target is a three-dimensional coordinate of the target in a world coordinate system, and the three-dimensional tracking result of the target includes a velocity and an acceleration of the target; The three-dimensional tracking result of the target is obtained according to the three-dimensional target detection result and the two-dimensional target motion trajectory, including: Determine the three-dimensional coordinates corresponding to the pixel position of the target on each frame of image according to the three-dimensional position of the target and the two-dimensional target motion trajectory; Obtain the three-dimensional coordinates of an origin in a world coordinate system, and connect the three-dimensional coordinates of the origin with the three-dimensional coordinates of the target corresponding to any one frame of image to obtain a position vector of the target corresponding to the frame of image; Obtain the acquisition times of adjacent two frames of image respectively, and obtain a time difference between the adjacent two frames of image according to the acquisition times of the adjacent two frames of image; Mathematically process the position vector of the target of adjacent two frames of image to obtain a position vector variable of the target; Mathematically process the position vector variable and the time difference to obtain the velocity and the acceleration of the target.

11. The method of claim 10, wherein, The three-dimensional tracking result of the target includes a heading angle and an angular velocity of the target; the method further includes: Obtain any one axis in the world coordinate system, and take the axis as a reference axis; Calculate an included angle between the position vector variable and the reference axis, and determine the obtained included angle as the heading angle of the target; Mathematically process the heading angle and the time difference to obtain the angular velocity of the target.

12. A target detection apparatus characterized by comprising: The target detection device is applied to a visual sensor, the visual sensor is arranged on a side pole in a road, and the target detection device includes: A first obtaining module is configured to obtain an image of a target scene; A first detection module is configured to process the image by using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result includes a pixel position and a category of a target; A first mapping processing module is configured to detect whether a radar sensor is invalid; if the radar sensor is invalid, a three-dimensional target detection result is obtained by processing the two-dimensional target detection result by using a preset mapping model; the mapping model includes a mapping relationship between a position of a pixel of an image and depth information of a point cloud of the radar sensor, the mapping model is determined according to historical image data and historical point cloud data at each time in the same scene, the depth information includes a distance between the target and the visual sensor, and the three-dimensional target detection result includes a size and / or a three-dimensional position of the target. The depth information acquisition module of the point cloud is configured to determine, according to the position of the pixel point, whether there is depth information corresponding to the position of the pixel point in the mapping model; if there is, the depth information corresponding to the position of the pixel point is acquired; if there is not, a plurality of fitting curves formed by each depth information in the mapping model on the image is determined, an extension line is determined based on the two-dimensional coordinate axes on the image, and two intersection points closest to the pixel point are selected from intersection points of the extension line and the plurality of fitting curves as a plurality of target pixel points; and the depth information corresponding to the position of each target pixel point is interpolated according to the distance between each target pixel point and the pixel point, to obtain the depth information corresponding to the position of the pixel point.

13. A target tracking device, characterized by The target tracking device is applied to a visual sensor arranged on a side pole in a road, and the target tracking device comprises: A second acquisition module is configured to acquire video data of a target scene; the video data comprises a plurality of images; A second detection module is configured to process the video data by using a target perception algorithm to obtain a two-dimensional target detection result; the two-dimensional target detection result comprises a pixel position and a category of a target; A two-dimensional tracking module is configured to process the two-dimensional target detection result by using an image tracking algorithm to obtain a two-dimensional target motion trajectory; A second mapping processing module is configured to detect whether a radar sensor is invalid; if the radar sensor is invalid, the two-dimensional target detection result is processed by using a preset mapping model to obtain a three-dimensional target detection result; the mapping model comprises a mapping relationship between a position of a pixel of an image and depth information of a point cloud of a radar sensor; the mapping model is determined according to historical image data and historical point cloud data at each time in the same scene; the depth information comprises a distance between the target and the visual sensor; and the three-dimensional target detection result comprises a size and / or a three-dimensional position of the target; The depth information acquisition module of the point cloud is configured to determine, according to the position of the pixel point, whether there is depth information corresponding to the position of the pixel point in the mapping model; if there is, the depth information corresponding to the position of the pixel point is acquired; if there is not, a plurality of fitting curves formed by each depth information in the mapping model on the image is determined, an extension line is determined based on the two-dimensional coordinate axes on the image, and two intersection points closest to the pixel point are selected from intersection points of the extension line and the plurality of fitting curves as a plurality of target pixel points; and the depth information corresponding to the position of each target pixel point is interpolated according to the distance between each target pixel point and the pixel point, to obtain the depth information corresponding to the position of the pixel point; A three-dimensional tracking module is configured to obtain a three-dimensional tracking result of the target according to the three-dimensional target detection result and the two-dimensional target motion trajectory.

14. A vision sensor, characterized by An apparatus comprising a camera, a memory storing a computer program, and a processor which, when executing the computer program, implements the steps of the method of any one of claims 1 to 11.

15. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program which, when executed by the processor, implements the steps of the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Intelligent target tracking trajectory recording method

    CN108090922A

  • A multi-camera data fusion method based on a spatial coordinate system

    CN109190508A

  • Article recognition pre-sorting system and method based on deep learning and robot

    CN111368852A