Image Tracking Processing Method, Device, Computer Equipment and Storage Medium
By preprocessing the current frame point cloud data and combining standard area images, the target tracking model is used to determine the target tracking area, which solves the accuracy and robustness problems of traditional visual tracking technology under the influence of image quality, and achieves a more stable target tracking effect.
Patent Information
- Application Number
- CN201980037486.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-30
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2039-12-30
AI Technical Summary
Traditional image-based visual tracking techniques are susceptible to image quality, especially when ambient light changes and target motion speed changes, resulting in less accuracy and robustness of tracking results.
By acquiring the current frame point cloud data for preprocessing, generating a projected image, and combining the standard area image corresponding to the standard frame point cloud data, the target tracking model is called to obtain the candidate area label of the candidate area, and finally determining the target tracking area.
Improve the accuracy and robustness of target tracking, reduce dependence on image quality, and enhance tracking stability when ambient light changes and target motion speed changes.
Smart Images

Figure CN113490965B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to an image tracking processing method, apparatus, computer device, storage medium, and vehicle. Background Art
[0002] Visual tracking refers to using computer technology to extract, identify, and track a target, obtain information such as the position of the target, and then perform subsequent processing and analysis. With the development of computer technology, visual tracking technology can be implemented in many application scenarios. For example, visual tracking technology can be applied to related fields such as the field of autonomous driving and the field of assisted driving.
[0003] In the traditional method, visual tracking technology usually performs target tracking based on images captured by devices such as cameras. However, the inventor has realized that the tracking result is easily affected by the image quality in the way of performing target tracking based on the captured images. Under the influence of factors such as environmental light changes and the target movement speed, the image quality is low, which in turn leads to low accuracy and robustness of the target tracking result. Summary of the Invention
[0004] According to various embodiments disclosed in the present application, there is provided an image tracking processing method, apparatus, computer device, storage medium, and vehicle.
[0005] An image tracking processing method includes:
[0006] Obtaining current frame point cloud data;
[0007] Preprocessing the current frame point cloud data to generate a projection image;
[0008] Obtaining a standard region image corresponding to the standard frame point cloud data;
[0009] Invoking a target tracking model to obtain a candidate region label corresponding to the candidate region based on the projection image and the standard region image; and
[0010] Determining a target tracking region corresponding to the current frame point cloud data according to the candidate region label.
[0011] An image tracking processing apparatus includes:
[0012] A point cloud acquisition module for obtaining current frame point cloud data;
[0013] A preprocessing module for preprocessing the current frame point cloud data to generate a projection image;
[0014] A standard image acquisition module for obtaining a standard region image corresponding to the standard frame point cloud data; and
[0015] A target tracking module, configured to call a target tracking model, obtain candidate region labels corresponding to candidate regions based on the projection image and the standard region image, and determine a target tracking region corresponding to the current frame point cloud data according to the candidate region labels.
[0016] A computer device includes a memory and one or more processors. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the one or more processors, the one or more processors perform the following steps:
[0017] Obtain the current frame point cloud data;
[0018] Preprocess the current frame point cloud data to generate a projection image;
[0019] Obtain a standard region image corresponding to the standard frame point cloud data;
[0020] Call a target tracking model, obtain candidate region labels corresponding to candidate regions based on the projection image and the standard region image; and
[0021] Determine a target tracking region corresponding to the current frame point cloud data according to the candidate region labels.
[0022] One or more non-volatile computer-readable storage media storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors perform the following steps:
[0023] Obtain the current frame point cloud data;
[0024] Preprocess the current frame point cloud data to generate a projection image;
[0025] Obtain a standard region image corresponding to the standard frame point cloud data;
[0026] Call a target tracking model, obtain candidate region labels corresponding to candidate regions based on the projection image and the standard region image; and
[0027] Determine a target tracking region corresponding to the current frame point cloud data according to the candidate region labels.
[0028] A vehicle includes steps of performing the above image tracking processing method.
[0029] Details of one or more embodiments of the present application are set forth in the following drawings and description. Other features and advantages of the present application will become apparent from the specification, drawings, and claims. Description of the Drawings
[0030] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0031] Figure 1 It is an application scenario diagram of an image tracking processing method according to one or more embodiments.
[0032] Figure 2 It is a schematic flowchart of an image tracking processing method according to one or more embodiments.
[0033] Figure 3 It is a schematic flowchart of the steps for obtaining a standard detection region corresponding to a standard frame image according to one or more embodiments.
[0034] Figure 4 It is a block diagram of an image tracking processing device according to one or more embodiments.
[0035] Figure 5 It is a block diagram of a computer device according to one or more embodiments. Specific Embodiments
[0036] In order to make the technical solutions and advantages of the present application more clearly understood, the following further details the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0037] The image tracking processing method provided by the present application can be applied to a variety of application environments. For example, it can be applied to Figure 1In the application environment of the autonomous driving shown, it may include a laser sensor 102 and a computer device 104. The computer device 104 can communicate with the laser sensor 102 according to the connection established with the laser sensor 102. A wired connection or a wireless connection can be established between the laser sensor 102 and the computer device 104. The laser sensor 102 can collect multiple frames of point cloud data of the surrounding environment. The computer device 104 can obtain the current frame of point cloud data collected by the laser sensor 102, and the computer device 104 can also obtain the preset current frame of point cloud data. The computer device 104 preprocesses the current frame of point cloud data to generate a projection image and obtains the standard region image corresponding to the standard frame of point cloud data. The computer device 104 calls the target tracking model and obtains the candidate region label corresponding to the candidate region based on the projection image and the standard region image. The computer device 104 determines the target tracking region corresponding to the current frame of point cloud data according to the candidate region label. The laser sensor 102 can be a laser sensor carried by an autonomous driving device, and specifically can include a lidar, a laser scanner, etc.
[0038] In one embodiment, as Figure 2 shown, an image tracking processing method is provided. Taking the method applied to Figure 1 the computer device 104 therein as an example for illustration, it includes the following steps:
[0039] Step 202, obtain the current frame of point cloud data.
[0040] The laser sensor can be carried by a device capable of autonomous driving. For example, it can be carried by an unmanned vehicle or by a vehicle including an autonomous driving model. The laser sensor can be used to collect environmental data within the visual range. Specifically, the laser sensor can emit detection signals, such as laser beams. The laser sensor compares the signals reflected back by the objects in the environment with the detection signals to obtain the surrounding environmental data. The environmental data collected by the laser sensor can specifically be point cloud data. Point cloud data refers to the collection of point data corresponding to multiple points on the surface of an object recorded in the form of points by scanning the object in the environment. Among them, multiple can specifically refer to two or more. The laser sensor can collect at a preset frequency to obtain multiple frames of point cloud data. The preset frequency can be preset according to actual needs. For example, it can be specifically set to 50 frames per second.
[0041] The point cloud data can be three-dimensional point cloud data, and each frame of point cloud data can include point data corresponding to multiple points. The point data can specifically include at least one of the three-dimensional coordinates corresponding to the points, the laser reflection intensity, and the color information, etc. Among them, the three-dimensional coordinates can be the coordinates of the points in the Cartesian coordinate system, specifically including the horizontal axis coordinate, the vertical axis coordinate, and the vertical axis coordinate of the points in the Cartesian coordinate system. The Cartesian coordinate system is a three-dimensional space coordinate system established with the position where the laser sensor is located as the origin, and the three-dimensional space coordinate system includes the horizontal axis (x-axis), the vertical axis (y-axis), and the vertical axis (z-axis). The three-dimensional space coordinate system established with the position where the laser sensor is located as the origin satisfies the right-hand rule.
[0042] The computer device can acquire the point cloud data. Specifically, the computer device can acquire the acquired point cloud data in real time when the laser sensor acquires each frame of point cloud data, or can acquire multiple frames of point cloud data acquired by the laser sensor after the laser sensor acquires multiple frames of point cloud data. The computer device can perform target tracking sequentially according to multiple frames of point cloud data in the time sequence of the point cloud data acquired by the laser sensor. The computer device can record the point cloud data at the start of target tracking or during target tracking as the current frame of point cloud data. The target can include organisms or non-organisms in the surrounding environment. The target can be moving or stationary. For example, the target can specifically include at least one of pedestrians, roadblocks, vehicles, and buildings, etc. It can be understood that when the computer device finishes tracking the current frame of point cloud data and starts to track the next frame of point cloud data, according to the acquisition order of the point cloud data, the current frame of point cloud data can be recorded as the previous frame of point cloud data, and the next frame of point cloud data is acquired and recorded as the current frame of point cloud data again.
[0043] Step 204, preprocess the current frame of point cloud data to generate a projection image.
[0044] The computer device can preprocess the acquired current frame of point cloud data, and the preprocessing can include at least one of multiple processing methods. Specifically, the preprocessing performed by the computer device on the current frame of point cloud data can specifically include at least one of processing methods such as data cleaning, point cloud segmentation, and point cloud projection. The computer device generates a projection image from the point cloud data with a large amount of discrete point data, effectively reducing the amount of data calculation and saving the computing resources of the computer device.
[0045] For example, the ways for a computer device to preprocess the current frame of point cloud data may include point cloud projection. Specifically, the computer device may obtain the point data corresponding to each of the multiple points in the current frame of point cloud data, and extract the three-dimensional coordinates corresponding to the points from the point data. The computer device may project the points in the current frame of point cloud data onto a plane according to the three-dimensional coordinates of the points, and record the image formed by the points projected on the plane as the projected image. The generated projected image is a two-dimensional image. For example, the computer device may project the points in the current frame of point cloud data onto the x-y plane where the horizontal axis and the vertical axis are located to obtain a top view of the point cloud, and the computer device may record the top view of the point cloud as the projected image.
[0046] The ways for a computer device to preprocess the current frame of point cloud data may also include data cleaning and point cloud projection. Specifically, the computer device may perform data cleaning on the current frame of point cloud data, and clean the abnormal point data from the multiple point data included in the current frame of point cloud data, so as to avoid the interference of abnormal point data on target tracking and ensure the accuracy of the tracking result. The computer device may perform point cloud projection according to the current frame of point cloud data after cleaning to obtain the projected image generated after projection.
[0047] The ways for a computer device to preprocess the current frame of point cloud data may also include point cloud segmentation and point cloud projection. Specifically, the computer device may divide the current frame of point cloud data into multiple sub-point clouds according to the point data, and generate a segmentation threshold corresponding to the sub-point cloud based on the point data included in each sub-point cloud. The computer device may segment the points in the corresponding sub-point cloud according to the segmentation threshold, and count the segmentation results corresponding to the multiple sub-point clouds to obtain the ground point set and the non-ground point set corresponding to the current frame of point cloud data. The computer device may project the points in the non-ground point set to generate a projected image. By segmenting the point cloud data, the interference of ground points on target tracking is excluded, thereby ensuring the accuracy of the tracking result. In one embodiment, the ways for a computer device to preprocess the current frame of point cloud data may further include data cleaning, point cloud segmentation, and point cloud projection.
[0048] Step 206, obtain the standard region image corresponding to the standard frame of point cloud data.
[0049] The standard frame of point cloud data can be used as a reference basis for target tracking. The computer device may perform target tracking on the current frame of point cloud data based on the standard frame of point cloud data. The standard frame of point cloud data may be one of various point cloud data. For example, the standard frame of point cloud data may be a frame of point cloud data determined by the user from multiple frames of point cloud data according to actual needs, or may also be the first frame of point cloud data among multiple frames of point cloud data collected by a laser sensor.
[0050] The computer device can obtain the standard region image corresponding to the standard frame point cloud data. The standard frame point cloud data can correspond to one or more standard region images. The standard region image refers to the image corresponding to the region where the target is located in the standard frame point cloud data. The standard region image can be an image of various shapes. For example, the standard region image can be rectangular or circular. The standard region image can be a part of the standard image corresponding to the standard frame point cloud data, and the standard image can be obtained by performing point cloud projection on the standard frame point cloud data.
[0051] The computer device can obtain the standard region image corresponding to the standard frame point cloud data in various ways. Specifically, the computer device can detect the standard frame point cloud data to obtain the standard region image corresponding to the standard frame point cloud data. The standard region image can also be preset by the user according to actual needs. For example, the computer device can receive the target selected by the user in advance to be tracked, and determine the standard region image corresponding to the target to be tracked. The computer device can obtain the standard region image corresponding to the standard frame point cloud data that is preset in advance.
[0052] Step 208: Invoke the target tracking model, and obtain the candidate region label corresponding to the candidate region based on the projection image and the standard region image.
[0053] The computer device can invoke the target tracking model, and perform tracking processing on the projection image according to the target tracking model to obtain the tracking region corresponding to the current frame point cloud data. The target tracking model can be pre-configured in the computer device. The target tracking model can be one of various deep learning models. For example, it can be one of various convolutional neural network models, deep belief network models, etc. The target tracking model can be obtained by training the deep learning model based on point cloud image samples.
[0054] The computer device can input the projection image generated through preprocessing and the standard region image corresponding to the standard frame point cloud data into the target tracking model, and perform operations on the projection image and the standard region image through the target tracking model to obtain the candidate region label corresponding to the candidate region output by the target tracking model. The candidate region refers to the region where the target may be located in the projection image. The candidate region can specifically include the position where the target may be located, the range of the region, and the shape, etc. The candidate region label refers to the marked label corresponding to the candidate region, and there is a unique association between the candidate region label and the candidate region. The candidate region label can include the region confidence or probability value indicating that the candidate region belongs to the true region of the target.
[0055] Step 210: Determine the target tracking region corresponding to the current frame point cloud data according to the candidate region label.
[0056] A computer device can obtain candidate region labels corresponding to multiple candidate regions, and determine a target tracking region corresponding to the current frame point cloud data based on the candidate region labels, so as to achieve target tracking. The target tracking region refers to the position region where the target is located in the current frame point cloud data estimated through tracking processing, and the target tracking region can be a target box corresponding to the target. Specifically, the computer device can use one of multiple algorithms to determine the target tracking region. For example, the computer device can use the maximum value algorithm to compare multiple candidate region labels with each other, and determine the candidate region corresponding to the candidate region label with the largest region confidence among the multiple candidate region labels as the target tracking region corresponding to the current frame point cloud data.
[0057] In one embodiment, the computer device can also use the non-maximum suppression algorithm (Non-Maximum Suppression, abbreviated as NMS) to screen the candidate region labels. Specifically, the computer device can perform multiple screenings on multiple candidate regions according to the non-maximum suppression algorithm based on the region confidence. Each screening clears the unselected candidate regions until the screening ends. The computer device can determine the candidate region corresponding to the screened candidate region label as the target tracking region corresponding to the current frame point cloud data, effectively improving the accuracy of determining the target tracking region from multiple candidate regions.
[0058] In this embodiment, the computer device preprocesses the acquired current frame point cloud data to generate a projection image, and performs tracking on the projection image. By processing the current frame point cloud data with a large amount of discrete point data to generate a projection image, the computational load of the computer device is effectively reduced, and the computing resources of the computer device are saved. The target tracking model is called to process the standard region image and the projection image corresponding to the standard frame point cloud data to obtain candidate region labels corresponding to the candidate regions, and the target tracking region is determined based on the candidate region labels, so as to achieve tracking of the target in the current frame point cloud data. Compared with the traditional method of target tracking based on images, the point cloud data collected by the laser sensor is not easily affected by factors such as environmental light changes and target movement speed, effectively improving the accuracy and robustness of target tracking.
[0059] In one embodiment, the steps of preprocessing the current frame point cloud data to generate a projection image include: obtaining a target tracking task; obtaining a corresponding image plane according to the target tracking task; projecting the points in the current frame point cloud data onto the image plane to obtain a projection image.
[0060] A computer device can obtain a target tracking task, which can be used to instruct the computer device and the laser sensor to perform target tracking. The target tracking task can be triggered according to the user's operation instruction, or can be automatically generated by the computer device according to actual needs. The target tracking task can carry a tracking task type. The tracking task type refers to the task type corresponding to the target tracking task, and the target tracking task can correspond to one of multiple task types.
[0061] The tracking task type can be used to represent multiple tracking scenarios. In different tracking scenarios, the requirements for point cloud projection can be different, and the tracking task type of the target tracking task can also be different. The computer device can obtain an image plane corresponding to the tracking task type according to the tracking task type. The image plane is a plane used to project the current frame of point cloud data to generate a projected image. In different tracking scenarios, the computer device can determine different planes as the image plane.
[0062] For example, when a vehicle equipped with a laser sensor is driving on a horizontal road surface and it is necessary to determine the distribution of the target in the horizontal plane where the vehicle is located, the computer device can determine the horizontal plane where the laser sensor is located, that is, the x-y plane formed by the horizontal axis and the vertical axis in the space coordinate system, as the image plane, without considering the vertical axis coordinate in the three-dimensional coordinates of the points. When a vehicle equipped with a laser sensor is driving on an uphill or downhill route, the computer device can determine the vertical plane corresponding to the laser sensor, that is, the y-z plane formed by the vertical axis and the vertical axis in the space coordinate system, as the image plane, without considering the horizontal axis coordinate in the three-dimensional coordinates of the points.
[0063] The computer device can project multiple points in the current frame of point cloud data, project the multiple points onto the image plane, and obtain multiple projected points in the image plane. The computer device can record the image corresponding to the multiple projected points in the image plane as a projected image, and the projected image is a two-dimensional image. The computer device can perform tracking based on the generated projected image to obtain a two-dimensional target tracking area in the projected image.
[0064] In one embodiment, the computer device can obtain multiple image planes, project the points in the current frame of point cloud data onto the multiple image planes respectively, and obtain multiple projected images. The computer device can perform tracking processing on the multiple projected images respectively to obtain the target tracking areas corresponding to the multiple projected images respectively. It can be understood that the target tracking area determined in the two-dimensional projected image is also two-dimensional. The computer device can synthesize the target tracking areas corresponding to the multiple projected images to generate a three-dimensional target tracking area corresponding to the current frame of point cloud data, so as to more accurately determine the position and size of the tracked target in the three-dimensional space, which is beneficial for the computer device to analyze and control autonomous driving according to the three-dimensional target tracking area.
[0065] In this embodiment, the computer device can determine the corresponding image plane according to the target tracking task, project the points in the current frame of point cloud data onto the image plane corresponding to the target tracking task to obtain a projected image, and reduce the dimension of the current frame of point cloud data, thereby reducing the data volume of the current frame of point cloud data. The computer device performs target tracking based on the generated projected image, and can utilize the image features in the projected image. Compared with the traditional method of performing Kalman filtering on point cloud data to achieve target tracking, the accuracy of target tracking based on point cloud data is effectively improved.
[0066] In one embodiment, the steps of obtaining the standard region image corresponding to the standard frame of point cloud data include: generating a standard frame image according to the standard frame of point cloud data; obtaining the standard detection region corresponding to the standard frame image; and intercepting the standard region image matching the standard detection region in the standard frame image.
[0067] The computer device can obtain the standard frame of point cloud data. The standard frame of point cloud data can be a frame of point cloud data determined by the user according to actual needs from multiple frames of point cloud data, or the first frame of point cloud data in the multiple frames of point cloud data collected by the laser sensor.
[0068] Specifically, the computer device can generate the standard frame image according to the standard frame of point cloud data in multiple ways. For example, the computer device can project the points in the standard frame of point cloud data and determine the image obtained by projection as the standard frame image. The method of obtaining the standard frame image by projecting the standard frame of point cloud data by the computer device can be similar to the method of generating the projected image according to the current frame of point cloud data in the above embodiment, so it will not be elaborated here. The computer device can also obtain the point data included in the standard frame of point cloud data, encode the points according to the point data to obtain the point features corresponding to each of the multiple points, and generate a feature map according to the point features corresponding to each of the multiple points. The computer device can denote the feature map generated according to the standard frame of point cloud data as the standard frame image.
[0069] The computer device can obtain the standard detection region corresponding to the standard frame image. The standard detection region can be used to represent the region where the target is located in the standard frame image and can be a part of the region range in the standard frame image. The standard detection region can be detected by the computer device according to the standard frame of point cloud data. Specifically, the computer device can perform target detection according to the standard frame of point cloud data to obtain the standard detection region. After the computer device generates the standard frame image according to the standard frame of point cloud data, it can also perform target detection according to the standard frame image to obtain the standard detection region. The standard detection region can specifically include the position, range, and region shape of the target in the standard frame of point cloud data. The computer device can obtain one standard detection region corresponding to the standard frame image, or multiple corresponding standard detection regions.
[0070] The computer device can intercept the standard region image in the standard frame image according to the standard detection region corresponding to the standard frame image, so as to obtain the standard region image corresponding to the standard detection region. The standard region image may include the target to be tracked, and the intercepted standard region image matches the size and shape of the standard region.
[0071] In this embodiment, the computer device generates a standard frame image based on the standard frame point cloud data, obtains the standard detection region corresponding to the standard frame image, and intercepts the standard region image matching the standard detection region in the standard frame image. The computer device can use the intercepted standard region image as the basis for target tracking to perform target tracking on the projection image. By generating the image, the depth features of the point cloud data are utilized, effectively improving the accuracy of target tracking.
[0072] In one embodiment, as Figure 3 shown, the steps of obtaining the standard detection region corresponding to the standard frame image include:
[0073] Step 302, rasterize the standard frame point cloud data to obtain a plurality of grids.
[0074] Step 304, extract the point features corresponding to the standard frame point cloud data in the plurality of grids to generate a point feature matrix.
[0075] Step 306, call the target detection model, input the point feature matrix into the target detection model, and obtain the point cloud detection region corresponding to the standard frame point cloud data.
[0076] Step 308, determine the standard detection region corresponding to the standard frame image according to the point cloud detection region.
[0077] The computer device can detect the target according to the standard frame point cloud data to obtain the standard detection region corresponding to the target. Specifically, the computer device can rasterize the standard frame point cloud data, divide the three-dimensional space corresponding to the standard frame point cloud data into a plurality of grids. The computer device can determine the grid to which the point belongs according to the three-dimensional coordinates of the point in the standard frame point cloud data.
[0078] The computer device can count the point data corresponding to the points in each grid, extract features of the points in each grid, and obtain the point features corresponding to the points. Specifically, the computer device can call a feature extraction model to extract the point features in the grid. The feature extraction model can be obtained by training with a large number of point cloud samples and point feature samples. The feature extraction model can be one of various neural network models. For example, the feature extraction model can be a convolutional neural network model, specifically the PointNet model. The computer device can input the point data in each grid into the feature extraction model, perform operations on the point data through the feature extraction model, and obtain the point features output by the feature extraction model. The computer device can count the point features corresponding to multiple points in the grid and generate a point feature matrix. The point feature matrix can be a three-dimensional matrix.
[0079] The computer device can call a target detection model to detect the targets in the standard frame point cloud data through the target detection model. The target detection model can be pre-configured in the computer device through training. The target detection model can be obtained by training based on a Convolutional Neural Networks (CNN) model. The target detection model can specifically include one of the YOLO model or the Mask RCNN model, etc. The computer device can input the generated point feature matrix into the target detection model, perform operations on the point feature matrix through the target detection model, and obtain the detection area output by the target detection model. The computer device can inverse rasterize the detection area output by the target detection model to obtain the point cloud detection area corresponding to the standard frame point cloud data.
[0080] Since the point cloud detection area is a three-dimensional detection area corresponding to the standard frame point cloud data, the computer can determine the standard detection area corresponding to the standard frame image according to the point cloud detection area. Specifically, the computer device can project the point cloud detection area onto the corresponding image plane in the way that the standard frame point cloud data projects to generate the standard frame image, and obtain the standard detection area corresponding to the standard frame image.
[0081] In one embodiment, when the computer device performs detection based on the standard frame image, it can obtain the two-dimensional detection area corresponding to the standard frame image. The computer device can directly record the two-dimensional detection area corresponding to the standard frame image as the standard detection area corresponding to the standard frame image.
[0082] In this embodiment, the computer device can call the target detection model to detect the point feature matrix corresponding to the standard frame point cloud data, and obtain the standard detection area corresponding to the standard frame image, so that the computer device can track the targets in the current frame point cloud data based on the standard detection area, effectively improving the accuracy of target tracking.
[0083] In one embodiment, the steps of invoking a target tracking model and obtaining candidate region labels corresponding to candidate regions based on a projection image and a standard region image include: extracting current image features corresponding to the projection image and standard image features corresponding to the standard region image; inputting the current image features and the standard image features into the target tracking model; and performing filtering processing on the current image features and the standard image features based on the target tracking model to obtain candidate region labels corresponding to multiple candidate regions output by the target tracking model.
[0084] The computer device can perform feature extraction on the projection image corresponding to the current frame point cloud data and the standard region image corresponding to the standard frame point cloud data to obtain the current image features corresponding to the projection image and the standard image features corresponding to the standard region image. Specifically, the computer device can sequentially extract the image features of the projection image and the standard region image in a single-threaded manner, or can extract the image features of the projection image and the standard region image in parallel in a multi-threaded manner. The computer device can invoke an image feature model to perform feature extraction on the projection image and the standard region image to obtain the current image features and the standard image features output by the image feature model. The image feature model can be a two-dimensional convolutional neural network model. In one embodiment, when the computer device extracts image features in parallel in a multi-threaded manner, the computer device can obtain a siamese network model corresponding to the image feature model and extract the features of the projection image and the standard region image in parallel.
[0085] The computer device can input the extracted current image features and standard image features into the target tracking model. The target tracking model can be one of various convolutional neural network models. For example, the target tracking model can specifically include a SiamMask model, a Siamese RPN (Region Proposal Network) model, etc. The computer device can process the current image features and the standard image features based on the target tracking model. Specifically, the target tracking model can perform convolutional filtering on the current image features and the standard image features, compare the current image features and the standard image features respectively, and obtain candidate region labels corresponding to each of the multiple candidate regions output by the target tracking model. The avatar image of the candidate region corresponds to the standard region image. In one embodiment, when the standard frame point cloud data corresponds to multiple standard region images, the computer can obtain siamese network models corresponding to multiple target tracking models, perform operations on the standard image features corresponding to the multiple standard region images, and obtain candidate regions corresponding to the multiple standard region images.
[0086] In this embodiment, the computer device operates on the current image features corresponding to the projection image and the standard image features corresponding to the standard region image by invoking the target tracking model, and obtains the candidate region labels corresponding to multiple candidate regions, making full use of the image features of the image corresponding to the point cloud data, and determining multiple candidate regions through the deep learning model. Compared with the tracking method of performing Kalman filtering on the point cloud, the accuracy of target tracking is effectively improved.
[0087] In one of the embodiments, the steps of filtering the current image features and the standard image features based on the target tracking model include: obtaining a historical feature matrix; adjusting the standard image features according to the historical feature matrix, and performing filtering processing according to the adjusted image features.
[0088] Before invoking the target tracking model to operate on the standard image features and the current image features, the computer device can also obtain a historical feature matrix. The historical feature matrix refers to the feature matrix generated by the computer device according to the historical image features corresponding to the historical target images in the historical point cloud data. The historical point cloud data may include the point cloud data including the target collected by the laser sensor before the current frame of point cloud data. The historical feature matrix may be generated from the image features corresponding to the target in multiple frames of historical point cloud data, and the historical feature matrix and the historical point cloud data may be stored in the corresponding memory of the computer device.
[0089] It can be understood that after the computer device finishes target tracking on the current frame of point cloud data, the current frame of point cloud data can be recorded as historical point cloud data. The computer device can adjust the historical feature matrix according to the image features of the target tracking area corresponding to the current frame of point cloud data, and continuously adjust the historical feature matrix corresponding to the target, effectively improving the accuracy and robustness of the historical feature matrix corresponding to the target.
[0090] The computer device can adjust the standard image features according to the obtained historical feature matrix. Specifically, the computer device can perform convolution processing on the historical feature matrix and the standard image features through the target tracking model to obtain the adjusted image features. The computer device can perform convolution filtering on the adjusted image features and the current image features to obtain the candidate region labels corresponding to multiple candidate regions.
[0091] In this embodiment, the computer device can obtain the historical feature matrix corresponding to the target to adjust the standard image features, and perform filtering processing according to the adjusted image features to obtain the candidate region labels corresponding to multiple candidate regions. By adjusting the standard image features through the historical feature matrix corresponding to the target, the adjusted image features can more accurately reflect the features of the target in the image, and through multiple frames of point cloud data in the historical time, the accuracy and robustness of target tracking are effectively improved.
[0092] It should be understood that although Figure 2-3 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-3 at least a part of the steps in
[0093] In one embodiment, as Figure 4 shown, an image tracking processing device is provided, including: a point cloud acquisition module 402, a preprocessing module 404, a standard image acquisition module 406, and a target tracking module 408, where:
[0094] The point cloud acquisition module 402 is configured to acquire current frame point cloud data.
[0095] The preprocessing module 404 is configured to preprocess the current frame point cloud data to generate a projection image.
[0096] The standard image acquisition module 406 is configured to acquire a standard region image corresponding to the standard frame point cloud data.
[0097] The target tracking module 408 is configured to call a target tracking model, obtain a candidate region label corresponding to a candidate region based on the projection image and the standard region image; and determine a target tracking region corresponding to the current frame point cloud data according to the candidate region label.
[0098] In one of the embodiments, the above-mentioned preprocessing module 404 is further configured to obtain a target tracking task; obtain a corresponding image plane according to the target tracking task; project the points in the current frame point cloud data onto the image plane to obtain a projection image.
[0099] In one of the embodiments, the above-mentioned standard image acquisition module 406 is further configured to generate a standard frame image according to the standard frame point cloud data; obtain a standard detection region corresponding to the standard frame image; and intercept a standard region image matching the standard detection region in the standard frame image.
[0100] In one embodiment, the above-mentioned standard image acquisition module 406 is further configured to rasterize the standard frame point cloud data to obtain a plurality of grids; extract the point features corresponding to the standard frame point cloud data in the plurality of grids to generate a point feature matrix; call a target detection model, input the point feature matrix into the target detection model, and obtain a point cloud detection area corresponding to the standard frame point cloud data; determine a standard detection area corresponding to the standard frame image according to the point cloud detection area.
[0101] In one embodiment, the above-mentioned target tracking module 408 is further configured to extract the current image features corresponding to the projected image and the standard image features corresponding to the standard area image; input the current image features and the standard image features into the target tracking model; perform filtering processing on the current image features and the standard image features based on the target tracking model to obtain candidate area labels corresponding to a plurality of candidate areas output by the target tracking model.
[0102] In one embodiment, the above-mentioned target tracking module 408 is further configured to obtain a historical feature matrix; adjust the standard image features according to the historical feature matrix, and perform filtering processing according to the adjusted image features.
[0103] In one embodiment, the candidate area label includes an area confidence level. The above-mentioned target tracking module 408 is further configured to screen a plurality of candidate areas according to the area confidence level; determine the screened candidate areas as the target tracking areas corresponding to the current frame point cloud data.
[0104] For the specific limitations of the image tracking processing device, reference may be made to the limitations on the image tracking processing method in the foregoing text, which will not be elaborated here. Each module in the above-mentioned image tracking processing device can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.
[0105] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 5As shown. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, computer-readable instructions, and a database. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the non-volatile storage medium. The database of the computer device is used to store image tracking processing data. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer-readable instructions are executed by the processor, an image tracking processing method is implemented.
[0106] Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0107] In one embodiment, a computer device is provided, including a memory and one or more processors. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the one or more processors, the one or more processors implement the steps in the above method embodiments when executed.
[0108] In one embodiment, one or more non-volatile computer-readable storage media storing computer-readable instructions are provided. When the computer-readable instructions are executed by one or more processors, the one or more processors implement the steps in the above method embodiments when executed.
[0109] In one embodiment, a vehicle is provided. The vehicle may specifically include an autonomous vehicle, an electric vehicle, a bicycle, an aircraft, etc. The vehicle includes the above computer device and can execute the steps in the above image tracking processing method embodiments.
[0110] The embodiments and implementation objects of the present invention are not limited to autonomous vehicles, electric vehicles, bicycles, aircraft, robots, etc., but also include simulation devices and test equipment related to these devices.
[0111] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. The computer-readable instructions can be stored in a non-volatile computer-readable storage medium. When the computer-readable instructions are executed, they can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0112] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0113] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. An image tracking and processing method, including: Obtaining current frame point cloud data; Obtaining a target tracking task; Obtaining a corresponding image plane according to the target tracking task; Projecting the points in the current frame point cloud data onto the image plane to obtain a projected image; Generating a standard frame image according to standard frame point cloud data; Performing rasterization processing on the standard frame point cloud data to obtain a plurality of grids; Extracting point features corresponding to the standard frame point cloud data in the plurality of grids to generate a point feature matrix; Invoking a target detection model, inputting the point feature matrix into the target detection model to obtain a point cloud detection area corresponding to the standard frame point cloud data; Determining a standard detection area corresponding to the standard frame image according to the point cloud detection area; Intercepting a standard area image matching the standard detection area in the standard frame image, where the standard area image is an image corresponding to the area where the target is located in the standard frame point cloud data; Extracting current image features corresponding to the projected image and standard image features corresponding to the standard area image; Inputting the current image features and the standard image features into a target tracking model, where the target tracking model is one of multiple convolutional neural network models; Performing filtering processing on the current image features and the standard image features based on the target tracking model to obtain candidate area labels corresponding to a plurality of candidate areas output by the target tracking model; and Determining a target tracking area corresponding to the current frame point cloud data according to the candidate area labels.
2. The method according to claim 1, wherein the target tracking task carries a tracking task type, and the method further includes: Obtaining an image plane corresponding to the tracking task type according to the tracking task type.
3. The method according to claim 1, wherein the method further includes: Determining the grid to which a point belongs according to the three-dimensional coordinates of the point in the standard frame point cloud data; Counting the point data corresponding to the points in each grid, performing feature extraction on the points in each grid to obtain point features corresponding to the points; and Counting a point matrix corresponding to multiple points in the grid to generate the point feature matrix.
4. The method according to claim 1, wherein the standard detection area includes the position, target range, and area shape of the target in the standard frame point cloud data.
5. The method according to claim 1, wherein the method further includes: Performing convolutional filtering on the current image features and the standard image features, comparing the current image features and the standard image features respectively to obtain candidate area labels corresponding to each of the multiple candidate areas output by the target tracking model.
6. The method according to claim 1, wherein the performing filtering processing on the current image features and the standard image features based on the target tracking model includes: Obtaining a historical feature matrix; and Adjusting the standard image features according to the historical feature matrix, and performing filtering processing according to the adjusted image features.
7. The method according to claim 1, wherein The candidate region label includes a region confidence level. Determining the target tracking region corresponding to the current frame point cloud data according to the candidate region label includes: Screening a plurality of the candidate regions according to the region confidence level; and Determining the screened candidate region as the target tracking region corresponding to the current frame point cloud data.
8. An image tracking processing device, comprising: A point cloud acquisition module, configured to acquire current frame point cloud data; A preprocessing module, configured to acquire a target tracking task; Acquiring a corresponding image plane according to the target tracking task; projecting points in the current frame point cloud data onto the image plane to obtain a projected image; A standard image acquisition module, configured to generate a standard frame image according to standard frame point cloud data; Performing rasterization processing on the standard frame point cloud data to obtain a plurality of grids; extracting point features corresponding to the standard frame point cloud data in the plurality of grids to generate a point feature matrix; Invoking a target detection model, and inputting the point feature matrix into the target detection model to obtain a point cloud detection region corresponding to the standard frame point cloud data; Determining a standard detection region corresponding to the standard frame image according to the point cloud detection region; intercepting a standard region image matching the standard detection region in the standard frame image, where the standard region image is an image corresponding to the region where the target is located in the standard frame point cloud data; and A target tracking module, configured to extract a current image feature corresponding to the projected image and a standard image feature corresponding to the standard region image; inputting the current image feature and the standard image feature into a target tracking model, where the target tracking model is one of a plurality of convolutional neural network models; performing filtering processing on the current image feature and the standard image feature based on the target tracking model to obtain candidate region labels corresponding to a plurality of candidate regions output by the target tracking model; and determining the target tracking region corresponding to the current frame point cloud data according to the candidate region labels.
9. The device according to claim 8, wherein, The target tracking module is further configured to: Acquire a historical feature matrix; and Adjust the standard image feature according to the historical feature matrix, and perform filtering processing according to the adjusted image feature.
10. A computer device, comprising a memory and one or more processors, where computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the one or more processors, the one or more processors perform the following steps: Acquire current frame point cloud data; Acquire a target tracking task; Acquire a corresponding image plane according to the target tracking task; Project points in the current frame point cloud data onto the image plane to obtain a projected image; Generate a standard frame image according to standard frame point cloud data; Perform rasterization processing on the standard frame point cloud data to obtain a plurality of grids; Extract point features corresponding to the standard frame point cloud data in the plurality of grids to generate a point feature matrix; Invoke a target detection model, and input the point feature matrix into the target detection model to obtain a point cloud detection region corresponding to the standard frame point cloud data; Determine the standard detection area corresponding to the standard frame image according to the point cloud detection area; Crop a standard area image that matches the standard detection area in the standard frame image, where the standard area image is the image corresponding to the area where the target is located in the standard frame point cloud data; Extract the current image features corresponding to the projection image and the standard image features corresponding to the standard area image; Input the current image features and the standard image features into a target tracking model, where the target tracking model is one of multiple convolutional neural network models; Perform filtering processing on the current image features and the standard image features based on the target tracking model to obtain candidate region labels corresponding to multiple candidate regions output by the target tracking model; and Determine the target tracking area corresponding to the current frame point cloud data according to the candidate region labels.
11. The computer device according to claim 10, wherein, The target tracking task carries a tracking task type, and when the processor executes the computer-readable instructions, the following steps are further executed: Obtain an image plane corresponding to the tracking task type according to the tracking task type.
12. The computer device according to claim 10, wherein, When the processor executes the computer-readable instructions, the following steps are further executed: Determine the grid to which a point belongs according to the three-dimensional coordinates of the point in the standard frame point cloud data; Count the point data corresponding to the points in each grid, extract features of the points in each grid to obtain point features corresponding to the points; and Count the point matrix corresponding to multiple points in the grid to generate the point feature matrix.
13. The computer device according to claim 12, wherein, The standard detection area includes the position where the target is located, the target range, and the area shape in the standard frame point cloud data.
14. The computer device according to claim 10, wherein, When the processor executes the computer-readable instructions, the following steps are further executed: Perform convolutional filtering on the current image features and the standard image features, compare the current image features and the standard image features respectively, and obtain candidate region labels corresponding to each of the multiple candidate regions output by the target tracking model.
15. The computer device according to claim 14, wherein, When the processor executes the computer-readable instructions, the following steps are further executed: Obtain a historical feature matrix; and Adjust the standard image features according to the historical feature matrix, and perform filtering processing according to the adjusted image features.
16. One or more non-volatile computer-readable storage media storing computer-readable instructions, when the computer-readable instructions are executed by one or more processors, cause the one or more processors to perform the following steps: Obtain the current frame point cloud data; Obtain a target tracking task; Obtain a corresponding image plane according to the target tracking task; Project the points in the current frame point cloud data onto the image plane to obtain a projection image; Generate a standard frame image according to the standard frame point cloud data; Perform rasterization processing on the standard frame point cloud data to obtain multiple grids; Extract the point features corresponding to the standard frame point cloud data in the multiple grids to generate a point feature matrix; Call a target detection model, input the point feature matrix into the target detection model, and obtain the point cloud detection area corresponding to the standard frame point cloud data; Determine the standard detection area corresponding to the standard frame image according to the point cloud detection area; Intercept a standard area image that matches the standard detection area in the standard frame image, and the standard area image is the image corresponding to the area where the target is located in the standard frame point cloud data; Extract the current image features corresponding to the projection image and the standard image features corresponding to the standard area image; Input the current image features and the standard image features into a target tracking model, and the target tracking model is one of multiple convolutional neural network models; Based on the target tracking model, perform filtering processing on the current image features and the standard image features to obtain candidate region labels corresponding to multiple candidate regions output by the target tracking model; and Determine the target tracking area corresponding to the current frame point cloud data according to the candidate region labels.
17. The storage medium according to claim 16, wherein, The target tracking task carries a tracking task type, and when the computer-readable instructions are executed by the processor, the following steps are further performed: Obtain an image plane corresponding to the tracking task type according to the tracking task type.
18. The storage medium according to claim 16, wherein, When the computer-readable instructions are executed by the processor, the following steps are further performed: Perform convolutional filtering on the current image features and the standard image features, compare the current image features and the standard image features respectively, and obtain candidate region labels corresponding to each of the multiple candidate regions output by the target tracking model.
19. The storage medium according to claim 18, wherein, When the computer-readable instructions are executed by the processor, the following steps are further performed: Obtain a historical feature matrix; and Adjust the standard image features according to the historical feature matrix, and perform filtering processing according to the adjusted image features.
20. A vehicle, including performing the image tracking processing method according to any one of claims 1-7.
Citation Information
Patent Citations
Object detection method, device, apparatus, storage medium and vehicle
CN109345510A
Target tracking method and target tracking device
CN110533693A
Three-dimensional object detection and tracking method based on streaming data
CN110570457A