Target tracking method, device, server and readable storage medium
By obtaining point cloud data scanned by multiple lidars, determining the three-dimensional spatial information of the target and comparing it with the predicted spatial information, the problem of poor target tracking effect when the occlusion exists in traditional technology is solved, and higher tracking accuracy is achieved.
Patent Information
- Application Number
- CN202010777637.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-05
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-08-05
AI Technical Summary
In traditional technology, the target tracking method has poor tracking effect when it has occlusion.
By obtaining point cloud data of multiple lidar scanning detection areas, the three-dimensional spatial information of each target is determined, and the three-dimensional spatial information at the current moment is compared with the predicted spatial information, and the corresponding identifier is matched to complete the target tracking.
This method reduces the impact caused by occlusion through the mutual complementation between multiple lidars, and fully considers the historical information of the target and improves the accuracy of the target tracking results through the use of predicted spatial information.
Smart Images

Figure CN114091561B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a target tracking method, device, server and readable storage medium. Background Art
[0002] In the current traffic field, it is usually necessary to track road targets, such as tracking whether vehicles are driving illegally, in order to reduce the pressure on traffic officers. With the continuous development of LiDAR technology, it is widely used in the target tracking process due to its advantages of high resolution, good concealment, and strong anti-active interference ability.
[0003] In traditional technology, a laser radar on one side of the road is used to collect point cloud data of a certain range at fixed time intervals, and target detection is performed on the point cloud data. The targets detected at adjacent moments are analyzed to achieve the purpose of tracking the target.
[0004] However, the target tracking method in traditional technology has poor tracking effect when there are occlusions. Summary of the invention
[0005] Based on this, it is necessary to propose a target tracking method, device, server and readable storage medium to address the problem that the target tracking method in traditional technology has poor tracking effect when there are obstructions.
[0006] A target tracking method, the method comprising:
[0007] Acquire point cloud data obtained by scanning a detection area with multiple laser radars; multiple laser radars are arranged at different positions in the detection area;
[0008] Determine the three-dimensional spatial information of each target in the detection area at the current moment based on the point cloud data obtained by multiple laser radar scans; the three-dimensional spatial information includes the location information and size information of the target;
[0009] The three-dimensional spatial information of each target in the detection area at the current moment is compared with the predicted spatial information of each target in the target set, and a corresponding identifier is determined for the target whose three-dimensional spatial information matches the predicted spatial information to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
[0010] In one embodiment, the three-dimensional spatial information of each target in the detection area at the current moment is determined based on the point cloud data obtained by scanning multiple laser radars, including:
[0011] A coordinate system where the first point cloud data is located is selected as a reference coordinate system from a plurality of point cloud data obtained by scanning a plurality of laser radars, and the second point cloud data is converted to the reference coordinate system where the first point cloud data is located according to a preset conversion matrix, and the converted second point cloud data is fused with the first point cloud data to obtain fused point cloud data; wherein the second point cloud data is other point cloud data in the plurality of point cloud data except the first point cloud data, and one point cloud data is obtained by one laser radar scan;
[0012] The fused point cloud data is processed for target detection to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0013] In one embodiment, the three-dimensional spatial information of each target in the detection area at the current moment is determined based on the point cloud data obtained by scanning multiple laser radars, including:
[0014] Perform target detection processing on the point cloud data of multiple laser radars respectively to obtain the three-dimensional spatial information of the target in each point cloud data;
[0015] A coordinate system where the first three-dimensional spatial information is located is selected as a reference coordinate system from multiple three-dimensional spatial information of multiple point cloud data, and second three-dimensional spatial information is converted to the reference coordinate system where the first three-dimensional spatial information is located according to a preset conversion matrix, and the converted second three-dimensional spatial information and the first three-dimensional spatial information are fused to obtain fused three-dimensional spatial information; wherein the second three-dimensional spatial information is other three-dimensional spatial information in the multiple three-dimensional spatial information corresponding to different point cloud data of the first three-dimensional spatial information, and one point cloud data corresponds to multiple three-dimensional spatial information;
[0016] The fused three-dimensional spatial information is de-redundanted to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0017] In one embodiment, the fused three-dimensional spatial information is de-redundantly processed to obtain the three-dimensional spatial information of each target in the detection area at the current moment, including:
[0018] The non-maximum suppression algorithm is used to remove redundancy from the fused three-dimensional spatial information to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0019] In one embodiment, comparing the three-dimensional spatial information of each target in the detection area at the current moment with the predicted spatial information of each target in the target set, and determining a corresponding identifier for the target whose three-dimensional spatial information matches the predicted spatial information, includes:
[0020] For each target corresponding to each three-dimensional space information at the current moment, identify the first feature of the target;
[0021] For each target corresponding to the predicted spatial information, identifying a second feature of the target;
[0022] If there is a target whose similarity between the first feature and the second feature is greater than the similarity threshold at the current moment, the identifier of the target corresponding to the second feature is used as the identifier of the target corresponding to the first feature.
[0023] In one embodiment, the method further comprises:
[0024] If there is a target whose similarity between the first feature and the second feature is not greater than the similarity threshold at the current moment, calculate the intersection-over-union ratio between the three-dimensional spatial information corresponding to the target whose similarity is not greater than the similarity threshold among the targets corresponding to the current moment and the candidate prediction spatial information; wherein the candidate prediction spatial information is the prediction spatial information of the target whose similarity is not greater than the similarity threshold in the target set;
[0025] If the intersection-over-union ratio is greater than the intersection-over-union ratio threshold, the identifier of the target corresponding to the candidate prediction space information is used as the identifier of the target corresponding to the three-dimensional space information.
[0026] In one embodiment, comparing the three-dimensional spatial information of each target in the detection area at the current moment with the predicted spatial information of each target in the target set, and determining a corresponding identifier for the target whose three-dimensional spatial information matches the predicted spatial information, includes:
[0027] A Kalman filter is used to predict the three-dimensional spatial information of the targets in the target set to obtain predicted spatial information of each target in the target set; wherein the identification of the target corresponding to the predicted spatial information corresponds to the identification of the target in the target set;
[0028] For each target at the current moment, the intersection-and-union ratio between the three-dimensional spatial information and all the predicted spatial information is calculated. If there is three-dimensional spatial information whose intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identifier of the target corresponding to the matched predicted spatial information is used as the identifier of the target corresponding to the three-dimensional spatial information.
[0029] In one embodiment, the method further comprises:
[0030] If there is three-dimensional spatial information whose intersection-and-union ratio is not greater than the intersection-and-union ratio threshold, identify the third feature of the first target and the fourth feature of the second target; wherein the first target is a target whose three-dimensional spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the targets corresponding to the current moment, and the second target is a target whose predicted spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the target set;
[0031] The similarity between the third feature and the fourth feature is calculated, and if the similarity is greater than a similarity threshold, the identifier of the second target is determined as the identifier of the first target.
[0032] In one embodiment, the method further comprises:
[0033] If there is a target with an undetermined identification at the current moment, a random identification is assigned to the target with an undetermined identification, and the target with an undetermined identification and the random identification are stored in a target set; wherein the random identification is different from identifications of other targets in the target set.
[0034] A target tracking device, characterized in that the device comprises:
[0035] An acquisition module is used to acquire point cloud data obtained by scanning a detection area by multiple laser radars; the multiple laser radars are arranged at different positions in the detection area;
[0036] A determination module is used to determine the three-dimensional spatial information of each target in the detection area at the current moment according to the point cloud data obtained by scanning multiple laser radars; the three-dimensional spatial information includes the position information and size information of the target;
[0037] A comparison module is used to compare the three-dimensional spatial information of each target in the detection area at the current moment with the predicted spatial information of each target in the target set, and determine the corresponding identification for the target whose three-dimensional spatial information matches the predicted spatial information to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
[0038] A server comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0039] Acquire point cloud data obtained by scanning a detection area with multiple laser radars; multiple laser radars are arranged at different positions in the detection area;
[0040] Determine the three-dimensional spatial information of each target in the detection area at the current moment based on the point cloud data obtained by multiple laser radar scans; the three-dimensional spatial information includes the location information and size information of the target;
[0041] The three-dimensional spatial information of each target in the detection area at the current moment is compared with the predicted spatial information of each target in the target set, and a corresponding identifier is determined for the target whose three-dimensional spatial information matches the predicted spatial information to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
[0042] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps:
[0043] Acquire point cloud data obtained by scanning a detection area with multiple laser radars; multiple laser radars are arranged at different positions in the detection area;
[0044] Determine the three-dimensional spatial information of each target in the detection area at the current moment based on the point cloud data obtained by multiple laser radar scans; the three-dimensional spatial information includes the location information and size information of the target;
[0045] The three-dimensional spatial information of each target in the detection area at the current moment is compared with the predicted spatial information of each target in the target set, and a corresponding identifier is determined for the target whose three-dimensional spatial information matches the predicted spatial information to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
[0046] The target tracking method, device, server and readable storage medium can obtain point cloud data obtained by scanning the detection area by multiple laser radars; then determine the three-dimensional spatial information of each target in the detection area at the current moment according to the point cloud data obtained by scanning the multiple laser radars, and the three-dimensional spatial information includes the position information and size information of the target; then compare the three-dimensional spatial information of each target in the detection area at the current moment with the predicted spatial information of each target in the target set, and determine the corresponding identification for the target whose three-dimensional spatial information matches the predicted spatial information to complete the target tracking. Among them, multiple laser radars are arranged in different directions of the detection area, so that the multiple laser radars can compensate each other, that is, the scanning blind area of one laser radar can be scanned by another laser radar to obtain the corresponding point cloud data, reducing the influence caused by the obstruction; in addition, the predicted spatial information is obtained by predicting the three-dimensional spatial information of the target in the target set, and the target set includes the target in the detection area at the previous moment, that is, the three-dimensional spatial information of the target at the current moment is predicted according to the three-dimensional spatial information of the target at the previous moment, and compared and matched with the real three-dimensional spatial information of the target at the current moment, thereby fully considering the historical information of the target to complete the target tracking process, which can greatly improve the accuracy of the target tracking result. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 An application scenario diagram of a target tracking method in an embodiment;
[0048] Figure 2 A schematic diagram of a target tracking method according to an embodiment of the present invention;
[0049] Figure 3 is a flow chart of a target tracking method in another embodiment;
[0050] Figure 4 is a flowchart of a target tracking method in yet another embodiment;
[0051] Figure 5 is a flowchart of a target tracking method in yet another embodiment;
[0052] Figure 6 is a flowchart of a target tracking method in yet another embodiment;
[0053] Figure 7 is a flowchart of a target tracking method in yet another embodiment;
[0054] Figure 8 is a structural block diagram of a target tracking device in one embodiment;
[0055] Fig. 9 FIG. 4 is a diagram showing the internal structure of a server in one embodiment.
[0056] Description of reference numerals:
[0057] 11: base station; 12: server. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0059] The target tracking method provided in the embodiment of the present application can be applied to Figure 1 In the application scenario shown. Among them, multiple base stations 11 are arranged in different directions of the detection area, such as the diagonal corners of a road intersection, and the scanned point cloud data is sent to the server 12; the server 12 can track the target in the detection area according to the point cloud data at different times. Optionally, the base station 11 may include sensors such as laser radar or millimeter wave radar, and the point cloud data is obtained by scanning by the laser radar or millimeter wave radar; the server 12 can be implemented by an independent server or a service cluster composed of multiple servers.
[0060] In one embodiment, Figure 2 As shown, a target tracking method is provided, which is applied to Figure 1 The server in the example is used to illustrate the specific process of the server tracking the target based on the point cloud data obtained by multiple laser radar scans. The method includes the following steps:
[0061] S101, obtaining point cloud data obtained by scanning a detection area with multiple laser radars; the multiple laser radars are arranged at different positions in the detection area.
[0062] Among them, multiple laser radars (or base station laser radars) are arranged at different positions in the detection area to scan the detection area from different angles. In this way, if the obstructed object cannot be scanned from the scanning angle of one laser radar, it can be compensated by the scanning angle of another laser radar. Multiple laser radars can continuously scan the detection area at fixed time intervals to obtain point cloud data at different times, and send each point cloud data to the server, so that the server can obtain the point cloud data obtained by multiple laser radars scanning the detection area.
[0063] S102, determining the three-dimensional spatial information of each target in the detection area at the current moment based on the point cloud data obtained by scanning multiple laser radars; the three-dimensional spatial information includes the position information and size information of the target.
[0064] Specifically, the server can select the point cloud data at the current moment from the acquired multiple point cloud data. Assuming that two laser radars A and B are arranged in the detection area, the server can select the point cloud data of A and the point cloud data of B at the current moment. Then the server performs target detection on the point cloud data of A and the point cloud data of B respectively, and determines the three-dimensional spatial information of each target in each point cloud data. The three-dimensional spatial information includes the position information and size information of the target; wherein the position information, i.e., the current geographical location of the target, can be represented by the latitude and longitude information in the geodetic coordinate system, and the size information can be represented by the size of the detection box that can surround the target, such as the length, width and height of the detection box.
[0065] Optionally, if the scanning areas of the two laser radars have no overlapping parts, the server can directly perform target detection on the point cloud data of A and the point cloud data of B respectively; if the scanning areas of the two laser radars have overlapping parts, the server can superimpose the point cloud data of the overlapping parts to make the point cloud density of the overlapping parts higher, and then perform the target detection process, which can improve the accuracy of the three-dimensional spatial information of the targets in the overlapping parts.
[0066] S103, comparing the three-dimensional spatial information of each target in the detection area at the current moment with the predicted spatial information of each target in the target set, determining a corresponding identifier for the target whose three-dimensional spatial information matches the predicted spatial information to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
[0067] The target tracking process is generally a process of associating the driving state of a target at the last moment (which may include location information, etc.) with the driving state at the current moment to obtain the entire driving state of the target. In this embodiment, the server stores the targets detected at the last moment and the three-dimensional spatial information corresponding to each target. Each target may be located in a target set, and the target set may be stored in a list form.
[0068] Specifically, the server can compare the three-dimensional spatial information of each target detected at the current moment with the predicted spatial information of each target in the target set, and the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, that is, the three-dimensional spatial information at the current moment predicted by the three-dimensional spatial information at the previous moment. If there is a target (a) at the current moment whose three-dimensional spatial information matches the predicted spatial information, the identifier of the target corresponding to the matched predicted spatial information can be used as the identifier of the target (a) at the current moment, thereby determining the position information of the target (a) at the previous moment and the position information at the current moment, and completing the tracking process of the target.
[0069] Optionally, the server can compare the position information of the target at the current moment with the position information in the predicted spatial information. If there are two targets with the same or similar position information, the size information between the two is compared; if the size information is also the same or similar, it can be considered that the target at the current moment and the target corresponding to the predicted spatial information are the same target, and an identifier is determined for the target at the current moment.
[0070] In the target tracking method provided by this embodiment, the server first obtains point cloud data obtained by scanning the detection area by multiple laser radars; then, based on the point cloud data obtained by scanning the multiple laser radars, the three-dimensional spatial information of each target in the detection area at the current moment is determined, and the three-dimensional spatial information includes the position information and size information of the target; then, the three-dimensional spatial information of each target in the detection area at the current moment is compared with the predicted spatial information of each target in the target set, and the corresponding identification is determined for the target whose three-dimensional spatial information matches the predicted spatial information, so as to complete the target tracking. Among them, multiple laser radars are arranged in different directions of the detection area, so that the multiple laser radars can compensate each other, that is, the scanning blind area of one laser radar can be scanned by another laser radar to obtain the corresponding point cloud data, reducing the influence caused by the obstruction; in addition, the predicted spatial information is obtained by predicting the three-dimensional spatial information of the target in the target set, and the target set includes the target in the detection area at the previous moment, that is, the three-dimensional spatial information of the target at the current moment is predicted based on the three-dimensional spatial information of the target at the previous moment, and compared and matched with the real three-dimensional spatial information of the target at the current moment, thereby fully considering the historical information of the target to complete the target tracking process, which can greatly improve the accuracy of the target tracking result.
[0071] Normally, different laser radars (or base station laser radars) have their own coordinate systems, so the point cloud data obtained by scanning may be in different coordinate systems, that is, in different spatial domains; the server needs to convert each point cloud data to the same coordinate system first, and then detect the three-dimensional spatial information of the target, so that the obtained three-dimensional spatial information can also be in the same coordinate system to reduce the error of the subsequent matching process. Optionally, such as Figure 3 As shown, the above S102 may include:
[0072] S201, select the coordinate system where the first point cloud data is located as the reference coordinate system from the multiple point cloud data obtained by multiple laser radar scans, convert the second point cloud data to the reference coordinate system where the first point cloud data is located according to a preset transformation matrix, and fuse the converted second point cloud data and the first point cloud data to obtain fused point cloud data; wherein the second point cloud data is the other point cloud data in the multiple point cloud data except the first point cloud data, and one point cloud data is obtained by one laser radar scan.
[0073] Among them, the server can select the coordinate system where the first point cloud data is located from the multiple point cloud data scanned by the above-mentioned multiple laser radars as the reference coordinate system, and convert the other point cloud data to the reference coordinate system, so that multiple point cloud data are located in the same coordinate system, and one laser radar usually scans and obtains one point cloud data at a time. Specifically, the server can convert the second point cloud data to the above-mentioned reference coordinate system according to a preset conversion matrix, and the second point cloud data is the other point cloud data in the multiple point cloud data except the first point cloud data. Optionally, the conversion matrix can characterize the relative relationship between the reference coordinate system and the coordinate system where the second point cloud data is located; optionally, the conversion matrix can be determined according to the iterative closest point algorithm (ICP) to convert the second point cloud data to the reference coordinate system where the first point cloud data is located. Then, the converted second point cloud data and the first point cloud data are fused to obtain fused point cloud data, and the fusion operation can be an overlay operation of the two point cloud data.
[0074] S202, performing target detection processing on the fused point cloud data to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0075] Specifically, after obtaining the fused point cloud data, the server can perform target detection processing on the fused point cloud data. Optionally, a deep learning-based target detection algorithm can be used to execute the target detection processing to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0076] In the target tracking method provided in this embodiment, the server can convert the second point cloud data to the reference coordinate system where the first point cloud data is located according to the preset conversion matrix, and fuse the converted second point cloud data and the first point cloud data to obtain fused point cloud data; then, target detection processing is performed on the fused point cloud data to obtain the three-dimensional spatial information of each target in the detection area at the current moment. By converting different point cloud data to the same coordinate system, the target detection processing is performed in the same spatial domain, which improves the accuracy of the target detection result, thereby improving the accuracy of the target tracking result.
[0077] Generally, the amount of point cloud data obtained by laser radar scanning is large. If all point cloud data are converted to coordinate system, the amount of calculation will inevitably increase. Therefore, in this embodiment, target detection can be performed first, and only the obtained three-dimensional space information can be converted to coordinate system to improve the calculation efficiency. Figure 4 As shown, the above S102 may include:
[0078] S301, performing target detection processing on the point cloud data of multiple laser radars respectively to obtain the three-dimensional spatial information of the target in each point cloud data.
[0079] Specifically, the server can first perform target detection processing on the point cloud data of each lidar respectively. Optionally, a deep learning-based target detection algorithm can be used to execute the target detection processing to obtain the three-dimensional spatial information of the target in each point cloud data.
[0080] S302, selecting a coordinate system where the first three-dimensional spatial information is located from multiple three-dimensional spatial information of multiple point cloud data as a reference coordinate system, converting the second three-dimensional spatial information to the reference coordinate system where the first three-dimensional spatial information is located according to a preset transformation matrix, and fusing the converted second three-dimensional spatial information and the first three-dimensional spatial information to obtain fused three-dimensional spatial information; wherein the second three-dimensional spatial information is other three-dimensional spatial information in the multiple three-dimensional spatial information that corresponds to different point cloud data from the first three-dimensional spatial information, and one point cloud data corresponds to multiple three-dimensional spatial information.
[0081] Among them, the server can select the coordinate system where the first three-dimensional spatial information is located from the above-mentioned multiple three-dimensional spatial information as the reference coordinate system, and convert other three-dimensional spatial information to the reference coordinate system, so that multiple three-dimensional spatial information is located in the same coordinate system, and one point cloud data usually corresponds to multiple three-dimensional spatial information, that is, the scene corresponding to one point cloud data includes multiple targets. Specifically, the server can convert the second three-dimensional spatial information to the above-mentioned reference coordinate system according to a preset conversion matrix. The second three-dimensional spatial information is other three-dimensional spatial information in the multiple three-dimensional spatial information corresponding to different point cloud data of the first three-dimensional spatial information, that is, the first three-dimensional spatial information and the second three-dimensional spatial information are obtained from different point cloud data. Optionally, the conversion matrix can characterize the relative relationship between the reference coordinate system and the coordinate system where the second three-dimensional spatial information is located; optionally, the conversion matrix can be determined according to the ICP algorithm to convert the second three-dimensional spatial information to the reference coordinate system where the first three-dimensional spatial information is located. Then, the converted second three-dimensional spatial information and the first three-dimensional spatial information are fused to obtain fused three-dimensional spatial information, and the fusion operation can be a union operation of two three-dimensional spatial information.
[0082] S303, performing redundancy removal processing on the fused three-dimensional spatial information to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0083] Specifically, for scenes where there are overlapping parts in the scanning areas of multiple base stations, there may be a target with multiple spatial information in the fused three-dimensional spatial information, that is, multiple base stations detect the target at the same time, then the server needs to perform de-redundancy processing on the scene so that each target corresponds to only one three-dimensional spatial information, that is, to obtain unique three-dimensional spatial information for each target in the detection area at the current moment. Optionally, the server can use a non-maximum suppression algorithm to perform de-redundancy processing on the fused three-dimensional spatial information to obtain the three-dimensional spatial information of each target in the detection area at the current moment. It can be understood that the best one (such as the one with the highest position information accuracy or the detection box size is the smallest box that can surround the target) is selected from multiple three-dimensional spatial information as the final three-dimensional spatial information.
[0084] In the target tracking method provided in this embodiment, the server first performs target detection processing on the point cloud data of multiple laser radars respectively to obtain the three-dimensional spatial information of the target in each point cloud data; then, according to the preset transformation matrix, the second three-dimensional spatial information is converted to the reference coordinate system where the first three-dimensional spatial information is located, and the converted second three-dimensional spatial information and the first three-dimensional spatial information are fused to obtain the fused three-dimensional spatial information; finally, the fused three-dimensional spatial information is de-redundantly processed to obtain the three-dimensional spatial information of each target in the detection area at the current moment. By converting different three-dimensional spatial information to the same coordinate system, each three-dimensional spatial information is placed in the same spatial domain to improve the accuracy of subsequent target tracking results; at the same time, only the three-dimensional spatial information is converted, which also improves the conversion efficiency; and the target detection processing is performed on one point cloud data, which reduces the amount of data processing compared to the execution on the fused point cloud data and improves the processing efficiency.
[0085] In one embodiment, the server compares the three-dimensional spatial information of each target with the predicted spatial information to determine the specific process of identifying the target in the detection area at the current moment. Figure 5 As shown, the above S103 may include:
[0086] S401, for a target corresponding to each three-dimensional space information at a current moment, identifying a first feature of the target.
[0087] S402: For each target corresponding to the predicted spatial information, identify a second feature of the target.
[0088] Specifically, for each target corresponding to the three-dimensional spatial information at the current moment, the server can identify the first feature of the target based on the deep learning target recognition algorithm, and for each target corresponding to the predicted spatial information, also identify the second feature of the target. Optionally, the server can also use a point cloud re-identification network to identify target features.
[0089] S403: If there is a target whose similarity between the first feature and the second feature is greater than a similarity threshold at the current moment, the identifier of the target corresponding to the second feature is used as the identifier of the target corresponding to the first feature.
[0090] Specifically, among all the targets corresponding to the current moment, if there is a target whose similarity between the first feature and the second feature is greater than the similarity threshold, that is, the target at the current moment exists in the target set, that is, the target was also scanned at the previous moment; then the server can use the identifier of the target corresponding to the second feature (the identifier of the target in the target set) as the identifier of the target corresponding to the first feature, that is, the identifier of the target at the current moment, thereby achieving the purpose of determining the identifier for the target at the current moment and associating it with the target at the previous moment.
[0091] Of course, among all the targets at the current moment, there must be targets whose similarity between the first feature and the second feature is not greater than the similarity threshold, that is, they fail the similarity match. Optionally, the server can calculate the intersection-and-union ratio between the three-dimensional spatial information corresponding to the target whose similarity is not greater than the similarity threshold in the target corresponding to the current moment and the candidate predicted spatial information, where the candidate predicted spatial information is the predicted spatial information of the target in the target set whose similarity is not greater than the similarity threshold, that is, calculate the intersection-and-union ratio of the spatial information of the target whose similarity fails the similarity match with the target in the target set at the current moment. If there is a case where the intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identifier of the target corresponding to the candidate predicted spatial information that meets this condition is used as the identifier of the target corresponding to the three-dimensional spatial information at the current moment.
[0092] In the target tracking method provided by this embodiment, the server identifies the first feature of the target corresponding to each three-dimensional spatial information at the current moment and the second feature of the target corresponding to each predicted spatial information. If there is a target whose similarity between the first feature and the second feature is greater than the similarity threshold at the current moment, the identifier of the target corresponding to the second feature is used as the identifier of the target corresponding to the first feature; if there is a target whose similarity between the first feature and the second feature is not greater than the similarity threshold at the current moment, the intersection-and-union ratio between the three-dimensional spatial information corresponding to the target whose similarity is not greater than the similarity threshold among the targets corresponding to the current moment and the candidate predicted spatial information is calculated. If the intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identifier of the target corresponding to the candidate predicted spatial information is used as the identifier of the target corresponding to the three-dimensional spatial information. Thus, by dual matching of the target feature and the intersection-and-union ratio of the three-dimensional spatial information, the corresponding identifier is determined for the target detected at the current moment, which can greatly improve the accuracy of the determined identifier, thereby improving the accuracy of the target tracking result.
[0093] In one embodiment, the server compares the three-dimensional spatial information of each target with the predicted spatial information to determine another specific process of identifying the target in the detection area at the current moment. Figure 6 As shown, the above S103 may include:
[0094] S501, using a Kalman filter to predict the three-dimensional spatial information of the targets in the target set, to obtain predicted spatial information of each target in the target set; wherein the identification of the target corresponding to the predicted spatial information corresponds to the identification of the target in the target set.
[0095] Specifically, for each target in the target set, the server uses a Kalman filter to predict its three-dimensional spatial information, and predicts the predicted spatial information of each target at the current moment. The identifier of the target corresponding to each predicted spatial information is the identifier of the target in the corresponding target set.
[0096] S502, for each target at the current moment, calculate the intersection-and-union ratio between the three-dimensional spatial information and all the predicted spatial information. If there is three-dimensional spatial information whose intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identifier of the target corresponding to the matched predicted spatial information is used as the identifier of the target corresponding to the three-dimensional spatial information.
[0097] Specifically, for each target detected at the current moment, the server calculates the intersection-and-union ratio between its three-dimensional spatial information and all predicted spatial information, which can be the degree of overlap of the target detection frame size; if there is three-dimensional spatial information whose intersection-and-union ratio is greater than the intersection-and-union ratio threshold (such as 90%), the identification of the target corresponding to the predicted spatial information that matches it will be used as the identification of the target corresponding to the three-dimensional spatial information.
[0098] Of course, among all the targets at the current moment, there must be three-dimensional spatial information whose intersection-and-union ratio is not greater than the intersection-and-union ratio threshold, that is, the intersection-and-union ratio matching fails. In this case, the server can identify the third feature of the first target and the fourth feature of the second target. The first target is the target whose three-dimensional spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold in the target corresponding to the current moment, and the second target is the target whose predicted spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold in the target set, that is, the target whose intersection-and-union ratio matching fails in the current moment and the target set. Optionally, the point cloud re-identification network can be used to extract the third feature and the fourth feature respectively. Then calculate the similarity between the third feature and the fourth feature. If there is a similarity greater than the similarity threshold, the identifier of the corresponding second target is used as the identifier of the matched first target.
[0099] The target tracking method provided in this embodiment is that for each target at the current moment, the server calculates the intersection-and-union ratio between the three-dimensional spatial information and all the predicted spatial information. If there is three-dimensional spatial information whose intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identification of the target corresponding to the matched predicted spatial information is used as the identification of the target corresponding to the three-dimensional spatial information; if there is three-dimensional spatial information whose intersection-and-union ratio is not greater than the intersection-and-union ratio threshold, the third feature of the first target and the fourth feature of the second target are identified, and the similarity between the third feature and the fourth feature is calculated. If the similarity is greater than the similarity threshold, the identification of the second target is determined as the identification of the first target. Thus, by also matching the target feature and the intersection-and-union ratio of the three-dimensional spatial information, the corresponding identification is determined for the target detected at the current moment, which can greatly improve the accuracy of the determined identification, thereby improving the accuracy of the target tracking result.
[0100] In one embodiment, there may be a target whose identification is not determined at the current moment, such as a target that has just entered the detection area and does not exist in the target set. In this case, the server may assign a random identification to the target whose identification is not determined, and store the target and the random identification in the target set, and the random identification is different from the identifications of other targets in the target set. Thus, each target in the target set can be used to match the target in the detection area at the next moment to determine the identification. Optionally, for the target in the target set, there may also be a situation where the target leaves the detection area at the next moment, in which case the server may remove the target that is no longer in the detection area from the target set.
[0101] In the actual application of the multi-base station system, the above target tracking process can also be applied to the environment perception process of the multi-base station system. The environment perception process of the multi-base station system is introduced in detail below, where the single-base station perception data (including point cloud data) collected by the roadside base station (including the lidar sensor) in the multi-base station system is used as an example for explanation:
[0102] A. Obtain the single base station perception data of each roadside base station respectively, and perform spatiotemporal synchronization processing on the single base station perception data of each roadside base station according to the calibration parameters of the multi-base station system.
[0103] Among them, the single base station perception data can be the collected data within the current detection range collected by the roadside base station, such as point cloud data or camera data. The server can obtain the collected single base station perception data from each roadside base station respectively. Because each roadside base station has its own base station coordinate system, the collected single base station perception data is in its own base station coordinate system; in order to make the obtained single base station perception data under the same reference, so as to obtain the perception information of the global scene under the same reference, the server needs to first perform spatiotemporal synchronization processing on the perception data of each single base station. Specifically, the server can perform spatiotemporal synchronization processing on the single base station perception data of each roadside base station according to the calibration parameters of the multi-base station system. Optionally, the server can align the perception data of each single base station to the same time and space according to the calibration parameters (the calibration parameters may include parameters such as translation vectors and rotation matrices).
[0104] B, based on the single base station perception data after spatiotemporal synchronization processing, the target detection results of each roadside base station are obtained.
[0105] Specifically, the server can perform target detection on the obtained single base station perception data after spatiotemporal processing, and obtain information such as the position, speed, heading angle, acceleration, and category (such as pedestrians, vehicles, etc.) of the target within the detection range of each roadside base station as the target detection result. Optionally, the server can perform target detection on the single base station perception data based on a deep learning algorithm (such as a neural network) to obtain the target detection result.
[0106] C. Map the target detection results of each roadside base station to the global scene to generate perception information under the global scene; wherein the global scene is determined based on the perception range of the multi-base station system.
[0107] Specifically, the target detection results of each roadside base station are based on a single roadside base station. In order to obtain the target detection results of the entire multi-base station system, the server can map each target detection result to the global scene, that is, map the target detection results of each roadside base station to global perception data to obtain perception information under the global scene. Among them, the global scene is determined based on the perception range of the multi-base station system, so the server can "mark" each target detection result to the global scene to obtain perception information under the global scene. Therefore, the multi-base station system is used to cover the detection range of the entire traffic scene, and the perception information of the entire global scene is obtained based on the single base station perception data of a single roadside base station, that is, the perception information of the entire traffic scene is obtained, which greatly improves the scope of the perception environment.
[0108] To facilitate understanding of the above process of performing spatiotemporal synchronization processing on the single base station sensing data of each roadside base station (also referred to as a base station) according to the calibration parameters of the multi-base station system, the following is a detailed description of the process. The process may include the following steps:
[0109] A1, use a measuring instrument to measure the longitude and latitude information of each roadside base station, and determine the initial calibration parameters based on the longitude and latitude information.
[0110] The base station has a measuring instrument inside that can measure its own longitude and latitude information, which is the positioning information of the base station in the geodetic coordinate system. Each base station has its own base station coordinate system, and usually different base stations have different base station coordinate systems. Therefore, the single base station perception data collected by different base stations are located in different base station coordinate systems (the following uses point cloud data as an example for explanation, the following point cloud data is single base station perception data, the first point cloud data is the first single base station perception data, and the point cloud data to be registered is the perception data to be registered).
[0111] Specifically, after measuring the longitude and latitude information of each base station using the above-mentioned measuring instrument, the server can determine the initial calibration parameters according to the longitude and latitude information of each base station, and the initial calibration parameters are used to roughly align the point cloud data collected by each base station. Optionally, the server can determine the distance between the base stations according to the longitude and latitude information of each base station, and determine the initial calibration parameters according to the distance between the base stations and its own base station coordinate system; wherein the initial calibration parameters may include the translation vector and rotation matrix required for alignment.
[0112] A2, using the initial calibration parameters to process the single base station perception data of each roadside base station to obtain the first single base station perception data corresponding to each roadside base station.
[0113] Specifically, the server can process the point cloud data of each base station according to the initial calibration parameters determined above, synchronize the point cloud data of each base station to the same space, and obtain the first point cloud data corresponding to each base station. Optionally, the same space can be the base station coordinate system space of a base station among the base stations, or it can be a reference coordinate system space selected by the server (such as the geodetic coordinate system). Optionally, assuming that the translation vector in the initial calibration parameters is T and the rotation matrix is R, the server can use the relationship containing P0*R+T to transform the point cloud data P0 of the base station to obtain the first point cloud data.
[0114] A3, according to preset conditions, select the perception data to be registered corresponding to each roadside base station from the first single base station perception data corresponding to each roadside base station, and use the preset registration algorithm to process the perception data to be registered to obtain the calibration parameters of the multi-base station system; the preset conditions are used to characterize the data range of the selected perception data to be registered.
[0115] Among them, since the above-mentioned coarse alignment process is performed based on the latitude and longitude information of the base station, the accuracy of the latitude and longitude information depends on the hardware factors of the base station itself. Therefore, in order to further improve the synchronization accuracy of the cloud data of each base station in the same space, this embodiment performs a fine alignment process on the cloud data of each base station.
[0116] Specifically, for the first point cloud data corresponding to each base station, the server can select the point cloud data to be registered corresponding to each base station from each first point cloud data according to preset conditions, and the preset conditions are used to characterize the data range of the selected point cloud data to be registered. Optionally, the data within the range of Xm (such as 10m) from the point cloud center in the first point cloud data can be selected as the point cloud data to be registered, that is, only point cloud data with a larger point cloud density is selected to reduce the amount of data in the registration process. The server then uses a preset registration algorithm to process the selected point cloud data to be registered to obtain calibration parameters for precise registration of the multi-base station system, and then uses the calibration parameters to register the data to be registered. Optionally, the above-mentioned preset registration algorithm can be an iterative closest point algorithm (Iterative Closest Point, ICP), or other types of point cloud registration algorithms, which are not limited in this embodiment. Therefore, for the point cloud data collected by multiple base stations, this embodiment determines the precise calibration parameters of the multi-base station system through the two processes of coarse alignment and fine alignment, and then aligns the point cloud data of the base stations according to the calibration parameters, thereby greatly improving the spatial synchronization of the point cloud data of multiple base stations.
[0117] In one embodiment, there is a certain overlapping area in the detection ranges of the above-mentioned multiple base stations, and the multiple base stations can detect a common target in the overlapping area. In order to improve the uniformity of the detected common target information, the server can select the point cloud data corresponding to the overlapping area for registration. The above-mentioned process of selecting the perception data to be registered corresponding to each roadside base station from the first single base station perception data corresponding to each roadside base station according to the preset conditions can include the following steps:
[0118] A31, determining the overlapping area between the base stations according to the detection range of each base station.
[0119] A32: For each base station, obtain point cloud data corresponding to the overlapping area from the first point cloud data as the point cloud data to be registered.
[0120] Specifically, through the detection range of each base station, the server can determine the overlapping area between the base stations. For example, assuming that the detection range of base station A and base station B are both circles with a radius of 50m, and the distance between base station A and base station B is 80m, it can be determined that the overlapping area of the detection range of base station A and the detection range of base station B is an area with a width of 20m.
[0121] Then, for each base station, the server can obtain the part of the point cloud data corresponding to the overlapping area from the first point cloud data as the point cloud data to be registered. Optionally, the server can delete the point cloud data of the non-overlapping area in the first point cloud data to obtain the point cloud data to be registered. By selecting the point cloud data corresponding to the overlapping area between base stations as the point cloud data to be registered, firstly, the amount of point cloud data during registration can be reduced and the registration efficiency can be improved. Secondly, the uniformity of the common target information within the detection range of the base station can be improved.
[0122] In one embodiment, the process of determining the initial calibration parameters according to the latitude and longitude information may include the following steps:
[0123] A11, obtaining original calibration parameters according to the latitude and longitude information of each base station.
[0124] A12, using common targets within the detection range of each base station to evaluate the original calibration parameters, and obtaining initial calibration parameters according to the evaluation results.
[0125] Specifically, the process of obtaining the original calibration parameters according to the longitude and latitude information of each base station can refer to the description of the above embodiment, which will not be repeated here. After obtaining the original calibration parameters, the server further evaluates the original calibration parameters to obtain calibration parameters with higher precision, thereby improving the accuracy of the rough registration result. After obtaining the original calibration parameters, the server can use the original calibration parameters to process the point cloud data of each base station, and then perform target detection on the processed point cloud data, and use the common target within the detection range of each base station to evaluate the above original calibration parameters to obtain the initial calibration parameters. Optionally, the server can calculate the distance from the common target to each base station respectively, and evaluate the original calibration parameters according to the difference of each distance. If the distance difference is less than the preset difference threshold, the original calibration parameters are used as the initial calibration parameters. If the distance error is not less than the difference threshold, the longitude and latitude information of each base station needs to be measured again using the measuring instrument, and the original calibration parameters are obtained again according to the longitude and latitude information, and the execution is repeated until the distance difference from the common target to each base station is less than the difference threshold. Optionally, the server may also evaluate the original calibration parameters according to the differences between the coordinates of the common target detected by each base station to obtain the initial calibration parameters.
[0126] In another achievable manner, the server may also obtain the detection frame of the common target within the detection range of each base station, and determine the overlap between the detection frames of the common target; if the overlap between the detection frames is greater than the overlap threshold, the original calibration parameters are used as the initial calibration parameters. Optionally, a target detection algorithm based on deep learning may be used to perform target detection on each processed point cloud data, and determine the detection frame of the common target within the detection range of each base station. The detection frame may be the smallest three-dimensional frame that can surround the target, and has information such as length, width, and height. Then, the overlap between the detection frames is determined based on the detection frames of the common target. If the overlap is greater than a preset overlap threshold (such as 90%), it means that the accuracy of the original calibration parameters obtained is already high, and the original calibration parameters can be used as the initial calibration parameters; if the overlap is not greater than the overlap threshold, it means that the accuracy of the original calibration parameters obtained is still low, and the longitude and latitude information of each base station needs to be measured again using the measuring instrument, and the original calibration parameters are re-obtained based on the longitude and latitude information, and this is repeated until the overlap between the detection frames of the common target is greater than the overlap threshold. Therefore, by performing the fine registration process under the premise of ensuring that the rough registration has a certain accuracy, the accuracy of point cloud registration can be further improved.
[0127] In one embodiment, the server may also use the latitude and longitude information of the target within the detection range of the base station and the latitude and longitude information of the base station to jointly determine the above original calibration parameters. The above A11 process may include:
[0128] A111, obtain the longitude and latitude information of the target within the detection range of each base station.
[0129] A112, determining the angle and distance between the base stations according to the latitude and longitude information of the base stations and the latitude and longitude information of the target.
[0130] Specifically, the longitude and latitude information of the target within the detection range of the base station can also be the position information in the geodetic coordinate system, which can be measured using a measuring instrument inside the base station; then the geodetic coordinate system is selected as the reference coordinate system, and the server determines the angle between the preset coordinate axis in each base station coordinate system and the reference direction in the geodetic coordinate system based on the longitude and latitude information of each base station, the longitude and latitude information of the target within the detection range of each base station, and the base station coordinate system of each base station, and determines the angle between each base station based on the angle between the preset coordinate axis in the coordinate system of each base station and the reference direction.
[0131] Exemplarily, the base station coordinate system may be a three-dimensional coordinate system including an X-axis, a Y-axis and a Z-axis, and the reference direction may be the true north direction. The server may determine the angle between the Y-axis in the base station coordinate system and the true north direction in the geodetic coordinate system. Assuming that the longitude of base station A is Aj and the latitude is Aw, and the longitude of the target is Bj and the latitude is Bw, the server may optionally determine the angle between the Y-axis in the base station coordinate system and the true north direction in the geodetic coordinate system. Of course, the server can also calculate a reference angle F based on the relationship The reference angle is calculated by using other relationship formulas. If the target is in the first quadrant of the base station coordinate system and the positive half axis of the Y axis, then the angle Azimuth between the Y axis and the due north direction in the base station coordinate system is F; if the target is in the second quadrant of the base station coordinate system, then Azimuth is 360°+A; if the target is in the third quadrant, the fourth quadrant and the negative half axis of the Y axis of the base station coordinate system, then Azimuth is 180°+A. Thus, the angle Azimuth1 between the Y axis in the base station A coordinate system and the due north direction in the geodetic coordinate system and the angle Azimuth2 between the Y axis in the base station B coordinate system and the due north direction in the geodetic coordinate system can be calculated. By performing the difference operation between the angle Azimuth1 and the angle Azimuth2, the angle ΔA between base station A and base station B is obtained as Azimuth1-Azimuth2.
[0132] In addition, the server can also determine the distance between two base stations according to the longitude and latitude information of each base station, such as by calculating the longitude difference between the two base stations and the latitude difference between the two base stations, and then calculating the distance between the two base stations according to the longitude and latitude information of each base station. The distance between the two base stations is determined by the distance formula, where ΔJ is the longitude difference and ΔW is the latitude difference; optionally, the server may also directly use ΔJ as the distance between the two base stations in the longitude direction and ΔW as the distance in the latitude direction.
[0133] A113, determining original calibration parameters according to the angles and distances between the base stations.
[0134] Specifically, the server can use the angle between the base stations as a rotation matrix, the distance between the base stations as a translation vector, and use the rotation matrix and translation vector as the original calibration parameters. Thus, the original calibration parameters are determined based on the latitude and longitude information of the base station and the latitude and longitude information of the target, which can improve the accuracy of the original calibration parameters obtained, thereby improving the spatial synchronization of the cloud data of multiple base station sites.
[0135] To facilitate understanding of the process of processing the point cloud data to be registered using the preset registration algorithm, this embodiment is explained using two base stations. Assuming that the point cloud data to be registered of one base station is the second point cloud data, and the point cloud data to be registered of the other base station is the third point cloud data, the process of processing the perception data to be registered using the preset registration algorithm to obtain the calibration parameters of the multi-base station system may include:
[0136] A33, obtaining matching point pairs in the second point cloud data and the third point cloud data according to the distance values between the point cloud points of the second point cloud data and the point cloud points of the third point cloud data.
[0137] Specifically, assuming that the second point cloud data is P0 and the third point cloud data is Q, for each point cloud point in the point cloud data P0, the point cloud point closest to the point cloud point P0 is searched from the point cloud data Q to form multiple point pairs.
[0138] A34, using an error function to calculate the mean square error of each point pair, determining the rotation transformation parameter corresponding to the minimum mean square error value, and using the rotation transformation parameter to process the second point cloud data and the third point cloud data to obtain the first candidate point cloud data and the second candidate point cloud data.
[0139] Specifically, each point pair includes a point cloud point of P0 and a point cloud point of Q (p i ,q i ), wherein the correspondences in the initial point pairs are not necessarily correct, and incorrect correspondences will affect the final registration results. In this embodiment, the direction vector threshold can also be used to eliminate incorrect point pairs. Then, the error function is used to calculate the mean square error of the multiple point pairs, and the rotation transformation parameters when the mean square error is the smallest are determined. The second point cloud data P0 is converted into the first candidate point cloud data P1 using the rotation transformation parameters. It should be noted that at this time, the third point cloud data Q does not need to be converted, and the third point cloud data Q is directly used as the second candidate point cloud data; optionally, the expression of the error function can be Where n is the number of point pairs, R is the rotation matrix in the rotation transformation parameters, and t is the translation vector in the rotation transformation parameters. The values of R and t that minimize the mean square error are currently determined, and according to p i '={Rp i +t,p i∈P0} converts the point cloud data P0 into P1.
[0140] A35, calculating the mean square error between the first candidate point cloud data and the second candidate point cloud data, if the mean square error is less than the error threshold, using the rotation transformation parameter as a calibration parameter of the multi-base station system.
[0141] Then, the mean square error between the first candidate point cloud data P1 and the second candidate point cloud data Q is calculated. Optionally, the mean square error can be calculated by The mean square error is calculated by the relationship between i 'For q i p at the same point pair i If the mean square error is less than the error threshold, the rotation conversion parameters obtained above are used as calibration parameters of the multi-base station system. If the mean square error is not less than the preset error, the point pairs between the point cloud data P1 and Q are determined again, and the process of calculating the mean square error of the point pairs is re-executed until the mean square error is less than the preset error or the number of iterations reaches the preset number. The calibration parameters of the fine alignment process obtained by iteration can greatly improve the accuracy of the obtained calibration parameters.
[0142] In one embodiment, after the server obtains the point cloud data to be registered corresponding to the above-mentioned base stations (such as the point cloud data corresponding to the overlapping area), it can also determine the data to be eliminated in the point cloud data to be registered whose data accuracy is not greater than the accuracy threshold, such as some data with unclear features, based on the data accuracy and accuracy threshold of the point cloud data to be registered, and eliminate the data to be eliminated from the point cloud data to be registered. The server can then process the point cloud data to be registered using a preset registration algorithm to obtain the calibration parameters of the multi-base station system. In this way, the data with higher accuracy in each point cloud data to be registered can be retained, providing high-precision data for the subsequent fine registration process, so as to further improve the accuracy of the point cloud registration results. Optionally, the server can also filter out ground points in the point cloud data to be registered, that is, filter out the ground point data in the point cloud data to be registered, so as to reduce the influence of the ground points on the data registration process.
[0143] In one embodiment, in addition to spatial synchronization of single base station sensing data of multiple base stations, time synchronization can also be achieved. Optionally, the process of time synchronization may include: receiving the base station time axis sent by each base station; synchronizing the time axis of each base station to the same time axis according to the base station time axis of each base station and the reference time axis. Specifically, first select a reference time axis, optionally, the reference time axis can be a GPS time axis; then calculate the time difference ΔT1, ΔT2, etc. between the base station time axis of each base station and the reference time axis. If two base stations are taken as an example, the difference between ΔT1 and ΔT2 is used as the time difference between the base station time axis of the first base station and the base station time axis of the second base station, then according to the time difference, the second base station can synchronize its own base station time axis to the base station time axis of the first base station. In this way, time synchronization between base stations is achieved.
[0144] In one embodiment, the specific process of obtaining the target detection results of each roadside base station based on the single base station sensing data after time-space synchronization processing is involved. Optionally, the above step B may include:
[0145] B1: If there is a perception overlap area among the roadside base stations, data enhancement processing is performed on the single base station perception data corresponding to the perception overlap area to obtain enhanced single base station perception data.
[0146] B2, use the target detection algorithm to process and enhance the single base station perception data to obtain the target detection results of each roadside base station.
[0147] Specifically, Figure 1 In the scene diagram shown, the detection range of the roadside base station will have a perception overlap area, so the single base station perception data collected by each roadside base station will also have overlapping data. For example, the detection area of base station A and base station B are both circular with a radius of 50m, and the distance between base station A and base station B is 80m. It can be determined that the perception overlap area width of the detection area of base station A and base station B is 20m, and the single base station perception data corresponding to the perception overlap area is the collected data corresponding to the 20m. The server can then perform data enhancement processing on this part of the single base station perception data to obtain enhanced single base station perception data. For example, densification processing is performed. If the single base station perception data is point cloud data, an interpolation algorithm can be used to increase the point cloud density of this part of the data to enhance the feature dimension of the target therein; if the single base station perception data is camera data (image data), a difference algorithm can be used to increase the pixel information dimension, thereby obtaining enhanced single base station perception data.
[0148] Then, the server can process the enhanced single base station perception data using a target detection algorithm, which can be a detection algorithm based on deep learning, such as an algorithm based on a neural network model, to obtain target detection results of each roadside base station after detecting the enhanced single base station perception data. By enhancing the single base station perception data, the accuracy of the target detection results can be greatly improved.
[0149] Optionally, when the above-mentioned single-base station perception data is point cloud data, each roadside base station can also share the target detection results in the perception overlap area detected by other base stations based on the enhanced single-base station perception data. For example, base station A detects a part of a target in the overlap area (such as the front of a car) based on the enhanced single-base station perception data, and base station B detects another part of the target (such as the body of the car) based on the enhanced single-base station perception data. Then base station B can share the front information of base station A, so that the obtained target detection results are complete, and the detection capability of base station B is also improved.
[0150] In one embodiment, the perception information in the global scene includes the target movement trajectory in the global scene, that is, the target tracking process is realized. Optionally, the above C may include: associating and matching the target detection result mapped to the global scene with the prior target detection result to obtain the target movement trajectory in the global scene; wherein the prior target detection result includes the target detection result corresponding to the moment before the current moment.
[0151] Specifically, the target detection result can include the position of the target at the current moment, so the previous target detection result also includes the position of the target at the moment before the current moment; the server can also assign a target identifier to the detected target to distinguish different targets, and the same target uses the same target identifier. Therefore, the server can associate the target detection result with the previous target detection result through the target identifier and the target's position to obtain the target movement trajectory in the global scene.
[0152] It should be noted that the server needs to first determine whether the target in the current target detection result and the target in the previous target detection result are the same target before assigning them the same target identifier to achieve the target tracking process. The following is a detailed description of the specific process of achieving target tracking:
[0153] In one embodiment, the target detection result may include the position of the target, the speed of the target and the heading angle of the target, and the previous target detection result also includes prediction information of the target; optionally, the above step C may include:
[0154] C1, based on the target detection results of each roadside base station and the relative positions between the roadside base stations, the position and direction of the corresponding target after a preset time period are calculated to obtain the prediction information of each target.
[0155] Specifically, the server can predict the position and direction of the target after a preset time (which can be multiple preset time periods) based on the position, speed and heading angle of the target at the current moment, as well as the relative positions between the roadside base stations. For example, the current time is 16:00:00, and the server predicts the target at 16:00:05, 16:00:10, 16:00:15, 16:00:20 and other ten subsequent time periods based on the distance and relative angle between base station A and base station B. It should be noted that the number of predicted subsequent time periods can be set according to the needs of the actual scenario. Optionally, the server can include The position of the target after the time interval Δt is calculated by the relationship of (X i , Y i ) is the longitude and latitude of the target at the current moment, V i is the speed of the target at the current moment, ψ i is the heading angle of the target at the current moment; according to the i +a i The velocity of the target at the subsequent moment after the Δt time interval is calculated by the relationship between a i is the acceleration of the target at the current moment.
[0156] In addition, each roadside base station will continuously collect data within a preset time period, and predict the target detection results collected at each moment, and the prediction information obtained at the latter moment will cover the prediction information obtained at the previous moment. For example, at 16:00:00, the prediction information of the target at 16:00:05, 16:00:10, 16:00:15, 16:00:20, etc. is predicted; if the target is still detected at 16:00:05, the prediction information of the target at 16:00:10, 16:00:15, 16:00:20, 16:00:25, etc. will continue to be predicted, and the newly predicted prediction information at 16:00:10, 16:00:15, 16:00:20 will cover the prediction information of the first prediction.
[0157] C2, according to the prediction information of each target, the target detection results in the global scene are associated and matched to obtain the target movement trajectory in the global scene.
[0158] Specifically, the server can match the prediction information of each target with the target detection result at the current moment. If they match, it means that the target is still in the detection area of the roadside base station at the current moment. The target identifier of the target corresponding to the prediction information is assigned to the target corresponding to the target detection result, and the moving trajectory of the target is obtained based on the position of the target at the previous moment and the position at the current moment.
[0159] Optionally, the server can also determine whether there are safety hazards in the global scenario based on the obtained prediction information; if there are safety hazards, a safety warning message is output. Optionally, the server can obtain the prediction information of multiple targets, and if the position information in the prediction information of multiple targets overlaps, it is determined that there are safety hazards in the global scenario. For example, if there is overlapping position information in the prediction information of two or more targets, it means that the two or more targets may collide, that is, there is a safety hazard, and a safety warning message can be output.
[0160] Optionally, the target detection result may also include size information of the target, and the process of achieving target tracking based on the target detection results (including three-dimensional spatial information) in the global scene and the prediction information of each target (including predicted spatial information) can be referred to the description in the above embodiments.
[0161] The above describes in detail the process in which the server tracks the target in the detection area to obtain the target movement trajectory in the entire multi-base station system. The following takes a roadside base station in the multi-base station system as an example to introduce the detection and tracking process of the roadside base station.
[0162] In one embodiment, the above step C2 may include:
[0163] C21, determining a target roadside base station from multiple roadside base stations based on the location information in the candidate prediction information; wherein the candidate prediction information is the prediction information of any one of the targets based on the current moment.
[0164] Specifically, the server can know where the target will reach based on the location information in the candidate prediction information, and can know which roadside base station the location is within the detection range of based on the location information and the detection range of the roadside base station, and then use the roadside base station as the target roadside base station.
[0165] C22, after a preset time, obtain the current single base station perception data of the target roadside base station, and perform target detection on the current single base station perception data to obtain the current target detection result of the target roadside base station.
[0166] C23, if the current target detection result matches the candidate prediction information, the target corresponding to the candidate prediction information is associated with the target in the current target detection result.
[0167] Specifically, after the target roadside base station is determined, the current single base station perception data of the target roadside base station after a preset time period can be obtained, and the current single base station perception data can be subjected to target detection to obtain the current target detection result. The target detection result is then matched with the above-mentioned candidate prediction information. The matching process can refer to the description of the above-mentioned embodiment (such as according to target features, detection frame intersection and union ratio, etc.). If the match is successful, the target corresponding to the candidate prediction information is associated with the target in the current target detection result, that is, the target identifier corresponding to the candidate prediction information is assigned to the target in the current target detection result.
[0168] Optionally, if the current target detection result does not match the candidate prediction information, the target roadside base station has not detected the target corresponding to the candidate prediction information, then it is determined whether the target corresponding to the current target detection result is a new target. For example, if the target has not been detected by the target roadside base station before, it is considered to be a new target, then the perception information of the new target is added to the perception information of the global scene to improve the comprehensiveness of the global scene perception information.
[0169] Optionally, the server can also obtain the location information in the candidate prediction information. If the target roadside base station does not detect the current target detection result corresponding to the location information, that is, the target roadside base station does not detect the target at the predicted location, indicating that the perception ability of the target roadside base station at this location is weak; then the server can determine the target subsequent time at which the target detection result matches the prediction information in the subsequent time, that is, determine the time when the target roadside base station detects the target; and then use the candidate prediction information corresponding to before the target subsequent time as the target detection result of the target roadside base station.
[0170] Exemplarily, for the current target detection result at 16:00:05, the server can match the detection result with the candidate prediction information. If the match is successful, the target corresponding to the candidate prediction information is the target detected by the target roadside base station at the current time, and the time (16:00:05) is the subsequent time of the target, that is, it can be considered that the target was detected by the target roadside base station at (16:00:05). If there is no matching pose data, it means that the target roadside base station has not detected the target at 16:00:05. The server then compares the current target detection result at 16:00:10 with the candidate prediction information at 16:00:10. If they match at this time, the target corresponding to the candidate prediction information is the target detected by the target roadside base station at the current time, and the time (16:00:10) is the target subsequent time; and the candidate prediction information before (16:00:10) is used as the target detection result of the target roadside base station, so that even if the target roadside base station does not detect the target, the corresponding target detection result can be obtained, which improves the perception ability of the target roadside base station. If there is still no matching pose data, the candidate prediction information at the next subsequent time is compared until the target subsequent time is determined.
[0171] In order to better understand the entire process of the above target tracking method, the method is introduced again in the form of an overall embodiment. Figure 7 As shown, the method includes:
[0172] S601, obtaining point cloud data obtained by scanning a plurality of laser radar detection areas;
[0173] S602, according to a preset transformation matrix, transforming the second point cloud data into a reference coordinate system where the first point cloud data is located, and fusing the transformed second point cloud data with the first point cloud data to obtain fused point cloud data;
[0174] S603, performing target detection processing on the fused point cloud data to obtain the three-dimensional spatial information of each target in the detection area at the current moment;
[0175] S604, for each target corresponding to the three-dimensional spatial information at the current moment, identifying a first feature of the target; for each target corresponding to the predicted spatial information, identifying a second feature of the target;
[0176] S605, determining whether the similarity between the first feature and the second feature at the current moment is greater than a target similarity threshold;
[0177] S606: If yes, use the identifier of the target corresponding to the second feature as the identifier of the target corresponding to the first feature;
[0178] S607, if not, calculating the intersection-over-union ratio between the three-dimensional spatial information corresponding to the target whose similarity is not greater than the similarity threshold among the targets corresponding to the current moment and the candidate prediction spatial information;
[0179] S608, if the intersection-over-union ratio is greater than the intersection-over-union ratio threshold, taking the identifier of the target corresponding to the candidate prediction space information as the identifier of the target corresponding to the three-dimensional space information;
[0180] S609: If the intersection-over-union ratio is not greater than the intersection-over-union ratio threshold, a random identifier is assigned to the target with an undetermined identifier.
[0181] The implementation process of each step can refer to the description of the above embodiment. The implementation principle and technical effect are similar and will not be repeated here.
[0182] It should be understood that although Figure 2-Figure 7 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2-Figure 7 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0183] In one embodiment, Figure 8 As shown, a target tracking device is provided, including: an acquisition module 21, a determination module 22 and a comparison module 23.
[0184] Specifically, the acquisition module 21 is used to acquire point cloud data obtained by scanning a detection area by multiple laser radars; the multiple laser radars are arranged at different positions in the detection area;
[0185] A determination module 22 is used to determine the three-dimensional spatial information of each target in the detection area at the current moment according to the point cloud data obtained by scanning multiple laser radars; the three-dimensional spatial information includes the position information and size information of the target;
[0186] The comparison module 23 is used to compare the three-dimensional spatial information of each target in the detection area at the current moment with the predicted spatial information of each target in the target set, and determine the corresponding identification for the target whose three-dimensional spatial information matches the predicted spatial information to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
[0187] The target tracking device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.
[0188] In one embodiment, the above-mentioned determination module 22 is specifically used to select the coordinate system where the first point cloud data is located from multiple point cloud data obtained by multiple laser radar scans as the reference coordinate system, and according to a preset transformation matrix, convert the second point cloud data to the reference coordinate system where the first point cloud data is located, and fuse the converted second point cloud data with the first point cloud data to obtain fused point cloud data; wherein the second point cloud data is other point cloud data in the multiple point cloud data except the first point cloud data, and one laser radar scan obtains one point cloud data; target detection processing is performed on the fused point cloud data to obtain three-dimensional spatial information of each target in the detection area at the current moment.
[0189] In one embodiment, the above-mentioned determination module 22 is specifically used to perform target detection processing on the point cloud data of multiple laser radars respectively to obtain the three-dimensional spatial information of the target in each point cloud data; select the coordinate system where the first three-dimensional spatial information is located from the multiple three-dimensional spatial information of the multiple point cloud data as the reference coordinate system, and according to a preset transformation matrix, transform the second three-dimensional spatial information to the reference coordinate system where the first three-dimensional spatial information is located, and fuse the converted second three-dimensional spatial information and the first three-dimensional spatial information to obtain fused three-dimensional spatial information; wherein the second three-dimensional spatial information is other three-dimensional spatial information in the multiple three-dimensional spatial information that corresponds to different point cloud data of the first three-dimensional spatial information, and one point cloud data corresponds to multiple three-dimensional spatial information; perform de-redundancy processing on the fused three-dimensional spatial information to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0190] In one embodiment, the determination module 22 is specifically configured to perform redundancy removal processing on the fused three-dimensional spatial information by using a non-maximum suppression algorithm to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0191] In one embodiment, the comparison module 23 is specifically used to identify the first feature of the target corresponding to each three-dimensional spatial information at the current moment; identify the second feature of the target corresponding to each predicted spatial information; if there is a target at the current moment whose similarity between the first feature and the second feature is greater than the similarity threshold, the identifier of the target corresponding to the second feature is used as the identifier of the target corresponding to the first feature.
[0192] In one embodiment, the comparison module 23 is also used to calculate the intersection-and-union ratio between the three-dimensional spatial information corresponding to the target whose similarity is not greater than the similarity threshold in the target corresponding to the current moment and the candidate prediction spatial information if there is a target whose similarity between the first feature and the second feature is not greater than the similarity threshold at the current moment; wherein the candidate prediction spatial information is the prediction spatial information of the target whose similarity is not greater than the similarity threshold in the target set; if the intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identifier of the target corresponding to the candidate prediction spatial information is used as the identifier of the target corresponding to the three-dimensional spatial information.
[0193] In one embodiment, the comparison module 23 is specifically used to use a Kalman filter to predict the three-dimensional spatial information of the targets in the target set to obtain the predicted spatial information of each target in the target set; wherein the identification of the target corresponding to the predicted spatial information corresponds to the identification of the target in the target set; for each target at the current moment, the intersection-and-union ratio between the three-dimensional spatial information and all the predicted spatial information is calculated, and if there is three-dimensional spatial information with an intersection-and-union ratio greater than an intersection-and-union ratio threshold, the identification of the target corresponding to the matched predicted spatial information is used as the identification of the target corresponding to the three-dimensional spatial information.
[0194] In one embodiment, the comparison module 23 is also used to identify the third feature of the first target and the fourth feature of the second target if there is three-dimensional spatial information whose intersection-and-union ratio is not greater than the intersection-and-union ratio threshold; wherein the first target is a target whose three-dimensional spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the targets corresponding to the current moment, and the second target is a target whose predicted spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the target set; calculate the similarity between the third feature and the fourth feature, and if the similarity is greater than the similarity threshold, determine the identifier of the second target as the identifier of the first target.
[0195] In one embodiment, the above-mentioned device also includes a random assignment module, which is used to assign a random identifier to the target with an undetermined identifier if there is a target with an undetermined identifier at the current moment, and store the target with an undetermined identifier and the random identifier in a target set; wherein the random identifier is different from the identifiers of other targets in the target set.
[0196] The target tracking device provided in this embodiment can execute the above method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.
[0197] For the specific definition of the target tracking device, please refer to the definition of the target tracking method above, which will not be repeated here. Each module in the above-mentioned target tracking device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0198] In one embodiment, a server is provided, the internal structure diagram of which can be as follows: Fig. 9 As shown. The server includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the server is used to provide computing and control capabilities. The memory of the server includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the server is used to store the point cloud data scanned by the laser radar and the three-dimensional spatial information of the target in the target set at the last moment. The network interface of the server is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a target tracking method is implemented.
[0199] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a partial structure related to the solution of the present application, and does not constitute a limitation on the server to which the solution of the present application is applied. The specific server may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0200] In one embodiment, a server is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the following steps are implemented:
[0201] Acquire point cloud data obtained by scanning a detection area with multiple laser radars; multiple laser radars are arranged at different positions in the detection area;
[0202] Determine the three-dimensional spatial information of each target in the detection area at the current moment based on the point cloud data obtained by multiple laser radar scans; the three-dimensional spatial information includes the location information and size information of the target;
[0203] The three-dimensional spatial information of each target in the detection area at the current moment is compared with the predicted spatial information of each target in the target set, and a corresponding identifier is determined for the target whose three-dimensional spatial information matches the predicted spatial information to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
[0204] The implementation principle and technical effects of the server provided in this embodiment are similar to those of the above method embodiments, and will not be repeated here.
[0205] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0206] A coordinate system where the first point cloud data is located is selected as a reference coordinate system from a plurality of point cloud data obtained by scanning a plurality of laser radars, and the second point cloud data is converted to the reference coordinate system where the first point cloud data is located according to a preset conversion matrix, and the converted second point cloud data is fused with the first point cloud data to obtain fused point cloud data; wherein the second point cloud data is other point cloud data in the plurality of point cloud data except the first point cloud data, and one point cloud data is obtained by one laser radar scan;
[0207] The fused point cloud data is processed for target detection to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0208] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0209] Perform target detection processing on the point cloud data of multiple laser radars respectively to obtain the three-dimensional spatial information of the target in each point cloud data;
[0210] A coordinate system where the first three-dimensional spatial information is located is selected as a reference coordinate system from multiple three-dimensional spatial information of multiple point cloud data, and second three-dimensional spatial information is converted to the reference coordinate system where the first three-dimensional spatial information is located according to a preset conversion matrix, and the converted second three-dimensional spatial information and the first three-dimensional spatial information are fused to obtain fused three-dimensional spatial information; wherein the second three-dimensional spatial information is other three-dimensional spatial information in the multiple three-dimensional spatial information corresponding to different point cloud data of the first three-dimensional spatial information, and one point cloud data corresponds to multiple three-dimensional spatial information;
[0211] The fused three-dimensional spatial information is de-redundanted to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0212] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0213] The non-maximum suppression algorithm is used to remove redundancy from the fused three-dimensional spatial information to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0214] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0215] For each target corresponding to each three-dimensional space information at the current moment, identify the first feature of the target;
[0216] For each target corresponding to the predicted spatial information, identifying a second feature of the target;
[0217] If there is a target whose similarity between the first feature and the second feature is greater than the similarity threshold at the current moment, the identifier of the target corresponding to the second feature is used as the identifier of the target corresponding to the first feature.
[0218] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0219] If there is a target whose similarity between the first feature and the second feature is not greater than the similarity threshold at the current moment, calculate the intersection-over-union ratio between the three-dimensional spatial information corresponding to the target whose similarity is not greater than the similarity threshold among the targets corresponding to the current moment and the candidate prediction spatial information; wherein the candidate prediction spatial information is the prediction spatial information of the target whose similarity is not greater than the similarity threshold in the target set;
[0220] If the intersection-over-union ratio is greater than the intersection-over-union ratio threshold, the identifier of the target corresponding to the candidate prediction space information is used as the identifier of the target corresponding to the three-dimensional space information.
[0221] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0222] A Kalman filter is used to predict the three-dimensional spatial information of the targets in the target set to obtain predicted spatial information of each target in the target set; wherein the identification of the target corresponding to the predicted spatial information corresponds to the identification of the target in the target set;
[0223] For each target at the current moment, the intersection-and-union ratio between the three-dimensional spatial information and all the predicted spatial information is calculated. If there is three-dimensional spatial information whose intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identifier of the target corresponding to the matched predicted spatial information is used as the identifier of the target corresponding to the three-dimensional spatial information.
[0224] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0225] If there is three-dimensional spatial information whose intersection-and-union ratio is not greater than the intersection-and-union ratio threshold, identify the third feature of the first target and the fourth feature of the second target; wherein the first target is a target whose three-dimensional spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the targets corresponding to the current moment, and the second target is a target whose predicted spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the target set;
[0226] The similarity between the third feature and the fourth feature is calculated, and if the similarity is greater than a similarity threshold, the identifier of the second target is determined as the identifier of the first target.
[0227] In one embodiment, when the processor executes the computer program, the processor further implements the following steps:
[0228] If there is a target with an undetermined identification at the current moment, a random identification is assigned to the target with an undetermined identification, and the target with an undetermined identification and the random identification are stored in a target set; wherein the random identification is different from identifications of other targets in the target set.
[0229] In one embodiment, a computer readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0230] Acquire point cloud data obtained by scanning a detection area with multiple laser radars; multiple laser radars are arranged at different positions in the detection area;
[0231] Determine the three-dimensional spatial information of each target in the detection area at the current moment based on the point cloud data obtained by multiple laser radar scans; the three-dimensional spatial information includes the location information and size information of the target;
[0232] The three-dimensional spatial information of each target in the detection area at the current moment is compared with the predicted spatial information of each target in the target set, and a corresponding identifier is determined for the target whose three-dimensional spatial information matches the predicted spatial information to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
[0233] The computer-readable storage medium provided in this embodiment has similar implementation principles and technical effects to those of the above method embodiments, and will not be described in detail here.
[0234] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0235] A coordinate system where the first point cloud data is located is selected as a reference coordinate system from a plurality of point cloud data obtained by scanning a plurality of laser radars, and the second point cloud data is converted to the reference coordinate system where the first point cloud data is located according to a preset conversion matrix, and the converted second point cloud data is fused with the first point cloud data to obtain fused point cloud data; wherein the second point cloud data is other point cloud data in the plurality of point cloud data except the first point cloud data, and one point cloud data is obtained by one laser radar scan;
[0236] The fused point cloud data is processed for target detection to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0237] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0238] Perform target detection processing on the point cloud data of multiple laser radars respectively to obtain the three-dimensional spatial information of the target in each point cloud data;
[0239] A coordinate system where the first three-dimensional spatial information is located is selected as a reference coordinate system from multiple three-dimensional spatial information of multiple point cloud data, and second three-dimensional spatial information is converted to the reference coordinate system where the first three-dimensional spatial information is located according to a preset conversion matrix, and the converted second three-dimensional spatial information and the first three-dimensional spatial information are fused to obtain fused three-dimensional spatial information; wherein the second three-dimensional spatial information is other three-dimensional spatial information in the multiple three-dimensional spatial information corresponding to different point cloud data of the first three-dimensional spatial information, and one point cloud data corresponds to multiple three-dimensional spatial information;
[0240] The fused three-dimensional spatial information is de-redundanted to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0241] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0242] The non-maximum suppression algorithm is used to remove redundancy from the fused three-dimensional spatial information to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
[0243] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0244] For each target corresponding to each three-dimensional space information at the current moment, identify the first feature of the target;
[0245] For each target corresponding to the predicted spatial information, identifying a second feature of the target;
[0246] If there is a target whose similarity between the first feature and the second feature is greater than the similarity threshold at the current moment, the identifier of the target corresponding to the second feature is used as the identifier of the target corresponding to the first feature.
[0247] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0248] If there is a target whose similarity between the first feature and the second feature is not greater than the similarity threshold at the current moment, calculate the intersection-over-union ratio between the three-dimensional spatial information corresponding to the target whose similarity is not greater than the similarity threshold among the targets corresponding to the current moment and the candidate prediction spatial information; wherein the candidate prediction spatial information is the prediction spatial information of the target whose similarity is not greater than the similarity threshold in the target set;
[0249] If the intersection-over-union ratio is greater than the intersection-over-union ratio threshold, the identifier of the target corresponding to the candidate prediction space information is used as the identifier of the target corresponding to the three-dimensional space information.
[0250] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0251] A Kalman filter is used to predict the three-dimensional spatial information of the targets in the target set to obtain predicted spatial information of each target in the target set; wherein the identification of the target corresponding to the predicted spatial information corresponds to the identification of the target in the target set;
[0252] For each target at the current moment, the intersection-and-union ratio between the three-dimensional spatial information and all the predicted spatial information is calculated. If there is three-dimensional spatial information whose intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identifier of the target corresponding to the matched predicted spatial information is used as the identifier of the target corresponding to the three-dimensional spatial information.
[0253] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0254] If there is three-dimensional spatial information whose intersection-and-union ratio is not greater than the intersection-and-union ratio threshold, identify the third feature of the first target and the fourth feature of the second target; wherein the first target is a target whose three-dimensional spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the targets corresponding to the current moment, and the second target is a target whose predicted spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the target set;
[0255] The similarity between the third feature and the fourth feature is calculated, and if the similarity is greater than a similarity threshold, the identifier of the second target is determined as the identifier of the first target.
[0256] In one embodiment, when the computer program is executed by a processor, the following steps are also implemented:
[0257] If there is a target with an undetermined identification at the current moment, a random identification is assigned to the target with an undetermined identification, and the target with an undetermined identification and the random identification are stored in a target set; wherein the random identification is different from identifications of other targets in the target set.
[0258] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0259] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0260] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A target tracking method, characterized in that: The method comprises: Acquire point cloud data obtained by scanning a detection area with multiple laser radars; the multiple laser radars are arranged at different positions of the detection area; If the scanning areas of two laser radars among the multiple laser radars have overlapping parts, the point cloud data of the overlapping parts are superimposed to obtain the point cloud data obtained by the superimposed multiple laser radar scans, so as to increase the point cloud density of the overlapping parts; Determine the three-dimensional spatial information of each target in each of the point cloud data according to the point cloud data obtained by scanning the plurality of laser radars; the three-dimensional spatial information includes the position information and size information of the target; For each target corresponding to the three-dimensional spatial information at the current moment, the first feature of the target is identified; for each target corresponding to the predicted spatial information, the second feature of the target is identified; if there is a target at the current moment whose similarity between the first feature and the second feature is greater than a similarity threshold, the identifier of the target corresponding to the second feature is used as the identifier of the target corresponding to the first feature to complete target tracking; The predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the last moment.
2. The method according to claim 1, characterized in that: Determining the three-dimensional spatial information of each target in the detection area at the current moment according to the point cloud data obtained by the multiple laser radar scans includes: The coordinate system where the first point cloud data is located is selected as the reference coordinate system from the multiple point cloud data obtained by the multiple laser radar scans, and the second point cloud data is converted to the reference coordinate system where the first point cloud data is located according to a preset conversion matrix, and the converted second point cloud data is fused with the first point cloud data to obtain fused point cloud data; wherein the second point cloud data is other point cloud data in the multiple point cloud data except the first point cloud data, and one point cloud data is obtained by one laser radar scan; The fused point cloud data is subjected to target detection processing to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
3. The method according to claim 1, characterized in that: Determining the three-dimensional spatial information of each target in the detection area at the current moment according to the point cloud data obtained by the multiple laser radar scans includes: Performing target detection processing on the point cloud data of the plurality of laser radars respectively to obtain three-dimensional spatial information of the target in each point cloud data; A coordinate system where the first three-dimensional spatial information is located is selected as a reference coordinate system from multiple three-dimensional spatial information of multiple point cloud data, and second three-dimensional spatial information is converted to the reference coordinate system where the first three-dimensional spatial information is located according to a preset conversion matrix, and the converted second three-dimensional spatial information is fused with the first three-dimensional spatial information to obtain fused three-dimensional spatial information; wherein the second three-dimensional spatial information is other three-dimensional spatial information in the multiple three-dimensional spatial information that corresponds to different point cloud data from the first three-dimensional spatial information, and one point cloud data corresponds to multiple three-dimensional spatial information; The fused three-dimensional spatial information is de-redundantly processed to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
4. The method according to claim 3, characterized in that The de-redundancy processing is performed on the fused three-dimensional spatial information to obtain the three-dimensional spatial information of each target in the detection area at the current moment, including: The non-maximum suppression algorithm is used to perform redundancy processing on the fused three-dimensional spatial information to obtain the three-dimensional spatial information of each target in the detection area at the current moment.
5. The method according to claim 1, characterized in that: The method further comprises: If there is a target whose similarity between the first feature and the second feature is not greater than the similarity threshold at the current moment, calculate the intersection-over-union ratio between the three-dimensional spatial information corresponding to the target whose similarity is not greater than the similarity threshold among the targets corresponding to the current moment and the candidate prediction spatial information; wherein the candidate prediction spatial information is the prediction spatial information of the target whose similarity is not greater than the similarity threshold in the target set; If the intersection-over-union ratio is greater than an intersection-over-union ratio threshold, the identifier of the target corresponding to the candidate prediction space information is used as the identifier of the target corresponding to the three-dimensional space information.
6. The method according to claim 1, characterized in that The comparing the three-dimensional spatial information of each target in the detection area at the current moment with the predicted spatial information of each target in the target set, and determining a corresponding identifier for the target whose three-dimensional spatial information matches the predicted spatial information, includes: A Kalman filter is used to predict the three-dimensional spatial information of the targets in the target set to obtain predicted spatial information of each target in the target set; wherein the identification of the target corresponding to the predicted spatial information corresponds to the identification of the target in the target set; For each target at the current moment, the intersection-and-union ratio between the three-dimensional spatial information and all the predicted spatial information is calculated. If there is three-dimensional spatial information whose intersection-and-union ratio is greater than the intersection-and-union ratio threshold, the identifier of the target corresponding to the matched predicted spatial information is used as the identifier of the target corresponding to the three-dimensional spatial information.
7. The method according to claim 6, characterized in that The method further comprises: If there is three-dimensional spatial information whose intersection-and-union ratio is not greater than the intersection-and-union ratio threshold, identify the third feature of the first target and the fourth feature of the second target; wherein the first target is a target whose three-dimensional spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the targets corresponding to the current moment, and the second target is a target whose predicted spatial information intersection-and-union ratio is not greater than the intersection-and-union ratio threshold among the target set; The similarity between the third feature and the fourth feature is calculated, and if the similarity is greater than a similarity threshold, the identifier of the second target is determined as the identifier of the first target.
8. The method according to claim 5 or 7, characterized in that: The method further comprises: If there is a target with an undetermined identification at the current moment, a random identification is assigned to the target with an undetermined identification, and the target with an undetermined identification and the random identification are stored in the target set; wherein the random identification is different from the identifications of other targets in the target set.
9. A target tracking device, characterized in that: The device comprises: An acquisition module is used to acquire point cloud data obtained by scanning a detection area by multiple laser radars; the multiple laser radars are arranged at different positions of the detection area; A determination module, configured to, when there is an overlapped portion in the scanning areas of two laser radars among the multiple laser radars, superimpose the point cloud data of the overlapped portion to obtain the point cloud data obtained by the superimposed multiple laser radar scans, so as to increase the point cloud density of the overlapped portion; and determine the three-dimensional spatial information of each target in each of the point cloud data according to the point cloud data obtained by the multiple laser radar scans; the three-dimensional spatial information includes the position information and size information of the target; A comparison module is used to identify the first feature of the target corresponding to each three-dimensional spatial information at the current moment; identify the second feature of the target corresponding to each predicted spatial information; if there is a target at the current moment whose similarity between the first feature and the second feature is greater than a similarity threshold, use the identifier of the target corresponding to the second feature as the identifier of the target corresponding to the first feature to complete target tracking; wherein the predicted spatial information is obtained by predicting the three-dimensional spatial information of the targets in the target set, and the target set includes the targets in the detection area at the previous moment.
10. A server comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Multi-target tracking method, system and device based on optical flow and Kalman filtering
CN110415277A
Data fusion method based on multiple laser radars
CN111413684A