An improved image multi-target cascade matching method and system
By constructing the target's feature vector F=[x,y,w,h,u,v,w/h,w*h] and weight vector W, the neural network extracts appearance features instead of the neural network, solving the problems of large calculation volume and high accuracy of the multi-objective tracking algorithm, and achieving efficient multi-objective matching, which is suitable for vehicle-mounted monocular camera perception.
Patent Information
- Application Number
- CN202311685600.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-11
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2043-12-11
AI Technical Summary
The existing multi-objective tracking algorithm has high calculation accuracy requirements and large data volume, which leads to large calculation volume and low efficiency, making it difficult to effectively apply in the field of vehicle-mounted monocular camera perception.
By constructing the target's feature vector F=[x,y,w,h,u,v,w/h,w*h], combined with the preset weight vector W, it replaces the appearance features extracted by the neural network, and uses the world coordinate system and image coordinate information to filter and match, reducing the calculation amount and improving the accuracy.
It reduces the computing volume, improves matching efficiency, expands the application scope, is suitable for general processors, and reduces application costs.
Smart Images

Figure CN117788855B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target tracking technology, and in particular to an improved image multi-target cascade matching method and system. Background Art
[0002] In the field of vehicle-mounted monocular camera perception, it is necessary to match and track perceived targets. With the development of neural networks, the DeepSort multi-target tracking algorithm uses convolutional neural networks (CNNs) to extract image features. Combined with the caching of historical features, it can effectively solve the problem of image target ID transformation.
[0003] Theoretically, the DeepSort multi-target tracking algorithm is effective and feasible. However, in practical applications, on the one hand, the vector accuracy requirements for CNN feature extraction are very high, and the accuracy of the board-side processor after general engineering is difficult to meet the requirements; on the other hand, CNN processes a large amount of data, which results in a large amount of algorithm calculation and low efficiency. Summary of the Invention
[0004] The embodiments of the present invention provide an improved image multi-target cascade matching method and system to solve the technical problems that the existing multi-target tracking algorithm has too high requirements for calculation accuracy and large amount of data to be processed, resulting in large amount of calculation and low efficiency.
[0005] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions:
[0006] In one aspect, an embodiment of the present invention provides an improved image multi-object cascade matching method, the improved image multi-object cascade matching method comprising:
[0007] Obtain a rectangular box for each target in the image, which can be detected by a neural network model or other means;
[0008] Calculating the world coordinates corresponding to the image coordinates of the rectangular frame;
[0009] Comparing the targets in the previous and next frames, predicting the location range of the target in the current frame based on the motion characteristics of the targets in the previous and next frames, and filtering the targets using the world coordinates of the rectangular frame corresponding to the targets, and filtering out targets whose distance difference exceeds a preset correlation threshold distance threshold;
[0010] For the filtered target, construct a feature vector of the target based on the image coordinates and size of the rectangular box corresponding to the target and the world coordinates of the rectangular box, combined with a preset weight vector;
[0011] The feature vector is used to replace the target appearance feature extracted by the neural network in the cascade matching algorithm, thereby improving the cascade matching algorithm and realizing multi-target matching based on the improved cascade matching algorithm.
[0012] Furthermore, calculating the world coordinates corresponding to the image coordinates of the rectangular frame includes:
[0013] Use the camera tool provided by MATLAB to calibrate the camera intrinsic parameters, and use the inverse perspective transform IPM to obtain the camera extrinsic parameters;
[0014] The image coordinates of the midpoint of the bottom edge of the rectangular frame are used as the ranging coordinates. The camera's intrinsic and extrinsic parameters are used to transform the ranging coordinates into world coordinates. The vehicle width ranging algorithm is then used to correct the coordinates of the rectangular frame to obtain the two-dimensional bird's-eye view coordinates of the ranging coordinates in the world coordinate system.
[0015] Furthermore, based on the motion characteristics of the targets in the previous and next frames, the position range of the target in the current frame is predicted, and the targets are screened using the world coordinates of the rectangular frame corresponding to the targets, and targets with a distance difference exceeding a preset correlation threshold are screened out, including:
[0016] Based on the target motion characteristics of the previous and next frames, the maximum distance threshold x_thres for the target's lateral motion and the maximum distance threshold y_thres for the target's longitudinal motion in the world coordinate are set;
[0017] Determine whether the absolute value of the difference between the horizontal coordinates in the two-dimensional bird's-eye view coordinates corresponding to the two rectangular boxes is greater than x_thres, and whether the absolute value of the difference between the vertical coordinates in the two-dimensional bird's-eye view coordinates corresponding to the two rectangular boxes is greater than y_thres; when the absolute value of the difference between the horizontal coordinates in the two-dimensional bird's-eye view coordinates corresponding to the rectangular boxes is greater than x_thres or the absolute value of the difference between the vertical coordinates in the two-dimensional bird's-eye view coordinates corresponding to the rectangular boxes is greater than y_thres, the target corresponding to the rectangular box is screened out.
[0018] Furthermore, the constructing of a feature vector of the target based on the image coordinates and size of the rectangular frame corresponding to the target and the world coordinates of the rectangular frame in combination with a preset weight vector includes:
[0019] Based on the image coordinates and size of the rectangular frame corresponding to the target and the world coordinates of the rectangular frame, a feature vector of the size and distance of the target is constructed: F = [x, y, w, h, u, v, w / h, w*h]; wherein x, y represent the horizontal and vertical coordinate values of the rectangular frame in the two-dimensional bird's-eye view coordinates; u, v represent the horizontal and vertical coordinate values of the upper left corner of the rectangular frame in the image coordinates; w represents the width of the rectangular frame in the image coordinates; h represents the height of the rectangular frame in the image coordinates;
[0020] Designing a weight vector having the same dimension as the feature vector of the size and distance of the target; wherein the weight vector is used to highlight the characteristics of the selected elements in the feature vector of the size and distance of the target;
[0021] Perform a dot multiplication operation on F and the weight vector to obtain the target feature vector.
[0022] Furthermore, the feature vector is used to replace the target appearance features extracted by the neural network in the cascade matching algorithm to improve the cascade matching algorithm, and multi-target matching is achieved based on the improved cascade matching algorithm, including:
[0023] Calculate the cosine similarity of the eigenvectors of all traces and trajectories, and construct a feature matrix based on the calculated cosine similarity of the eigenvectors; wherein the traces refer to the detected targets; the trajectories refer to the tracked targets in the library; and the feature matrix is used to represent the similarity between all traces and trajectories;
[0024] Obtain the three-point vectors of the rectangular frame corresponding to the image target, use the three-point vectors to calculate the sum of the three-point Euclidean distances between the point trace and the trajectory, and construct a distance matrix with the sum of the calculated three-point Euclidean distances; wherein the three-point vectors are expressed as: [u, v, u+0.5*w, v+0.5*h, u+w, v+h];
[0025] Processing the feature matrix and the distance matrix;
[0026] The processed distance matrix and feature matrix are used for cascade matching. When performing cascade matching, the trajectories that have been successfully matched are given priority. The matching result matrix of points and trajectories is obtained by cascade matching, and the unmatched points and trajectories are output.
[0027] Furthermore, processing the feature matrix and the distance matrix includes:
[0028] Design a similarity threshold and filter out targets whose calculated similarity is less than the similarity threshold;
[0029] Design a Euclidean distance threshold to filter out targets whose calculated distance is greater than the Euclidean distance threshold.
[0030] Furthermore, the Euclidean distance threshold is a coefficient raised to the power of w.
[0031] Furthermore, after obtaining a matching result matrix of points and trajectories by using cascade matching and outputting unmatched points and trajectories, the optimized image multi-object matching method further includes:
[0032] For unmatched points and trajectories, the overlap ratio between the rectangular boxes corresponding to the target is calculated to obtain a cost matrix; then, based on the cost matrix, the optimal match is obtained through the Hungarian matching algorithm;
[0033] The cascade matching result and the optimal match are combined to obtain the final target matching result.
[0034] On the other hand, the present invention also provides an improved image multi-object cascade matching system, the improved image multi-object cascade matching system comprising:
[0035] Data processing module for:
[0036] Get the rectangular box of each target in the image;
[0037] Calculating the world coordinates corresponding to the image coordinates of the rectangular frame;
[0038] Comparing the targets in the previous and next frames, predicting the location range of the target in the current frame based on the motion characteristics of the targets in the previous and next frames, and filtering the targets using the world coordinates of the rectangular frame corresponding to the targets, and filtering out targets whose distance difference exceeds a preset correlation threshold distance threshold;
[0039] For the filtered target, construct a feature vector of the target based on the image coordinates and size of the rectangular box corresponding to the target and the world coordinates of the rectangular box, combined with a preset weight vector;
[0040] Target matching module, used to:
[0041] The feature vector is used to replace the target appearance feature extracted by the neural network in the cascade matching algorithm, thereby improving the cascade matching algorithm and realizing multi-target matching based on the improved cascade matching algorithm.
[0042] Furthermore, the data processing module is specifically used to:
[0043] Obtain a rectangular box for each target in the image, which can be detected by a neural network model or other means;
[0044] Use the camera tool provided by MATLAB to calibrate the camera intrinsic parameters, and use the inverse perspective transform IPM to obtain the camera extrinsic parameters;
[0045] The image coordinates of the midpoint of the bottom edge of the rectangular frame are used as the ranging coordinates, and the ranging coordinates are transformed into world coordinates using the intrinsic and extrinsic parameters of the camera. The coordinates of the rectangular frame are then corrected using the vehicle width ranging algorithm to obtain the two-dimensional bird's-eye view coordinates of the ranging coordinates in the world coordinate system;
[0046] Based on the target motion characteristics of the previous and next frames, the maximum distance threshold x_thres for the target's lateral motion and the maximum distance threshold y_thres for the target's longitudinal motion in the world coordinate are set;
[0047] Determine whether the absolute value of the difference between the horizontal coordinates in the two-dimensional bird's-eye view coordinates corresponding to the two rectangular boxes is greater than x_thres, and whether the absolute value of the difference between the vertical coordinates in the two-dimensional bird's-eye view coordinates corresponding to the two rectangular boxes is greater than y_thres; when the absolute value of the difference between the horizontal coordinates in the two-dimensional bird's-eye view coordinates corresponding to the rectangular boxes is greater than x_thres or the absolute value of the difference between the vertical coordinates in the two-dimensional bird's-eye view coordinates corresponding to the rectangular boxes is greater than y_thres, filter out the targets corresponding to the rectangular boxes;
[0048] Based on the image coordinates and size of the rectangular frame corresponding to the target and the world coordinates of the rectangular frame, a feature vector of the size and distance of the target is constructed: F = [x, y, w, h, u, v, w / h, w*h]; wherein x, y represent the horizontal and vertical coordinate values of the rectangular frame in the two-dimensional bird's-eye view coordinates; u, v represent the horizontal and vertical coordinate values of the upper left corner of the rectangular frame in the image coordinates; w represents the width of the rectangular frame in the image coordinates; h represents the height of the rectangular frame in the image coordinates;
[0049] Designing a weight vector having the same dimension as the feature vector of the size and distance of the target; wherein the weight vector is used to highlight the characteristics of the selected elements in the feature vector of the size and distance of the target;
[0050] Perform a dot multiplication operation on F and the weight vector to obtain the target feature vector;
[0051] The target matching module is specifically used for:
[0052] Calculate the cosine similarity of the eigenvectors of all traces and trajectories, and construct a feature matrix based on the calculated cosine similarity of the eigenvectors; wherein the traces refer to the detected targets; the trajectories refer to the tracked targets in the library; and the feature matrix is used to represent the similarity between all traces and trajectories;
[0053] Obtain the three-point vectors of the rectangular frame corresponding to the image target, use the three-point vectors to calculate the sum of the three-point Euclidean distances between the point trace and the trajectory, and construct a distance matrix with the sum of the calculated three-point Euclidean distances; wherein the three-point vectors are expressed as: [u, v, u+0.5*w, v+0.5*h, u+w, v+h];
[0054] Processing the feature matrix and the distance matrix;
[0055] Cascade matching is performed using the processed distance matrix and feature matrix. When performing cascade matching, the previously successfully matched trajectories are prioritized. The matching result matrix of points and trajectories is obtained using cascade matching, and the unmatched points and trajectories are output.
[0056] The processing of the feature matrix and the distance matrix includes:
[0057] Design a similarity threshold and filter out targets whose calculated similarity is less than the similarity threshold;
[0058] Design a Euclidean distance threshold to filter out targets whose calculated distance is greater than the Euclidean distance threshold;
[0059] The Euclidean distance threshold is a coefficient raised to the power of w;
[0060] After obtaining the matching result matrix of points and trajectories by cascade matching and outputting the unmatched points and trajectories, the target matching module is further configured to:
[0061] For unmatched points and trajectories, the overlap ratio between the rectangular boxes corresponding to the target is calculated to obtain a cost matrix; then, based on the cost matrix, the optimal match is obtained through the Hungarian matching algorithm;
[0062] The cascade matching result and the optimal match are combined to obtain the final target matching result.
[0063] On the other hand, the present invention further provides an electronic device, comprising a processor and a memory; the memory stores at least one instruction, and the instruction is loaded and executed by the processor to implement the above method.
[0064] In yet another aspect, the present invention further provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, and the instruction is loaded and executed by a processor to implement the above method.
[0065] The beneficial effects brought about by the technical solution provided by the present invention include at least:
[0066] Inspired by the DeepSort multi-target tracking algorithm, this paper proposes an improved image multi-target cascade matching method. Compared to DeepSort, which relies on deep neural network training data to extract image feature vectors, this method, on the one hand, suffers from large feature dimensionality and high computational complexity; on the other hand, the feature vectors' high precision places high demands on processors, making it difficult for conventional computers to accurately extract features and resulting in high application costs. The multi-target matching method provided by this paper does not rely on CNN network computations. The image target features are constructed using the bird's-eye view features [xy] and the coordinate size features of the target rectangular box to construct the target feature vector F = [x, y, w, h, u, v, w / h, w*h], significantly reducing dimensionality and computational complexity. Furthermore, the use of the world coordinate system [xy] and target motion feature information can filter out irrelevant targets, further increasing the probability of matching. The design of the weight vector W allows for specific features to be highlighted, ensuring satisfactory accuracy while enabling standard processors to meet requirements. Therefore, it has a wider range of applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0068] Figure 1 is a flow chart of an improved image multi-object cascade matching method according to an embodiment of the present invention;
[0069] Figure 2 4 is a flow chart of an improved cascade matching algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION
[0070] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present invention. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other in the absence of conflict.
[0071] Furthermore, it should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the associated drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such processes, methods, products, or apparatus.
[0072] First embodiment
[0073] In view of the problem that in the existing DeepSort algorithm, image feature extraction generally uses CNN network extraction, and the appearance feature extraction of CNN network has a large amount of calculation, low efficiency, and high calculation accuracy requirements, which is not conducive to implementation and popularization, this embodiment provides an improved image multi-target cascade matching method, which can be implemented by an electronic device, which can be a terminal or a server. The execution process of this method is as follows Figure 1 As shown, the following steps are included:
[0074] S1, obtain the rectangular box of each target in the image;
[0075] In this embodiment, the upper left corner coordinate [uv], width w, and height h are used to represent a rectangular box, referred to as bbox. The bbox can be obtained by a neural network model or other means.
[0076] S2, calculating the world coordinates corresponding to the image coordinates of the rectangular frame;
[0077] Specifically, in this embodiment, the process of obtaining the world coordinates is as follows: using the camera tool provided by MATLAB to calibrate the camera intrinsic parameters, and using the inverse perspective transformation IPM to obtain the camera extrinsic parameters; using the image coordinates of the midpoint of the bottom edge of the rectangular frame as the ranging coordinates, using the intrinsic and extrinsic parameters of the camera to transform the ranging coordinates into world coordinates, and then using the vehicle width ranging algorithm to correct the coordinates of the rectangular frame to obtain the two-dimensional bird's-eye view coordinates [xy] of the ranging coordinates in the world coordinate system.
[0078] S3, comparing the targets in the previous and next frames, predicting the position range of the target in the current frame based on the motion characteristics of the targets in the previous and next frames, and screening the targets using the world coordinates of the rectangular frame corresponding to the targets, and screening out targets whose distance difference exceeds a preset correlation threshold distance threshold;
[0079] Specifically, in this embodiment, the process of screening targets is as follows: using the target motion characteristics of the previous and next frames, the maximum distance thresholds for the target's lateral and longitudinal motion in world coordinates are estimated. A maximum distance threshold x_thres for the target's lateral motion and a maximum distance threshold y_thres for the target's longitudinal motion are designed. A determination is made as to whether (|x1-x2|>x_thres or |y1-y2|>y_thres) is satisfied. Targets that meet these conditions are then screened out to roughly eliminate targets with large distance differences, reduce mismatched targets, and increase the probability of accurate matching. Furthermore, in this embodiment, image coordinates can be taken into consideration to design an adaptive threshold, which is used to filter out irrelevant targets to increase matching accuracy.
[0080] S4, for the filtered target, constructing a feature vector of the target based on the image coordinates and size of the rectangular frame corresponding to the target and the world coordinates of the rectangular frame, in combination with a preset weight vector;
[0081] Specifically, this embodiment obtains the target size and distance feature vector F = [x, y, w, h, u, v, w / h, w*h] based on the size characteristics of the image target bbox and the characteristics of the two-dimensional bird's-eye view coordinates [xy] in the world coordinates. It should be noted that because the horizontal distance and vertical distance of the target in the world coordinates are significantly different, in order to highlight the characteristics of x and y, a weight vector W of dimension [1x8] is designed. The weight ratio parameters can be adjusted according to the eight-dimensional features of the target [x, y, w, h, u, v, w / h, w*h]. This design can make the feature distinction between the targets more obvious, the similarity difference is large, and it is easy to distinguish. The feature vector feas of each target is then obtained by multiplying the F vector by the W vector.
[0082] S5, using the feature vector to replace the target appearance feature extracted by the neural network in the cascade matching algorithm, to improve the cascade matching algorithm, and to achieve multi-target matching based on the improved cascade matching algorithm.
[0083] Specifically, in this embodiment, Figure 2 As shown, the process of implementing cascade matching is as follows:
[0084] S51: Calculate the feature vector of each detected target (hereinafter referred to as a point) and the feature vector of each tracked target (hereinafter referred to as a track) using the above method. Calculate the cosine similarity of all points and tracks to obtain the feature matrix, feature. Since the feature vector is 1x8 in dimension, the computation is simple and easy.
[0085] S52, on the other hand, take the three-point vector [u, v, u+0.5*w, v+0.5*h, u+w, v+h] of the image target bbox, and use this vector to calculate the sum of the three-point Euclidean distances between the point trace and the trajectory to obtain the distance matrix uvDist;
[0086] In step S53, because the distance of the target in the world coordinate system is inversely proportional to the bbox width w, a Euclidean distance threshold uv_thres is designed. The threshold is a coefficient raised to the power of w. Therefore, uv_thres can be adaptively adjusted based on the size of the bbox. Objects with uvDist greater than uvThres are filtered out, reducing mismatches and increasing the probability of accurate matches.
[0087] S54: The processed Euclidean distance matrix uvDist and feature matrix feature are input into cascade matching. Cascade matching prioritizes matching stable trajectories to ensure the trajectory continuity of the stable target. Cascade matching generates a matching matrix for point and trajectory matching. Unmatched points unMatchMeas and trajectories unMatchTracks are also output.
[0088] S55: For the unmatched points unMatchMeas and tracks unMatchTracks, calculate the IoU between the two bboxes to obtain the cost matrix costMatrix. Then, use the Hungarian matching method to obtain the optimal match M.
[0089] S56, combining the above two levels of matching and M to obtain the final target matching result.
[0090] Next, the method of this embodiment is applied to the field of monocular camera perception, and its implementation process is described. Specifically, the implementation process of this method in the field of monocular camera perception is as follows:
[0091] 1) Use the built-in camera tool in MATLAB to calibrate the intrinsic parameters of the monocular camera, and use the inverse perspective transform (IPM) to obtain the extrinsic parameters of the vehicle-mounted camera.
[0092] 2) Each target is represented in the image by a rectangular box (denoted by its upper left corner coordinates [uv], width w, and height h, referred to as a bbox). The midpoint of the bbox's base is used as the ranging coordinates [u0, v0]. Using intrinsic and extrinsic parameters, the bbox image coordinates [u0, v0] are transformed to world coordinates [xy]. The vehicle width ranging method is then used to correct the bbox's longitudinal x coordinates. This results in the calculated 2D bird's-eye view coordinates [xy] in the world coordinate system.
[0093] 3) For motor vehicle targets, set x_thres = 15m, y_thres = 2m; for non-motor vehicle targets, set x_thres = 10m, y_thres = 1.5m; take the absolute value of the difference between the point trace and the track target distance, and for targets with x_err > x_thres or y_err > y_thres, determine the cosine similarity = 0.
[0094] 4) For targets with x_err <= x_thres or y_err <= y_thres, obtain the target size and distance feature vector F = [x, y, w, h, u, v, w / h, w*h] based on the size characteristics of the image target bbox and the characteristics of the two-dimensional bird's-eye view coordinates [xy] in the world coordinates.
[0095] 5) Design the weight vector W to have dimensions [1x8]. Adjust the weight ratio parameters based on the eight target features [x, y, w, h, u, v, w / h, w*h]. In this case, for non-motorized vehicles, W = [100, 100, 1, 1, 10, 10, 1, 1]; for motor vehicles, W = [100, 1, 1, 1, 10, 10, 1e4, 1]. This design makes the features between the targets more distinct and increases the degree of familiarity. Then, dot-product the F vector with the W vector to obtain the feature vector feas for each target.
[0096] 6) Using the above method, calculate the feature vector of each detected target (hereafter referred to as a point) and the feature vector of each tracked target (hereafter referred to as a track). Calculate the cosine similarity of all points and tracks to obtain the feature matrix, feature. Since the feature vector is 1x8 in dimension, the computation is minimal and simple.
[0097] 7) The feature matrix feature represents the similarity between all points and trajectories, and the similarity threshold f_thres is set to 0.99. For targets where feature(i,j)>f_thres, the corresponding indicator matrix B2(i,j) is set to 1; otherwise, B2(i,j) is set to 0.
[0098] 8) On the other hand, take the three-point vector [u, v, u+0.5*w, v+0.5*h, u+w, v+h] of the image target bbox, and use this vector to calculate the sum of the three-point Euclidean distances between the point and the trajectory to obtain the distance matrix uvDist.
[0099] 9) Design uv_thres = 1.1 ^ (0.1 * w), with the threshold limit 1.0 < uv_thres < 2.5. Set the target Euclidean distance with uv_Dist > uv_thres to 1e4, which is the default maximum distance. Set the indicator matrix B1. If uv_Dist(i,j) > uv_thres, then the element B1(i,j) in the indicator matrix is 0; otherwise, B1(i,j) = 1.
[0100] 10) The processed Euclidean distance matrix uvDist and the feature matrix feature are input into cascade matching. Cascade matching preferentially matches stable trajectories to ensure the continuity of the trajectories of stable targets. Use cascade matching to obtain the matching result matching matrix for point tracks and trajectories. Additionally, output the unmatched point tracks unMatchMeas and trajectories unMatchTracks. As Figure 2 shown.
[0101] 11) For the unmatched point tracks unMatchMeas and trajectories unMatchTracks, calculate the overlap rate IoU result between the two bboxes to obtain the cost matrix costMatrix. Then, obtain the optimal matching M through the Hungarian matching method.
[0102] 12) Combine the above two-level matchings matching and M to obtain the final target matching result MM.
[0103] 13) Process the trajectories according to MM. If a match is found, update the target in the trajectory library through Kalman filtering. If a trajectory is unmatched, delete the trajectory target if the conditions are met, and update the trajectory target through Kalman prediction if the conditions are not met. For unmatched point tracks, add them as new temporary trajectories to the trajectory library. As Figure 1 shown.
[0104] In summary, this embodiment provides an improved method for multi-target cascade matching of images. This method first converts the image coordinates of the target into world coordinates, compares the targets in the front and back frames, and roughly filters out the targets with a large difference in distance using the world coordinate information to increase the matching accuracy. Then, considering the image coordinates, an adaptive threshold is designed to filter out some irrelevant targets to increase the matching accuracy. Finally, considering the distance characteristics of the target in the world coordinates and the size characteristics of the bbox in the image coordinates, these characteristics are extracted and the preset weight vector is used to highlight the characteristics of the target to obtain the feature vector of the target. Using the above feature vector to replace the appearance features extracted by the CNN network can avoid using neural networks to extract features, thereby improving the calculation efficiency, reducing the algorithm cost, and having good feasibility.
[0105] Moreover, it should be noted that, through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, or of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the method described in the embodiment of the present invention.
[0106] Second embodiment
[0107] This embodiment provides an improved image multi-object cascade matching system, which is used to implement the above-mentioned embodiments and preferred implementations. The details that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated. The improved image multi-object cascade matching system includes the following modules:
[0108] Data processing module for:
[0109] Get the rectangular box of each target in the image;
[0110] Calculating the world coordinates corresponding to the image coordinates of the rectangular frame;
[0111] Comparing the targets in the previous and next frames, predicting the location range of the target in the current frame based on the motion characteristics of the targets in the previous and next frames, and filtering the targets using the world coordinates of the rectangular frame corresponding to the targets, and filtering out targets whose distance difference exceeds a preset correlation threshold distance threshold;
[0112] For the filtered target, construct a feature vector of the target based on the image coordinates and size of the rectangular box corresponding to the target and the world coordinates of the rectangular box, combined with a preset weight vector;
[0113] Target matching module, used to:
[0114] The feature vector is used to replace the target appearance feature extracted by the neural network in the cascade matching algorithm, thereby improving the cascade matching algorithm and realizing multi-target matching based on the improved cascade matching algorithm.
[0115] It should be noted that the improved image multi-target cascade matching system of this embodiment corresponds to the improved image multi-target cascade matching method of the above-mentioned first embodiment; wherein, the functions implemented by each functional module in the improved image multi-target cascade matching system of this embodiment correspond one-to-one to each process step in the improved image multi-target cascade matching method of the above-mentioned first embodiment; therefore, they will not be repeated here.
[0116] In addition, it should be noted that the above-mentioned modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned modules are all located in the same processor; or the above-mentioned modules are located in different processors in any combination.
[0117] Third embodiment
[0118] This embodiment provides an electronic device comprising a processor and a memory; wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment described above. The electronic device may vary significantly due to different configurations or performance, and may comprise one or more processors and one or more memories, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the method of the first embodiment.
[0119] Fourth embodiment
[0120] This embodiment provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the above-described method. The computer-readable storage medium may be a ROM, random access memory, CD-ROM, magnetic tape, floppy disk, or optical data storage device. The instructions stored therein can be loaded by a processor in a terminal to execute the method of the first embodiment.
[0121] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.
[0122] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0123] In the several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0124] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0125] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0126] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program codes.
[0127] Finally, it should be noted that the above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art could make numerous improvements and modifications without departing from the principles of the present invention, and such improvements and modifications should also be considered within the scope of protection of the present invention. Therefore, the appended claims are intended to include the preferred embodiments and all variations and modifications that fall within the scope of the preferred embodiments.
Claims
1. An improved image multi-target cascade matching method, characterized in that: include: Get the rectangular box of each target in the image; Calculating the world coordinates corresponding to the image coordinates of the rectangular frame; Comparing the targets in the previous and next frames, predicting the location range of the target in the current frame based on the motion characteristics of the targets in the previous and next frames, and screening the targets using the world coordinates of the rectangular box corresponding to the targets, and filtering out targets whose distance difference exceeds a preset correlation threshold; For the filtered target, construct a feature vector of the target based on the image coordinates and size of the rectangular box corresponding to the target and the world coordinates of the rectangular box, combined with a preset weight vector; The feature vector is used to replace the target appearance feature extracted by the neural network in the cascade matching algorithm, thereby improving the cascade matching algorithm and realizing multi-target matching based on the improved cascade matching algorithm.
2. The improved image multi-object cascade matching method according to claim 1, characterized in that: The calculating the world coordinates corresponding to the image coordinates of the rectangular frame includes: Use the camera tool provided by MATLAB to calibrate the camera intrinsic parameters, and use the inverse perspective transform IPM to obtain the camera extrinsic parameters; The image coordinates of the midpoint of the bottom edge of the rectangular frame are used as the ranging coordinates. The camera's intrinsic and extrinsic parameters are used to transform the ranging coordinates into world coordinates. The vehicle width ranging algorithm is then used to correct the coordinates of the rectangular frame to obtain the two-dimensional bird's-eye view coordinates of the ranging coordinates in the world coordinate system.
3. The improved image multi-object cascade matching method according to claim 2, characterized in that: The method of predicting the position range of the target in the current frame based on the motion characteristics of the targets in the previous and next frames, screening the targets using the world coordinates of the rectangular frame corresponding to the targets, and eliminating targets whose distance difference exceeds a preset correlation threshold, includes: Based on the target motion characteristics of the previous and next frames, the maximum distance threshold x_thres for the target's lateral motion and the maximum distance threshold y_thres for the target's longitudinal motion in the world coordinate are set; Determine whether the absolute value of the difference between the horizontal coordinates in the two-dimensional bird's-eye view coordinates corresponding to the two rectangular boxes is greater than x_thres, and whether the absolute value of the difference between the vertical coordinates in the two-dimensional bird's-eye view coordinates corresponding to the two rectangular boxes is greater than y_thres; when the absolute value of the difference between the horizontal coordinates in the two-dimensional bird's-eye view coordinates corresponding to the rectangular boxes is greater than x_thres or the absolute value of the difference between the vertical coordinates in the two-dimensional bird's-eye view coordinates corresponding to the rectangular boxes is greater than y_thres, the target corresponding to the rectangular box is screened out.
4. The improved image multi-object cascade matching method according to claim 2, characterized in that: The step of constructing a feature vector of the target based on the image coordinates and size of the rectangular frame corresponding to the target and the world coordinates of the rectangular frame, combined with a preset weight vector, includes: Based on the image coordinates and size of the rectangular frame corresponding to the target and the world coordinates of the rectangular frame, a feature vector of the size and distance of the target is constructed: F = [x, y, w, h, u, v, w / h, w*h]; wherein x, y represent the horizontal and vertical coordinate values of the rectangular frame in the two-dimensional bird's-eye view coordinates; u, v represent the horizontal and vertical coordinate values of the upper left corner of the rectangular frame in the image coordinates; w represents the width of the rectangular frame in the image coordinates; h represents the height of the rectangular frame in the image coordinates; Designing a weight vector having the same dimension as the feature vector of the size and distance of the target; wherein the weight vector is used to highlight the characteristics of the selected elements in the feature vector of the size and distance of the target; Perform a dot multiplication operation on F and the weight vector to obtain the target feature vector.
5. The improved image multi-object cascade matching method according to claim 4, characterized in that: The feature vector is used to replace the target appearance feature extracted by the neural network in the cascade matching algorithm to improve the cascade matching algorithm, and multi-target matching is achieved based on the improved cascade matching algorithm, including: Calculate the cosine similarity of the eigenvectors of all traces and trajectories, and construct a feature matrix based on the calculated cosine similarity of the eigenvectors; wherein the traces refer to the detected targets; the trajectories refer to the tracked targets in the library; and the feature matrix is used to represent the similarity between all traces and trajectories; Obtain the three-point vectors of the rectangular frame corresponding to the image target, use the three-point vectors to calculate the sum of the three-point Euclidean distances between the point trace and the trajectory, and construct a distance matrix with the sum of the calculated three-point Euclidean distances; wherein the three-point vectors are expressed as: [u, v, u+0.5*w, v+0.5*h, u+w, v+h]; Processing the feature matrix and the distance matrix; The processed distance matrix and feature matrix are used for cascade matching. When performing cascade matching, the trajectories that have been successfully matched are given priority. The matching result matrix of points and trajectories is obtained by cascade matching, and the unmatched points and trajectories are output.
6. The improved image multi-object cascade matching method according to claim 5, characterized in that: Processing the feature matrix and the distance matrix includes: Design a similarity threshold and filter out targets whose calculated similarity is less than the similarity threshold; Design a Euclidean distance threshold to filter out targets whose calculated distance is greater than the Euclidean distance threshold.
7. The improved image multi-object cascade matching method according to claim 6, characterized in that: The Euclidean distance threshold is a coefficient raised to the power of w.
8. The improved image multi-object cascade matching method according to claim 5, characterized in that: After obtaining a matching result matrix of points and trajectories by using cascade matching and outputting unmatched points and trajectories, the optimized image multi-object matching method based on cascade matching further includes: For unmatched points and trajectories, the overlap ratio between the rectangular boxes corresponding to the target is calculated to obtain a cost matrix; then, based on the cost matrix, the optimal match is obtained through the Hungarian matching algorithm; The cascade matching result and the optimal match are combined to obtain the final target matching result.
9. An improved image multi-target cascade matching system, characterized in that: include: Data processing module for: Get the rectangular box of each target in the image; Calculating the world coordinates corresponding to the image coordinates of the rectangular frame; Comparing the targets in the previous and next frames, predicting the location range of the target in the current frame based on the motion characteristics of the targets in the previous and next frames, and filtering the targets using the world coordinates of the rectangular frame corresponding to the targets, and filtering out targets whose distance difference exceeds a preset correlation threshold distance threshold; For the filtered target, construct a feature vector of the target based on the image coordinates and size of the rectangular box corresponding to the target and the world coordinates of the rectangular box, combined with a preset weight vector; Target matching module, used to: The feature vector is used to replace the target appearance feature extracted by the neural network in the cascade matching algorithm, thereby improving the cascade matching algorithm and realizing multi-target matching based on the improved cascade matching algorithm.
10. The improved image multi-object cascade matching system according to claim 9, characterized in that: The data processing module is specifically used for: Get the rectangular box of each target in the image; Use the camera tool provided by MATLAB to calibrate the camera intrinsic parameters, and use the inverse perspective transform IPM to obtain the camera extrinsic parameters; The image coordinates of the midpoint of the bottom edge of the rectangular frame are used as the ranging coordinates, and the ranging coordinates are transformed into world coordinates using the intrinsic and extrinsic parameters of the camera. The coordinates of the rectangular frame are then corrected using the vehicle width ranging algorithm to obtain the two-dimensional bird's-eye view coordinates of the ranging coordinates in the world coordinate system; Based on the target motion characteristics of the previous and next frames, the maximum distance threshold x_thres for the target's lateral motion and the maximum distance threshold y_thres for the target's longitudinal motion in the world coordinate are set; Determine whether the absolute value of the difference between the horizontal coordinates in the two-dimensional bird's-eye view coordinates corresponding to the two rectangular boxes is greater than x_thres, and whether the absolute value of the difference between the vertical coordinates in the two-dimensional bird's-eye view coordinates corresponding to the two rectangular boxes is greater than y_thres; when the absolute value of the difference between the horizontal coordinates in the two-dimensional bird's-eye view coordinates corresponding to the rectangular boxes is greater than x_thres or the absolute value of the difference between the vertical coordinates in the two-dimensional bird's-eye view coordinates corresponding to the rectangular boxes is greater than y_thres, filter out the targets corresponding to the rectangular boxes; Based on the image coordinates and size of the rectangular frame corresponding to the target and the world coordinates of the rectangular frame, a feature vector of the size and distance of the target is constructed: F = [x, y, w, h, u, v, w / h, w*h]; wherein x, y represent the horizontal and vertical coordinate values of the rectangular frame in the two-dimensional bird's-eye view coordinates; u, v represent the horizontal and vertical coordinate values of the upper left corner of the rectangular frame in the image coordinates; w represents the width of the rectangular frame in the image coordinates; h represents the height of the rectangular frame in the image coordinates; Designing a weight vector having the same dimension as the feature vector of the size and distance of the target; wherein the weight vector is used to highlight the characteristics of the selected elements in the feature vector of the size and distance of the target; Perform a dot multiplication operation on F and the weight vector to obtain the target feature vector; The target matching module is specifically used for: Calculate the cosine similarity of the eigenvectors of all traces and trajectories, and construct a feature matrix based on the calculated cosine similarity of the eigenvectors; wherein the traces refer to the detected targets; the trajectories refer to the tracked targets in the library; and the feature matrix is used to represent the similarity between all traces and trajectories; Obtain the three-point vectors of the rectangular frame corresponding to the image target, use the three-point vectors to calculate the sum of the three-point Euclidean distances between the point trace and the trajectory, and construct a distance matrix with the sum of the calculated three-point Euclidean distances; wherein the three-point vectors are expressed as: [u, v, u+0.5*w, v+0.5*h, u+w, v+h]; Processing the feature matrix and the distance matrix; Cascade matching is performed using the processed distance matrix and feature matrix. When performing cascade matching, the previously successfully matched trajectories are prioritized. The matching result matrix of points and trajectories is obtained using cascade matching, and the unmatched points and trajectories are output. The processing of the feature matrix and the distance matrix includes: Design a similarity threshold and filter out targets whose calculated similarity is less than the similarity threshold; Design a Euclidean distance threshold to filter out targets whose calculated distance is greater than the Euclidean distance threshold; The Euclidean distance threshold is a coefficient raised to the power of w; After obtaining the matching result matrix of points and trajectories by cascade matching and outputting the unmatched points and trajectories, the target matching module is further configured to: For unmatched points and trajectories, the overlap ratio between the rectangular boxes corresponding to the target is calculated to obtain a cost matrix; then, based on the cost matrix, the optimal match is obtained through the Hungarian matching algorithm; The cascade matching result and the optimal match are combined to obtain the final target matching result.
Citation Information
Patent Citations
Multi-target identification tracking method and identification tracking system
CN114419343A
Multi-target tracking method, equipment and medium
CN114638855A