Multi-monitoring equipment active tracking method and device based on road network behavior rule prediction
By constructing a spatiotemporal graph convolution model, the spatial and temporal behavior patterns of targets in the monitoring road network are predicted, and the problems of target loss and path loss in low-density cross-spectrum tracking are solved, and precise tracking and efficient prediction of targets in the monitoring road network are achieved.
Patent Information
- Application Number
- CN202510122015.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The prior art has problems of target loss and path loss or fragmentation in low-density cross-scope tracking, resulting in the need of manual intervention to predict target moving routes, which is time-consuming and labor-intensive.
By using road network data, external parameters of monitoring equipment, target file gathering data and external time feature data, a spatiotemporal graph convolution model is built to predict the spatiotemporal behavior of the target in the monitoring road network, and combining target detection technology and image coordinates to road network meteoral and latitude coordinate conversion technology, fast and accurate active tracking of sparse multi-monitoring equipment is achieved.
Accurate tracking of targets in the monitoring road network is achieved, manual intervention is reduced, and tracking efficiency and accuracy is improved.
Smart Images

Figure CN119991747A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and geographic information systems, and in particular to a method and device for active tracking of multiple monitoring devices based on prediction of road network behavior rules. Background Art
[0002] In the urban public security monitoring system, real-time monitoring and tracking of moving targets on the road surface (such as people, non-motor vehicles, motor vehicles and other small-scale objects) helps city managers better understand and analyze the activity patterns and events of urban personnel, which plays an important role and significance in the fields of behavior analysis, criminal investigation analysis, traffic management, etc. In order to achieve this goal, cross-camera linkage technology between multiple monitoring devices is usually used to ensure smooth tracking of the target between different cameras, so as to obtain all-round and multi-angle target information for in-depth analysis and processing.
[0003] At present, in the field of computer vision and geographic information systems, there are two main methods for cross-lens tracking of multi-monitoring targets: high-density cross-lens tracking and low-density cross-lens tracking. Among them, the former is the current mainstream method, which is usually suitable for areas with high-density coverage of monitoring equipment. It realizes multi-monitoring cross-lens linkage without blind spots through technologies such as boundary overlap relay. However, the monitoring area in real scenarios is usually low-density, and the application of such methods is very limited and has poor practicality. For low-density cross-lens tracking technology, due to the independence and lack of linkage between low-density monitoring devices, the target often loses tracking after switching the lens, resulting in the missing and fragmented path of the tracking target. Generally, it is necessary to use external signals such as satellite positioning to assist tracking, or rely on algorithms to search all monitoring devices in the area in the entire domain, which requires a large amount of equipment or computing power. In most scenarios, cross-lens identification of targets with low-density multi-monitoring devices still needs to rely on manual intervention to summarize, generalize and predict the target's movement route, which is time-consuming and labor-intensive, and brings huge challenges to the current target tracking technology that relies on modern urban monitoring systems. Summary of the invention
[0004] In order to solve the problems existing in the prior art, an embodiment of the present invention provides a method and device for active tracking of multiple monitoring devices based on the prediction of road network behavior laws. The method and device utilize multi-source data encoding consisting of road network data, monitoring device external parameters, target cluster data and external time feature data to monitor the heterogeneous correlations and spatiotemporal dependencies between road network devices. A spatiotemporal graph convolution model consisting of a graph convolutional network and a temporal neural network is used to construct a target spatiotemporal behavior law model under the monitoring road network. Target detection technology and image coordinate to road network longitude and latitude coordinate conversion technology are used to combine the target's moving direction and speed to assist in the realization of fast and accurate sparse active tracking and trajectory characterization of multiple monitoring devices.
[0005] In order to solve any of the above technical problems, the specific technical solutions of the present invention are as follows:
[0006] The embodiment of the present invention provides a method for active tracking of multiple monitoring devices based on prediction of road network behavior rules, comprising:
[0007] Receive target image features and initial monitoring points of the tracked target;
[0008] Determine continuous video image frames containing the target image features from the video image of the initial monitoring point as target video image frames, and determine the position information of the tracking target in each target video image frame;
[0009] Calculate the movement information of the tracked target in the monitoring road network where the initial monitoring point is located according to the position information of the tracked target in each target video image frame;
[0010] Acquire the current multi-directional traffic data of the monitoring road network, and use the pre-built spatiotemporal convolution model to calculate the current multi-directional traffic data and the movement information to obtain the predicted monitoring point position where the tracking target will appear in the monitoring road network at the next moment;
[0011] Calculate the predicted arrival time of the tracking target at each predicted monitoring point according to the distance between each predicted monitoring point and the initial monitoring point and the movement information;
[0012] For each predicted monitoring point, the target image feature is searched in the video image corresponding to the predicted arrival time, and the predicted monitoring point where the target image feature is found is used as the actual monitoring point.
[0013] Furthermore, calculating the movement information of the tracked target in the monitoring road network where the initial monitoring point is located according to the position information of the tracked target in each target video image frame further includes:
[0014] For each target video image frame, convert the position information from the image coordinate system of the target video image frame to the longitude and latitude coordinate system of the monitoring road network where the initial monitoring point is located, to obtain the longitude and latitude coordinates of the tracking target;
[0015] The longitude and latitude coordinates of the tracked target corresponding to all target video image frames are fitted in the order of the target video image frames to obtain the movement information.
[0016] Furthermore, the step of constructing the spatiotemporal convolution model includes:
[0017] Acquire information parameters of each monitoring point in the monitoring road network, and calculate the spatial dependence characteristics of the monitoring road network according to the information parameters of each monitoring point;
[0018] Obtaining historical traffic data of each monitoring point, and calculating the time-dependent characteristics of the monitoring road network based on the historical traffic data of each monitoring point;
[0019] Performing feature fusion on the space-dependent features and the time-dependent features to obtain fused features;
[0020] The fused features are input into a two-layer graph convolution network for graph convolution calculation to obtain spatial dynamic correlation features of each monitoring point;
[0021] Acquiring external time characteristic data of the monitored road network;
[0022] The external time feature data and the spatial dynamic correlation features of each monitoring point are input into a recurrent neural network for calculation to obtain the spatiotemporal convolution model.
[0023] Furthermore, the external time feature data and the spatial dynamic correlation features of each monitoring point are input into a recurrent neural network for calculation, and the spatiotemporal convolution model is obtained, further comprising:
[0024] Each time step of the spatial dynamic correlation feature is connected through a flattening layer and a fully connected layer and input into an encoding layer;
[0025] The external time feature data is normalized and fully connected, and input into each time step of the decoding layer;
[0026] The recurrent neural network is trained to obtain the spatiotemporal convolution model.
[0027] Furthermore, the method further comprises:
[0028] Acquiring current external time characteristic data of the monitored road network;
[0029] Calculating the current multi-directional traffic data and the movement information using a pre-built spatiotemporal convolution model to obtain a predicted monitoring point where the tracking target will appear at the next moment in the monitoring road network further includes:
[0030] The current multi-directional traffic data and the current external time feature data are input into the spatiotemporal convolution model for calculation to obtain the predicted monitoring point where the tracking target will appear in the monitored road network at the next moment.
[0031] Furthermore, the method further comprises:
[0032] Calculate the probability of the tracking target appearing at each predicted monitoring point at the next moment;
[0033] For each predicted monitoring point, searching for the target image feature in the video image corresponding to the predicted arrival time further includes:
[0034] According to the order of probability values from large to small, the target image features are searched for the predicted monitoring points in the video image corresponding to the predicted arrival time.
[0035] Furthermore, after finding the predicted monitoring point of the target image feature as the actual monitoring point, the method further includes:
[0036] Taking the actual monitoring point as the initial monitoring point, repeatedly performing the step of determining continuous video image frames having the target image feature from the video image of the initial monitoring point as target video image frames, until the target image feature does not appear in the video images corresponding to all the predicted monitoring points;
[0037] The movement information corresponding to all the initial monitoring points of the tracking target is fitted to obtain the complete movement information of the tracking target.
[0038] Furthermore, searching for the target image feature in the video image corresponding to the predicted arrival time further includes:
[0039] Determine the corresponding time window based on the predicted arrival time;
[0040] The target image feature is searched in the video image corresponding to the time window.
[0041] On the other hand, an embodiment of the present invention further provides a multi-monitoring device active tracking device based on road network behavior law prediction, comprising:
[0042] A tracking target information receiving unit, used to receive target image features and initial monitoring points of the tracking target;
[0043] A tracking target searching unit, used to determine continuous video image frames having the target image features from the video image of the initial monitoring point as target video image frames, and determine the position information of the tracking target in each target video image frame;
[0044] A tracking target movement calculation unit, used to calculate the movement information of the tracking target in the monitoring road network where the initial monitoring point is located according to the position information of the tracking target in each target video image frame;
[0045] A predicted monitoring point calculation unit is used to obtain the current multi-directional flow data of the monitoring road network, and use a pre-built spatiotemporal convolution model to calculate the current multi-directional flow data and the movement information to obtain the predicted monitoring point where the tracking target will appear in the monitoring road network at the next moment;
[0046] A predicted arrival time calculation unit, used to calculate the predicted arrival time of the tracking target to each predicted monitoring point according to the distance between each predicted monitoring point and the initial monitoring point and the movement information;
[0047] The actual monitoring point determination unit is used to search for the target image feature in the video image corresponding to the predicted arrival time for each predicted monitoring point, and use the predicted monitoring point where the target image feature is found as the actual monitoring point.
[0048] On the other hand, an embodiment of the present specification further provides a computer device, including a memory, a processor, and a computer program stored in the memory, and when the processor executes the computer program, the above method is implemented.
[0049] The beneficial effects of the embodiments of the present invention are as follows:
[0050] The embodiments of this specification pre-construct a spatiotemporal convolution model for predicting the movement of targets in a monitoring network. When tracking a target, the location information of the target is first determined, and then the movement information of the target is calculated based on the location information, and the current multi-directional traffic data of the monitoring network is obtained. The pre-constructed spatiotemporal convolution model is used to calculate the current multi-directional traffic data and the movement information to obtain the predicted monitoring point where the tracking target will appear at the next moment in the monitoring network. The predicted arrival time of the tracking target at each predicted monitoring point is calculated based on the distance between each predicted monitoring point and the initial monitoring point and the movement information. For each predicted monitoring point, the target image feature is searched in the video image corresponding to the predicted arrival time, and the predicted monitoring point where the target image feature is found is used as the actual monitoring point. The embodiments of this specification implement accurate tracking of targets in a monitoring network. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0052] Figure 1It is a schematic diagram of a process of a multi-monitoring device active tracking method based on road network behavior law prediction in an embodiment of the present invention;
[0053] Figure 2 It is a schematic diagram of the structure of a multi-monitoring equipment active tracking device based on road network behavior law prediction in an embodiment of the present invention;
[0054] Figure 3 FIG. 2 is a schematic diagram showing the structure of a computer device in an embodiment of the present invention.
[0055]
Description of reference numerals
[0056] 201. Tracking target information receiving unit;
[0057] 202. Tracking target search unit;
[0058] 203, tracking target movement calculation unit;
[0059] 204. Prediction monitoring point calculation unit;
[0060] 205. predicted arrival time calculation unit;
[0061] 206. Actual monitoring point determination unit;
[0062] 302. Computer equipment;
[0063] 304. Processing equipment;
[0064] 306. Storage resources;
[0065] 308, driving mechanism;
[0066] 310, input / output module;
[0067] 312. Input devices;
[0068] 314. Output devices;
[0069] 316. Presentation equipment;
[0070] 318. Graphical user interface;
[0071] 320, network interface;
[0072] 322. Communication link;
[0073] 324. Communication bus. DETAILED DESCRIPTION
[0074] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0075] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, device, product or equipment that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0076] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0077] In order to solve the problems existing in the prior art, an embodiment of the present invention provides an active tracking method for multiple monitoring devices based on the prediction of road network behavior laws. The method uses multi-source data consisting of road network data, monitoring device external parameters, target archive data and external time feature data to encode the heterogeneous correlations and spatiotemporal dependencies between monitoring road network devices. A spatiotemporal graph convolution model consisting of a graph convolution network and a temporal neural network is used to construct a target spatiotemporal behavior law model under the monitoring road network. Target detection technology and image coordinate to road network longitude and latitude coordinate conversion technology are used to combine the target's moving direction and speed to assist in achieving fast and accurate sparse active tracking and trajectory characterization of multiple monitoring devices.
[0078] Figure 1 The figure shows a flow chart of a method for active tracking of multiple monitoring devices based on prediction of road network behavior rules according to an embodiment of the present invention. This figure describes the process of tracking a target in a monitoring road network, but it may include more or fewer operating steps based on conventional or non-creative work. The sequence of steps listed in the embodiment is only one way of executing the steps among many sequences, and does not represent the only execution sequence. When the system or device product is executed in practice, it may be executed in the sequence or in parallel according to the method shown in the embodiment or the accompanying drawings. Specifically, Figure 1As shown, the method may include:
[0079] Step 101: receiving target image features and initial monitoring points of a tracking target;
[0080] Step 102: determining continuous video image frames having the target image features from the video image of the initial monitoring point as target video image frames, and determining the position information of the tracked target in each target video image frame;
[0081] In this step, for the target to be tracked, firstly, according to the input target image features and the initial monitoring point, target detection and other technologies are used to perform feature comparison in the video image of the initial monitoring point, and the target is continuously located and tracked to obtain the specific position of the target in the video image (anchor frame coordinates B);
[0082] The anchor frame coordinate B is expressed as: B = [x B ,y B ,w B ,h B ].
[0083] Among them, (x B ,y B ) is the bottom coordinate of the target anchor box, w B Indicates the width of the target anchor box, h B Indicates the height of the target anchor box.
[0084] Step 103: Calculating the movement information of the tracked target in the monitoring road network where the initial monitoring point is located according to the position information of the tracked target in each target video image frame;
[0085] In an embodiment of the present specification, calculating the movement information of the tracked target in the monitoring road network where the initial monitoring point is located according to the position information of the tracked target in each target video image frame further includes:
[0086] For each target video image frame, convert the position information from the image coordinate system of the target video image frame to the longitude and latitude coordinate system of the monitoring road network where the initial monitoring point is located, to obtain the longitude and latitude coordinates of the tracking target;
[0087] The longitude and latitude coordinates of the tracked target corresponding to all target video image frames are fitted in the order of the target video image frames to obtain the movement information.
[0088] Specifically, for the video image of the detected target, the latitude and longitude coordinates of the target in the video image on the road network map are calculated using the image coordinate to road network longitude and latitude coordinate conversion method. Here, the field of view area coordinate mapping method is used. That is, first, the closest monitoring distance V of the monitoring device in the vertical field of view direction is calculated by the height, depression angle, horizontal field of view angle h and vertical field of view angle v of the monitoring device. min And the maximum monitoring distance V max , and then pass the bottom coordinate (x B ,y B ) Calculate the inclination angle θ and distance d of the target in the road network coordinate system relative to the monitoring device s Finally, according to the longitude lat, latitude lon, and direction of the monitoring device And the radius R of the earth to calculate the actual latitude and longitude of the target in the road network coordinate system;
[0089] The calculation formula of the inclination angle θ is: θ = x B ·hh / 2;
[0090] Relative distance d s Calculation formula:
[0091] Longitude lat′ calculation formula:
[0092] Latitude lon′ calculation formula:
[0093] Step 104: obtaining the current multi-directional traffic data of the monitoring road network, and calculating the current multi-directional traffic data and the movement information using a pre-built spatiotemporal convolution model to obtain a predicted monitoring point position where the tracking target will appear in the monitoring road network at the next moment;
[0094] In the embodiment of this specification, the step of constructing the spatiotemporal convolution model includes:
[0095] Acquire information parameters of each monitoring point in the monitoring road network, and calculate the spatial dependence characteristics of the monitoring road network according to the information parameters of each monitoring point;
[0096] Obtaining historical traffic data of each monitoring point, and calculating the time-dependent characteristics of the monitoring road network based on the historical traffic data of each monitoring point;
[0097] Performing feature fusion on the space-dependent features and the time-dependent features to obtain fused features;
[0098] The fused features are input into a two-layer graph convolution network for graph convolution calculation to obtain spatial dynamic correlation features of each monitoring point;
[0099] Acquiring external time characteristic data of the monitored road network;
[0100] The external time feature data and the spatial dynamic correlation features of each monitoring point are input into a recurrent neural network for calculation to obtain the spatiotemporal convolution model.
[0101] Specifically, the monitoring equipment is first calibrated, and the installation height, pitch angle, orientation, and longitude and latitude coordinates of the monitoring equipment are recorded. According to the road network data and the longitude and latitude coordinates of the monitoring equipment, the non-Euclidean road network distance and connectivity between the monitoring devices are obtained.
[0102] Then the spatial topological structure of the monitoring network is defined as an undirected graph G = (V, E, M), where represents the set of monitoring points in the monitoring road network, N represents the number of monitoring points, E∈e ij represents a set of edges from monitoring point i to point j, is the adjacency matrix, which represents the correlation weights between monitoring points. The larger the weight, the higher the correlation between the points.
[0103] Then, the spatial dependency of the monitoring network is established according to the road network distance, road network connectivity and orientation between the monitoring points, which are represented as the spatial distance graph G 1 =(V, E, M 1 ), spatial structure diagram G 2 =(V, E, M 2 ) and the spatial orientation graph G 3 =(V, E, M 3 ). Among them, the spatial distance graph G 1 The actual distance d between monitoring points ij The inverse of the spatial distance related adjacency matrix M 1 The weight K 1 , spatial structure correlation graph G 2 The neighborhood connectivity between monitoring points is used as the spatial structure related adjacency matrix M 2 The weight K 2 , space orientation graph G 3 The horizontal angle θ between the monitoring points ij (The maximum difference is 180 degrees) as the spatial orientation related adjacency matrix M 3 The weight K 3 ;
[0104] Weight K 1 Calculation formula: K 1 =1 / d ij ;
[0105] Weight K 2 Calculation formula:
[0106] Weight K 3 Calculation formula: K 3 =1-sinθ ij ;
[0107] Then, the time dependency of the monitoring network is established based on the target data of each monitoring point, and the target flow data of each monitoring point is aggregated every ten minutes (based on the average target speed of 1.2 meters per second and the average distance of the monitoring point of 800 meters) through a one-hour sliding window to obtain the historical flow time series. In addition, considering the directionality of the movement of the target between the monitoring points, for the target flow data of each cycle, according to the monitoring points where these targets appear in the next cycle (i.e., the outflow direction), the existing flow data is further grouped to obtain a multi-directional flow time series matrix.
[0108] Then, based on the external time feature data composed of working days, holidays, time, weather and temperature, the temporal autocorrelation, periodicity and trend modeling of the monitoring network are further enhanced. t Indicates the current day of the week, holiday item h t Indicates whether it is a holiday, the time item p t Indicates that the current time is a specific time of the day (each hour is a time period), and the weather item w t In terms of classification items, there are eight types of weather: sunny, cloudy, overcast, light rain, showers, fog and thunderstorm. The temperature item c t Indicates the current temperature in degrees Celsius;
[0109] Work Itemd t Formula: d t =r, r=1, 2, ..., 7;
[0110] Holiday Items t formula:
[0111] Time period item t Formula: t =s,s=1,2,...,24;
[0112] Weather Items t Formula: w t = k, k = 1, 2, ..., 8;
[0113] Then the adjacency matrix is standardized. First, the identity matrix I is added to the existing adjacency matrix M. N , to prevent the information of the graph nodes from being omitted, and in the identity matrix I NA trainable hyperparameter λ is added before. The larger the λ, the higher the importance of the node features itself, and the self-transfer adjacency matrix is obtained. Then use the degree matrix (Adjacency Matrix The degree of the adjacency matrix The rows and columns of the matrix are weighted averaged, so that nodes with low degrees have a greater influence on their neighbors, while nodes with high degrees have a smaller influence on their neighbors because their influence is dispersed to more neighbors, and finally the standardized adjacency matrix is obtained.
[0114] Self-transitive adjacency matrix Calculation formula:
[0115] Normalized adjacency matrix Calculation formula:
[0116] Then the normalized adjacency matrix Repeat the expansion L times to make it consistent with the multi-directional traffic timing matrix The graphs have the same dimension, and the feature fusion operation (matrix multiplication) is performed on multiple graphs respectively, and then input into the two-layer graph convolution network for graph convolution calculation to extract the spatial dynamic correlation between the monitored road network points. After the graph convolution calculation is completed, the spatial distance graph G 1 , spatial structure diagram G 2 and the spatial orientation graph G 3 The graph convolutional network output O i Perform weighted fusion, where the hyperparameter β of weighted fusion i It can also be learned;
[0117] Graph convolutional network calculation formula:
[0118] Among them, W 0 and W 1 is the convolution kernel; X represents the input feature, that is, the multi-directional traffic time series matrix.
[0119] Output graph weighted fusion calculation formula: O = β 1 O 1 +β 2 O 2 +β 3 O 3 ;
[0120] Among them, O represents the output after fusion.
[0121] Then, a recurrent neural network with an encoding and decoding structure (LSTM long short-term memory network is selected here) is used to establish the temporal dynamic correlation of the monitoring road network. First, each time step of the weighted fusion output graph O is connected through the flatten layer flatten(·) and the fully connected layer FCN(·), and input into the LSTM encoding layer, whose feature size is [L, N×N]. Then, the external time feature data E = [d, h, o, t] is normalized by norm(·) and fully connected, and input into each time step of the LSTM decoding layer. The LSTM neural network and the graph convolution network together form a spatiotemporal graph convolution model to model the spatiotemporal behavior of targets under the monitoring road network;
[0122] LSTM neural network calculation formula:
[0123] f(O, E) = LSTM decoder (LSTM encoder (FCN(flatten(O))))+norm(FCN(E))).
[0124] When calculating the predicted monitoring point, obtaining the current external time characteristic data of the monitoring road network;
[0125] Calculating the current multi-directional traffic data and the movement information using a pre-built spatiotemporal convolution model to obtain a predicted monitoring point where the tracking target will appear at the next moment in the monitoring road network further includes:
[0126] The current multi-directional traffic data and the current external time feature data are input into the spatiotemporal convolution model for calculation to obtain the predicted monitoring point where the tracking target will appear in the monitored road network at the next moment.
[0127] Step 105: Calculating the predicted arrival time of the tracking target at each predicted monitoring point according to the distance between each predicted monitoring point and the initial monitoring point and the movement information;
[0128] In the embodiment of the present specification, the movement information includes the movement direction and the movement speed, so the predicted arrival time of the tracking target to each predicted monitoring point can be calculated according to the distance and the movement speed.
[0129] Step 106: For each predicted monitoring point, the target image feature is searched in the video image corresponding to the predicted arrival time, and the predicted monitoring point where the target image feature is found is used as the actual monitoring point.
[0130] Specifically, searching for the target image feature in the video image corresponding to the predicted arrival time further includes:
[0131] Determine the corresponding time window based on the predicted arrival time;
[0132] The target image feature is searched in the video image corresponding to the time window.
[0133] In the embodiments of this specification, the time window may be an empirical value or a test value, such as 3 minutes, which is not limited in the embodiments of this specification. For example, the target image feature is searched in the video image corresponding to the time window of 3 minutes before and after the predicted arrival time (a total of 6 minutes).
[0134] In the embodiment of this specification, in order to speed up the search, the method further includes:
[0135] Calculate the probability of the tracking target appearing at each predicted monitoring point at the next moment;
[0136] For each predicted monitoring point, searching for the target image feature in the video image corresponding to the predicted arrival time further includes:
[0137] According to the order of probability values from large to small, the target image features are searched for the predicted monitoring points in the video image corresponding to the predicted arrival time.
[0138] In the embodiment of this specification, the spatiotemporal convolution model can predict the target flow probability under the current monitoring point, and weight this probability by the actual moving direction of the target (the probability of the monitoring point in the same direction as the target is greater), and obtain the monitoring point retrieval priority where the final target may appear at the next moment. The retrieval priority indicates the probability of the predicted monitoring point appearing at the tracking target at the next moment. This improves the search efficiency.
[0139] Furthermore, after finding the predicted monitoring point of the target image feature as the actual monitoring point, the method further includes:
[0140] Taking the actual monitoring point as the initial monitoring point, repeatedly performing the step of determining continuous video image frames having the target image feature from the video image of the initial monitoring point as target video image frames, until the target image feature does not appear in the video images corresponding to all the predicted monitoring points;
[0141] The movement information corresponding to all the initial monitoring points of the tracking target is fitted to obtain the complete movement information of the tracking target.
[0142] The embodiments of this specification pre-construct a spatiotemporal convolution model for predicting the movement of targets in a monitoring network. When tracking a target, the position information of the target is first determined, and then the movement information of the target is calculated based on the position information. The current multi-directional traffic data of the monitoring network is obtained based on the target data. The pre-constructed spatiotemporal convolution model is used to calculate the current multi-directional traffic data and the movement information to obtain the predicted monitoring point where the tracking target will appear at the next moment in the monitoring network. The predicted arrival time of the tracking target at each predicted monitoring point is calculated based on the distance between each predicted monitoring point and the initial monitoring point and the movement information. For each predicted monitoring point, the target image feature is searched in the video image corresponding to the predicted arrival time, and the predicted monitoring point where the target image feature is found is used as the actual monitoring point. The embodiments of this specification implement accurate tracking of targets in a monitoring network.
[0143] Based on the same inventive concept, the embodiment of the present invention also provides a multi-monitoring device active tracking device based on road network behavior law prediction, such as Figure 2 As shown, including:
[0144] The tracking target information receiving unit 201 is used to receive the target image features and initial monitoring points of the tracking target;
[0145] A tracking target searching unit 202 is used to determine continuous video image frames having the target image features from the video image of the initial monitoring point as target video image frames, and determine the position information of the tracking target in each target video image frame;
[0146] A tracking target movement calculation unit 203, configured to calculate movement information of the tracking target in the monitoring road network where the initial monitoring point is located according to the position information of the tracking target in each target video image frame;
[0147] The predicted monitoring point calculation unit 204 is used to obtain the current multi-directional flow data of the monitoring road network, and use the pre-built spatiotemporal convolution model to calculate the current multi-directional flow data and the movement information to obtain the predicted monitoring point where the tracking target will appear in the monitoring road network at the next moment;
[0148] The predicted arrival time calculation unit 205 is used to calculate the predicted arrival time of the tracking target to each predicted monitoring point according to the distance between each predicted monitoring point and the initial monitoring point and the movement information;
[0149] The actual monitoring point determination unit 206 is used to search for the target image feature in the video image corresponding to the predicted arrival time for each predicted monitoring point, and use the predicted monitoring point where the target image feature is found as the actual monitoring point.
[0150] The beneficial effects obtained by the above-mentioned device are consistent with the beneficial effects obtained by the above-mentioned method, and are not described in detail in the embodiments of the present invention.
[0151] like Figure 3 The structure diagram of the computer device of the embodiment of the present invention is shown. The apparatus in the present invention can be the computer device in the embodiment, and executes the method of the present invention described above. The computer device 302 may include one or more processing devices 304, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The computer device 302 may also include any storage resource 306, which is used to store any kind of information such as code, settings, data, etc. Non-limitingly, for example, the storage resource 306 may include any one or more combinations of the following: any type of RAM, any type of ROM, flash memory device, hard disk, optical disk, etc. More generally, any storage resource can use any technology to store information. Further, any storage resource can provide volatile or non-volatile retention of information. Further, any storage resource can represent a fixed or removable component of the computer device 302. In one case, when the processing device 304 executes an associated instruction stored in any storage resource or a combination of storage resources, the computer device 302 can perform any operation of the associated instruction. The computer device 302 also includes one or more drive mechanisms 308 for interacting with any storage resources, such as a hard disk drive mechanism, an optical disk drive mechanism, and the like.
[0152] The computer device 302 may also include an input / output module 310 (I / O) for receiving various inputs (via input devices 312) and for providing various outputs (via output devices 314). A specific output mechanism may include a presentation device 316 and an associated graphical user interface (GUI) 318. In other embodiments, the input / output module 310 (I / O), input device 312, and output device 314 may not be included, and the computer device 302 may be used as a computer device in a network. The computer device 302 may also include one or more network interfaces 320 for exchanging data with other devices via one or more communication links 322. One or more communication buses 324 couple the components described above together.
[0153] The communication link 322 may be implemented in any manner, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 322 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0154] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program implements the above method when executed by a processor.
[0155] An embodiment of the present invention further provides a computer-readable instruction, wherein when a processor executes the instruction, the program therein causes the processor to execute the above method.
[0156] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0157] It should also be understood that in the embodiments of the present invention, the term "and / or" is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in the present invention generally indicates that the associated objects before and after are in an "or" relationship.
[0158] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0159] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0160] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, or it can be an electrical, mechanical or other form of connection.
[0161] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present invention.
[0162] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0163] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0164] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of the present invention should not be understood as a limitation on the present invention.
Claims
1. A multi-monitoring device active tracking method based on road network behavior law prediction, characterized in that: include: Receive target image features and initial monitoring points of the tracked target; Determine continuous video image frames containing the target image features from the video image of the initial monitoring point as target video image frames, and determine the position information of the tracking target in each target video image frame; Calculate the movement information of the tracked target in the monitoring road network where the initial monitoring point is located according to the position information of the tracked target in each target video image frame; Acquire the current multi-directional traffic data of the monitoring road network, and use the pre-built spatiotemporal convolution model to calculate the current multi-directional traffic data and the movement information to obtain the predicted monitoring point position where the tracking target will appear in the monitoring road network at the next moment; Calculate the predicted arrival time of the tracking target at each predicted monitoring point according to the distance between each predicted monitoring point and the initial monitoring point and the movement information; For each predicted monitoring point, the target image feature is searched in the video image corresponding to the predicted arrival time, and the predicted monitoring point where the target image feature is found is used as the actual monitoring point.
2. The method according to claim 1, characterized in that Calculating the movement information of the tracked target in the monitoring road network where the initial monitoring point is located according to the position information of the tracked target in each target video image frame further includes: For each target video image frame, convert the position information from the image coordinate system of the target video image frame to the longitude and latitude coordinate system of the monitoring road network where the initial monitoring point is located, to obtain the longitude and latitude coordinates of the tracking target; The longitude and latitude coordinates of the tracked target corresponding to all target video image frames are fitted in the order of the target video image frames to obtain the movement information.
3. The method according to claim 1, characterized in that The steps of constructing the spatiotemporal convolution model include: Acquire information parameters of each monitoring point in the monitoring road network, and calculate the spatial dependence characteristics of the monitoring road network according to the information parameters of each monitoring point; Obtaining historical traffic data of each monitoring point, and calculating the time-dependent characteristics of the monitoring road network based on the historical traffic data of each monitoring point; Performing feature fusion on the space-dependent features and the time-dependent features to obtain fused features; The fused features are input into a two-layer graph convolution network for graph convolution calculation to obtain spatial dynamic correlation features of each monitoring point; Acquiring external time characteristic data of the monitored road network; The external time feature data and the spatial dynamic correlation features of each monitoring point are input into a recurrent neural network for calculation to obtain the spatiotemporal convolution model.
4. The method according to claim 3, characterized in that Inputting the external time feature data and the spatial dynamic correlation features of each monitoring point into a recurrent neural network for calculation, and obtaining the spatiotemporal convolution model further includes: Each time step of the spatial dynamic correlation feature is connected through a flattening layer and a fully connected layer and input into an encoding layer; The external time feature data is normalized and fully connected, and input into each time step of the decoding layer; The recurrent neural network is trained to obtain the spatiotemporal convolution model.
5. The method according to claim 3, characterized in that: The method further comprises: Acquiring current external time characteristic data of the monitored road network; Calculating the current multi-directional traffic data and the movement information using a pre-built spatiotemporal convolution model to obtain a predicted monitoring point where the tracking target will appear at the next moment in the monitoring road network further includes: The current multi-directional traffic data and the current external time feature data are input into the spatiotemporal convolution model for calculation to obtain the predicted monitoring point where the tracking target will appear in the monitored road network at the next moment.
6. The method according to claim 5, characterized in that The method further comprises: Calculate the probability of the tracking target appearing at each predicted monitoring point at the next moment; For each predicted monitoring point, searching for the target image feature in the video image corresponding to the predicted arrival time further includes: According to the order of probability values from large to small, the target image features are searched for the predicted monitoring points in the video image corresponding to the predicted arrival time.
7. The method according to claim 1, characterized in that After taking the predicted monitoring point of the target image feature as the actual monitoring point, the method further includes: Taking the actual monitoring point as the initial monitoring point, repeatedly performing the step of determining continuous video image frames having the target image feature from the video image of the initial monitoring point as target video image frames, until the target image feature does not appear in the video images corresponding to all the predicted monitoring points; The movement information corresponding to all the initial monitoring points of the tracking target is fitted to obtain the complete movement information of the tracking target.
8. The method according to claim 1, characterized in that Searching for the target image feature in the video image corresponding to the predicted arrival time further includes: Determine the corresponding time window based on the predicted arrival time; The target image feature is searched in the video image corresponding to the time window.
9. An active tracking device for multiple monitoring devices based on prediction of road network behavior rules, characterized in that: include: A tracking target information receiving unit, used to receive target image features and initial monitoring points of the tracking target; A tracking target searching unit, used to determine continuous video image frames having the target image features from the video image of the initial monitoring point as target video image frames, and determine the position information of the tracking target in each target video image frame; A tracking target movement calculation unit, used to calculate the movement information of the tracking target in the monitoring road network where the initial monitoring point is located according to the position information of the tracking target in each target video image frame; A predicted monitoring point calculation unit is used to obtain the current multi-directional flow data of the monitoring road network, and use a pre-built spatiotemporal convolution model to calculate the current multi-directional flow data and the movement information to obtain the predicted monitoring point where the tracking target will appear in the monitoring road network at the next moment; A predicted arrival time calculation unit, used to calculate the predicted arrival time of the tracking target to each predicted monitoring point according to the distance between each predicted monitoring point and the initial monitoring point and the movement information; The actual monitoring point determination unit is used to search for the target image feature in the video image corresponding to the predicted arrival time for each predicted monitoring point, and use the predicted monitoring point where the target image feature is found as the actual monitoring point.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Intelligent tracking method and intelligent tracking device for suspected target based on GIS (Geographic Information System)
CN104679864A
Trajectory generation method and device for query object
CN107315755A
Subway pedestrian flow network fusion method based on video pedestrian identification, and pedestrian flow prediction method
CN112541440A
Multi-camera multi-target trajectory prediction method based on traffic flow and graph convolution
CN115984334A
Vehicle track prediction method and system based on ReID and graph convolutional network
CN116824520A