Video Behavior Prediction Model Training Method, Prediction Method, Device, and Storage Medium

By constructing and processing undirected graph data of video images and updating parameters using Kalman filtering and smoothing models, the problem of training models when video image data is incomplete or tagged is chaotic is solved, and accurate prediction of video behavior is achieved.

CN117218164BActive Publication Date: 2025-07-01ZHUHAI ZHONGKE HUIZHI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310827785.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-06
Publication Date
2025-07-01
Estimated Expiration
2043-07-06

AI Technical Summary

Technical Problem

In the prior art, it is difficult to train an accurate video behavior prediction model when video image data is incomplete or tagged.

Method used

By determining the marking nodes of each frame of the image in the video, undirected graph data is constructed, and the marking node sequence is rearranged using the arrangement matrix to generate an observation image data sequence. Then, the data sequence is input into the Kalman filtered prediction sub-model and the Kalman smoothing sub-model, and the prediction parameters are updated until converge, and a trained video behavior prediction model is obtained.

Benefits of technology

In the case of incomplete video image data or chaotic marking, an accurate video behavior prediction model can be trained to achieve accurate prediction of video behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117218164B_ABST
    Figure CN117218164B_ABST
Patent Text Reader

Abstract

The present invention discloses a video image behavior prediction method, system, device and storage medium, relating to the field of artificial intelligence technology. The training method includes constructing undirected graph data for each frame of image according to the labeled nodes of each frame of image; arranging the order of the labeled nodes of the undirected graph data according to the permutation matrix to obtain observed image data; inputting the sequence of observed image data based on the video time sequence into the Kalman filter prediction sub-model to obtain a sequence of predicted image data, and correcting the sequence of predicted image data according to the sequence of observed image data; inputting the sequence of predicted image data into the Kalman smoothing sub-model to obtain a sequence of smoothed image data; updating the prediction parameters and correction parameters of the prediction sub-model according to the sequence of smoothed image data and the sequence of observed image data until the prediction parameters and correction parameters converge, so as to obtain a video behavior prediction model. The present invention can train an accurate video behavior prediction model when the video image data is incomplete or the labels are chaotic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, system, device and storage medium for predicting video image behaviors. Background Art

[0002] Image features are now commonly used for modeling and predicting object trajectories. In many instances, image algorithms are used for tracking and reasoning about the movement behaviors and trajectories between people or between people and objects. Another application of image feature learning is multi-object tracking, which is used to match the objects detected in different frames and add annotations.

[0003] The prediction of image time series is commonly used for modeling the variable relationships in multivariate time series data, such as predicting traffic conditions, and recognizing actions and gestures. Some current methods for predicting the real image structure across time are based on using Gaussian process regression, neural network models, etc. on image and network data. These methods require all nodes of the image to be accurately registered and marked without confusion during the training process. When the image data is incomplete or the markings are confused, the model cannot accurately predict image behaviors. Summary of the Invention

[0004] The present invention aims to at least solve one of the technical problems existing in the prior art. For this purpose, the present invention provides a method, system, device and storage medium for training a video behavior prediction model, aiming to train an accurate video behavior prediction model when the video image data is incomplete or the markings are confused.

[0005] On the one hand, an embodiment of the present invention provides a method for training a video behavior prediction model, including the following steps:

[0006] Determine the marked nodes of each frame image in the video, and construct undirected graph data for each frame image according to the marked nodes;

[0007] Rearrange the order of the marked nodes of the undirected graph data according to the permutation matrix to obtain observed image data;

[0008] Arrange the observed image data corresponding to the images in the order of video time to obtain an observed image data sequence;

[0009] Input the observed image data sequence into a Kalman filter prediction sub-model to obtain a predicted image data sequence, and correct the predicted image data sequence according to the observed image data sequence;

[0010] Input the predicted image data sequence into a Kalman smoothing sub-model to obtain a smoothed image data sequence;

[0011] Update the prediction parameters and correction parameters of the Kalman filter prediction sub-model according to the smoothed image data sequence and the observed image data sequence until the prediction parameters and correction parameters converge, and obtain a trained video behavior prediction model.

[0012] According to some embodiments of the present invention, the determining the labeled nodes of each frame of image in the video and constructing the undirected graph data of each frame of image according to the labeled nodes includes the following steps:

[0013] Use a bounding box to label the objects in each frame of image as labeled nodes, and annotate the bounding box;

[0014] Use the linear interpolation method to fill the bounding boxes of all frames in the video;

[0015] Construct the undirected graph data of each frame of image according to all the bounding boxes in the image, where the undirected graph data includes an edge matrix and an attribute matrix, the edge matrix is mapped from an adjacency matrix representing the relationship between labeled nodes, and the attribute matrix is used to characterize the coordinate positions of the labeled nodes.

[0016] According to some embodiments of the present invention, the reordering the labeled node order of the undirected graph data according to the permutation matrix to obtain the observed image data includes the following steps:

[0017] Obtain a set of permutation matrices, where the set of permutation matrices includes the n×n permutation matrix P corresponding to all the undirected graph data, n is the number of labeled nodes in the undirected graph data, and the permutation matrix is a permutation matrix;

[0018] Map the edge matrix and the attribute matrix of the undirected graph data according to the permutation matrix corresponding to each undirected graph data to reorder the labeled node order, and obtain the observed image data.

[0019] According to some embodiments of the present invention, the inputting the observed image data sequence into the Kalman filter prediction sub-model to obtain the predicted image data sequence and correcting the predicted image data sequence according to the observed image data sequence includes the following steps:

[0020] Input the observed image data in the observed image data sequence into the dynamic system of the Kalman filter prediction sub-model in turn for prediction to obtain the predicted image data, and obtain the predicted image data sequence according to multiple predicted image data;

[0021] Use the observed image data at the moment corresponding to the predicted image data to correct the predicted image data, and obtain the predicted image data sequence according to multiple corrected predicted image data.

[0022] According to some embodiments of the present invention, the dynamic system is expressed as:

[0023]

[0024] Among them, w represents the edge matrix, v represents the attribute matrix, (e) and (n) respectively represent the models regarding edges and nodes, B and C both represent the parameters of the dynamic system, and u t is the system noise of a random standard normal distribution.

[0025] According to some embodiments of the present invention, the mapping representation of the permutation matrix in the undirected graph data is:

[0026] (P, X) = (P, (w, v)) = (P * w, Pv);

[0027] Among them, X represents the undirected graph data, w represents the edge matrix, v represents the attribute matrix, and P represents the permutation matrix.

[0028] According to some embodiments of the present invention, the method for training the video behavior prediction model further includes the following steps:

[0029] When comparing two consecutive images with the number of labeled nodes being n1 and n2 respectively in the video, empty nodes are introduced into the undirected graph data of the two images so that the total number of labeled nodes in the undirected graph data of the two images is both n1 + n2;

[0030] The true labeled nodes of the undirected graph data of one image are paired with the empty nodes of the undirected graph data of the other image to obtain a pairing result, and the pairing result represents the birth of new labeled nodes or the deletion of old labeled nodes.

[0031] On the other hand, embodiments of the present invention also provide a video behavior prediction method, including the following steps:

[0032] Obtain the video data to be predicted;

[0033] Preprocess the video data to obtain an undirected graph data sequence;

[0034] Input the undirected graph data sequence into the video behavior prediction model as described in the previous embodiments to obtain a video behavior prediction result.

[0035] On the other hand, embodiments of the present invention also provide an electronic device, including:

[0036] At least one processor;

[0037] At least one memory for storing at least one program;

[0038] When the at least one program is executed by the at least one processor, the at least one processor implements the method for training the video behavior prediction model or the video behavior prediction method as described in the previous embodiments.

[0039] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions for causing a computer to execute the video behavior prediction model training method or the video behavior prediction method as described in the previous embodiments.

[0040] At least one of the above technical solutions of the present invention has the following advantages or beneficial effects: First, determine the labeled nodes of each frame image in the video, and construct the undirected graph data of each frame image according to the labeled nodes. Then, rearrange the order of the labeled nodes of the undirected graph data according to the permutation matrix to obtain the observed image data, so as to reduce the situation of chaotic labeling of the labeled nodes in the image. Arrange the observed image data corresponding to the images in the order of video time to obtain an observed image data sequence. Then, input the observed image data sequence into the Kalman filter prediction sub-model to obtain a predicted image data sequence, and correct the predicted image data sequence according to the observed image data sequence. Then, input the predicted image data sequence into the Kalman smoothing sub-model to obtain a smoothed image data sequence. Update the prediction parameters and correction parameters of the Kalman filter prediction sub-model according to the smoothed image data sequence and the observed image data sequence until the prediction parameters and correction parameters converge, and obtain a trained video behavior prediction model. The present invention can still train an accurate video behavior prediction model even when the video image data is incomplete or the labeling is chaotic, so as to achieve accurate prediction of video behavior. Description of the Drawings

[0041] Figure 1 is a flowchart of the video behavior prediction model training method provided by an embodiment of the present invention;

[0042] Figure 2 is a schematic diagram of the undirected graph provided by an embodiment of the present invention;

[0043] Figure 3 is a schematic diagram of the undirected graph node rearrangement process provided by an embodiment of the present invention;

[0044] Figure 4 is a schematic diagram of the image prediction and prediction parameter update process provided by an embodiment of the present invention;

[0045] Figure 5 is a schematic diagram of the video behavior prediction process provided by an embodiment of the present invention;

[0046] Figure 6 is a schematic diagram of the change of edge attributes and node attributes over time provided by an embodiment of the present invention;

[0047] Figure 7 is a schematic diagram of the change of edge attribute error and node attribute error over time provided by an embodiment of the present invention;

[0048] Figure 8 It is a comparison schematic diagram of true values, observed values, and predicted values under the first set of simulation data provided by an embodiment of the present invention;

[0049] Figure 9 It is a comparison schematic diagram of true values, observed values, and predicted values under the second set of simulation data provided by an embodiment of the present invention;

[0050] Figure 10 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0051] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0052] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, left, right, etc., is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0053] In the description of the present invention, if the first, second, etc. are described, they are only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0054] An embodiment of the present invention provides a method for training a video behavior prediction model. The training method in the embodiments of the present application can be applied to a terminal, a server, or software running on a terminal or a server. The terminal can be a tablet computer, a laptop computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0055] Referring to Figure 1 , the method for training a video behavior prediction model according to an embodiment of the present invention includes but is not limited to step S110, step S120, step S130, step S140, step S150, and step S160.

[0056] Step S110: Determine the marked nodes of each frame image in the video, and construct undirected graph data for each frame image according to the marked nodes;

[0057] Step S120: Rearrange the order of the marked nodes of the undirected graph data according to the permutation matrix to obtain the observed image data;

[0058] Step S130: Arrange the observed image data corresponding to the images in the order of video time to obtain a sequence of observed image data;

[0059] Step S140: Input the sequence of observed image data into the Kalman filter prediction sub-model to obtain a sequence of predicted image data, and correct the sequence of predicted image data according to the sequence of observed image data;

[0060] Step S150: Input the sequence of predicted image data into the Kalman smoothing sub-model to obtain a sequence of smoothed image data;

[0061] Step S160: Update the prediction parameters and correction parameters of the Kalman filter prediction sub-model according to the sequence of smoothed image data and the sequence of observed image data until the prediction parameters and correction parameters converge, and obtain a trained video behavior prediction model.

[0062] According to some embodiments of the present invention, in step S110, the step of determining the marked nodes of each frame image in the video and constructing undirected graph data for each frame image according to the marked nodes includes, but is not limited to, the following steps:

[0063] Step S210: Use a bounding box to label the objects in each frame image as marked nodes, and annotate the bounding box;

[0064] Step S220: Use the linear interpolation method to fill in the bounding boxes of all frames in the video;

[0065] Step S230: Construct undirected graph data for each frame image according to all the bounding boxes in the image, where the undirected graph data includes an edge matrix and an attribute matrix, the edge matrix is mapped from an adjacency matrix representing the relationship between the marked nodes, and the attribute matrix is used to characterize the coordinate positions of the marked nodes.

[0066] Specifically, in the preprocessing stage of the images in the embodiments of the present invention, the deTR software can be used to label the objects in each frame of the scene, and further use the davinci Resolve video editor to annotate the bounding boxes of the key frames. Then, the linear interpolation method is used to fill in the bounding boxes of all frames in the video, and the undirected graph data corresponding to each frame image in the video is obtained according to all the bounding boxes in the image. Refer to Figure 2 , the association relationship of the marked nodes in the image is represented as an undirected graph.Figure 2 The circles marked as 1, 2, 3, and 4 respectively represent marked nodes. The two marked nodes connected by each line represent an associated relationship. The thicker the line, the greater the associated relationship.

[0067] In the embodiments of the present invention, two sets of variables are used to represent an undirected graph with n nodes. The two sets of variables are an adjacency matrix and a node attribute vector respectively. Among them, the node attributes include the coordinate positions (x, y) of the nodes in the image, and the adjacency matrix represents the attributes of the edges. Each element in the adjacency matrix represents the interaction between node i and node j. Assume that {A ij} is a symmetric matrix with diagonal elements being 0 and having degrees of freedom in . Therefore, the upper triangular matrix part of the adjacency matrix A can be replaced by a long vector: . That is, there is a one-to-one relationship between the adjacency matrix A and the edge matrix w. φ(A) = w represents the mapping from A to w, that is, φ -1 (w) = A. The attribute matrix in the embodiments of the present invention, that is, the node attribute vector Each element v in the attribute matrix i represents the attribute of the i-th node of the undirected graph.

[0068] Furthermore, in the embodiments of the present invention, X is used to represent the undirected graph data. X = (w, v) ∈ χ, representing the space χ n is a Euclidean space using the standard Euclidean metric. The distance in χ is defined as: for any X (1) ≡ (w (1) , v (1) ), the distance between X (2) ≡ (w (2) , v (2) ) is as follows:

[0069]

[0070] where λ > 0 is the weight related to the second term, and d x is the weighted metric of the combination of the edges and the node attributes. For the set of all graphs with different marked nodes, the present invention is defined as

[0071] According to some embodiments of the present invention, in step S120, the step of reordering the marked node order of the undirected graph data according to the permutation matrix to obtain the observed image data includes, but is not limited to, the following steps:

[0072] Step S310: Obtain a set of permutation matrices, where the set of permutation matrices includes n×n permutation matrices P corresponding to all undirected graph data, n being the number of labeled nodes in the undirected graph data, and the permutation matrix being a permutation matrix;

[0073] Step S320: Map the edge matrix and attribute matrix of the undirected graph data according to the permutation matrix corresponding to each undirected graph data to rearrange the order of the labeled nodes, and obtain the observed image data.

[0074] Specifically, in the embodiment of the present invention, the labeled nodes in the undirected graph data are rearranged to solve the labeling problem that may be caused by different numbers of nodes in two associated graphs. Define the set of permutation matrices P n as the set of all n×n permutation matrices P, and the permutation matrix P ∈ P n is a permutation matrix, that is, there is only one 1 in each row and each column of the permutation matrix, and all other elements are 0. Set the permutation matrix P to rearrange the order of the labeled nodes in the image, so that the nodes corresponding to the target in the undirected graph data corresponding to each frame of image will not be confused when used as the observed data subsequently. Exemplarily, referring to Figure 3 , the target in the current frame of image is the labeled node 2, but it is possible that in the next frame of image, the target is considered to be the labeled node 3. However, if through the mapping of P, if the node corresponding to the target in each image is 1 at the beginning, it will always be 1 afterwards. The mapping of v→Pv will be rearranged according to the elements in P, thereby changing the sorting of the nodes in the graph, where the sorting in the node attributes will not change. At the same time, the corresponding adjacency matrix changes to A→PAP T , that is, the corresponding edge matrix w becomes φ(Pφ -1 (w)P T ), and the embodiment of the present invention can use P*w to represent the changed w.

[0075] Furthermore, P n acts on each permutation matrix P in the set of labeled permutation matrices on χ n to map each undirected graph data corresponding to χ n The process is expressed as:

[0076] (P,X)=(P,(w,v))=(P*w,Pv)∈χ;

[0077] where X represents the undirected graph data, w represents the edge matrix, v represents the attribute matrix, and P represents the permutation matrix.

[0078] After the undirected graph data is mapped, the image space is defined as the quotient space G n =χ n / P n. In linear algebra, the quotient of a vector space V with respect to a subspace N is the vector space obtained by "collapsing" N to zero, and the resulting vector space is called the quotient space, denoted as V / N. The elements of the quotient space in the embodiments of the present invention represent the permutation orbits of images, denoted as [X] = (P*w, Pv)|P ∈ P n . The distance between any two undirected graphs of size n in the quotient space is expressed as:

[0079]

[0080] Further, with respect to P ∈ P n Regarding the problem of optimizing wireless graph data, the embodiments of the present invention use the Umeyama algorithm for optimization. Through this algorithm, image data with variable numbers of labeled nodes can be processed, such as Figure 2 The derivation process of undirected graph permutation. The Umeyama algorithm is expressed as:

[0081]

[0082] For q i and p i in the above formula being the predicted nodes in the image, for the embodiments of the present invention, when t = 0, no translation is performed on the nodes in the image. Where c is the gain coefficient of the coefficient R to be obtained, the above formula changes to:

[0083]

[0084] The output parameter of this algorithm formula is cR, which is the above-mentioned permutation matrix P.

[0085] Finding the minimum value, the final image space is the union of spaces of different sizes:

[0086] According to some embodiments of the present invention, the method for training a video behavior prediction model in the embodiments of the present invention further includes but is not limited to the following steps:

[0087] Step S410, when comparing two consecutive images with the numbers of labeled nodes being n1 and n2 respectively in the video, introduce empty nodes into the undirected graph data of the two images so that the total numbers of labeled nodes in the undirected graph data of the two images are both n1 + n2;

[0088] Step S420, pair the real labeled nodes of the undirected graph data of one image with the empty nodes of the undirected graph data of the other image to obtain a pairing result, and the pairing result represents the birth of new labeled nodes or the deletion of old labeled nodes.

[0089] Specifically, when comparing the undirected graph data corresponding to two images with the number of labeled nodes being n1 and n2 respectively, it is necessary to introduce empty nodes for the two undirected graph data respectively, so that the total number of nodes in both undirected graph data becomes n1 + n2, and then use the above Umeyama algorithm for two-point spatial matching. The pairing result of the real nodes of one undirected graph data and the empty nodes of the other undirected graph data can result in a new node being born or an old node being deleted in two consecutive frames of images.

[0090] According to some embodiments of the present invention, in step S140, the step of inputting the observed image data sequence into the Kalman filter prediction sub-model to obtain the predicted image data sequence and correcting the predicted image data sequence according to the observed image data sequence includes the following steps:

[0091] Step S510, input the observed image data in the observed image data sequence into the dynamic system of the Kalman filter prediction sub-model in sequence for prediction to obtain the predicted image data, and obtain the predicted image data sequence according to multiple predicted image data;

[0092] Step S520, correct the predicted image data with the observed image data at the moment corresponding to the predicted image data, and obtain the predicted image data sequence according to multiple corrected predicted image data.

[0093] Specifically, the purpose of Bayesian time series image data analysis is to predict X = x t ∈χ: t = 1, 2,..., T from the observation data Y = y t : t = 1,..., T. Embodiments of the present invention use the Kalman filter prediction sub-model for prediction and adopt the Kalman smoothing algorithm to update the prediction sub-model parameters. The prediction sub-model parameters θ = {B {l} , C {l} , E {l} , F {l} : l = e, n}, where l = e represents the edge and l = n represents the node. In the Kalman filter prediction sub-model, a dynamic system S for image prediction tracking and a dynamic system O for observed data are established as follows:

[0094]

[0095]

[0096] Among them, (e) and (n) respectively represent the models regarding the edge and the node, P t+1 ∈P n is a random permutation matrix, and B, C, E, and F are all prediction sub-model parameter matrices that need to be trained. B and E are respectively the transition matrices of the real sequence and the observation sequence, Q = CC' and Λ = FF' are covariance matrices. and is the system noise of a random standard normal distribution. In the embodiments of the present invention, it is assumed that obeys a multivariate normal distribution with a mean of 0 and the same covariance. W t+1 is a matrix used to encode by discarding real nodes or adding false nodes. x t = w t , v t represents the undirected graph data corresponding to the image at time t, represents the noisy observation data at time t.

[0097] Exemplarily, with reference to Figure 4 , the image prediction and prediction parameter update of the embodiments of the present invention are described as follows:

[0098] S1. Input the observation image data sequence Y = {y t : t = 1, 2,..., T}, initialize the first item of the predicted image data sequence , that is and initialize K 1|1 = I;

[0099] S2. Starting from the second frame of the image, use the Umeyama algorithm to rearrange the labeled nodes of the observation image data to obtain a new observation image data sequence. The rearrangement process is expressed as:

[0100]

[0101]

[0102] S3. Input the observation image data sequence from time t = 1 to t = T into the Kalman filter prediction sub-model for state prediction and measurement update to obtain the predicted image data sequence;

[0103] The state prediction process is expressed as:

[0104]

[0105] K T+1|t = B (l) K t|t B (l)′ + Q (l) Q (l) = C (l) C (l)′ ;

[0106] Among them, K t+1|t is the mean square error of the prior prediction x^(t + 1).

[0107] The measurement update process is expressed as:

[0108] S t = W {t] E (l) K (t+1|t ) * E (l) W′ t + W t A (l) W_t′;

[0109]

[0110]

[0111] K t+1|t+1 = K t+1|t - G t W t E (l) K t+1|t Λ (l) = F (l) F (l)′ ;

[0112] S4. Smooth the predicted image data obtained at each moment in the predicted image data sequence from time t = T - 1 to t = 1 according to the Kalman smoothing algorithm to obtain a smoothed image data sequence The specific process is as follows:

[0113]

[0114]

[0115]

[0116] S5. Update the predicted sub - model parameters based on the parameter prediction of Kalman smoothing according to the observed image data sequence and the smoothed image data sequence Until the predicted sub - model parameters converge, the update of the predicted sub - model parameters is expressed as:

[0117]

[0118]

[0119]

[0120]

[0121] Among them, B is the state transition matrix of the actual value, and E is the state transition matrix of the observed value.

[0122] Refer to Figure 5, during the prediction process of the video behavior prediction model according to the embodiments of the present invention, if a new target (target object) is added, the new target needs to be processed through the above node arrangement process, and then enter the Kalman prediction filtering sub-model to obtain the predicted association relationship network as the output result.

[0123] The following uses simulated data to illustrate the effect of the video behavior prediction model according to the embodiments of the present invention.

[0124] To simulate the real situation, the embodiments of the present invention generate a set of sorted image sequence data, with a total of n = 10 nodes and t = 5000 time points. The node attribute v t contains the position information of the target in the scene (m = 2), and the edge attribute w t is a scalar representing the strength of the connection between nodes (p = 1). To simulate the entire image time series, the embodiments of the present invention first randomly generate a v0 in a unit circle, and then generate a w0 in a continuous uniform distribution. Subsequently, the parameters of the model are designed such that the elements in M follow a standard normal distribution, k n = n(n - 1) / 2, B (n) = I 2m , Q (n) = 100I 2n , Λ (n) = 10I 2n , E (e) = 3M + I 2n , the elements in M follow a normal distribution of N(0, 0.01).

[0125] The embodiments of the present invention use the L 2 norm to compare the evolution of images. As Figure 6 shown, the upper part is ‖w t - w t+1 ‖, which is the change of the edge attribute at different time points, and the lower part is ‖v t - v t+1 ‖, which is the change of the node over time. The larger the value, the greater the change.

[0126] As Figure 7 shown, the embodiments of the present invention compare the observed values and the true values v t , w t of the error, and also use the L 2 norm for calculation. The upper part of the picture is the error between the observed edge attribute and the true edge attribute, and the lower part is the error between the observed node and the true node.

[0127] As Figure 8As shown, for each time t, the embodiment of the present invention shows the true value x t (blue), the observed value y t (pink), and the predicted value (green). The thickness of each line is positively correlated with the value of the edge attribute.

[0128] Furthermore, the embodiment of the present invention modifies the above first set of simulation data to obtain a second set of simulation data, that is, randomly selects some time nodes based on the first set of simulation data and randomly removes 2-3 nodes in the selected time period, so as to simulate the detection of loss or the entry or exit of the target over time. This data is data with lost and incorrect nodes. Figure 9 The color of represents and Figure 8 The meaning of the edge line and the color is the same. At the time point t = 239, it can be seen that each node of the true value (blue) is connected to other nodes. However, there are 3 empty nodes in the observed value (pink) (these three empty nodes are not connected to other nodes), representing the missing of the observed value at this moment, and the predicted value (green) of the embodiment of the present invention can solve the problem of these three empty nodes.

[0129] To better compare the training method of the embodiment of the present invention, the embodiment of the present invention uses other methods for time series prediction, that is, uses training with multiple iterations, such as recurrent neural network (RNN), gated recurrent unit (GRU), and Transformer model. The embodiment of the present invention uses the online learning method for training, that is, only one step of prediction for the next graph is performed in each iteration. The Kalman filter and Kalman smoothing methods belong to this method. Finally, the embodiment of the present invention tests static training and uses the median filter for prediction. The comparison results of different methods are as follows:

[0130] Table 1 Comparison results of training methods

[0131]

[0132] As can be seen from Table 1, the Kalman filter / smoothing method used in the embodiment of the present invention has excellent performance, with smaller errors than other methods both in the data without missing values and in the data with missing values. Compared with the deep learning black box model, the method proposed based on the Kalman filter in the embodiment of the present invention will be more controllable, and the dynamic mapping matrix obtained in the prediction process of this method can help users better understand the change process of nodes between different time periods and better observe the influencing factors of data on nodes.

[0133] On the other hand, the embodiment of the present invention also provides a video behavior prediction method, including the following steps:

[0134] Obtain video data to be predicted;

[0135] Preprocess the video data to obtain an undirected graph data sequence;

[0136] Input the undirected graph data sequence into the video behavior prediction model as in the previous embodiments to obtain the video behavior prediction result.

[0137] Refer to Figure 10 , Figure 10 , which is a schematic diagram of an electronic device provided by an embodiment of the present invention. The electronic device according to the embodiment of the present invention includes one or more control processors and a memory. Figure 10 In

[0138] , a control processor and a memory are taken as an example. Figure 10 In

[0139] , the control processor and the memory can be connected by a bus or other means.

[0140] Those skilled in the art can understand that Figure 10 , the device structure shown in

[0141] does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine some components, or arrange different components.

[0142] The non-transitory software program and instructions required to implement the video behavior prediction model training method or the video behavior prediction method applied to the electronic device in the above embodiments are stored in the memory. When executed by the control processor, the video behavior prediction model training method or the video behavior prediction method applied to the electronic device in the above embodiments is executed.

[0143] Those of ordinary skill in the art will understand that all or some of the steps and systems disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes but is not limited to RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those of ordinary skill in the art that communication media typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery media.

[0144] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the relevant technical field.

Claims

1. A method for training a video behavior prediction model, characterized in that, Including the following steps: Determine the marked nodes of each frame image in the video, and construct undirected graph data for each frame image according to the marked nodes. The undirected graph data includes an edge matrix, and the edge matrix is used to characterize the association relationship between the marked nodes in the image; Map the edge matrix of the undirected graph data according to the permutation matrix to adjust the association relationship between the marked nodes in the image, and obtain the observed image data, specifically including: obtaining a set of permutation matrices, where the set of permutation matrices includes the n×n permutation matrix P corresponding to all undirected graph data, n is the number of marked nodes in the undirected graph data, and the permutation matrix is a permutation matrix; mapping the edge matrix of the undirected graph data according to the permutation matrix corresponding to each undirected graph data to re-arrange the marked node order, and obtaining the observed image data; Arrange the observed image data corresponding to the images in the order of video time to obtain a sequence of observed image data; Input the sequence of observed image data into the Kalman filter prediction sub-model to obtain a sequence of predicted image data, and correct the sequence of predicted image data according to the sequence of observed image data; Input the sequence of predicted image data into the Kalman smoothing sub-model to obtain a sequence of smoothed image data; Update the prediction parameters and correction parameters of the Kalman filter prediction sub-model according to the sequence of smoothed image data and the sequence of observed image data until the prediction parameters and correction parameters converge, and obtain a trained video behavior prediction model.

2. The video behavior prediction model training method according to claim 1, wherein The step of determining the marked nodes of each frame image in the video and constructing undirected graph data according to the marked nodes includes the following steps: Use bounding boxes to label the objects in each frame image as marked nodes, and annotate the bounding boxes; Use the linear interpolation method to fill the bounding boxes of all frames in the video; Construct undirected graph data for each frame image according to all the bounding boxes in the image, where the undirected graph data includes an edge matrix and an attribute matrix, the edge matrix is mapped from an adjacency matrix representing the relationship between the marked nodes, and the attribute matrix is used to characterize the coordinate positions of the marked nodes.

3. The video behavior prediction model training method according to claim 2, characterized in that The step of inputting the sequence of observed image data into the Kalman filter prediction sub-model to obtain a sequence of predicted image data and correcting the sequence of predicted image data according to the sequence of observed image data includes the following steps: Input the observed image data in the sequence of observed image data into the dynamic system of the Kalman filter prediction sub-model for prediction to obtain predicted image data, and obtain a sequence of predicted image data according to multiple predicted image data; Correct the predicted image data with the observed image data at the moment corresponding to the predicted image data, and obtain a sequence of predicted image data according to multiple corrected predicted image data.

4. The video behavior prediction model training method according to claim 3, wherein The dynamic system is expressed as: Among them, w represents the edge matrix, v represents the attribute matrix, (e) and (n) respectively represent the models regarding edges and nodes, B and C both represent the parameters of the dynamic system, and u t is the system noise of a random standard normal distribution.

5. The video behavior prediction model training method according to claim 2, characterized in that The mapping of the permutation matrix on the undirected graph data is expressed as: (P,X)=(P,(w,v))=(P*w,Pv); where X represents the undirected graph data, w represents the edge matrix, v represents the attribute matrix, and P represents the permutation matrix.

6. The video behavior prediction model training method according to claim 1, characterized in that The video behavior prediction model training method further includes the following steps: When comparing two consecutive images in a video with the number of labeled nodes being n1 and n2 respectively, null nodes are introduced into the undirected graph data of the two images so that the total number of labeled nodes in the undirected graph data of the two images is both n1 + n2; Pair the true labeled nodes of the undirected graph data of one image with the null nodes of the undirected graph data of the other image to obtain a pairing result, which indicates the birth of new labeled nodes or the deletion of old labeled nodes.

7. A video behavior prediction method, characterized in that, It includes the following steps: Obtain the video data to be predicted; Preprocess the video data to obtain an undirected graph data sequence; Input the undirected graph data sequence into the video behavior prediction model as claimed in claim 1 to obtain a video behavior prediction result.

8. An electronic device, characterized in that, It includes: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, it enables at least one of the processors to implement the video behavior prediction model training method as claimed in any one of claims 1 to 6 or the video behavior prediction method as claimed in claim 7.

9. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to implement the video behavior prediction model training method as claimed in any one of claims 1 to 6 or the video behavior prediction method as claimed in claim 7.

Citation Information

Patent Citations

  • Pedestrian abnormal behavior detection method and device suitable for inspection vehicle

    CN112149618A

  • Coal mine personnel behavior detection method and device, and storage medium

    CN114038067A