Cluster agent motion prediction method based on rotation isovariant convolutional neural network

By introducing rotational denaturation into the CNN model, replacing the convolution operator and designing the model architecture of encoder, converter and decoder, the problem of poor prediction performance in the existing technology under rotational is solved, and more efficient and accurate spatio-temporal sequence prediction is achieved.

CN119941787AInactive Publication Date: 2025-05-06TONGJI UNIV

Patent Information

Application Number
CN202411935379.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art ignores isodenativity under geometric transformation in spatiotemporal sequence prediction, resulting in poor prediction performance of the model under rotational isovariable.

Method used

Integrate rotation and other degeneration into the pure CNN model framework, and by replacing the traditional convolution operator with rotation and equal variable convolution operator, the model architecture of encoder, converter and decoder is designed to enhance the model's adaptability to rotation changes.

Benefits of technology

The model's ability to characterize spatiotemporal data is improved, prediction error is reduced, prediction accuracy and robustness are enhanced, and the high parallel efficiency of CNN is maintained.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941787A_ABST
    Figure CN119941787A_ABST
Patent Text Reader

Abstract

The invention discloses a cluster agent motion prediction method based on a rotation isovariant convolutional neural network, and the method comprises the steps: abstracting large-scale cluster motion prediction into a density map sequence prediction task, namely, predicting a future frame density map through an observed historical frame density map; a prediction model is constructed, the prediction model comprises an encoder, a converter and a decoder, and the encoder and the decoder are in jump connection; the converter is a rotation equivariant converter of a converter in a SimVP model, a SimVPv2 model and a TAU model, and a traditional convolution operator is replaced by an equivariant convolution operator. According to the invention, a more accurate prediction result is realized in the fields of multi-agent cluster motion trail prediction and the like, and an effect of considering both training efficiency and prediction accuracy is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular to a cluster intelligent body motion prediction method based on a rotation equivariant convolutional neural network. Background Art

[0002] Swarm behavior refers to the collective behavior dynamics exhibited by a group of intelligent agents (such as flocks of birds, schools of fish, and even self-driving vehicles) through interaction with neighboring individuals and their environment without the control of a central regulator. In recent years, significant progress has been made in the swarm application of autonomous systems, involving cluster control, formation target tracking, and robot collaboration. Expanding this concept, a swarm can be viewed as a distributed multi-agent system that is able to self-organize and operate in an orderly manner. In a broad sense, population flow, traffic flow, temperature changes, and even particle motion can be regarded as swarm behavior. Predicting the motion trajectory of a multi-agent swarm is crucial to revealing the motion dynamics behind it, which helps unmanned systems make more informed decisions and ensure the safety of operations.

[0003] Since there are usually a large number of agents in a group, it is more practical and meaningful to regard the group as a whole and focus on its overall spatial distribution and motion pattern. Using density maps to predict group motion is an effective solution because it can intuitively show distribution and dynamic changes. However, when predicting the density map sequence of agents, the model needs to have the ability to deeply understand the potential information in spatial and temporal dimensions, which adds complexity to the spatiotemporal prediction task. The key to designing an effective model is to capture the spatial and temporal correlation of features at the same time. From the perspective of video prediction, many methods have tried to address this challenge through complex module designs or training strategies. Some methods based on recurrent neural networks (RNNs) predict future frames in an autoregressive manner, but may face problems of insufficient parallel capabilities and increased time complexity. At present, there are also models that only use convolutional neural networks (CNNs) to provide a more efficient training process while maintaining performance.

[0004] In previous studies of spatiotemporal sequence prediction, an often overlooked but crucial principle is that the prediction model needs to maintain equivariance under geometric transformations. If equivariance is destroyed in the initial stage, the entire model will lose this property. The concept of equivariance is closely related to geometry and symmetry, and many natural phenomena exhibit equivariance due to their underlying geometric properties. In the field of machine learning, a large amount of data also has equivariance, such as Euclidean coordinates and graph structured data. EGNN introduces a simple equivariant information transfer method in the field of graph neural networks without relying on high-order representations that are computationally expensive. ReDet proposes a rotation equivariant detector for aerial target detection, which simultaneously encodes rotation equivariance and invariance to adapt to images of randomly rotated aerial targets. EqMotion proposes a motion prediction model that theoretically ensures equivariance of sequence-to-sequence motion under Euclidean transformations and includes an invariant interaction reasoning module to enhance the robustness of model predictions. Incorporating the principle of equivariance into the design of neural network models can bring three benefits: (i) enhancing the network's robustness to specific geometric transformations and ensuring stability; (ii) improving the network's generalization ability under transformed data, increasing flexibility and data efficiency; and (iii) the network contains a large number of parameter sharing mechanisms, reducing space complexity and improving parameter efficiency.

[0005] In spatiotemporal sequence prediction tasks where both input and output are video data, it is important to consider the model's rotation and translation variability in the spatial dimension to improve the model's ability to extract features from data. CNN methods are superior to RNNs in terms of parallel computing and training speed, which makes them attractive in terms of low training cost and high real-time performance, but their ability to model temporal relationships is generally inferior to RNN methods. In existing CNN methods, although the convolution operation has translation variability, it mainly considers translation transformations and ignores rotation variability. Summary of the invention

[0006] In view of the shortcomings of the prior art, the purpose of the present invention is to provide a cluster agent motion prediction method based on a rotation equivariant convolutional neural network, which integrates rotation equivariance into the model framework of a pure CNN, fully utilizes the speed advantage of CNN in processing spatiotemporal data, and enhances its adaptability to rotation changes, thereby achieving more accurate prediction results in the fields of multi-agent cluster motion trajectory prediction, and achieving the effect of taking into account both training efficiency and prediction accuracy. In order to achieve the above-mentioned purpose and other advantages according to the present invention, a cluster agent motion prediction method based on a rotation equivariant convolutional neural network is provided, comprising:

[0007] Abstract large-scale cluster motion prediction as a density map sequence prediction task, that is, predicting future frame density maps through observed historical frame density maps;

[0008] Constructing a prediction model, wherein the prediction model includes an encoder, a converter, and a decoder, and the encoder and the decoder are jump-connected;

[0009] The converter is a rotational equivariant converter of the converter in the SimVP, SimVPv2 and TAU models, replacing the traditional convolution operator with an equivariant convolution operator;

[0010] When designing the converter, the channel dimension and the time dimension are coupled through a linear layer to create a new dimension, so that the original four-dimensional data is converted into three-dimensional data; and the end of the converter remaps and splits the coupled dimension into the channel and time dimensions;

[0011] The encoder and decoder are both stacked by a series of basic convolution modules consisting of convolution, batch normalization and activation functions.

[0012] The encoder extracts the spatial features of the input sequence frame by frame, and helps the model to mine features of different granularities through downsampling operations. The converter receives the encoding results of the encoder as input, further extracts spatial features, and mines the temporal relationship in the sequence, converting historical features into features of the predicted target for output. The decoder restores the hidden features of the predicted sequence obtained by the converter into a density map sequence of the same size as the input data, including necessary upsampling operations. The jump connection between the encoder and the decoder has two specific functions: one is the role of the reference residual connection, which is used to stabilize the training process; the other is to make up for the information loss that may be caused by the downsampling operation in the encoder.

[0013] In the settings of the converter, SimVP, SimVPv2 and TAU respectively design three types of structure extraction features containing spatiotemporal correlation. On this basis, the present invention replaces the traditional convolution operator with a rotation equivariant convolution operator to enhance the model's ability to characterize spatiotemporal data and effectively reduce the prediction error. This improvement enables the model to better adapt to geometric transformations such as rotation, and improves the accuracy and robustness of the prediction.

[0014] The introduction of rotational equivariance results in an extra dimension in the features. S . Corresponds to S The feature map of the group elements. The traditional convolution can be viewed as a special case where S=1 , corresponding to the trivial representation. This integration allows for richer feature representations, enabling the model to analyze features in multiple directions, thereby enhancing its ability to capture complex patterns and improving robustness to input variations.

[0015] Compared with the prior art, the present invention has the following beneficial effects: using a pure CNN network to capture spatial and temporal correlations is more efficient than a recurrent neural prediction network. Integrating rotational equivariance into the convolutional baseline model can not only improve data efficiency, but also enhance the model's adaptability to rotational changes. The method of the present invention improves prediction performance by using a rotational equivariant convolution operator, while enhancing the stability of the training process. Compared with the baseline method, the method of the present invention shows superiority in performance, which is mainly reflected in the improvement of prediction accuracy and the reduction of model parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A model architecture diagram of an encoder-converter-decoder of a cluster agent motion prediction method based on a rotation equivariant convolutional neural network according to the present invention;

[0017] Figure 2 A structural diagram of an encoder and a decoder of a cluster agent motion prediction method based on a rotation equivariant convolutional neural network according to the present invention;

[0018] Figure 3 It is a principle diagram of the action of the equivariant convolution operator on the C4 group of the cluster intelligent agent motion prediction method based on the rotation equivariant convolutional neural network according to the present invention;

[0019] Figure 4 Three types of feasible basic modules of converters with rotation equivariance and corresponding structural diagrams of the swarm intelligence agent motion prediction method based on rotation equivariant convolutional neural network according to the present invention;

[0020] Figure 5 This is a qualitative comparison diagram of the swarm-motion dataset between the method of the swarm agent motion prediction method based on the rotation equivariant convolutional neural network according to the present invention and the SimVP baseline method. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0022] Predicting the future motion trajectories of multiple intelligent agents is crucial for a deep understanding of their motion laws and dynamic mechanisms. Considering the group as a whole to focus on its spatial distribution and motion, and performing density map sequence prediction is an effective way to capture these features, but capturing spatial and temporal correlations at the same time increases the complexity of the task. Existing models in the field of video prediction often ignore equivariance under geometric transformations. The technical problem to be solved by the present invention is to explore the feasibility of integrating rotational equivariance into the framework of a pure CNN model, overcome the shortcomings of traditional spatiotemporal sequence prediction models, and propose a cluster motion prediction method based on rotational equivariant convolutional neural networks to extract density map features and obtain potential cluster motion features, aiming to maintain the high parallel efficiency of CNN while improving model performance.

[0023] Reference Figure 1 , a cluster agent motion prediction method based on rotation equivariant convolutional neural network, comprising:

[0024] Abstracting large-scale cluster motion prediction as a density map sequence prediction task, that is, predicting future frame density maps through observed historical frame density maps; when dealing with large-scale cluster motion prediction, predicting the spatial distribution and evolution trend of the cluster as a whole is more valuable than accurately predicting the trajectory of each agent. By converting the spatial distribution of the cluster into a density map sequence, the value of each pixel represents the number of agents, which can capture the spatial distribution characteristics of the group, such as scale, shape, and center of mass, and provide a more intuitive representation for motion prediction.

[0025] Specifically, the problem is abstracted into a density graph sequence prediction task, that is, using the observed history T h Frame density map to predict future T f Frame density map. The historical density map sequence can be represented as a variable The future density map sequence that needs to be predicted can be represented as a variable Where C, H, and W represent the number of channels, height, and width of the density map, respectively. For this problem, the density map has only one channel, which is used to represent the size of the density value, that is, C = 1. The core of the task is to build a deep learning model The input is the observed density map sequence X, and the output is the predicted output that is as close as possible to the true output Y. The model parameters are optimized using the following formula:

[0026]

[0027] Among them, Θ * represents the optimal model parameters, It is a loss function that measures the difference between the predicted value and the true value. By minimizing the loss function, the best model parameters can be found, so that the accuracy of model prediction is maximized.

[0028] Construct a prediction model, the prediction model includes an encoder, a converter and a decoder, and the encoder and the decoder are jump-connected; Figure 1 As shown in the figure, the encoder extracts the spatial features of the input sequence frame by frame, and helps the model to mine features of different granularities through downsampling operations. The converter receives the encoding result of the encoder as input, further extracts spatial features, and mines the temporal relationship in the sequence, converting the historical features into the features of the predicted target for output. The decoder restores the hidden features of the predicted sequence obtained by the converter into a density map sequence of the same size as the input data, including the necessary upsampling operations. The jump connection between the encoder and the decoder has two specific functions: one is the role of the reference residual connection, which is used to stabilize the training process; the other is to make up for the information loss that may be caused by the downsampling operation in the encoder.

[0029] Moreover, the encoder and decoder are stacked by a series of basic convolution modules consisting of convolution, batch normalization and activation functions. Each convolution module can be expressed as:

[0030]

[0031] Among them, x (j) denotes the input of the j-th basic convolutional block, σ(·) denotes the activation function, BatchNorm(·) denotes the batch normalization operation, and Conv(·) denotes the convolutional layer.

[0032] In the encoder, every two basic convolution blocks are followed by a 2×2 average / maximum pooling layer for downsampling; in the decoder, every two basic convolution blocks are followed by an interpolation upsampling layer. The structure diagram of the encoder and decoder is shown in the figure. Figure 2 shown.

[0033] The converter is a rotational equivariant converter of the converters in the SimVP, SimVPv2 and TAU models, which replaces the traditional convolution operator with an equivariant convolution operator; when designing the converter, the channel dimension and the time dimension are first coupled through a linear layer to create a new dimension, which converts the original four-dimensional data into three-dimensional data. Such a conversion allows the use of the 2D-CNN method as a converter to process spatiotemporal features. At the end of the converter, the coupled dimensions are remapped and split into channel and time dimensions. In the setting of the converter, SimVP, SimVPv2 and TAU respectively design three types of structural extraction features containing spatiotemporal correlations. On this basis, the present invention enhances the model's ability to characterize spatiotemporal data and effectively reduces prediction errors by replacing the traditional convolution operator with a rotational equivariant convolution operator. This improvement enables the model to better adapt to geometric transformations such as rotation, thereby improving the accuracy and robustness of the prediction.

[0034] Assume f,ψ: Respectively represent the mapping function and the convolution function, f and ψ are decomposed into and Then the traditional convolution operator (★) can be expressed as:

[0035]

[0036] If ψ i (xy) is replaced by the group representation corresponding to the translation transformation Then formula (3) can be further transformed into:

[0037]

[0038] Ignoring the boundary effect, the traditional convolution operator satisfies translation equivariance:

[0039]

[0040] That is, the result of convolution followed by translation is the same as the result of translation followed by convolution.

[0041] We can generalize the above example to a more general case, that is, derive the equivariant convolution operator. Assume f,ψ: Represent the mapping function and the convolution operation function respectively, where G represents the group corresponding to the considered geometric transformation, then the equivariant convolution operator It can be expressed as:

[0042]

[0043] Among them, g, represents two elements in the group G,

[0044] The above formula generalizes the equivariance of the traditional convolution operator from the translation group to the wider general group G. Formula (6) can ensure that on the group G, the order of convolution operation and transformation operation can be interchanged, that is:

[0045]

[0046] When processing video data, in addition to considering the spatial translation transformation, the cyclic group C is also considered. n (rotation) and dihedral group D n (Rotation and mirroring). Since the coordinates of image pixels are all discrete integer values, only when n=1, 2, 4, the result of image rotation corresponds to the pixel value in the original image one by one, without involving interpolation operations. Therefore, the present invention mainly designs rotation equivariant convolution modules on C4 group and D4 group.

[0047] The affine group on an affine space is the group consisting of all reversible affine transformations in the space. Specifically, given an affine space consisting of a vector space X, its affine group G can be described as the semidirect product between X and the general linear group GL(X) ( ):

[0048]

[0049] For any g1=(x1,h1),g2=(x2,h2)∈G, where x1,x2∈X and h1,h2∈GL(X), their group product is defined as:

[0050]

[0051] In practical applications, the vector space that is most often concerned is Commonly used affine groups include the Euclidean group and special Euclidean groups Where O(d) and SO(d) represent the orthogonal group and the special orthogonal group respectively.

[0052] If the group G is an affine group, that is, Where H corresponds to GL(X) in equation (8), then the group representation of group G can be decoupled into two parts as follows:

[0053]

[0054] in, In the model designed by the present invention, the affine group considered is or From the perspective of image transformation, the translation, rotation and mirroring of the image are taken into account (for mirroring transformation, C n Contains, D n Therefore, this property allows the processing of transformations to be divided into two categories: translation and non-translation transformations, and to be processed separately.

[0055] The properties of affine groups allow equivariant convolution on their H-groups to be transformed into a two-stage process. The first stage consists of exactly |H| sub-processes, each corresponding to a unique group element, to preserve equivariance on the group H. Each sub-process involves a traditional convolution, maintaining translation equivariance. Formally, equivariant convolution can be decomposed into |H| traditional convolution operators, where the input signal f is transformed into each h-transform filter There are the following processes:

[0056]

[0057] Taking C4 group as an example, Figure 3It shows how the equivariant convolution operator works on this group. The four colors on the left correspond to the feature channels under four transformations on the C4 group (no rotation, 90° clockwise rotation, 180° rotation, and 270° clockwise rotation). The right side shows four convolution kernels whose weights are not shared. The right side of the equation shows the features obtained after calculation by the equivariant convolution operator.

[0058] To incorporate rotational equivariance into the model, regular convolutions in each block are replaced with equivariant convolution operators. Throughout the network, each convolution operation is modified to maintain rotational equivariance to the D4 group. Since the input corresponds to a trivial representation, rotational equivariant convolutions are used to transform features from trivial to regular representation at the beginning of the encoder. At the end of the decoder, features are reversely transformed from trivial to regular representation. Since BatchNorm and activation layers are performed element-wise, the original module already has equivariance.

[0059] The introduction of rotation equivariance results in an additional dimension S for the feature, and the shape becomes (T, S, C, H, W), corresponding to the feature map of S group elements. The traditional convolution can be regarded as a special case where S = 1, corresponding to a trivial representation. This integration allows for richer feature representations, enabling the model to analyze features in multiple directions, thereby enhancing its ability to capture complex patterns and improving robustness to input changes.

[0060] Figure 4 Three feasible converter designs are given, which are the rotational equivariant versions of converters in the SimVP, SimVPv2, and TAU models. The prefix "Re" in each module indicates that the general convolution operator in the original module is converted into an equivariant convolution operator. For other existing model designs, it is also possible to improve them by replacing the traditional convolution operator with an equivariant convolution operator.

[0061] In the training phase, the prediction results are used The mean square error MSE of the true result Y is used as the loss function to train the model, that is:

[0062]

[0063] To quantitatively evaluate the prediction results of the model, standard evaluation indicators such as MSE, MAE, and RMSE can be used to predict the density map sequence. If the value of the image of the prediction task is the value of the standard image, indicators such as SSIM and PSNR can also be used to measure the similarity between images. The calculation method of the standard evaluation indicators is as follows:

[0064]

[0065]

[0066]

[0067]

[0068]

[0069] in, and μ Y They represent the average values ​​of the predicted sequence and the true sequence respectively. and σ Y Represent the standard deviation of the predicted sequence and the true sequence, The covariance of the predicted sequence and the true sequence is shown in Table 1. c1 and c2 are pre-set constants. MAX Represents the maximum pixel value of the image. It is worth noting that, following the commonly used calculation method in the field of video prediction, the error of each image, not each pixel, is obtained when calculating MSE, MAE, and RMSE here. Therefore, only the time step T is averaged after summing.

[0070] Example 1

[0071] This example conducts a cluster motion density map prediction experiment on the Swarm-Motion dataset. The Swarm-Motion dataset is a simulated dataset built based on the Boids mathematical model. Each sample contains 200 to 800 moving individuals. At the initial moment, all individuals are evenly distributed in a given area. As time goes by, clusters of various shapes are formed and move.

[0072] The present invention is used to perform cluster motion density map prediction tasks on the Swarm-Motion dataset. The specific implementation steps are as follows:

[0073] Step 1: Select frames 151 to 160 in the Swarm-Motion dataset as the input observation sequence, and frames 161 to 170 as the future sequence that the model needs to predict. After converting the motion of all individuals into a density map sequence, a 3×3 Gaussian kernel is used for smoothing. The resolution of each image is 80×80. The dataset contains 5,000 samples, with a training set, validation set, and test set ratio of 85%:5%:15%.

[0074] Step 2: Build a prediction model:

[0075] First, we build a basic prediction network of encoder-converter-decoder structure, and replace all convolution operators in the entire network with equivariant convolution operators. For the three different converters in SimVP, SimVPv2 and TAU, we build three corresponding rotation equivariant models.

[0076] Step 3: Train the model based on the training set obtained in step 1, using the Adam optimizer and the OneCycle learning rate change strategy with an initial learning rate of 2e-4 and a maximum learning rate of 1e-3. The training batches for all prediction methods are set to 16, and a total of 200 epochs are trained. After each epoch of training, the prediction effect is tested on the validation set, and the model with the smallest error on the validation set is selected as the final training result.

[0077] Step 4: Input the cluster motion density map sequence obtained in step 1 into the trained prediction network to obtain the predicted density map sequence, and use the corresponding evaluation index to evaluate the prediction results. The test results of different methods on the Swarm-Motion dataset are shown in Table 5:

[0078] Table 5 Quantitative results of the proposed method and other comparative methods on the Swarm-Motion dataset

[0079]

[0080]

[0081] As shown in Table 5, the prediction accuracy of the proposed method under the three types of converter settings is better than the baseline method and other prediction methods in terms of quantitative values. Figure 5 As shown in the figure, the prediction effect in the short term is satisfactory. As time goes by, the prediction of the shape and position of each cluster by the method of the present invention is generally better than that of the baseline method, and the prediction of the cluster outline and high-density distribution area is relatively accurate.

[0082] Step 5: Conduct ablation experiments and consider different transformation groups C n and D n The corresponding models are trained based on the settings. The quantitative results of each model are shown in Table 6-2.

[0083] Table 6-2 Quantitative results comparison between the proposed method and the SimVP method when considering different groups on the Swarm-Motion dataset

[0084]

[0085]

[0086] As can be seen from Table 6-2, as the number of group elements increases, the number of parameters decreases proportionally and the performance tends to improve. However, for models with higher-order groups, the computational complexity increases due to the increased complexity of the kernel rotation operation.

[0087] The proposed method is compared with the SimVP method for data enhancement. The input data is randomly rotated in space. The data of each sample has a probability p of being rotated by an integer multiple of 90° during training. The quantitative experimental results are shown in Table 5.

[0088] Table 6-3 Quantitative results comparison of the proposed method with the SimVP method with data enhancement on the Swarm-Motion dataset

[0089]

[0090] As can be seen from Table 6-2, data enhancement during baseline model training is helpful to improve the accuracy of prediction tasks, but when the enhancement probability reaches a certain ratio, the performance enhancement effect reaches a bottleneck. The method proposed in this invention is equivariant as a whole, so the performance improvement is better than the baseline model with data enhancement.

[0091] Example 2

[0092] This example conducts experiments on the Moving MNIST dataset to verify the performance of the proposed method in the video prediction task. The Moving MNIST dataset is a video version of the classic handwritten digit MNIST dataset. It is a simulated dataset. Each sample contains a series of image sequence data, and each image contains two moving handwritten digits. This dataset records the movement of two handwritten digits and is a widely recognized benchmark public dataset in the field of video prediction.

[0093] The present invention is used to perform video prediction tasks on the Moving MNIST dataset. The specific implementation steps are as follows:

[0094] Step 1: Each sample in the Moving MNIST dataset contains 20 frames of images. The first 10 frames are selected as the input observation sequence, and the last 10 frames are selected as the future sequence for the model to predict. The resolution of each image is 64×64. The dataset contains a total of 10,000 training samples and 10,000 test samples.

[0095] Step 2: Build a prediction model:

[0096] First, we build a basic prediction network of encoder-converter-decoder structure, and replace all convolution operators in the entire network with equivariant convolution operators. For the three different converters in SimVP, SimVPv2 and TAU, we build three corresponding equivariant versions of the model.

[0097] Step 3: Train the model based on the training set obtained in step 1, using the Adam optimizer and the OneCycle learning rate change strategy with an initial learning rate of 2e-4 and a maximum learning rate of 1e-3. The training batch size for all prediction methods is set to 16, and a total of 200 epochs are trained.

[0098] Step 4: Input the image sequence obtained in step 1 into the trained prediction network, obtain the predicted image sequence, and use the corresponding evaluation index to evaluate the prediction results. The test results of different methods on the Moving MNIST dataset are shown in Table 6-4:

[0099] Table 6-4 Quantitative results of the method of the present invention and other comparative methods on the Moving MNIST dataset

[0100]

[0101]

[0102] As shown in Table 6-4, the prediction accuracy of the method of the present invention under the settings of three types of converters is better than that of the baseline method and other prediction methods in terms of quantitative values.

[0103] The test results of the above embodiments demonstrate the effectiveness and superiority of the present invention in cluster motion prediction and video prediction.

[0104] The number of devices and processing scales described here are used to simplify the description of the present invention, and the application, modification and variation of the present invention are obvious to those skilled in the art. Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and the implementation mode, and they can be fully applied to various fields suitable for the present invention. For those familiar with the art, other modifications can be easily realized, so without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the legends shown and described here.

Claims

1. A cluster agent motion prediction method based on rotation equivariant convolutional neural network, characterized in that: include: Abstract large-scale cluster motion prediction as a density map sequence prediction task, that is, predicting future frame density maps through observed historical frame density maps; Constructing a prediction model, wherein the prediction model includes an encoder, a converter, and a decoder, and the encoder and the decoder are jump-connected; The converter is a rotational equivariant converter of the converter in the SimVP, SimVPv2 and TAU models, replacing the traditional convolution operator with an equivariant convolution operator; When designing the converter, the channel dimension and the time dimension are coupled through a linear layer to create a new dimension, so that the original four-dimensional data is converted into three-dimensional data; and the end of the converter remaps and splits the coupled dimension into the channel and time dimensions; The encoder and decoder are both stacked by a series of basic convolution modules consisting of convolution, batch normalization and activation functions.

2. The method for predicting motion of swarm agents based on rotation equivariant convolutional neural network according to claim 1, characterized in that: In the encoder, every two basic convolutional blocks are followed by a 2×2 average / max pooling layer for downsampling.

3. The method for predicting motion of swarm agents based on rotation equivariant convolutional neural network according to claim 1, characterized in that: In the decoder, two basic convolution blocks are followed by an interpolation upsampling layer; and at the end of the decoder, features are reversely converted from regular representation to trivial representation.

4. The method for predicting motion of swarm agents based on rotation equivariant convolutional neural network according to claim 1, characterized in that: The traditional convolution operator is replaced by an equivariant convolution operator, that is, the conventional convolution in each block is replaced by an equivariant convolution operator; in the entire network, each convolution operation is modified to maintain the rotation equivariance to the D4 group. Since the input corresponds to a trivial representation, the rotation equivariant convolution is used to convert the features from the trivial representation to the regular representation at the beginning of the encoder.

5. The method for predicting swarm intelligence motion based on rotation equivariant convolutional neural network according to claim 1, characterized in that: Quantitatively evaluate the prediction results of the prediction model, that is, predict the density map sequence obtained by using the MSE, MAE and RMSE standard evaluation indicators; When the value of the image of the prediction task is the value of the standard image, the SSIM and PSNR indicators can also be used to measure the similarity between images.

Citation Information

Patent Citations

  • Multi-scale feature fusion method based on target detection

    CN114694003A

  • Bird-eye view target direction prediction method based on group equivariant

    CN116188933A

  • Pedestrian trajectory prediction method based on rotation isovariant convolutional neural network

    CN117173633A

Cited By

  • Intelligent agent trajectory prediction method and device, equipment and medium

    CN121858926A

  • An agent trajectory prediction method, device, equipment and medium

    CN121858926B