Class-independent motion prediction network system and method combining space-time and frequency domain characteristics, equipment and medium

By combining the class-independent motion prediction network system with space-time and frequency domain features, using HTSIM and FSFM modules, the problem of inability to fully capture space-time dependencies and dynamic features in the prior art is solved, and higher motion prediction accuracy and robustness are achieved.

CN120086575APending Publication Date: 2025-06-03TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510226258.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing motion prediction methods cannot fully capture space-time dependencies and dynamic features in complex dynamic scenarios, especially when predicting targets with high-speed moving, their accuracy is insufficient.

Method used

A class-independent motion prediction network system combining space-time and frequency domain features is designed. Through the hierarchical space-time integration module (HTSIM) and frequency-space fusion module (FSFM), the modeling of multi-scale space-time dependencies and the deep fusion of frequency and spatial characteristics is realized.

Benefits of technology

It significantly improves the accuracy and robustness of class-independent motion prediction, and can more effectively capture space-time dependencies and dynamic features, especially in high-speed targets and complex dynamic scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086575A_ABST
    Figure CN120086575A_ABST
Patent Text Reader

Abstract

The invention provides a class-independent motion prediction network system and method combining space-time and frequency domain features, equipment and a medium, and belongs to the field of automatic driving. The problem that the space-time dependency relationship and the dynamic characteristics cannot be fully captured in a complex dynamic scene in the moving target prediction is solved; comprising construction of a feature coding module, a hierarchical space-time integration module, a frequency-space fusion module and a feature decoding module, and the HTSIM fuses space-time information step by step through a multi-level space-time feature interactive learning mechanism, and can capture a long-term space-time dependency relationship and local dynamic change, so that a complex motion mode is modeled more accurately; fSFM combines the characteristics of a frequency domain and a space domain, multiband decomposition is carried out on input characteristics through two-dimensional discrete wavelet transform, and the capturing capability of dynamic characteristics is enhanced through deep fusion of the characteristics of the frequency domain and the space domain; the method is suitable for real-time track prediction tasks in intelligent traffic systems such as automatic driving and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of autonomous driving technology, and particularly to a class-agnostic motion prediction network system and method, device, and medium that combine spatio-temporal and frequency-domain features. Background Art

[0002] With the continuous development of autonomous driving technology, accurately predicting the motion trajectories of traffic participants in dynamic scenarios has become a key challenge for achieving safe and efficient autonomous driving. Existing motion prediction methods mainly handle object detection, object tracking, and trajectory prediction tasks independently. However, these methods have certain limitations in dynamic and complex scenarios, especially when facing high-speed moving targets, they often fail to accurately capture the dynamic characteristics of the targets.

[0003] In the field of autonomous driving, perception methods based on bird's-eye view (BEV, Bird’s Eye View) representation have received extensive attention due to their clear spatial structure and suitability for combination with 2D convolutional operations. Such methods can effectively extract spatial features and reduce computational complexity, but they still have deficiencies in capturing spatio-temporal dependencies and dynamic characteristics. Especially in high-speed target prediction and complex dynamic scenarios, the performance of the model is often unsatisfactory.

[0004] In recent years, class-agnostic motion prediction methods have attracted extensive attention. Different from traditional class-based motion prediction methods, class-agnostic methods directly extract dynamic features from sensor data (such as point cloud or image representation), avoiding dependence on target categories and being able to better handle unseen target categories. However, the main challenge faced by class-agnostic methods is how to capture feature changes in the spatio-temporal dimension to support generalization prediction of unknown target categories. Summary of the Invention

[0005] To solve the problem that motion target prediction in the prior art cannot fully capture spatio-temporal dependencies and dynamic characteristics in complex dynamic scenarios, this application proposes a class-agnostic motion prediction network system and method, device, and medium that combine spatio-temporal and frequency-domain features. By designing an innovative hierarchical spatio-temporal integration module (HTSIM) and a frequency-space fusion module (FSFM), this application can more effectively model multi-scale spatio-temporal dependencies and deeply fuse frequency and spatial features, thereby improving the accuracy and robustness of class-agnostic motion prediction.

[0006] The technical solution adopted in this application is as follows: A class-agnostic motion prediction network system that combines spatio-temporal and frequency-domain features, including a feature encoding module, a hierarchical spatio-temporal integration module, a frequency-space fusion module, and a feature decoding module. Among them, the hierarchical spatio-temporal integration module gradually fuses spatio-temporal information through a multi-level spatio-temporal feature interaction learning mechanism to capture long-term spatio-temporal dependencies and local dynamic changes. The frequency-space fusion module is used to connect the encoder and the decoder;

[0007] The hierarchical spatio-temporal integration module includes two channels. The inputs of the two channels are the previous-scale feature and the current-scale feature respectively. The two inputs are respectively passed through 3D global max pooling to aggregate temporal information, and then grouped. After that, the previous-scale feature is compressed to align its dimension with the current-scale feature. After dimension alignment, the features of the two channels are respectively processed by wavelet convolution and two-dimensional selective scanning and then input into the hierarchical interaction learning mechanism to perform multi-scale fusion on the interaction between the current-scale feature and the previous-scale feature among different-level features. The weight map output by the hierarchical interaction learning mechanism and the current-scale feature are multiplied element by element and the grouped channels are restored to complete the fusion, obtaining the final output feature of this module;

[0008] The frequency-space fusion module includes a frequency branch module, a space branch module, and a fusion module based on two-dimensional discrete wavelet transform. The frequency-domain branch module combines wavelet convolution and an attention mechanism, extracts directional features through pooling operations and splices them, and then completes the decomposition of high-frequency features and low-frequency features through wavelet convolution. A dynamic spatial attention mechanism is introduced to optimize important features, and GroupNorm is used to improve training stability. The space branch module uses stacked MambaBlock modules to recursively model long-range dependencies and global semantic information. The fusion module based on two-dimensional discrete wavelet transform uses two-dimensional discrete wavelet transform to deconstruct the input feature into a low-frequency component and a high-frequency component. The low-frequency component is used to extract global information, and the high-frequency component is used to capture local spatial detail changes. The low-frequency component and the spatial feature are integrated through splicing and normalization to generate a long-range dependence representation. The high-frequency component strengthens the detail information through element-wise addition and is spliced and normalized with the spatial feature to ensure the consistency of the high-frequency component. Finally, it is further optimized through an efficient channel attention module and the spatial resolution is restored through bilinear interpolation to generate a multi-frequency feature representation.

[0009] The feature encoding module includes five stages, which are the downsampling module, encoder 1, encoder 2, encoder 3, and encoder 4 in sequence. The downsampling module consists of two 3×3 convolutions and one 7×7 convolution; the 4 encoders are respectively stacked by 3, 4, 6, and 3 residual blocks; each residual block is composed of two 3×3 convolution operations and a residual connection; and 3D convolution is used to adjust the time dimension between every two encoders.

[0010] The feature decoding module includes five stages, namely decoder 1, decoder 2, decoder 3, decoder 4 and its three output heads. The three output heads are classification loss, motion prediction loss and state estimation loss respectively. Each of the 4 decoders has two inputs, and the scales of the two inputs differ by a factor of two. The decoder uses two-dimensional linear interpolation to perform scale transformation on the small-scale features, and then processes them through two 3×3 convolutions.

[0011] Each of the three output heads uses a 3×3 convolution with a stride of 1 and a 1×1 convolution, and the resulting output dimensions are 5×H×W, 20×2×H×W and 2×H×W respectively, where H is the height and W is the width.

[0012] The hierarchical interaction learning mechanism first uses the efficient channel attention mechanism to adjust the feature relationship between channels, and then extracts global information through global average pooling to generate a spatio-temporal attention weight matrix. Then, using matrix multiplication and dynamic weighting methods, key feature patterns are strengthened while redundant information is suppressed, so as to ensure that important spatio-temporal features are effectively modeled.

[0013] The high-frequency components include horizontal high-frequency components, vertical high-frequency components and diagonal high-frequency components.

[0014] A class-agnostic motion prediction method combining spatio-temporal and frequency-domain features includes the following steps:

[0015] Step 1: Construct a dataset, divide it into a training set and a test set, and perform data preprocessing;

[0016] Step 2: Construct a model of the class-agnostic motion prediction network system;

[0017] Step 3: Model training;

[0018] Step 4: Result prediction: Input the preprocessed test set data into the trained model to obtain the prediction results of the moving target, including the category, displacement and motion state of the target.

[0019] During model training, the cross-entropy loss function is used to optimize category prediction and state estimation, the Smooth L1 loss function is used to optimize motion trajectory prediction, the residual blocks are initialized with the weights of ResNet-34 pre-trained on the ImageNet-1K dataset, the Adam optimizer is used during the training process, and the learning rate scheduling strategy is used to accelerate convergence, and the final model parameters are saved.

[0020] An electronic device includes: a processor and a memory; the memory stores the computer execution instructions of the class-agnostic motion prediction network system; the processor executes the computer execution instructions stored in the memory.

[0021] A computer-readable storage medium stores computer-executable instructions of the class-agnostic motion prediction network system, and the computer-executable instructions are executed by a processor.

[0022] The beneficial effects of this application compared with the prior art are as follows:

[0023] (1) Through the innovative design that combines spatio-temporal integration and frequency-spatial fusion, this application significantly improves the accuracy of class-agnostic motion prediction. HTSIM can capture spatio-temporal dependencies over a long time range and local dynamic changes simultaneously through hierarchical spatio-temporal feature fusion, avoiding the problem of losing fine-grained information due to feature compression in existing methods. FSFM realizes the deep fusion of frequency-domain and spatial-domain features through two-dimensional discrete wavelet transform (2D-DWT), significantly enhancing the modeling ability for high-speed moving targets and complex dynamic scenes.

[0024] (2) This application designs a hierarchical spatio-temporal integration module (HTSIM). By gradually fusing spatio-temporal features, this module overcomes the problem of dynamic information loss caused by 3D max pooling in traditional methods. HTSIM can effectively capture multi-scale spatio-temporal dependencies, especially having superior processing capabilities for high-dynamic scenes and high-speed targets, ensuring the accuracy and continuity of prediction results.

[0025] (3) This application designs a frequency-spatial fusion module (FSFM). By deeply fusing frequency-domain and spatial-domain features, it further enhances the feature representation ability in the motion prediction task. The introduction of 2D-DWT technology enables the model to capture high-frequency local details and low-frequency global information of the target, thus providing more comprehensive information support when dealing with complex dynamic targets.

[0026] (4) This application adopts a class-agnostic motion prediction strategy, which does not depend on the target category and can adapt to unseen target categories. This design makes the model of this application have stronger generalization ability. Especially when facing diverse target types in a dynamic environment, it can maintain high accuracy and stability.

[0027] (5) Through the collaborative optimization of spatio-temporal features and frequency-spatial features, this application optimizes the traditional feature fusion strategy. The combined use of HTSIM and FSFM not only enhances the modeling ability for spatio-temporal changes but also effectively improves the ability to capture different motion patterns in dynamic scenes, especially having significant advantages for multi-target prediction in complex traffic environments.

[0028] (6) The present application has been optimized in terms of computational efficiency and real-time performance. By designing an efficient network structure and feature fusion module, the model can achieve high computational efficiency while ensuring high accuracy. It is applicable to real-time motion prediction tasks in autonomous driving systems and meets the high requirements for real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The present application will be further described below with reference to the accompanying drawings:

[0030] Figure 1 It is a schematic flow chart of the class-agnostic motion prediction method proposed in an embodiment of the present application;

[0031] Figure 2 It is a schematic diagram of the model structure of the class-agnostic motion prediction network system proposed in an embodiment of the present application;

[0032] Figure 3 It is a schematic diagram of the structure of the hierarchical spatio-temporal integration module (HTSIM) proposed in an embodiment of the present application;

[0033] Figure 4 It is a schematic diagram of the structure of the frequency-space fusion module (FSFM) proposed in an embodiment of the present application;

[0034] Figure 5 It is an example diagram of the point cloud BEV conversion of the nuScenes dataset used in an embodiment of the present application;

[0035] Figure 6 It is an example diagram of the motion prediction labels in the nuScenes dataset used in an embodiment of the present application;

[0036] Figure 7 It is a motion target prediction result diagram generated by using the method of the present application;

[0037] Figure 8 It is a motion prediction result diagram changing with time generated by using the method of the present application. SPECIFIC EMBODIMENTS

[0038] As Figure 1-8As shown in the figure, in order to further improve the accuracy and robustness of class-agnostic motion prediction, this application proposes an innovative method that combines spatio-temporal integration and frequency-spatial fusion, which is achieved by improving the existing spatio-temporal feature extraction and dynamic feature capture mechanisms. Specifically, this application designs a Hierarchical Temporal-Spatial Integrated Module (HTSIM) and a Frequency-Spatial Fusion Module (FSFM), which enhances the ability to model spatio-temporal dependencies, effectively captures high-frequency and low-frequency dynamic features, and significantly enhances the feature representation ability of the model.

[0039] The class-agnostic motion prediction network system that combines spatio-temporal and frequency-domain features proposed in this application mainly includes a feature encoding module, a hierarchical spatio-temporal integration module, a frequency-spatial fusion module, and a feature decoding module. The HTSIM captures spatio-temporal dependencies within a long time range by gradually fusing spatio-temporal features of different time scales and spatial scales, and retains fine-grained information of local dynamic changes. This module can effectively retain dynamic features and improve the accuracy and continuity of motion prediction when dealing with scenarios such as high-speed moving targets, complex backgrounds, and multi-target interactions. The FSFM decomposes features into multiple frequency bands through two-dimensional discrete wavelet transform (2D-DWT) and deeply fuses frequency-domain features with spatial-domain features. This fusion enables the model to simultaneously extract high-frequency local details and low-frequency global information, significantly enhancing the model's adaptability to dynamic changes, especially performing excellently in multi-target prediction tasks in high-speed motion and complex traffic scenarios. The class-agnostic motion prediction method of the present invention does not depend on the target category, can effectively adapt to unseen target categories, and has strong generalization ability and robustness. Through multi-scale spatio-temporal feature fusion and frequency-spatial feature optimization, the present invention significantly improves the accuracy and computational efficiency of the class-agnostic motion target prediction task, and is applicable to real-time trajectory prediction tasks in intelligent transportation systems such as autonomous driving.

[0040] This application also proposes a class-agnostic motion prediction method, and the specific implementation steps are as follows:

[0041] Step 1: Obtain the original point cloud data and its corresponding target position, displacement and other information through the nuScenes dataset, divide the training set and the test set, and perform data preprocessing. Both the training set and the test set contain the original LIDAR point cloud data and its corresponding motion prediction labels. The division of the dataset ensures the generalization ability of the model in different scenarios.

[0042] Step 2: Construct a class-agnostic motion prediction network model: Design and construct multiple key modules, including a feature encoding module, a Hierarchical Temporal-Spatial Integrated Module (HTSIM), and a Frequency-Spatial Fusion Module (FSFM). The overall model architecture is as shown in Figure 2 as follows. The following are the specific steps for building the network:

[0043] The specific steps for building the model of the class-agnostic motion prediction network system are as follows:

[0044] Step 2.1: Construct a feature encoding module: The feature encoding module consists of five stages, namely a downsampling module, Encoder 1, Encoder 2, Encoder 3, and Encoder 4 in sequence. The downsampling module consists of two 3×3 convolutions and one 7×7 convolution; the 4 encoders are respectively stacked by 3, 4, 6, and 3 residual blocks; each residual block is composed of two 3×3 convolution operations and a residual connection. And 3D convolution is used to adjust the time dimension between two adjacent encoders. The input raw data first extracts low-level features through convolution operations and gradually reduces the spatial resolution of the data through pooling operations. Then, high-level features are further extracted through the convolution operations of the 4 encoders while retaining the spatial structure information of the data. The feature encoding module can effectively capture the low-level features of moving objects and provide an effective basis for subsequent spatio-temporal feature fusion.

[0045] Step 2.2: Construct a Hierarchical Temporal-Spatial Integrated Module (HTSIM): As shown in Figure 3 the input of HTSIM consists of the current-scale feature and the previous-scale feature, which provide deep semantic information and fine-grained spatial features respectively. This module first uses 3D global max pooling to aggregate temporal information, thereby capturing broad semantic dependencies and fine spatial details at different scales. In the feature enhancement stage, HTSIM first performs channel grouping to reduce the parameter complexity while improving the processing efficiency. To ensure the dimensional consistency between different scales, the feature A of the previous scale i-1 is compressed through an image embedding block (PatchEmbedding) to align its dimension with the current-scale feature A i . After alignment, A i and A i-1All need to go through wavelet convolution (WTConv) and two-dimensional selective scanning (SS2D) processing. Wavelet convolution expands the receptive field and enhances the ability to recognize motion features through multi-frequency response features; two-dimensional selective scanning extracts global context information using multi-directional selective scanning. This process effectively enhances the model's ability to capture long-range dependencies while maintaining computational efficiency. This application also proposes a hierarchical interaction learning mechanism for the current scale feature A i and the previous scale feature A i-1 The interaction between different hierarchical features further realizes multi-scale fusion. Specifically, first, the ECA efficient channel attention mechanism is used to adjust the feature relationship between channels, and then global average pooling is used to extract global information to generate a spatio-temporal attention weight matrix. Next, matrix multiplication and dynamic weighting methods are used to strengthen the key feature patterns while suppressing redundant information, thereby ensuring that important spatio-temporal features are effectively modeled. During the fusion process, the fine-grained spatial information of the previous scale and the long-range global dependencies of the current scale are efficiently integrated. Finally, the weight map output by the hierarchical interaction learning mechanism Sigmoid and the current scale feature A i are fused by element-wise multiplication and restoring the grouped channels to obtain the final output feature, which not only retains local details and global dependencies but also further improves the network's expression ability and robustness for complex dynamic scenarios.

[0046] Through the hierarchical fusion of multi-level features by four hierarchical spatio-temporal integration modules (HTSIM), this application can efficiently extract and fuse spatio-temporal features at different time scales and spatial scales. Each HTSIM gradually improves the model's prediction ability for moving targets by layer-by-layer fusing local and global spatio-temporal information. Especially when dealing with complex dynamic scenarios, the hierarchical fusion method can effectively reduce the problem of information loss, enabling the model to capture the fine-grained motion features of the target while maintaining the spatio-temporal dependencies over a long time range. This design not only improves the model's performance in high-speed and complex scenarios but also effectively enhances the ability to model the interaction relationships between targets, thereby improving the prediction accuracy and robustness.

[0047] Step 2.3: Construct a frequency-space fusion module (FSFM): The frequency-space fusion module (FSFM) is an important bridge connecting the encoder and the decoder, aiming to achieve efficient feature fusion and expression enhancement through joint modeling of the frequency domain and the spatial domain. The core of the module design lies in combining two-dimensional discrete wavelet transform (2D-DWT) to capture both the frequency domain dependencies and spatial context relationships of the input features. The structure of FSFM is as Figure 4As shown in (a), it mainly consists of three parts: a frequency-domain branch module, a spatial branch module, and a fusion module based on two-dimensional discrete wavelet transform, aiming to enhance the model's ability to model multi-scale and multi-directional features in complex scenes through the deep fusion of frequency-domain and spatial features.

[0048] The frequency-spatial fusion module effectively captures multi-frequency features and spatial context dependencies through the joint modeling of the frequency-domain branch and the spatial branch. The frequency-domain branch module combines wavelet convolution and attention mechanism. First, it extracts directional features through horizontal pooling and vertical pooling, and splices them to form a joint representation for modeling multi-directional dependencies. Subsequently, the joint representation completes the decomposition of high-frequency and low-frequency features through wavelet convolutions (1×1, 5×5, and 7×7) to enhance local details and global context expression. At the same time, a dynamic spatial attention mechanism is introduced to weight important features, and group normalization (GroupNorm) is used to improve the training stability and feature robustness of the model; the spatial branch module uses stacked MambaBlock modules to capture long-range dependencies and global semantic information through recursive modeling. The frequency-spatial fusion module highlights key regions through a dynamic weighting mechanism and combines residual connections to retain the original features, ensuring the integrity and stability of global modeling. The two branch modules fully fuse high-frequency and low-frequency information, as well as local and global context features through joint modeling, significantly improving the feature expression ability and showing higher perception and prediction accuracy in complex scenes.

[0049] The structure of the fusion module based on two-dimensional discrete wavelet transform (WMFF) is as Figure 4 shown in (b). The core of the design is to achieve the deep fusion of frequency-domain and spatial features through two-dimensional discrete wavelet transform to enhance the hierarchy and diversity of feature expression. Combining frequency decomposition and spatial modeling, the module can deconstruct the input features into different frequency components in the frequency domain while retaining their spatial structure information, thus achieving a comprehensive modeling of details and global context in complex scenes. The WMFF module uses the Haar wavelet transform to decompose the input features in the frequency domain and extracts multiple frequency components through the low-frequency filter L and the high-frequency filter H. The low-frequency filter L is used to extract global information, while the high-frequency filter H focuses on capturing local detail changes. Through the filtering operation, the low-frequency component LL (global information), the horizontal high-frequency component LH, the vertical high-frequency component HL, and the diagonal high-frequency component HH (detail information) are generated, respectively capturing the spatial detail changes of the input features in different directions, thus providing rich multi-scale and multi-directional feature expressions for the model: the low-frequency component LL is used to capture global information, while the horizontal high-frequency component LH, the vertical high-frequency component HL, and the diagonal high-frequency component HH are mainly used to capture spatial detail changes. The specific calculation formulas of these components are as follows:

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056] wherein: represents a filtering operation, through which the input features are explicitly decomposed in the frequency domain, enabling the global information (low frequency) and local details (high frequency) to be modeled and optimized separately.

[0057] In the fusion stage, the module processes and combines the frequency-domain features and spatial features respectively:

[0058] 1) Low-frequency component fusion: The low-frequency component LL represents the global context information. The low-frequency components from the frequency-domain features and the low-frequency components of the spatial features are concatenated along the channel dimension and normalized to integrate the global dependencies and semantic relationships, generating a unified long-range dependency representation.

[0059] 2) High-frequency component fusion: The high-frequency components LH, HL, and HH represent fine-grained edges and local variations. The module combines the high-frequency components by element-wise addition to enhance the expression of detail information. Subsequently, the high-frequency components from the frequency-domain features and the high-frequency components of the spatial features are concatenated along the channel dimension and normalized to balance the scale differences of different feature components, ensuring the consistency and robustness of the high-frequency information during the fusion process.

[0060] 3) Fusion refinement: The fused low-frequency and high-frequency features first combine the high-frequency and low-frequency information by element-wise addition. Subsequently, a mapping layer composed of an efficient channel attention (ECA) module further adjusts the channel relationships of the features. The fused features are upsampled through a bilinear interpolation operation to restore the spatial resolution. Finally, an efficient feature integration is completed through a multi-layer perceptron (MLP) and a residual connection, generating a rich and highly expressive multi-frequency feature representation.

[0061] The Frequency-Spatial Fusion Module (FSFM) effectively improves the model's performance in dynamic scenarios by deeply fusing the features in the frequency domain and the spatial domain. FSFM first decomposes the input features into multiple frequency bands through the two-dimensional discrete wavelet transform (2D-DWT), thereby capturing dynamic information at different scales in the frequency domain. Then, FSFM combines the frequency-domain features with the spatial-domain features and uses convolutional operations to fuse the information of both, further enhancing the ability to predict the motion of targets. This module can simultaneously extract the high-frequency details and low-frequency global information of targets. Especially when dealing with high-speed moving targets or background interference in complex scenarios, it can provide more accurate motion predictions. Through this deep fusion of frequency-spatial features, FSFM effectively overcomes the drawback of traditional methods that are prone to losing fine-grained information when dealing with complex dynamic targets, improves the adaptability of the model to dynamic changes and the prediction ability for complex environments, and shows significant advantages especially in the multi-target prediction task in complex traffic scenarios.

[0062] Step 2.4: Construct a feature decoding module: The feature decoding module consists of five stages, successively including decoder 1, decoder 2, decoder 3, decoder 4 and its three output heads. The three output heads are the classification loss, the motion prediction loss, and the state estimation loss respectively. Each of the 4 decoders has two inputs, and the scales of the two inputs differ by a factor of two. The decoder uses two-dimensional linear interpolation to perform scale transformation on the smaller-scale features, and then processes them using two 3×3 convolutions. Each of the three output heads uses a 3×3 convolution with a stride of 1 and a 1×1 convolution to obtain the results, and the three output dimensions obtained are (5×H×W), (20×2×H×W), and (2×H×W).

[0063] Specifically, the decoder gradually restores the spatial resolution and fuses the features from the encoder and other modules by concatenating and upsampling the two input features x and y. First, the input feature y is upsampled to the same spatial resolution as x through bilinear interpolation, and then the two are concatenated in the channel dimension. Next, after being processed by two 3×3 convolutions and batch normalization, the features are gradually processed and fused to output the final feature map. This decoder can efficiently fuse multi-scale features, enhance the expressive ability of motion target prediction, while retaining fine-grained spatial information, and improve the accuracy and robustness of the network.

[0064] Step 3: Model Training: The training program is built based on the PyTorch framework. The Adam algorithm is selected as the optimizer, and the loss function adopts a multi-task learning framework, combining cross-entropy loss and Smooth L1 loss to optimize the accuracy and generalization ability of the network in motion prediction. The loss function consists of three main parts: classification loss, motion prediction loss, and state estimation loss. Both the classification loss and the state estimation loss use the cross-entropy loss function to ensure the accurate classification of targets and the accurate estimation of states, while the motion prediction loss is optimized by the Smooth L1 loss to improve the prediction of motion trajectories and enhance the accuracy of trajectories. The Smooth L1 loss has strong robustness and can effectively handle outliers in regression tasks. Through this balanced loss function design, the network can achieve optimization among classification accuracy, trajectory prediction, and state estimation, thus significantly improving the overall performance and prediction ability of the model in complex dynamic scenarios.

[0065] Step 4: Result Prediction: The preprocessed test set data is input into the trained class-agnostic motion prediction network model to obtain the prediction results of moving targets, including the category, displacement, and motion state of the targets.

[0066] The formula for the cross-entropy loss function is as follows:

[0067]

[0068] In the formula: x is the input of the loss function, class represents the true category label of the data, is the exponential sum of all category prediction values, and x[class] is the prediction value corresponding to the true category;

[0069] The formula for the Smooth L1 loss is as follows:

[0070]

[0071] In the formula: x' is the input value for calculating the loss.

[0072] The calculation formula for the network loss function is obtained as: L = L cls + L motion + L state , where L is the network loss function, L cls is the classification loss function, L motion is the motion prediction loss function, and L state is the state estimation loss function.

[0073] This application also proposes an electronic device, including: a processor and a memory; the memory stores the computer execution instructions of the class-agnostic motion prediction network system; the processor executes the computer execution instructions stored in the memory.

[0074] The present application also proposes a computer-readable storage medium storing computer-executable instructions of the class-agnostic motion prediction network system, and the computer-executable instructions are executed by a processor.

[0075] The following further describes the present application according to specific embodiments.

[0076] This embodiment conducts systematic experiments based on the large-scale autonomous driving dataset nuScenes, which contains multi-modal data such as continuous point clouds and images, providing rich visual and geometric information for multi-modal learning in autonomous driving scenarios. The dataset is divided into a training set, a validation set, and a test set, containing 500, 100, and 250 scenes respectively, comprehensively covering diverse environmental and target object distribution characteristics. The LiDAR sensors in the nuScenes dataset collect point cloud data at a frequency of 20 Hz, capable of capturing a complete 360° horizontal field of view and accompanied by high-precision annotation information to support the accurate perception and prediction of dynamic objects. In the training and testing phases, this embodiment selects to use only LiDAR point cloud data, generating pseudo-map features from bird's-eye view (BEV) at a sampling rate of 20 Hz as network inputs to efficiently model spatial relationships and dynamic changes in 2D space.

[0077] When evaluating the model performance, non-empty grid cells are divided into three categories according to their velocity characteristics: static (velocity ≤ 0.2 m / s), low-speed motion (velocity ≤ 5 m / s), and high-speed motion (velocity > 5 m / s). For each velocity category, by calculating the Euclidean distance L2 between the predicted displacement and the actual displacement within the next 1 second, the average and median of the model prediction errors are reported, which serve as key metrics for measuring the accuracy of motion prediction. Further, to comprehensively evaluate the performance of the model in classification tasks, this embodiment introduces two metrics: Overall Accuracy (OA) and Mean Category Accuracy (MCA), corresponding to the classification accuracy formulas for all non-empty cells and five specific categories respectively as follows:

[0078]

[0079]

[0080] where: n is the number of categories, C ij represents the elements in the confusion matrix, where C ii is the number of correct classifications of the i-th category, and C ij (when i ≠ j) represents the number of times the i-th category is misclassified as the j-th category.

[0081] Table 1 below shows the motion prediction performance of the method of this application in different speed ranges on the nuScenes dataset. For static targets (Static), the method of this application has a mean displacement error of 0.0267 meters and a median error of 0 meters, indicating that the network has extremely high accuracy in predicting static targets. For low-speed targets (Speed≤5m / s), the mean displacement error is 0.2277 meters and the median error is 0.0949 meters, indicating that the model can accurately predict the trajectories of low-speed targets. In the prediction of high-speed targets (Speed>5m / s), the network shows strong robustness, with a mean displacement error of 0.8359 meters and a median error of 0.6156 meters. Despite the fast target speed, the prediction results still maintain high accuracy. Overall, the average time of the model is 56.7 milliseconds, indicating its strong real-time prediction ability in dynamic scenarios.

[0082] Table 1 Motion Prediction at Different Speeds under the nuScenes Dataset

[0083]

[0084] Table 2 below shows the performance of the classification auxiliary task of the method of this application on the nuScenes dataset. The classification accuracy indicates the recognition ability of the method of this application for different target categories. For the background (Bg) category, the classification accuracy is 97.07%, showing very high accuracy; for the vehicle (Vehicle) category, the accuracy is 92.91%, demonstrating the superior performance of the model in traffic target detection. The classification accuracies of pedestrians (Ped) and bicycles (Bike) are relatively low, 82.53% and 31.96% respectively, which is due to the large variability of the target categories themselves, and the model's recognition of these targets needs to be further improved. For the other (Others) category, the accuracy is 74.81%, showing good recognition ability for non-standard targets. The multi-class average accuracy (MCA) is 75.86% and the overall accuracy (OA) is 96.40%, proving the efficiency and accuracy of the method of this application in the multi-class target classification task.

[0085] Table 2 Classification Auxiliary Task under the nuScenes Dataset

[0086]

[0087] As Figure 5 、 Figure 6 and Figure 7 shown, the application and results of this application in the motion target prediction task are specifically demonstrated. Specifically, Figure 5Disclosed is a method for converting radar point cloud data into a bird's-eye view (BEV) map, which converts three-dimensional radar point cloud data into a two-dimensional spatial representation, enabling the network to more efficiently model spatial relationships and dynamic changes. Figure 6 Disclosed is a label map for the moving target prediction task, which provides the ground truth motion trajectory labels corresponding to the input BEV map for learning the position and motion state of the target during network training. Figure 7 Disclosed is the prediction result map of the method of this application, which shows the predicted trajectories of the network on the test set, proving that the network model can accurately predict the motion path and position of the moving target, especially in complex dynamic scenarios, and can effectively capture the motion characteristics of the target.

[0088] As Figure 8 shown, this figure shows the change trend of the mean error over time in different speed scenarios, including three cases: static, slow, and fast. Specifically, the blue curve in the figure represents the error change in the static scenario, showing that the model has high accuracy and stability when dealing with static targets; the orange curve represents the error change in the slow-moving scenario, and the model can accurately capture the slow motion characteristics of the target; the green curve represents the error change in the fast-moving scenario, demonstrating the effectiveness of the model in predicting the motion of fast targets and its ability to better adapt to the changes in dynamic scenarios.

[0089] This application discloses an innovative class-agnostic moving target prediction method that integrates spatio-temporal integration and deep fusion technology of frequency-spatial features. By designing a hierarchical spatio-temporal integration module (HTSIM) and a frequency-spatial fusion module (FSFM), this application effectively improves the accuracy and robustness of moving target prediction in complex dynamic scenarios. This method uses the joint modeling of multi-scale spatio-temporal features and frequency-spatial features to enhance the model's adaptability to multiple targets in a dynamic environment while capturing the target motion trajectories.

[0090] The present application trains the model through the nuScenes dataset, ensuring the diversity of data and the representativeness of training. The present application also constructs a network architecture with spatio-temporal integration and frequency-space fusion capabilities. HTSIM ensures that the model can capture long-range spatio-temporal dependencies and local dynamic changes through multi-level spatio-temporal feature fusion, while FSFM enhances the prediction accuracy of the target motion trajectory by deeply fusing the features of the frequency domain and the spatial domain through two-dimensional discrete wavelet transform (2D-DWT). During the model training process, the present application adopts a loss function that combines cross-entropy loss and Smooth L1 loss, effectively balancing the classification accuracy, motion trajectory prediction, and target state estimation, thereby improving the overall performance and prediction ability of the network model. The present application can accurately predict the motion trajectory of the target in complex scenarios, especially having significant advantages in high-speed targets and complex backgrounds. Compared with the prior art, the uniqueness of the present application lies in its deep fusion of spatio-temporal features and frequency-space features, which not only improves the accuracy of moving target prediction but also enhances the adaptability and robustness of the model to dynamic scenarios. Through multi-scale feature fusion and efficient loss function design, the present application shows stronger generalization ability and precision in complex environments.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A class-independent motion prediction network system combining spatiotemporal and frequency domain features, characterized by: It includes feature encoding module, hierarchical spatiotemporal integration module, frequency-space fusion module and feature decoding module. The hierarchical spatiotemporal integration module integrates spatiotemporal information step by step through a multi-level spatiotemporal feature interactive learning mechanism to capture long-term spatiotemporal dependencies and local dynamic changes. The frequency-space fusion module is used to connect the encoder and decoder. The hierarchical spatiotemporal integration module includes two channels. The inputs of the two channels are the previous scale features and the current scale features, respectively. The two inputs are respectively subjected to 3D global maximum pooling to aggregate time information and then grouped. After that, the previous scale features are compressed to align the dimensions with the current scale features. After the dimensions are aligned, the features of the two channels are respectively processed by wavelet convolution and two-dimensional selective scanning and input into the hierarchical interactive learning mechanism. The interaction between the current scale features and the previous scale features at different levels is multi-scale fused. The weight map output by the hierarchical interactive learning mechanism is fused with the current scale features by element-by-element multiplication and the grouping channel is restored to obtain the final output features of the module. The frequency-space fusion module includes a frequency branch module, a space branch module and a fusion module based on two-dimensional discrete wavelet transform. The frequency branch module combines wavelet convolution and attention mechanism, extracts directional features through pooling operation and splices them, and then completes the decomposition of high-frequency features and low-frequency features through wavelet convolution, introduces dynamic spatial attention mechanism, optimizes important features, and improves training stability through GroupNorm; the space branch module adopts stacked MambaBlock modules to recursively model long-distance dependencies and global semantic information; the fusion module based on two-dimensional discrete wavelet transform uses two-dimensional discrete wavelet transform to deconstruct input features into low-frequency components and high-frequency components. The low-frequency components are used to extract global information, and the high-frequency components are used to capture local spatial detail changes. The low-frequency components and spatial features are integrated through splicing and normalization to generate long-range dependency representations; the high-frequency components enhance the detail information through element-by-element addition, and are spliced ​​and normalized with the spatial features to ensure the consistency of the high-frequency components; finally, it is further optimized through the efficient channel attention module, and the spatial resolution is restored through bilinear interpolation to generate multi-frequency feature representations.

2. The class-independent motion prediction network system combining spatiotemporal and frequency domain features according to claim 1, characterized in that: The feature encoding module consists of five stages, namely downsampling module, encoder 1, encoder 2, encoder 3 and encoder 4. The downsampling module consists of two 3×3 convolutions and one 7×7 convolution. The four encoders are composed of 3, 4, 6, and 3 residual blocks stacked respectively. Each residual block consists of two 3×3 convolution operations and residual connections. 3D convolution is used between encoders to adjust the time dimension.

3. The class-independent motion prediction network system combining spatiotemporal and frequency domain features according to claim 2, characterized in that: The feature decoding module consists of five stages, namely decoder 1, decoder 2, decoder 3, decoder 4 and their three output heads. The three output heads are classification loss, motion prediction loss and state estimation loss respectively. Each of the four decoders includes two inputs, and the scales of the two inputs differ by one times. The decoder uses two-dimensional linear interpolation to scale the small-scale features, and then processes them through two 3×3 convolutions.

4. The class-independent motion prediction network system combining spatiotemporal and frequency domain features according to claim 3, characterized in that: All three output heads use a 3×3 convolution with a stride of 1 and a 1×1 convolution, and the output dimensions are 5×H×W, 20×2×H×W, and 2×H×W, respectively, where H is the height and W is the width.

5. The class-independent motion prediction network system combining spatiotemporal and frequency domain features according to claim 1, characterized in that: The hierarchical interactive learning mechanism first uses an efficient channel attention mechanism to adjust the feature relationship between channels, and then extracts global information through global average pooling to generate a spatiotemporal attention weight matrix; then, it uses matrix multiplication and dynamic weighting to enhance key feature patterns while suppressing redundant information, thereby ensuring that important spatiotemporal features are effectively modeled.

6. The class-independent motion prediction network system combining spatiotemporal and frequency domain features according to claim 1, characterized in that: The high frequency components include horizontal high frequency components, vertical high frequency components and diagonal high frequency components.

7. A class-independent motion prediction method combining spatiotemporal and frequency domain features, characterized in that: The following steps are involved: Step 1: Construct a data set, divide it into a training set and a test set, and perform data preprocessing; Step 2: construct a model of the class-independent motion prediction network system as described in any one of claims 1 to 6; Step 3: Model training; Step 4: Result prediction: Input the preprocessed test set data into the trained model to obtain the prediction results of the moving target, including the target category, displacement and motion state.

8. The class-independent motion prediction method combining spatiotemporal and frequency domain features according to claim 7, characterized in that: During model training, the cross entropy loss function is used to optimize category prediction and state estimation, and the Smooth L1 loss function is used to optimize motion trajectory prediction. The residual block is initialized with the ResNet-34 weights pre-trained on the ImageNet-1K dataset. The Adam optimizer is used in the training process, and the convergence is accelerated through the learning rate scheduling strategy, and the final model parameters are saved.

9. An electronic device, characterized in that: include: A processor and a memory; the memory stores computer-executable instructions of the class-independent motion prediction network system according to any one of claims 1 to 6; The processor executes computer-executable instructions stored in the memory.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions of the class-independent motion prediction network system according to any one of claims 1 to 6, and the computer-executable instructions are executed by a processor.

Citation Information

Cited By

  • Method and system for predicting service life of electric energy meter production line bearing, equipment and medium

    CN120449721A

  • Premonition capture and effectiveness evaluation method and system for monitoring aviation unsafe events in real time

    CN120496369A

  • Target detection method based on high-resolution distributed optical fiber sensing data

    CN120974098A

  • Self-adaptive multi-scale earthquake high-resolution processing method fusing time-frequency characteristics

    CN121276597A