Intelligent mining equipment-oriented spatial-temporal feature aggregation point cloud semantic segmentation method

Through the method of spatiotemporal feature aggregation, the problems of point cloud sparseness and fuzzy terrain boundary features are solved, and higher semantic segmentation accuracy and lower computational cost are achieved, which is suitable for field mining environments in intelligent mining.

CN120148034APending Publication Date: 2025-06-13ZHONGBEI UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510190855.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The inherent sparsity of point clouds and blurred terrain boundary features in the environment of field mining have limited the accuracy of three-dimensional semantic segmentation in intelligent mining.

Method used

The method of spatiotemporal feature aggregation is adopted. By establishing a three-dimensional semantic segmentation model based on spatiotemporal feature aggregation, point cloud data is collected using lidar sensors, segmented into current frame data and historical frame data, and dynamically balanced through feature encoding, spatiotemporal feature extraction and gated units, aggregating spatiotemporal features, and finally semantic segmentation is performed.

Benefits of technology

It effectively reduces the impact of point cloud sparseness on semantic segmentation, improves the accuracy of recognition of terrain boundary features, reduces calculation costs, and improves the efficiency and scalability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148034A_ABST
    Figure CN120148034A_ABST
Patent Text Reader

Abstract

The invention discloses a spatial-temporal feature aggregation point cloud semantic segmentation method for intelligent mining equipment, and belongs to the technical field of intelligent mining and intelligent robots. Aiming at the problem that the recognition difficulty is increased due to the inherent sparsity of point clouds and the characteristic that terrain boundary characteristics in a field mining area environment are fuzzy, the method introduces spatio-temporal context information through a three-dimensional semantic segmentation model based on spatio-temporal characteristic aggregation, and effectively relieves the problem. Specifically, a space-time encoder is designed to extract space-time information in point cloud historical multi-frame data, and a current scanning frame is fused to enhance feature expression; a spatio-temporal feature selection state space module is proposed to more fully utilize and model long-term spatio-temporal features while maintaining lower time and spatial computation overhead. The method shows remarkable advantages in performance and applicability, and a technical path which is excellent in performance and friendly to resources is provided for semantic segmentation tasks of related scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent mining and intelligent robotics, and particularly relates to a spatio-temporal feature aggregation point cloud semantic segmentation method for intelligent mining equipment. Background Art

[0002] In recent years, the intelligent construction of mining areas has developed rapidly. Introducing artificial intelligence technology into mining operations not only shows great potential in ensuring safety, but also can significantly improve production efficiency and the stability of operations. In this process, point cloud data, as the main source of three-dimensional coordinate information of the environmental space, can comprehensively reflect the geometric characteristics of the terrain and scenes in the mining area and is the core input data for intelligent construction. Based on three-dimensional semantic segmentation of point clouds, pixel-level parsing of the environment can be achieved, accurately identifying key information such as terrain structures and object distributions. Accurate three-dimensional point cloud semantic segmentation has become an important part of the construction of smart mining areas.

[0003] However, the inherent sparsity of point clouds still poses a challenge to the accuracy of semantic segmentation. Especially in the wild mining area environment, due to the often unclear features of the ground and the boundaries of mining area stockpiles, the point cloud sparsity further exacerbates the fuzziness of terrain boundary features, making it more difficult to accurately segment the boundaries. Aggregating temporal data and making full use of long-term spatio-temporal features are considered an effective means to alleviate the problem of point cloud sparsity, and it is usually processed through the global modeling ability of the attention mechanism. However, this method brings a quadratic computational cost, limiting its efficiency and scalability in practical applications. Summary of the Invention

[0004] Aiming at the problem that the inherent sparsity of point clouds and the fuzzy terrain boundary features in the wild mining area environment exacerbate the difficulty of recognition, the present invention provides a spatio-temporal feature aggregation point cloud semantic segmentation method for intelligent mining equipment.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] A spatio-temporal feature aggregation point cloud semantic segmentation method for intelligent mining equipment, comprising the following steps:

[0007] Step 1, establish a three-dimensional semantic segmentation model based on spatio-temporal feature aggregation, then use a lidar sensor to collect point cloud data, divide the collected point cloud data into current frame data and historical frame data, and then input it into the data input unit of the model;

[0008] Further, the specific steps of the said Step 1 are:

[0009] The data input unit in the above-mentioned step 1 adopts a two-branch structure and inputs historical frame data and current frame data respectively. For the input of historical frame data, continuous point cloud frames within a specified time period are taken for processing;

[0010] Historical frame data and current frame data X t are respectively represented as:

[0011]

[0012] X t = S t

[0013] where S t-i represents the point cloud data of the (t - i)th frame, and concat(.) is the concatenation operation; for the point cloud frame at a certain moment t is generally composed of N t points p k = [x k , y k , z k , r k T [x k , y k , z k T is the three-dimensional coordinate of the point cloud, and [r k is the lidar reflection intensity; historical frame data is formed by concatenating and aggregating multiple consecutive frame data between the point cloud frame S t at time t and the point cloud frame S t-i at time (t - i). Usually, i is determined by comprehensive evaluation of hardware conditions and model performance; current frame data X t is represented by the point cloud frame S t .

[0014] Step 2: Perform preliminary feature encoding on the current frame data and historical frame data respectively to obtain the current feature and sequence feature;

[0015] Furthermore, the specific steps of the above-mentioned step 2 are as follows:

[0016] Encode the sequence frame data and the current frame data X t respectively through the following formula:

[0017] Y = σ(Norm(SpC(X)))

[0018] where SpC is three-dimensional sparse convolution, Norm is batch normalization operation, σ is a variable and is set to the RELU function; after encoding, the sequence feature X SF ​​and the current feature X CF , where Y is the encoding result, and after encoding, sequence features X SF and the current feature X CF are obtained respectively.

[0019] This step aims to extract low-level features from the point cloud data that are meaningful for the semantic segmentation task, and prepare for subsequent spatio-temporal feature learning.

[0020] Step 3: Mine and learn the potential spatio-temporal cues in the sequence features. Through the spatio-temporal feature extraction module, extract the long-term spatio-temporal features from the historical frame data to obtain the long-term spatio-temporal dependencies in the historical frame data. The specific operation is as follows: For the sequence feature X SF , first cross-correlate the features in the sequence through a linear layer and map them to a high-dimensional space to promote information interaction of long-term dependencies; subsequently, the features processed by the linear layer are input into the MambaBlock to further mine and capture potential spatio-temporal cues. The formula is as follows:

[0021]

[0022] where, represents the intermediate feature generated during the calculation process; X SF represents the sequence feature; Norm is the batch normalization operation; X STF represents the long-term spatio-temporal feature; Linear is the linear layer, and σ is a variable set to the GELU function.

[0023] Step 4: For the long-term spatio-temporal feature and the current feature, use a gated unit for dynamic balance, and use the model to adjust the dependence degree between the current frame data and the historical frame data, so as to obtain the aggregated spatio-temporal feature;

[0024] The gated unit first generates a weight representation Gate through a linear layer and the Sigmoid function, and then the long-term spatio-temporal feature X STF and the current feature X CF are respectively multiplied by different weight representations, and the weighted features are fused by element-wise addition to obtain the effectively aggregated long-term spatio-temporal feature X STAF , and the formula is as follows:

[0025] Gate = Sigmoid(Linear(X CF ))

[0026] X STAF = X CF × Gate + X STF × (1 - Gate)

[0027] In the formula, Linear represents the linear layer.

[0028] Through the gating mechanism, while retaining the key historical spatio-temporal clues, the model can flexibly adjust the weights of each spatio-temporal feature in the final decision-making, improving the quality of the overall feature representation.

[0029] Step 5: Use the spatio-temporal feature selection state space module to further model the long-term dependencies in the spatio-temporal features;

[0030] Furthermore, the specific steps of Step 5 are as follows:

[0031] Step 5.1: Design a spatio-temporal encoder to further model its long-term dependencies. The spatio-temporal encoder consists of a convolutional layer and a Mamba block. The convolutional layer continuously expands the receptive field, promoting information interaction between local features with long-term dependencies. After each layer of convolution, it passes through the Mamba block for global context modeling. The process is as follows:

[0032] X E-L atent = σ(BatchNorm(SpC(X STAF )))

[0033] XEO = MambaBlock(XE-Latent)

[0034] where X E-Latent is the temporary feature generated during the process, SpC is the three-dimensional sparse convolution, σ is a variable set to the RELU function, and X EO is the feature generated by the spatio-temporal encoder; The Mamba block consists of a basic Mamba model, introducing a parameterized state space model SelectiveSSM with a selection mechanism, transforming the time-invariant system into a linear time-varying system. The formula is as follows:

[0035] SelectiveSSM(x t ) = y t

[0036]

[0037] where x t is the input representation, obtained by linear mapping from the feature X EO generated by the spatio-temporal encoder, is the latent state updated at time step t, the parameter N is the state size, and y t is the output representation; The discrete matrix is updated over time Δ through the zero-order hold rule:

[0038]

[0039] where and are model parameters; I represents the identity matrix;

[0040] Step 5.2, use the spatio-temporal decoder to reconstruct the features obtained by the spatio-temporal encoder using successive transposed convolutional layers, and at the same time add Mamba blocks to prevent the loss of key spatio-temporal context information:

[0041] XD-Latent = MambaBlock(XEO)

[0042] XDO = σ(BatchNorm(SpInvC(XD-Latent)))

[0043] where SpInvC is three-dimensional sparse transposed convolution, σ is a variable, set to the RELU function; X D-Latent represents the intermediate result in the feature reconstruction process; X DO is the feature generated by the spatio-temporal decoder; BatchNorm is the batch normalization operation.

[0044] Through the spatio-temporal encoder and spatio-temporal decoder, it is ensured that the model can capture complex time series feature information while maintaining computational efficiency.

[0045] Step 6, send the features output by the spatio-temporal decoder into the semantic segmentation head, assign accurate semantic labels to each point cloud data point, and obtain the final point cloud semantic segmentation result.

[0046] Compared with the prior art, the present invention has the following advantages:

[0047] By deeply mining and encoding the latent spatio-temporal clues in historical data and modeling the long-term dependencies in spatio-temporal features, the spatio-temporal information in the data is more fully expressed. At the same time, this method effectively maintains low time and space computational overheads. Considering the high requirements of the wild mining area environment for computing resources, real-time performance, and accuracy, this method shows significant advantages in terms of performance and applicability, providing a technical solution that not only has excellent performance but also can efficiently utilize resources for 3D semantic segmentation tasks in related scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a schematic diagram of spatio-temporal feature encoding and aggregation;

[0049] Figure 2 is a schematic diagram of the spatio-temporal feature selection state space module. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0050] To gain a deep understanding of the present invention, we will provide a comprehensive and detailed description of it. However, the present invention has multiple implementation manners and is not limited to the specific examples listed herein. The presentation of these examples aims to deepen the overall understanding of the disclosed content of the present invention.

[0051] A spatio-temporal feature aggregation point cloud semantic segmentation method for intelligent mining equipment includes the following steps:

[0052] Step 1: Establish a three-dimensional semantic segmentation model based on spatio-temporal feature aggregation. Then, use a lidar sensor to collect point cloud data, divide the collected point cloud data into current frame data and historical frame data, and then input it into the data input unit of the model;

[0053] Further, the specific steps of Step 1 are as follows:

[0054] The data input unit in Step 1 adopts a two-branch structure and inputs historical frame data and current frame data respectively. For the input of historical frame data, continuous point cloud frames within a specified time period are processed;

[0055] Historical frame data and current frame data X t are respectively represented as:

[0056]

[0057] X t = S t

[0058] where S t-i represents the point cloud data of the (t - i)th frame, and concat(.) is the concatenation operation; for the point cloud frame at a certain moment t is generally composed of N t points p k = [x k , y k , z k , r k T [x k , y k , z k T is the three-dimensional coordinate of the point cloud, and [r k is the lidar reflection intensity; historical frame data is formed by concatenating and aggregating multiple consecutive frame data between the point cloud frame S t at the tth moment and the point cloud frame S t-i at the (t - i)th moment. i is usually determined by comprehensively evaluating the hardware conditions and model performance; current frame data X t is represented by the point cloud frame S t .​​

[0059] Step 2: Perform preliminary feature encoding on the current frame data and historical frame data respectively to obtain the current feature and sequence feature;

[0060] Further, the specific steps of Step 2 are as follows:

[0061] Encode the sequence frame data and the current frame data X t respectively through the following formula:

[0062] Y = σ(Norm(SpC(X)))

[0063] where SpC is 3D sparse convolution, Norm is batch normalization operation, σ is a variable set as the RELU function; after encoding, the sequence feature X SF and the current feature X CF are obtained respectively. Y is the encoding result. After encoding, the sequence feature X SF and the current feature X CF are obtained respectively.

[0064] The purpose of this step is to extract low-level features meaningful for the semantic segmentation task from the point cloud data and prepare for subsequent spatio-temporal feature learning.

[0065] Step 3: Mine and learn the potential spatio-temporal clues in the sequence feature. Through the spatio-temporal feature extraction module, extract the long-term spatio-temporal features in the historical frame data and obtain the long-term spatio-temporal dependency relationship in the historical frame data; the specific operation is as follows: For the sequence feature X SF , first correlate the features in the sequence through a linear layer and map them to a high-dimensional space to promote the information interaction of long-term dependencies; subsequently, the features processed by the linear layer are input into the MambaBlock to further mine and capture potential spatio-temporal clues. The formula is as follows:

[0066]

[0067] where, represents the intermediate feature generated during the calculation process; X SF represents the sequence feature; Norm is batch normalization operation; X STF represents the long-term spatio-temporal feature; Linear is a linear layer, and σ is set as the GELU function.

[0068] Step 4: For the long-term spatio-temporal feature and the current feature, use a gated unit for dynamic balance, and use the model to adjust the dependency degree between the current frame data and the historical frame data, so as to obtain the aggregated spatio-temporal feature;

[0069] The gating unit first generates a weight representation Gate through a linear layer and a Sigmoid function, and then the long-term spatio-temporal feature X STF and the current feature X CF are respectively multiplied by different weight representations, and the weighted features are fused by element-wise addition to obtain an effectively aggregated long-term spatio-temporal feature X STAF , as shown in the following formula:

[0070] Gate = Sigmoid(Linear(X CF ))

[0071] X STAF = X CF × Gate + X STF × (1 - Gate)

[0072] In the formula, Linear represents the linear layer.

[0073] Through the gating mechanism, while retaining the key historical spatio-temporal clues, the model can flexibly adjust the weights of each spatio-temporal feature in the final decision, improving the quality of the overall feature representation.

[0074] Step 5: Use the spatio-temporal feature selection state space module to further model the long-term dependencies in the spatio-temporal features;

[0075] Further, the specific steps of step 5 are as follows:

[0076] Step 5.1: Design a spatio-temporal encoder to further model its long-term dependencies. The spatio-temporal encoder consists of a convolutional layer and a Mamba block. The convolutional layer continuously expands the receptive field to promote information interaction between local features with long-term dependencies. After each layer of convolution, it passes through the Mamba block for global context modeling; the process is as follows:

[0077] X E-L atent = σ(BatchNorm(SpC(X STAF )))

[0078] XEO = MambaBlock(XE-Latent)

[0079] Among them, X E-Latent is the temporary feature generated during the process, SpC is the three-dimensional sparse convolution, σ is a variable, set to the RELU function, X EO is the feature generated by the spatio-temporal encoder; the Mamba block consists of a basic Mamba model, introducing a parameterized state space model SelectiveSSM with a selection mechanism to transform the time-invariant system into a linear time-varying system, and the formula is as follows:

[0080] SelectiveSSM(x t ) = y t

[0081] y t = Ch t ,

[0082] where x t is the input representation, the feature X generated by the spatio-temporal encoder EO obtained through linear mapping, is the latent state updated at time step t, with parameter N being the state size, and y t is the output representation; the discrete matrix is updated over time Δ according to the zero-order hold rule:

[0083]

[0084] where and are model parameters; I represents the identity matrix;

[0085] Step 5.2, use the spatio-temporal decoder to reconstruct the features obtained by the spatio-temporal encoder using consecutive transposed convolutional layers, while adding Mamba blocks to prevent the loss of key spatio-temporal context information:

[0086] XD-Latent = MambaBlock(XEO)

[0087] XDO = σ(BatchNorm(SpInvC(XD-Latent)))

[0088] where SpInvC is three-dimensional sparse transposed convolution, σ is a variable set to the RELU function; X D-Latent represents the intermediate result during the feature reconstruction process; X DO is the feature generated by the spatio-temporal decoder; BatchNorm is the batch normalization operation.

[0089] Through the spatio-temporal encoder and spatio-temporal decoder, it is ensured that the model can capture complex time series feature information while maintaining computational efficiency.

[0090] Step 6, send the features output by the spatio-temporal decoder into the semantic segmentation head to assign accurate semantic labels to each point cloud data point and obtain the final point cloud semantic segmentation result.

[0091] The semantic segmentation method proposed in this paper was evaluated on the MWD-L field mining area dataset and the SemanticKITTI public dataset. The MWD-L field mining area dataset was collected by our team in a certain mining area on-site, and the SemanticKITTI dataset is a large-scale outdoor point cloud public dataset for autonomous driving tasks.

[0092] Table 1 Comparison of results of different methods on the MWD-L dataset

[0093]

[0094] Table 1 shows the comparison results of the evaluation metrics of the method proposed in this paper and some mainstream semantic segmentation methods on the MWD-L dataset. The proposed method is better than the recent methods Retro-FPN and WaffleIron, indicating that using spatio-temporal features can effectively improve the model performance. Compared with the Cylinder3D method that directly stacks historical frame data, the proposed method improves by 3.9% in mIoU, demonstrating the superiority of the spatio-temporal clue encoding mechanism in this paper. PointNet++ enhances the learning of local features but lacks in-depth modeling of the dependence between local and global features. In contrast, the proposed method constructs a more comprehensive feature dependence relationship from local to global, achieving a 4.0% performance improvement. In addition, the improved model improves by 1.3% compared with SphereFormer, further indicating that the method proposed in this paper enhances the model's utilization of long-sequence spatio-temporal information, fully reflecting the role of spatio-temporal information in improving the model performance in the field mining area scenario.

[0095] Experiments were conducted on the outdoor dataset SemanticKITTI. As shown in Table 2, the method proposed in this paper is superior to all comparison methods in semantic segmentation metrics. Specifically, by strengthening the ability to extract and aggregate spatio-temporal information, the improved model improves by 1.8% based on SphereFormer, verifying the enhancing effect of long-term spatio-temporal features on the model performance. Compared with the Two-streamMOS and LiDAR-IMU-GNSS methods that integrate point cloud and image as inputs, the proposed method improves by 3.3% and 2.2% respectively. This advantage is mainly attributed to the difficulty of feature alignment caused by the dimensional inconsistency between point cloud and image, which may lead to the loss of feature information in the alignment mode of multi-modal fusion methods.

[0096] Table 2 Comparison of results of different methods on the SemanticKITTI dataset

[0097]

[0098] The content not described in detail in the specification of the present invention belongs to the prior art well-known to those skilled in the art. Although the illustrative specific embodiments of the present invention have been described above for the understanding of those skilled in the art of the present technology, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. A spatiotemporal feature aggregation point cloud semantic segmentation method for intelligent mining equipment, characterized in that: The following steps are involved: Step 1: Establish a 3D semantic segmentation model based on spatiotemporal feature aggregation, then use a lidar sensor to collect point cloud data, divide the collected point cloud data into current frame data and historical frame data, and then input them into the data input unit of the model; Step 2, respectively perform preliminary feature encoding on the current frame data and the historical frame data to obtain the current feature and the sequence feature; Step 3: mining and learning the potential spatiotemporal clues in the sequence features, extracting the long-term spatiotemporal features in the historical frame data through the spatiotemporal feature extraction module, and obtaining the long-term spatiotemporal dependencies in the historical frame data; Step 4: For long-term spatiotemporal features and current features, a gated unit is used to dynamically balance them, and a model is used to adjust the degree of dependence between the current frame data and the historical frame data, thereby obtaining aggregated spatiotemporal features. Step 5, use the spatiotemporal feature selection state space module to further model the long-term dependencies in the spatiotemporal features; In step 6, the features output by the spatiotemporal decoder in the model are fed into the semantic segmentation head, a semantic label is assigned to each point cloud data point, and the final point cloud semantic segmentation result is obtained.

2. The method for semantic segmentation of spatiotemporal feature aggregation point cloud for intelligent mining equipment according to claim 1 is characterized in that: The data input unit in step 1 adopts a two-branch structure to input historical frame data and current frame data respectively. For the input of historical frame data, continuous point cloud frames within a specified time period are taken for processing; Historical frame data and the current frame data X t Respectively expressed as: X t =S t Among them, S t-i is the point cloud data of (ti) frame, concat(.) is the connection operation; for the point cloud frame at a certain time t Generally, N t Points p k =[x k ,y k ,z k ,r k ] T Composition, [x k ,y k ,z k ] T is the three-dimensional coordinate of the point cloud, [r k ] is the laser radar reflection intensity; historical frame data The point cloud frame S at time t t At time (ti) point cloud frame S t-i The continuous multi-frame data is connected and aggregated, i is usually determined after comprehensive evaluation of hardware conditions and model performance; the current frame data X t That is, the point cloud frame S t express.

3. The method for semantic segmentation of spatiotemporal feature aggregation point cloud for intelligent mining equipment according to claim 2 is characterized in that: The step 2 performs preliminary feature encoding on the current frame data and the historical frame data respectively, and the specific operations to obtain the current features and sequence features are: Sequence frame data and the current frame data X t They are encoded by the following formulas: Y=σ(Norm(SpC(X))) Among them, SpC is a three-dimensional sparse convolution, Norm is a batch normalization operation, σ is a variable, set to the RELU function; after encoding, the sequence features X are obtained respectively SF and the current feature X CF , Y is the encoding result, and after encoding, the sequence features X are obtained respectively SF and the current feature X CF .

4. The method for semantic segmentation of spatiotemporal feature aggregation point cloud for intelligent mining equipment according to claim 3 is characterized in that: The step 3 is to mine and learn the potential spatiotemporal clues in the sequence features, extract the long-term spatiotemporal features in the historical frame data through the spatiotemporal feature extraction module, and obtain the long-term spatiotemporal dependency relationship in the historical frame data. The specific process is: For the sequence feature X SF First, the features in the sequence are correlated and mapped to a high-dimensional space through a linear layer to promote information interaction of long-term dependencies; then, the features processed by the linear layer are input into MambaBlock to further mine and capture potential spatiotemporal clues. The formula is as follows: in, is the intermediate feature generated during the calculation process; X SF is the sequence feature; Norm is the batch normalization operation; X STF is the long-term spatiotemporal feature; Linear is the linear layer, σ is a variable, and is set to the GELU function.

5. The method for semantic segmentation of spatiotemporal feature aggregation point cloud for intelligent mining equipment according to claim 4 is characterized in that: In step 4, for long-term spatiotemporal features and current features, a gate control unit is used to dynamically balance the long-term spatiotemporal features and the current features, and a model is used to adjust the degree of dependence between the current frame data and the historical frame data, so as to obtain the aggregated spatiotemporal features. The specific process is as follows: The gate control unit first generates a weight representation Gate through a linear layer and a Sigmoid function, and then the long-term spatiotemporal feature X STF With the current feature X CF Multiply them with different weight representations respectively, and fuse the weighted features by element-by-element addition to obtain the effectively aggregated long-term spatiotemporal feature X STAF , the announcement is as follows: Gate=Sigmoid(Linear(X CF )) X STAF =X CF ×Gate+X STF ×(1-Gate) Where Linear is the linear layer.

6. The method for semantic segmentation of spatiotemporal feature aggregation point cloud for intelligent mining equipment according to claim 5, characterized in that: In step 5, the spatiotemporal feature selection state space module is used to further model the long-term dependency in the spatiotemporal features. The specific steps are as follows: Step 5.1, design a spatiotemporal encoder to further model its long-term dependency. The spatiotemporal encoder consists of convolutional layers and Mamba blocks. Each convolutional layer is passed through a Mamba block for global context modeling. The process is as follows: X E-Latent =σ(BatchNorm(SpC(X STAF ))) X EO =MambaBlock(X E-Latent ) Among them, X E-Latent is a temporary feature generated in the process, SpC is a three-dimensional sparse convolution, σ is a variable, set to RELU function, X EO The features generated by the spatiotemporal encoder; the Mamba block consists of the basic Mamba model, which introduces the parameterized state space model SelectiveSSM of the selection mechanism to transform the time-invariant system into a linear time-varying system. The formula is as follows: SelectiveSSM(x t )=y t the t =Ch t , Among them, x t is the input representation, the feature X generated by the spatiotemporal encoder EO Through linear mapping, is the potential state updated at time step t, parameter N is the state size, y t is the output representation; discrete matrix Updated over time Δ by the zero-order hold rule: in, and is the model parameter; I is the unit matrix; Step 5.2, use the spatiotemporal decoder to reconstruct the features obtained by the spatiotemporal encoder using continuous deconvolution layers, and add Mamba blocks to prevent the loss of key spatiotemporal context information: X D-Latent =MambaBlock(X EO ) X DO =σ(BatchNorm(SpInvC(X D-Latent ))) Among them, SpInvC is a three-dimensional sparse deconvolution, σ is a variable, set to RELU function; X D-Latent is the intermediate result in the feature reconstruction process; X DO Features generated by the spatiotemporal decoder; BatchNorm is the batch normalization operation.