Video Mama superheat degree identification method
By preprocessing and feature embedding of fireeye video, combining bidirectional Mamba module and nonlinear Fourier module, the problem of low recognition efficiency and accuracy of VideoMamba model in the overheat recognition task in the aluminum electrolytic industry is solved, and a more efficient overheat recognition and stable training process is achieved.
Patent Information
- Application Number
- CN202510455466.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
The existing VideoMamba model has low recognition efficiency and accuracy in overheating recognition tasks in the aluminum electrolytic industry, and there are problems with convergence difficulties in the training process.
By preprocessing fireeye videos, fireeye image blocks are generated and feature embeddings, learning vectors and position embeddings are introduced, and deep feature extraction and classification are used for bidirectional Mamba module and nonlinear Fourier module to enhance information interaction between channels and model stability.
The recognition efficiency and accuracy of the model are improved, and the problem of training convergence difficulties is solved, making it more suitable for video classification tasks.
Smart Images

Figure CN120339708A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image recognition technology, and in particular, to a method for recognizing the superheat degree of VideoMamba. Background Art
[0002] The superheat degree is defined as the difference between the electrolyte temperature and the primary crystallization temperature. Among them, the electrolyte temperature can be measured by a sensor, while the primary crystallization temperature needs to be obtained through experiments. However, existing experimental methods are difficult to meet the requirements of rapidity and accuracy for superheat degree in industrial measurement. In the aluminum electrolysis industrial site, workers mainly rely on the visual inspection method to judge the superheat degree state according to the color change and fluctuation of the electrolyte. However, due to the lack of a unified judgment system, individual experience differences and environmental interferences such as smoke, residues, high temperature, etc., the observation results are inconsistent, and continuous monitoring is difficult to achieve. In addition, manual observation depends on the inheritance of experience and is difficult to be systematized, which affects the stability and operation efficiency of the aluminum electrolysis industry.
[0003] When applying VideoMamba to the recognition of the superheat degree of an aluminum electrolysis cell, the detection effect fails to meet the expectation. By analyzing its model structure, it is found that this model lacks information interaction between channels, which limits the feature extraction ability. In addition, from the training results of the FireEye video, there is a problem of difficult convergence in the training process of the model. These factors jointly lead to a low accuracy rate of VideoMamba in the superheat degree recognition task.
[0004] It can be seen that there is an urgent need for a method for recognizing the superheat degree of VideoMamba with high recognition efficiency and accuracy. Summary of the Invention
[0005] In view of this, embodiments of the present disclosure provide a method for recognizing the superheat degree of VideoMamba, which at least partially solves the problem of poor recognition efficiency and accuracy in the prior art.
[0006] Embodiments of the present disclosure provide a method for recognizing the superheat degree of VideoMamba, including:
[0007] Step 1, preprocess the FireEye video to generate a thumbnail corresponding to each frame of the FireEye image, and divide it into multiple FireEye image blocks;
[0008] Step 2, perform feature embedding on the FireEye image blocks and thumbnails to generate a video feature sequence;
[0009] Step 3, introduce a learnable vector at the head of the video feature sequence as a classification token vector;
[0010] Step 4: Introduce learnable spatial position embeddings and temporal position embeddings into the video feature sequence, and encode all the image patches in the embedded video feature sequence to generate Fire-Eye image patch feature embedding vectors containing spatio-temporal features, and construct an input sequence based on this.
[0011] Step 5: Input the input sequence into the bidirectional Mamba modules of multiple consecutive fusion channel attention mechanisms to extract the deep features of the Fire-Eye video and form a deep feature sequence.
[0012] Step 6: Input the deep feature sequence into a learnable non-linear Fourier module to enable channel interaction of the sequence features in the frequency domain.
[0013] Step 7: Extract the classification tokens in the deep feature sequence after channel interaction, convert them into corresponding overheat categories, and finally output the classification result of the Fire-Eye video.
[0014] According to a specific implementation manner of the embodiments of the present disclosure, the step of preprocessing the Fire-Eye video includes:
[0015] Project the Fire-Eye video using a 3D convolution of 1×16×16 to generate L non-overlapping spatio-temporal patches
[0016] where t = T, and L = t×h×w, where H and W respectively represent the height and width of the Fire-Eye video frames.
[0017] According to a specific implementation manner of the embodiments of the present disclosure, the expression of the input sequence is
[0018] X = [X cls , X] + p s + p t
[0019] where X cls is a learnable classification token, and p s and p t are respectively learnable temporal position embeddings and spatial position embeddings.
[0020] According to a specific implementation manner of the embodiments of the present disclosure, step 5 specifically includes:
[0021] Step 5.1: Input the input sequence X into the bidirectional Mamba module to obtain the output feature X B ;
[0022] Step 5.2: After layer normalization of the output feature X B , input it into the channel attention mechanism module again to obtain the feature input X CAM;
[0023] Step 5.3, fuse the output feature X B and the feature input X CAM to obtain the deep feature sequence X BC finally output by the Mamba module, where the expression of the deep feature sequence X BC is
[0024] X BC = X CAM + X B = CAM(Norm(X B )) + X B
[0025] where CAM represents the channel attention mechanism and Norm represents layer normalization.
[0026] According to a specific implementation manner of the embodiments of the present disclosure, the step 6 specifically includes:
[0027] Step 6.1, use the discrete Fourier transform to convert the deep feature sequence into a first frequency-domain feature sequence, where the expression of the discrete Fourier transform is
[0028]
[0029] where DFT represents the discrete Fourier transform, X (l) represents the image frame at the l-th time node, χ is the corresponding frequency-domain feature sequence obtained after the deep feature sequence X undergoes the discrete Fourier transform, and χ R and χ I are the real part and the imaginary part of the first frequency-domain feature sequence χ respectively;
[0030] Step 6.2, input the first frequency-domain feature sequence χ, the frequency-domain complex weight matrix W, and the frequency-domain complex bias B into the frequency-domain channel matrix multiplication layer to obtain a second frequency-domain feature sequence h
[0031]
[0032] where FCMM represents the frequency-domain channel matrix multiplication layer, and Re(h) and Im(h) are the real part and the imaginary part of the frequency-domain output h respectively;
[0033] Step 6.3, use a non-linear activation function to convert the second frequency-domain feature sequence h into a third frequency-domain feature sequence y
[0034]
[0035] where σ represents the non-linear activation function, y R and y IThey are the real part and the imaginary part of the third frequency-domain feature sequence y, respectively;
[0036] Step 6.4: Input the third frequency-domain feature sequence y, the frequency-domain complex weight matrix and the frequency-domain complex deviation into the frequency-domain channel matrix multiplication layer to obtain the fourth frequency-domain feature sequence
[0037]
[0038] wherein, and are the real part and the imaginary part of the fourth frequency-domain feature sequence respectively;
[0039] Step 6.5: Use the inverse discrete Fourier transform to convert the fourth frequency-domain feature sequence into the time-domain feature sequence Z (l)
[0040]
[0041] wherein, IDFT represents the inverse discrete Fourier transform;
[0042] Step 6.6: Integrate the time-domain feature sequence Z (l) to generate the overall output corresponding to the deep feature sequence
[0043] According to a specific implementation manner of an embodiment of the present disclosure, the calculation formula of the frequency-domain channel matrix multiplication layer is
[0044]
[0045] wherein, IDFT represents the inverse discrete Fourier transform, C b represents the number of blocks in the channel, C d represents the number of sub-channels in each block, and the channels are arranged in the manner of C = C b × C d × C d , b << d to form a block diagonal matrix. After the input X and the weight matrix are calculated by the frequency-domain channel matrix multiplication, a mixed feature vector Y is generated.
[0046] According to a specific implementation manner of an embodiment of the present disclosure, the step of converting it into the corresponding superheat degree category includes:
[0047] Pass the classification token through the fully connected layer to obtain the currently predicted Huoyan video category.
[0048] The VideoMamba superheat degree recognition solution in the embodiments of the present disclosure includes: Step 1, preprocess the Huoyan video to generate thumbnails corresponding to each frame of Huoyan images and divide them into multiple Huoyan image blocks; Step 2, perform feature embedding on the Huoyan image blocks and thumbnails to generate a video feature sequence; Step 3, introduce a learnable vector at the head of the video feature sequence as a classification token vector; Step 4, introduce learnable spatial position embedding and temporal position embedding into the video feature sequence, and encode all the image blocks in the embedded video feature sequence to generate Huoyan image block feature embedding vectors containing spatio-temporal features, and construct an input sequence based on this; Step 5, input the input sequence into a bidirectional Mamba module with multiple consecutive fusion channel attention mechanisms to extract deep features of the Huoyan video and form a deep feature sequence; Step 6, input the deep feature sequence into a learnable non-linear Fourier module to enable channel interaction of the sequence features in the frequency domain; Step 7, extract the classification token in the deep feature sequence after channel interaction and convert it into the corresponding superheat degree category, and finally output the classification result of the Huoyan video.
[0049] The beneficial effects of the embodiments of the present disclosure are as follows: Through the solution of the present disclosure, based on the VideoMamba framework with a selective scanning mechanism, redundant information can be effectively removed and key features can be retained, thereby improving the data processing efficiency. By integrating the channel attention mechanism and the bidirectional Mamba module, this method realizes the collaborative processing of the forward and reverse state space models (SSMs) for visual sequences, enhancing the spatial perception ability of the model. In addition, introducing the channel attention mechanism can strengthen the information interaction between channels and effectively extract channel dependence relationships. On this basis, after the bidirectional Mamba module based on the channel attention mechanism, a learnable non-linear Fourier module is further introduced to enable channel interaction of the input sequence in the frequency domain, thereby improving the convergence and stability of the model. The present invention effectively solves the problems of low classification accuracy and difficult training convergence in the Huoyan video classification task of the original VideoMamba model, making it more suitable for the video classification task. Description of the Drawings
[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 It is a schematic flowchart of a VideoMamba superheat degree recognition method provided by the embodiments of the present disclosure;
[0052] Figure 2It is the overall framework diagram of a VideoMamba superheat recognition method provided by an embodiment of the present disclosure;
[0053] Figure 3 It is the structural diagram of a channel attention mechanism module provided by an embodiment of the present disclosure;
[0054] Figure 4 It is the structural diagram of a learnable non - linear Fourier module provided by an embodiment of the present disclosure. Detailed implementation manners
[0055] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0056] The following uses specific specific examples to illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0057] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device can be implemented and this method can be practiced using other structures and / or functionality in addition to one or more of the aspects described herein.
[0058] It also should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure in a schematic manner. The diagrams only show the components related to the present disclosure, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.
[0059] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.
[0060] Embodiments of the present disclosure provide a method for identifying the superheat degree of VideoMamba. This method can be applied to the process of identifying the superheat degree of aluminum electrolytic cells in industrial scenarios.
[0061] See Figure 1 , which is a schematic flowchart of a method for identifying the superheat degree of VideoMamba provided by an embodiment of the present disclosure. As Figure 1 shown, the method mainly includes the following steps:
[0062] Step 1: Preprocess the FireEye video to generate a thumbnail corresponding to each frame of the FireEye image, and divide it into multiple FireEye image blocks;
[0063] Specifically, the overall framework diagram of the method of the present disclosure is as Figure 2 shown. Preprocess the FireEye video to generate a thumbnail corresponding to each frame of the FireEye image, and divide it into several image blocks. The preprocessing method is as follows: Use a 3D convolution of 1×16×16 to project the input video to generate L non-overlapping spatio-temporal patches where t = T, and L = t×h×w. H and W respectively represent the height and width of the FireEye video frame.
[0064] Step 2: Perform feature embedding on the FireEye image blocks and thumbnails to generate a video feature sequence;
[0065] Specifically, perform feature embedding on the obtained FireEye image blocks and thumbnails to generate corresponding feature embedding vectors to represent the feature information of each image block and thumbnail
[0066] Step 3: Introduce a learnable vector at the head of the video feature sequence as a classification token vector;
[0067] Step 4: Introduce learnable spatial position embedding and temporal position embedding into the video feature sequence, and encode all the image blocks in the embedded video feature sequence to generate a FireEye image block feature embedding vector containing spatio-temporal features, and construct an input sequence accordingly;
[0068] Specifically, use the learnable temporal position embedding and spatial position embedding to encode all the image blocks in Step 3 to fuse the spatio-temporal position information, generate a FireEye image block feature embedding vector containing spatio-temporal features, and construct the input sequence of the subsequent model. The specific structure of this input sequence is as follows:
[0069] X = [X cls , X] + p s + p t
[0070] where X cls is a learnable classification token, and p s and p t are learnable temporal position embedding and spatial position embedding respectively.
[0071] Step 5: Input the input sequence into a series of bidirectional Mamba modules with a fusion channel attention mechanism to extract the deep features of the Huoyan video and form a deep feature sequence;
[0072] Specifically, in implementation, input the input sequence obtained in step (4) into a series of bidirectional Mamba modules with a fusion channel attention mechanism to extract the deep features of the Huoyan video. The bidirectional Mamba module processes the flattened visual sequence using forward and backward state space models (SSMs) to enhance the model's spatial perception ability; meanwhile, the channel attention mechanism strengthens the information interaction between different channels and extracts the dependency relationships between channels to improve the feature expression ability. The bidirectional Mamba module with a fusion channel attention mechanism is as shown Figure 2 on the right. The calculation process of this step can be expressed by the following formula:
[0073] X BC = X CAM + X B = CAM(Norm(X B )) + X B
[0074] (5-1) Input the input sequence X generated in step (4) into the bidirectional Mamba module to obtain the output feature X B ;
[0075] (5-2) After layer normalization of the feature sequence X B from step (5-1), input it into the channel attention mechanism module again to obtain the feature input X CAM . The structure of the channel attention mechanism module is as shown Figure 3 on the right, and it dynamically weights each channel through a series of operations such as convolution, GeLU activation, and global average pooling, thereby improving the model's ability to allocate importance to features;
[0076] (5-3) Finally, fuse the feature output X B from step (5-1) and the feature output X CAM from step (5-2) to obtain the final output X BC of this module.
[0077] Step 6: Input the deep feature sequence into the learnable non-linear Fourier module to enable channel interaction of the sequence features in the frequency domain;
[0078] In specific implementation, input the feature sequence output in step (5) into the learnable non-linear Fourier module to enable channel interaction of the sequence features in the frequency domain, thereby improving the convergence and discriminative ability of the model. The structure of the learnable non-linear Fourier module is as Figure 4 shown.
[0079] (6-1) Use the discrete Fourier transform to convert the time-domain feature sequence obtained in step (5) into a frequency-domain feature sequence. The calculation expression is transformed as follows:
[0080]
[0081] where DFT represents the discrete Fourier transform, X (l) represents the image frame at the l-th time node, and χ is the corresponding frequency-domain feature sequence obtained after the discrete Fourier transform of the time-domain feature sequence X. χ R and χ I are the real and imaginary parts of the frequency-domain feature sequence χ respectively.
[0082] (6-2) Input the frequency-domain feature sequence χ obtained in step (6-1) and the frequency-domain complex weight matrix frequency-domain complex bias into the frequency-domain channel matrix multiplication layer to obtain the frequency-domain feature sequence h. This process can be expressed by the following formula:
[0083]
[0084] Re(h) and Im(h) are the real and imaginary parts of the frequency-domain output h respectively. FCMM represents the frequency-domain channel matrix multiplication layer, and its specific operation method can be expressed by the following formula:
[0085]
[0086] where IDFT represents the inverse discrete Fourier transform, C b represents the number of blocks in the channel, and C d represents the number of sub-channels in each block. The channels are arranged in the form of C = C b ×C d ×C d (b << d) to form a block diagonal matrix. After the calculation of the frequency-domain channel matrix multiplication of the input X and the weight matrix , a mixed feature vector Y is generated.
[0087] (6-3) Using a non-linear activation function, the frequency-domain feature sequence h in step (6-2) is transformed into a frequency-domain feature sequence y. This process can be represented by the following formula:
[0088]
[0089] where σ represents the non-linear activation function, and y R and y I are the real and imaginary parts of the frequency-domain feature sequence y, respectively.
[0090] (6-4) The frequency-domain feature sequence y obtained from step (6-3) is once again input, together with the frequency-domain complex weight matrix frequency-domain complex bias into the frequency-domain channel matrix multiplication layer to obtain the frequency-domain feature sequence This process can be represented by the following formula:
[0091]
[0092] where and are the real and imaginary parts of the frequency-domain feature sequence respectively.
[0093] (6-5) Using the inverse discrete Fourier transform, the frequency-domain feature sequence obtained in step (6-4) is transformed into a time-domain feature sequence Z (l) . This process can be represented by the following formula:
[0094]
[0095] where IDFT represents the inverse discrete Fourier transform.
[0096] (6-6) Finally, the output Z (l) across L time nodes is integrated to generate the overall output
[0097] Step 7: Extract the classification token from the deep feature sequence after channel interaction, convert it into the corresponding superheat degree category, and finally output the classification result of the Huoyan video.
[0098] Specifically, when implementing, extract the head vector, i.e., the classification token, from the feature sequence obtained in step 6, use the fully connected layer to convert it into the corresponding superheat degree category, and finally output the classification result of the Huoyan video.
[0099] During the actual training process, the present invention adopts a pre-training strategy. First, the channel attention mechanism module is pre-trained independently to obtain better weight parameters, enabling it to effectively extract channel dependencies. Subsequently, during the training process of a VideoMamba superheat recognition method that combines the channel attention mechanism and the Fourier transform, the pre-trained weight file is directly loaded, thereby improving the convergence speed and classification performance of the model without significantly increasing the consumption of computing resources.
[0100] The VideoMamba superheat recognition method provided in this embodiment can effectively remove redundant information and retain key features based on the VideoMamba framework with a selective scanning mechanism, thereby improving the data processing efficiency. By fusing the channel attention mechanism and the bidirectional Mamba module, this method realizes the collaborative processing of forward and reverse state space models (SSMs) for visual sequences, enhancing the model's spatial perception ability. In addition, introducing the channel attention mechanism can strengthen the information interaction between channels and effectively extract channel dependencies. On this basis, after the bidirectional Mamba module based on the channel attention mechanism, a learnable non-linear Fourier module is further introduced to enable channel interaction of the input sequence in the frequency domain, thereby improving the convergence and stability of the model. The present invention effectively solves the problems of low classification accuracy and difficult training convergence of the original VideoMamba model in the Huoyan video classification task, making it more suitable for video classification tasks.
[0101] It should be understood that each part of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof.
[0102] As described above, the above are only specific embodiments of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present disclosure should be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. A method for identifying the superheat degree of VideoMamba, characterized in that, Including: Step 1: Preprocess the Huoyan video to generate thumbnails corresponding to each frame of the Huoyan image and divide them into multiple Huoyan image patches; Step 2: Perform feature embedding on the Huoyan image patches and thumbnails to generate a video feature sequence; Step 3: Introduce a learnable vector at the head of the video feature sequence as a classification token vector; Step 4: Introduce learnable spatial position embeddings and temporal position embeddings into the video feature sequence, and encode all the image patches in the embedded video feature sequence to generate Huoyan image patch feature embedding vectors containing spatio-temporal features, and construct an input sequence accordingly; Step 5: Input the input sequence into a bidirectional Mamba module with multiple consecutive fused channel attention mechanisms to extract deep features of the Huoyan video and form a deep feature sequence; Step 6: Input the deep feature sequence into a learnable non-linear Fourier module to enable channel interaction of the sequence features in the frequency domain; Step 7: Extract the classification token in the deep feature sequence after channel interaction and convert it into the corresponding superheat degree category, and finally output the classification result of the Huoyan video.
2. The method according to claim 1, characterized in that, The steps of preprocessing the Huoyan video include: Project the Huoyan video using a 1×16×16 3D convolution to generate L non - overlapping spatio - temporal patches where t = T, and L = t × h × w, where H and W respectively represent the height and width of the Huoyan video frame.
3. The method according to claim 2, wherein The expression of the input sequence is X = [X cls , X] + p s + p t Among them, X cls is a learnable classification token, p s and p t are learnable temporal position embedding and spatial position embedding respectively.
4. The method according to claim 3, wherein Specifically, Step 5 includes: Step 5.1, the input sequence X is input into the bidirectional Mamba module to obtain the output feature X B ; Step 5.2, input the output feature X B After layer normalization, input it into the channel attention mechanism module again to obtain the feature input X CAM ; Step 5.3, fuse the output feature X B and the feature input X CAM to obtain the deep feature sequence X which is the final output of the Mamba module BC , where the expression of the deep feature sequence X BC is X BC = X CAM + X B = CAM(Norm(X B )) + X B Where, CAM represents the channel attention mechanism, and Norm represents layer normalization.
5. The method according to claim 4, wherein Specifically, Step 6 includes: Step 6.1: Use the discrete Fourier transform to convert the deep feature sequence into a first frequency domain feature sequence, where the expression of the discrete Fourier transform is where DFT represents the discrete Fourier transform, X (l) represents the image frame at the l-th time node, χ is the corresponding frequency domain feature sequence obtained after the discrete Fourier transform of the deep feature sequence X, χ R and χ I are the real part and the imaginary part of the first frequency domain feature sequence χ, respectively; Step 6.2: Input the first frequency domain feature sequence χ, the frequency domain complex weight matrix W, and the frequency domain complex bias B into the frequency domain channel matrix multiplication layer to obtain a second frequency domain feature sequence h FCMM(χ (l) ,W,B) = Re(h) l + jIm(h) l Where, FCMM represents the frequency domain channel matrix multiplication layer, and Re(h) and Im(h) are the real part and the imaginary part of the frequency domain output h respectively; Step 6.3: Use a non-linear activation function to convert the second frequency domain feature sequence h into a third frequency domain feature sequence y where σ represents a non-linear activation function, and y R and y I are the real and imaginary parts of the third frequency domain feature sequence y, respectively; Step 6.4: Input the third frequency domain feature sequence y, the frequency domain complex weight matrix W, and the frequency domain complex bias B into the frequency domain channel matrix multiplication layer to obtain a fourth frequency domain feature sequence Z Among them, Z R and Z I are respectively the real part and the imaginary part of the fourth frequency domain feature sequence Z; Step 6.5, use the inverse discrete Fourier transform to transform the fourth frequency domain feature sequence Z into a time domain feature sequence Z (l) Where, IDFT represents the inverse discrete Fourier transform; Step 6.6, integrating the time-domain feature sequence Z (l) Generating the overall output corresponding to the deep feature sequence 6. The method according to claim 5, characterized in that, The calculation formula of the frequency domain channel matrix multiplication layer is Among them, IDFT represents the inverse discrete Fourier transform, C b represents the number of blocks in the channel, C d represents the number of sub-channels in each block. The channel is arranged in the form of C = C b ×C d ×C d , where b << d, to form a block diagonal matrix. After the input X and the weight matrix are calculated through the multiplication of the frequency domain channel matrix, a mixed feature vector Y is generated.
7. The method according to claim 6, wherein The steps of converting it into the corresponding superheat degree category include: Pass the classification token through a fully connected layer to obtain the currently predicted Huoyan video category.