Indoor human body posture detection method based on light and radio fusion

By integrating WiFi CSI and visible light signals for posture detection, the method addresses accuracy and privacy issues in existing technologies, offering robust and accurate posture estimation in complex indoor settings.

CN120316614APending Publication Date: 2025-07-15SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510383864.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing human posture detection methods have problems such as incomplete feature extraction, poor data fusion effect, and difficulty in dealing with multimodal signals in complex indoor environments. Especially in environments with significant multipath effects, the traditional methods are not robust enough.

Method used

Using a method based on light and radio fusion, the WiFi CSI signal and visible light intensity signal are preprocessed, the CSI features and light intensity characteristics are extracted, and the cross-modal fusion is performed. The attention mechanism is used to determine the feature representation after the cross-modal fusion, and the posture recognition is finally completed.

Benefits of technology

It realizes the accuracy and reliability of human posture estimation and personnel density analysis in complex environments, protects user privacy, and does not require users to wear equipment for contactless monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316614A_ABST
    Figure CN120316614A_ABST
Patent Text Reader

Abstract

The invention discloses an indoor human body posture detection method based on light and radio fusion, and belongs to the technical field of posture detection, and the method comprises the following steps: S1, carrying out the preprocessing of a WiFi CSI signal and a visible light intensity signal collected indoors; s2, extracting a CSI feature of the preprocessed WiFi CSI signal and a light intensity feature of the visible light intensity signal; s3, fusing the CSI features and the light intensity features, and determining feature representation after cross-modal fusion; and S4, according to the feature representation after cross-modal fusion, obtaining a posture probability vector, and completing posture recognition. According to the method, the accuracy of human body posture estimation and personnel density analysis can be effectively improved, and particularly in a dynamic and variable indoor environment, the posture recognition efficiency and reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of attitude detection, and particularly relates to an indoor human body attitude detection method based on the fusion of light and radio. Background Art

[0002] Human body pose recognition has wide demands and high application value in society. In the field of medical monitoring, activity trajectory and pose recognition technologies can monitor unaccompanied patients or the elderly in real time, judge whether an accident has occurred to the target, and accurately locate and call for help in time after discovery, ensuring that they can receive help and treatment in the first time; in the field of home security, by deploying human body abnormal pose recognition and detection technologies, the household can monitor the dynamics at home in real time, discover and warn of any suspicious activities in time, and significantly enhance home security protection; in the field of public security, this technology can monitor the flow of people in public places in real time, discover abnormal behaviors or potential security threats in time, and prevent the occurrence of security accidents such as stampedes and fights. In summary, human body pose recognition has significant application value.

[0003] Traditional detection methods mainly include detection methods based on built-in sensors of wearable devices or smart devices, detection methods based on computer vision, and detection methods based on radar. First, although human activity recognition based on sensors can obtain high accuracy, users must always carry these devices, and sensors need to be attached to each body part to be sensed, which is very inconvenient in some application scenarios. In addition, the limited lifespan of sensors also increases the user's usage cost. Second, the perception effect of vision-based perception methods is not ideal under poor lighting conditions, and when the perception target is blocked, there are monitoring blind spots and effective target information cannot be obtained. More importantly, using cameras for human behavior activity recognition will cause serious privacy problems. Although researchers have also proposed using depth cameras and thermal infrared cameras to overcome privacy problems, the cost of the devices limits their wide application. Finally, although radar signals are not affected by light and have strong penetration ability, they can effectively protect user privacy while achieving high-precision perception. However, this perception method requires dedicated wireless transmitters and receivers, which limits the wide use of this technology in people's daily lives.

[0004] Existing human body pose detection based on CSI signals or light intensity signals also has problems. Traditional non-machine learning methods using CSI signals rely on manually designed features (such as signal amplitude mean and Doppler frequency shift, etc.), and it is difficult to effectively capture complex time-frequency domain dynamic features. Especially in an environment with significant multipath effects, the feature robustness is insufficient. And traditional image processing methods based on light intensity signals (such as background subtraction and contour detection) need to preset rules (such as human body proportion models), and it is difficult to handle occlusion or complex poses. Summary of the Invention

[0005] In order to solve the problems of incomplete feature extraction, poor data fusion effect and difficulty in processing multi-modal signals in complex indoor environments existing in the existing analysis methods based on a single signal source, the present invention proposes an indoor human posture detection method based on the fusion of light and radio.

[0006] The technical solution of the present invention is: an indoor human posture detection method based on the fusion of light and radio includes the following steps:

[0007] S1. Preprocess the WiFi CSI signal and visible light intensity signal collected indoors;

[0008] S2. Extract the CSI features of the preprocessed WiFi CSI signal and the light intensity features of the visible light intensity signal;

[0009] S3. Fuse the CSI features and light intensity features to determine the feature representation after cross-modal fusion;

[0010] S4. According to the feature representation after cross-modal fusion, obtain the posture probability vector to complete posture recognition.

[0011] Further, in S1, the method for preprocessing the WiFi CSI signal is: separate the real part and the imaginary part of the WiFi CSI signal, and determine the amplitude and phase of the subcarriers;

[0012] In S1, the method for preprocessing the visible light intensity signal is: perform moving average filtering on the visible light intensity signal.

[0013] Further, S2 includes the following sub-steps:

[0014] S21. Use a sliding convolution kernel to extract the frequency response correlation characteristics of adjacent subcarriers in the preprocessed WiFi CSI signal;

[0015] S22. Use a pooling layer to perform temporal downsampling on the frequency response correlation characteristics;

[0016] S23. Based on the temporal downsampling result, use a convolutional layer to fuse the real part, imaginary part, amplitude, and phase of the subcarriers;

[0017] S24. Use a pooling layer to perform spatial downsampling on the fusion result;

[0018] S25. Use a spatio-temporal fully connected layer to perform multi-dimensional mapping on the spatial downsampling result to obtain the CSI features of the WiFi CSI signal;

[0019] S26. Use a convolutional layer to extract the local time series of the preprocessed visible light intensity signal;

[0020] S27. Coarsely abstract the local time series using a pooling layer;

[0021] S28. Based on the coarsely abstracted result, compress the channel dimension of the local time series using a depthwise convolutional layer;

[0022] S29. Use a global average pooling layer to expand the compressed result and align the expanded compressed result with the CSI features of the WiFi CSI signal to obtain the light intensity features of the visible light intensity signal.

[0023] Further, S3 includes the following sub-steps:

[0024] S31. Use the CSI features as queries, and the light intensity features as keys and values;

[0025] S32. Calculate the correlation between the queries and keys and determine the attention weights;

[0026] S33. Determine the feature representation after cross-modal fusion according to the attention weights.

[0027] Further, in S31, the query Q CSI has the following expression:

[0028]

[0029] In the formula, X CSI represents the feature matrix of the WiFi CSI signal, N represents the first time step, W Q represents the first learning weight matrix, R represents the real number space, and d Q represents the dimension of the first learning weight matrix;

[0030] The key K Light has the following expression:

[0031]

[0032] In the formula, X Light represents the feature matrix of the light intensity signal, W K represents the second learning weight matrix, M represents the spatial dimension, and d K represents the dimension of the second learning weight matrix;

[0033] The value V Light has the following expression:

[0034]

[0035] In the formula, W V represents the third learning weight matrix, and d V represents the dimension of the third learning weight matrix.

[0036] Further, in S32, the calculation formula for the correlation between the query and the key textAttention(Q CSI ,K Light ) is as follows:

[0037]

[0038] In the formula, Q CSI represents the query, K Light represents the key, H represents the transpose operation of the matrix, and d Q represents the dimension of the first learning weight matrix;

[0039] In S31, the calculation formula for the attention weight textAttn_weights is:

[0040]

[0041] In the formula, Softmax(·) represents the activation function.

[0042] Further, in S33, the expression for the feature representation textCross-Attention_Output after cross-modal fusion is:

[0043] textCross-Attention_Output = textAttn_weights·V Light ;

[0044] In the formula, textAttn_weights represents the attention weight, and V Light represents the value.

[0045] Further, S4 includes the following sub-steps:

[0046] S41. Perform global max pooling on the feature representation after cross-modal fusion to obtain a global feature vector;

[0047] S42. Obtain class scores based on the global feature vector;

[0048] S43. Convert the class scores into a multi-class behavior probability distribution;

[0049] S44. Obtain a pose probability vector based on the multi-class behavior probability distribution to complete pose recognition.

[0050] Further, in S41, the expression for the global feature vector h is:

[0051]

[0052] In the formula, T represents the time step, and F fusionIt represents the latest expression form of the feature representation after cross-modal fusion, t represents the time step number, d represents the feature dimension number, GlobalMaxPool(·) represents the global max pooling function, and D represents the feature dimension.

[0053] Further, in S42, the calculation formula for the class score z is:

[0054] z = Wh + b;

[0055] In the formula, W represents the learnable weight matrix, h represents the global feature vector, and b represents the bias vector;

[0056] In S44, the calculation formula for the pose probability vector p is:

[0057] p = Softmax(W·GlobalMaxPool(F fusion ));

[0058] In the formula, Softmax(·) represents the activation function, W represents the learnable weight matrix, F fusion represents the latest expression form of the feature representation after cross-modal fusion, b represents the bias vector, and GlobalMaxPool(·) represents the global max pooling function.

[0059] The beneficial effects of the present invention are:

[0060] (1) Compared with the traditional camera solution, the present invention does not involve image data, fundamentally protecting user privacy; compared with the solution based on light intensity signals, the present invention reduces the influence of ambient light intensity changes on the accuracy rate; compared with the wearable device solution, the present invention does not require the user to wear any device, realizing non-contact monitoring; compared with the solution based on WiFi CSI signals, the present invention improves the robustness in complex environments through multi-modal signal fusion;

[0061] (2) The present invention can effectively improve the accuracy of human pose estimation and personnel density analysis, especially in dynamic and changing indoor environments, improving the efficiency and reliability of pose recognition. Brief Description of the Drawings

[0062] Figure 1 is a flowchart of the indoor human pose detection method based on the fusion of light and radio;

[0063] Figure 2 is a schematic diagram of the model of the CSI processing branch in the present invention;

[0064] Figure 3 is a schematic diagram of the model of the light intensity processing branch in the present invention;

[0065] Figure 4Schematic diagram of the model for cross-modal feature fusion and pose classification in the present invention. Detailed implementation manners

[0066] The embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0067] As Figure 1 shown, the present invention provides an indoor human pose detection method based on the fusion of light and radio, including the following steps:

[0068] S1. Preprocess the WiFi CSI signals and visible light intensity signals collected indoors;

[0069] S2. Extract the CSI features of the preprocessed WiFi CSI signals and the intensity features of the visible light intensity signals;

[0070] S3. Fuse the CSI features and the intensity features to determine the feature representation after cross-modal fusion;

[0071] S4. According to the feature representation after cross-modal fusion, obtain the pose probability vector to complete pose recognition.

[0072] In the embodiments of the present invention, in S1, the method for preprocessing the WiFi CSI signals is: separating the real part and the imaginary part of the WiFi CSI signals and determining the amplitude and phase of the subcarriers;

[0073] In S1, the method for preprocessing the visible light intensity signals is: performing moving average filtering on the visible light intensity signals.

[0074] Deploy multiple sensor nodes in the target indoor environment to synchronously collect WiFi channel state information (CSI) and visible light intensity signals. For the CSI part, the real part and the imaginary part are separated, and the amplitude and phase of the subcarriers are calculated. For the light intensity part, environmental noise is eliminated through dynamic splitting filtering. Finally, ensure the spatio-temporal consistency of the multi-modal data to provide a high signal-to-noise ratio input for subsequent processing.

[0075] In the embodiments of the present invention, S2 includes the following sub-steps:

[0076] S21. Use a sliding convolution kernel to extract the frequency response correlation characteristics of adjacent subcarriers in the preprocessed WiFi CSI signals;

[0077] S22. Use a pooling layer to perform temporal downsampling on the frequency response correlation characteristics;

[0078] S23. Based on the temporal downsampling result, use a convolutional layer to fuse the real part, the imaginary part, the amplitude, and the phase of the subcarriers;

[0079] S24. Perform spatial downsampling on the fusion result using a pooling layer;

[0080] S25. Perform multi-dimensional mapping on the spatially downsampled result using a spatio-temporal fully connected layer to obtain the CSI features of the WiFi CSI signal;

[0081] S26. Extract the local time series of the preprocessed visible light intensity signal using a convolutional layer;

[0082] S27. Coarsely abstract the local time series using a pooling layer;

[0083] S28. Based on the coarsely abstracted result, compress the channel dimension of the local time series using a depthwise separable convolutional layer;

[0084] S29. Use a global average pooling layer to expand the compressed result and align the expanded compressed result with the CSI features of the WiFi CSI signal to obtain the intensity features of the visible light intensity signal.

[0085] WiFi CSI signal: After completing real-imaginary separation and amplitude-phase feature calculation in the preprocessing layer, first perform cross-subcarrier 2D convolution along the subcarrier dimension to capture the frequency response correlation characteristics between adjacent subcarriers by sliding the convolutional kernel. Subsequently, compress the time series resolution through max pooling in the time direction to reduce high-frequency noise interference. Then, use wide-kernel cross-feature convolution to fuse the complementary information of the four channels of real part, imaginary part, amplitude, and phase. Finally, reduce the redundant spatial dimension of the sensor array through spatial pooling operation. Ultimately, map the multi-dimensional features to a 256-dimensional time series vector through a spatio-temporal fully connected layer and wait for cross-modal interaction.

[0086] Light intensity signal: After the original light intensity signal suppresses the ambient light mutation noise through dynamic differential filtering, first extract the local time series pattern and maintain time continuity through long-range time convolution. Subsequently, use an average pooling layer to coarsely abstract the time series to enhance the anti-pulse interference ability. Then, compress the channel dimension while retaining the time series dynamics through depthwise separable convolution. Finally, expand the compressed features to a 128-dimensional time series vector through global average pooling and linear interpolation reconstruction to achieve time series alignment with the CSI branch.

[0087] In the embodiment of the present invention, S3 includes the following sub-steps:

[0088] S31. Use the CSI features as the query, and the intensity features as the key and value;

[0089] S32. Calculate the correlation between the query and the key and determine the attention weights;

[0090] S33. Determine the feature representation after cross-modal fusion according to the attention weights.

[0091] In the embodiment of the present invention, in S31, the query Q CSI has the following expression:

[0092]

[0093] In the formula, X CSI represents the feature matrix of the WiFi CSI signal, N represents the first time step, W Q represents the first learning weight matrix, R represents the real number space, and d Q represents the dimension of the first learning weight matrix;

[0094] The key K Light has the following expression:

[0095]

[0096] In the formula, X Light represents the feature matrix of the light intensity signal, W K represents the second learning weight matrix, M represents the spatial dimension, and d K represents the dimension of the second learning weight matrix;

[0097] The value V Light has the following expression:

[0098]

[0099] In the formula, W V represents the third learning weight matrix, and d V represents the dimension of the third learning weight matrix.

[0100] and are the learning weight matrices, corresponding to the transformations of the query, key, and value respectively.

[0101] In the embodiment of the present invention, in S32, the correlation between the query and the key textAttention(Q CSI , K Light ) is calculated as follows:

[0102]

[0103] In the formula, Q CSI represents the query, K Light represents the key, H represents the transpose operation on the matrix, and d Q represents the dimension of the first learning weight matrix;

[0104] In S31, the calculation formula of the attention weight textAttn_weights is:

[0105]

[0106] In the formula, Softmax(·) represents the activation function.

[0107] In the embodiment of the present invention, in S33, the expression of the feature representation textCross - Attention_Output after cross - modal fusion is:

[0108] textCross - Attention_Output = textAttn_weights·V Light ;

[0109] In the formula, textAttn_weights represents the attention weight, and V Light represents the value.

[0110] In the embodiment of the present invention, S4 includes the following sub - steps:

[0111] S41. Perform global max - pooling processing on the feature representation after cross - modal fusion to obtain a global feature vector;

[0112] S42. Obtain the class score according to the global feature vector;

[0113] S43. Convert the class score into a multi - class behavior probability distribution;

[0114] S44. Obtain the pose probability vector according to the multi - class behavior probability distribution to complete pose recognition.

[0115] In the embodiment of the present invention, in S41, the expression of the global feature vector h is:

[0116]

[0117] In the formula, T represents the time step, and F fusion represents the latest expression form of the feature representation after cross - modal fusion, t represents the time step number, d represents the feature dimension number, GlobalMaxPool(·) represents the global max - pooling function, and D represents the feature dimension.

[0118] In the embodiment of the present invention, in S42, the calculation formula of the class score z is:

[0119] z = Wh + b;

[0120] In the formula, W represents the learnable weight matrix, h represents the global feature vector, and b represents the bias vector;

[0121] In S44, the calculation formula of the pose probability vector p is:

[0122] p = Softmax(W·GlobalMaxPool(Ffusion ) + b);

[0123] Wherein, Softmax(·) represents an activation function, W represents a learnable weight matrix, F fusion represents the latest expression form of the feature representation after cross-modal fusion, b represents a bias vector, and GlobalMaxPool(·) represents a global max pooling function.

[0124] In the embodiment of the present invention, as Figure 2 shown, in the CSI data input layer, four sensors are used for CSI data, and each sensor records the real and imaginary parts of 32 subcarriers, including data of T groups of time. In the preprocessing layer, the amplitude and phase of each subcarrier are first calculated. Now each subcarrier has four features, namely the real part, the imaginary part, the amplitude and the phase, that is, each sensor has 32 * 4 features per unit time; in the C1 layer, a 3×1 convolution kernel slides along the subcarrier dimension for a total of 32 channels, while extracting the correlation between subcarriers, the time dimension is kept intact. In the S2 layer, the pooling layer performs time downsampling with a pooling kernel of 2. In the C3 layer, a 3*4 convolution kernel is used for convolution operation to fuse the features of the four dimensions of the real part, the imaginary part, the amplitude and the phase. In the pooling layer, a 2*2 pooling kernel is used for spatial downsampling. In the F5 layer, the fully connected layer flattens the spatial and feature dimensions, waiting for cross-modal fusion.

[0125] In the embodiment of the present invention, as Figure 3 shown, in the light intensity data input layer, four sensors are used for light intensity data, and each sensor records one light intensity value, including data of T groups of time. In the preprocessing layer, dynamic differential filtering is performed, and median filtering is performed using a sliding window of size 10. In the C1 layer, the time expansion convolution is performed with a convolution kernel of size 5, and the padding is same. In the S2 layer, the pooling layer performs time downsampling with a pooling kernel of 2. In the C3 layer, the time alignment convolution is performed with a convolution kernel of size 5, and the padding is same. In the S4 layer, statistical features are extracted through temporal global average to eliminate the influence of local fluctuations on classification. In the F5 layer, time interpolation expansion is performed, and linear interpolation is performed after the fully connected layer to align the time dimension of the CSI branch.

[0126] In the embodiment of the present invention, as Figure 4 shown, CSI is used as the Query, light intensity is used as the Key and Value, and the outputs are concatenated and then fused; the pooling layer performs temporal global max pooling; the fully connected layer + Softmax function is used for pose classification, and the probability distribution vectors of different classes are output.

[0127] Those of ordinary skill in the art will realize that the embodiments described herein are provided to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on these technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.

Claims

1. An indoor human body posture detection method based on the fusion of light and radio, characterized in that, It includes the following steps: S1. Preprocess the WiFi CSI signals and visible light intensity signals collected indoors; S2. Extract the CSI features of the preprocessed WiFi CSI signals and the intensity features of the visible light intensity signals; S3. Fuse the CSI features and the intensity features to determine the feature representation after cross-modal fusion; S4. Obtain the pose probability vector based on the feature representation after cross-modal fusion to complete pose recognition.

2. The indoor human body posture detection method based on the fusion of light and radio according to claim 1, wherein In S1, the method for preprocessing the WiFi CSI signals is: separate the real part and the imaginary part of the WiFi CSI signals, and determine the amplitude and phase of the subcarriers; In S1, the method for preprocessing the visible light intensity signals is: perform moving average filtering on the visible light intensity signals.

3. The indoor human body posture detection method based on the fusion of light and radio according to claim 1, wherein S2 includes the following sub-steps: S21. Use a sliding convolution kernel to extract the frequency response correlation characteristics of adjacent subcarriers in the preprocessed WiFi CSI signals; S22. Use a pooling layer to perform temporal downsampling on the frequency response correlation characteristics; S23. Based on the temporal downsampling results, use a convolutional layer to fuse the real part, the imaginary part, the amplitude, and the phase of the subcarriers; S24. Use a pooling layer to perform spatial downsampling on the fusion results; S25. Use a spatio-temporal fully connected layer to perform multi-dimensional mapping on the spatial downsampling results to obtain the CSI features of the WiFi CSI signals; S26. Use a convolutional layer to extract the local time series of the preprocessed visible light intensity signals; S27. Use a pooling layer to perform coarse-grained abstraction on the local time series; S28. Based on the coarse-grained abstraction results, use a depthwise separable convolutional layer to compress the channel dimension of the local time series; S29. Use a global average pooling layer to expand the compression results and align the expanded compression results with the CSI features of the WiFi CSI signals to obtain the intensity features of the visible light intensity signals.

4. The indoor human body posture detection method based on the fusion of light and radio according to claim 1, characterized in that, S3 includes the following sub-steps: S31. Use the CSI features as queries, and use the intensity features as keys and values; S32. Calculate the correlation between the queries and the keys, and determine the attention weights; S33. Determine the feature representation after cross-modal fusion according to the attention weights.

5. The indoor human body posture detection method based on the fusion of light and radio according to claim 4, characterized in that In the S31, the query Q CSI has the following expression: where X CSI represents the feature matrix of the WiFi CSI signal, N represents the first time step, and W Q represents the first learning weight matrix, R represents the real number space, and d Q represents the dimension of the first learning weight matrix; The key K Light has the following expression: In the formula, X Light represents the feature matrix of the optical intensity signal, W K represents the second learning weight matrix, M represents the spatial dimension, d K represents the dimension of the second learning weight matrix; The value V Light has the following expression: Where, W V represents the third learning weight matrix, and d V represents the dimension of the third learning weight matrix.

6. The indoor human body posture detection method based on the fusion of light and radio according to claim 4, characterized in that In S32, the calculation formula for querying the correlation between and keys textAttention(Q CSI ,K Light ) is as follows: Where Q CSI represents a query, K Light represents a key, H represents the transpose operation on the matrix, and d Q represents the dimension of the first learning weight matrix; In S31, the calculation formula for the attention weight textAttn_weights is: In the formula, Softmax(·) represents the activation function.

7. The indoor human body posture detection method based on the fusion of light and radio according to claim 4, characterized in that In S33, the expression for the feature representation textCross-Attention_Output after cross-modal fusion is: textCross-Attention_Output = textAttn_weights · V Light ; where twxtAttn_weeights represents the attention weights, and V Light represents the value.

8. The indoor human body posture detection method based on the fusion of light and radio according to claim 1, characterized in that S4 includes the following sub-steps: S41. Perform global max pooling on the feature representation after cross-modal fusion to obtain the global feature vector; S42. Obtain the class scores based on the global feature vector; S43. Convert the class scores into a multi-class behavior probability distribution; S44. Obtain the pose probability vector according to the multi-class behavior probability distribution to complete pose recognition.

9. The indoor human body posture detection method based on the fusion of light and radio according to claim 8, characterized in that, In S41, the expression for the global feature vector h is: where T represents the time step, and F fusion represents the latest expression form of the feature representation after cross-modal fusion, t represents the time step number, d represents the feature dimension number, GlobalMaxPool(·) represents the global maximum pooling function, and D represents the feature dimension.

10. The indoor human body posture detection method based on the fusion of light and radio according to claim 8, characterized in that In S42, the calculation formula for the class score z is: z = Wh + b; In the formula, W represents the learnable weight matrix, h represents the global feature vector, and b represents the bias vector; In S44, the calculation formula for the attitude probability vector p is as follows: p = Softmax(W·GlobalMaxPool(F fusion ) + b); Wherein, Softmax(·) represents an activation function, W represents a learnable weight matrix, F fusion represents the latest expression form of the feature representation after cross-modal fusion, b represents a bias vector, and GlobalMaxPool(·) represents a global max pooling function.