User behavior recognition method, electronic equipment, program product and storage medium

By obtaining user video data in RGB, infrared and depth modes, extracting and fusing spatiotemporal features, and generating fusion features using cross-modal attention weights, the problems of high-cost equipment and environmental impact in the prior art are solved, and accurate user behavior recognition under different lighting conditions is achieved.

CN119942634APending Publication Date: 2025-05-06CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411897078.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art faces the problems of high-cost multimodal data acquisition equipment and environmental impacts in driver behavior recognition, and the single-modal recognition accuracy is insufficient, making it difficult to achieve accurate recognition in the case of light changes and lack of three-dimensional information.

Method used

By obtaining user video data from three modalities: RGB, infrared and depth, extracting spatiotemporal features and performing feature fusion, using cross-modal attention weights for weighted fusion, and generating fusion features to identify user behavior.

Benefits of technology

It reduces the cost of the acquisition equipment, improves the accuracy and robustness of the identification, and can effectively identify user behavior under different lighting conditions, making up for the shortcomings of single modal identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942634A_ABST
    Figure CN119942634A_ABST
Patent Text Reader

Abstract

The invention discloses a user behavior recognition method, electronic equipment, a program product and a storage medium. The user behavior identification method comprises the following steps: acquiring user video data in multiple modes, wherein the user video data in multiple modes comprises user video data in an RGB mode, user video data in an infrared mode and user video data in a depth mode; extracting spatio-temporal features from the user video data of each mode; performing feature fusion on the spatio-temporal features of the user video data of the multiple modes to obtain fusion features; and identifying the user behavior based on the fused features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a user behavior recognition method, electronic equipment, program product and storage medium. Background Art

[0002] In the field of human behavior recognition, multiple data modalities such as audio and visual data can be used to organically integrate these different modal data to enhance the recognition effect. Similarly, in the field of driver behavior recognition, multiple modal data are used to achieve accurate driving behavior analysis. These modal data include the driver's physiological characteristics, vehicle driving data, surrounding environment information, and visual features. This method helps to overcome the problem of insufficient accuracy of single modality recognition. However, this method still faces some challenges, such as the high cost of some modal data acquisition equipment and susceptibility to environmental influences. Summary of the invention

[0003] In view of this, embodiments of the present invention provide a user behavior identification method, an electronic device, a program product, and a storage medium.

[0004] The technical solution of the embodiment of the present invention is achieved as follows:

[0005] On the one hand, an embodiment of the present invention provides a method for identifying user behavior, including:

[0006] Acquire user video data in multiple modalities, wherein the user video data in multiple modalities includes user video data in RGB modality, user video data in infrared modality, and user video data in depth modality;

[0007] Extracting spatiotemporal features from user video data of each modality;

[0008] Performing feature fusion on the spatiotemporal features of the user video data of the multiple modalities to obtain fused features;

[0009] The user behavior is identified based on the fusion features.

[0010] In the above solution, after acquiring the user video data in multiple modes, the method further includes:

[0011] Divide the user video data of each modality into multiple video clips;

[0012] Extracting video frames from multiple video clips respectively and combining them into a video frame sequence;

[0013] Correspondingly, the step of extracting spatiotemporal features from user video data of each modality includes:

[0014] Spatiotemporal features are extracted from the video frame sequences of each modality.

[0015] In the above solution, after extracting video frames from the multiple video clips and combining them into a video frame sequence, the method further includes:

[0016] Feature enhancement processing is performed on the infrared modality video frame sequence and the depth modality video frame sequence.

[0017] In the above solution, the feature enhancement processing of the infrared modality video frame sequence and the depth modality video frame sequence includes:

[0018] performing contrast enhancement, Gaussian filtering, and dilation and corrosion processing on the video frame sequence of the infrared modality in sequence;

[0019] The video frame sequence of the depth modality is sequentially subjected to power rate transformation, foreground enhancement, and dilation and erosion processing.

[0020] In the above solution, the step of extracting spatiotemporal features from user video data of each modality includes:

[0021] The spatial and temporal features are extracted from the user video data of each modality via a residual convolutional neural network.

[0022] In the above solution, the feature fusion of the spatiotemporal features of the user video data of the multiple modes to obtain the fusion features includes:

[0023] For any two modal spatiotemporal features, the spatiotemporal features of one modality are used as queries, and the spatiotemporal features of the other modality are used as keys and values ​​to calculate the cross-modal attention weights.

[0024] Based on the cross-modal attention weights, the spatiotemporal features of the corresponding two modalities are weightedly fused to obtain features to be fused; any two modalities among the multiple modalities correspond to one feature to be fused respectively;

[0025] All the features to be fused are concatenated to obtain the fused features.

[0026] In the above solution, identifying the user behavior based on the fusion feature includes:

[0027] The fusion feature is input into a classification model to obtain user behavior information output by the classification model, and the loss function of the classification model is a cross entropy loss function.

[0028] On the other hand, an embodiment of the present application further provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the above-mentioned user behavior identification method.

[0029] On the other hand, an embodiment of the present invention provides an electronic device, including a processor and a memory, which are interconnected, wherein the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute the steps of the user behavior identification method provided by the embodiment of the present invention.

[0030] On the other hand, an embodiment of the present invention provides a computer-readable storage medium, including: the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the user behavior identification method provided in the embodiment of the present invention are implemented.

[0031] The embodiment of the present application obtains user video data of multiple modes, and the user video data of multiple modes includes user video data of RGB mode, user video data of infrared mode and user video data of depth mode. Spatiotemporal features are extracted from the user video data of each mode, and the spatiotemporal features of the user video data of multiple modes are feature fused to obtain fused features, and user behavior is identified based on the fused features. The multimodal data obtained by the embodiment of the present application are all video data, and the required acquisition equipment is single, without the need to use multiple types of acquisition equipment, and the cost of the required acquisition equipment is low. The user video data of infrared mode and depth mode obtained can make up for the problem that the user video data of RGB mode is affected by lighting and lacks three-dimensional information, and can improve the information richness of multimodal fusion features. And by combining multimodal fusion with computer vision technology, accurate identification of user behavior can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic diagram of an implementation flow of a user behavior identification method provided by an embodiment of the present invention;

[0033] Figure 2 is a schematic diagram of segmented sampling provided by an embodiment of the present invention,

[0034] Figure 3 It is a schematic diagram of a processing flow for enhancing a video frame in an infrared mode provided by an embodiment of the present invention;

[0035] Figure 4 is a pixel value histogram of an IR original image provided by an embodiment of the present invention;

[0036] Figure 5 A pixel value histogram after contrast enhancement of an IR original image is provided in an embodiment of the present invention;

[0037] Figure 6 is a schematic diagram of a processing flow for enhancing a video frame in a depth mode provided by an embodiment of the present invention;

[0038] Figure 7 is a schematic diagram of an original depth image provided by an embodiment of the present invention;

[0039] Figure 8 is a schematic diagram of a depth image after power rate transformation provided by an embodiment of the present invention;

[0040] Fig. 9 It is a structural diagram of InceptionV1 provided by an embodiment of the present invention;

[0041] Fig.10 It is a structural diagram of InceptionV2 provided by an embodiment of the present invention;

[0042] Fig.11 is a schematic diagram of a three-dimensional identity mapping residual block structure provided by an embodiment of the present invention;

[0043] Fig.12 It is a schematic diagram of a basic framework of an SR3D network provided by an embodiment of the present invention;

[0044] Fig.13 is a schematic diagram of an SR-Inc network structure provided by an embodiment of the present invention;

[0045] Fig.14 is a schematic diagram of a calculation process of a self-attention mechanism provided by an embodiment of the present invention;

[0046] Fig.15 It is a schematic diagram of a calculation process of a multi-head attention mechanism provided by an embodiment of the present invention;

[0047] Fig.16 is a structural schematic diagram of a multi-modal fusion module based on multi-head attention provided by an embodiment of the present invention;

[0048] Fig.17 is a schematic diagram of a driving behavior recognition algorithm architecture based on multimodal visual fusion provided by an embodiment of the present invention;

[0049] Fig.18 is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0051] The main goal of multimodal fusion methods is to fuse different types of data modalities together to improve the accuracy of behavior recognition. For example, in the field of human behavior recognition, researchers use multiple data modalities such as audio and visual data, and apply technologies such as deep learning to organically fuse information from these different modalities to enhance recognition. Similarly, in the field of driver behavior recognition, multiple modal data are also used, including physiological characteristics, vehicle data, and visual features. Researchers use different methods such as recursive neural networks, deep learning, and attention mechanisms to effectively fuse these data modalities to accurately identify the driver's status.

[0052] In the field of driver behavior recognition, researchers use multiple modal data to achieve accurate driving behavior analysis. These modal data include the driver's physiological characteristics, vehicle driving data, surrounding environment information, and visual features. For example, a related technology proposes a multimodal fusion recursive neural network (Matrix Factorization Recurrent Neural Network, MFRNN), which combines the heart rate collected by the wrist tracker with the modal features such as eye opening and mouth opening collected by the RGBD (RGB+Depth Map) camera, and fully considers the connection between fatigue features in the time dimension. Another related technology designs a distracted driving monitoring method that integrates driver images and sensor data (such as speed, engine speed, Global Positioning System (GPS) data, etc.), and uses a long short-term memory network (Long Short-Term Memory, LSTM)-recurrent neural network (Recurrent Neural Network, RNN) model to process these two modal data. Another related technology proposes a distracted driving behavior recognition method based on multimodal deep learning, introduces federated learning to solve the problem of insufficient data, and realizes the fusion of multiple modal features such as driver images and audio. Another related technology uses an automatically constructed deep neural network to achieve fine-grained detection of various types of driver distraction states by fusing visual information, vehicle dynamics information, and physiological signals.

[0053] At present, the behavior recognition methods based on multimodal fusion mainly focus on fusing the visual modality with data from other modalities to improve the accuracy of behavior recognition. This method helps to overcome the problem of insufficient accuracy of single modality recognition. However, this method still faces some challenges, such as the high cost of some modality data acquisition equipment and susceptibility to environmental influences.

[0054] Compared with the traditional sensor-based acquisition method, the driving behavior recognition method based on a single RGB modality has significant advantages in practicality and recognition cost. However, the input data of this method is RGB video. Although RGB video data has the characteristics of rich information and easy acquisition, it also has some limitations. It is easily disturbed by factors such as lighting, occlusion, and background changes. In addition, RGB images project three-dimensional information onto a two-dimensional plane, so they lack depth information, which may limit accurate spatial perception and behavior analysis in some cases.

[0055] Compared with the RGB modality, the imaging principle of the infrared (IR) modality is based on infrared rays, which enables it to work normally in low-light environments, thus making up for the problem of poor image quality of RGB in poor lighting conditions. However, relying solely on depth modality data for complex human-object interaction behavior recognition has certain challenges. This is because the performance of depth modality data is weak in extracting appearance information such as color and texture, and cannot provide rich detail features, which may have an adverse effect on the performance of the algorithm.

[0056] There is a problem in the process of fusing features of different modalities, that is, the correlation between modalities is ignored, resulting in the failure to fully utilize the complementary information between different modalities. This may cause redundancy in information between different modalities, thus affecting the performance of the algorithm. Commonly used fusion methods, such as simple concatenation or weighting, only directly combine the feature vectors of different modalities without considering the correlation between different modalities.

[0057] In addition, distinguishing between similar actions is also a big challenge. For example, a series of similar action categories, such as drinking water and eating, are very similar in appearance and motion characteristics after removing category-related items. Therefore, it is difficult to accurately classify these similar actions. In this case, only by focusing on the key features within the modality can these similar driving behaviors be effectively distinguished. This may require the use of more advanced feature extraction and classification methods, as well as a deeper understanding of the subtle differences in driving behaviors to obtain more accurate classification results.

[0058] In view of the shortcomings of the above-mentioned related technologies, an embodiment of the present invention provides a user behavior recognition method, which combines multimodal fusion and computer vision technology to solve the difficulty of identifying the driver's distracted driving behavior during vehicle driving, so as to achieve accurate recognition of the driver's distracted driving behavior. In order to illustrate the technical solution of the present invention, a specific embodiment is used for description below.

[0059] refer to Figure 1 , Figure 11 is a schematic diagram of an implementation flow of a user behavior recognition method provided by an embodiment of the present invention. The user behavior recognition method is applied to an electronic device. The user behavior recognition method includes:

[0060] S101, acquiring user video data in multiple modalities, where the user video data in multiple modalities includes user video data in RGB modality, user video data in infrared modality, and user video data in depth modality.

[0061] In one embodiment, video data in three modes, RGB, IR and depth, may be obtained, wherein the user video data in the RGB mode may be acquired through an RGB camera, the user video data in the IR mode may be acquired through an IR camera, and the user video data in the depth mode may be acquired through a depth camera.

[0062] Depth cameras collect depth information mainly through the following technologies:

[0063] Structured light technology: by projecting a specific coded pattern (such as stripes, dots, etc.) onto the surface of an object, the depth of the object is calculated based on the deformed pattern received by the camera.

[0064] Time of flight (ToF) technology: measures the time it takes for light to be emitted and reflected by an object back to the camera, thereby calculating the distance of the object, that is, the depth.

[0065] Binocular stereo vision technology: Use two or more cameras to shoot the same object from different angles, and obtain depth information by calculating the parallax of corresponding points in the image, just like human eyes perceive distance through parallax.

[0066] Among them, user video data of multiple modes may not be required to be collected in the same time period.

[0067] In this embodiment, infrared (IR) and depth (Depth) modalities are used to compensate for the problem that the RGB modality is affected by lighting and lacks three-dimensional information.

[0068] Since the embodiment of the present application only needs to obtain user video data and does not need to collect other modal data such as user physiological data, audio data, vehicle driving data, surrounding environment data, etc., the cost of the acquisition equipment required by the embodiment of the present application is low and the acquisition equipment is easy to deploy.

[0069] S102, extracting spatiotemporal features from user video data of each modality.

[0070] Videos are usually composed of multiple consecutive frames, and the feature information contained in each frame is closely related to the adjacent frames. 3D Convolutional Neural Networks (3DCNN) can be used to extract spatiotemporal features in user video data to provide a more comprehensive description of user behavior in the video.

[0071] For example, the operation formula of 3DCNN is shown in formula (1):

[0072]

[0073] Among them, xyz is the pixel position, represents the spatiotemporal characteristics of the output, f() represents the activation function, and s represents the input image feature information. represents the weight, b ij represents the bias, and (H, W, D) represents the size of the 3D convolution kernel.

[0074] S103, performing feature fusion on the spatiotemporal features of the user video data of the multiple modalities to obtain fused features.

[0075] Modal fusion is the core of multimodal methods, and the effect of fusion determines the performance of behavior recognition.

[0076] In one embodiment, the spatiotemporal features of multiple modalities may be directly concatenated to obtain fused features.

[0077] In another embodiment, a multi-head attention mechanism can be combined to perform modal fusion to obtain fusion features. Cross-modal attention can make full use of complementary information between modalities to achieve feature enhancement within the modality and information interaction between different modalities.

[0078] The process of cross-modal attention can be divided into the following steps:

[0079] 1. Feature extraction: Extract spatiotemporal features of user video data in multiple modalities.

[0080] 2. Attention calculation: When considering the association between the two modalities, calculate cross-modal attention. This can be achieved by calculating the similarity or correlation between the two modalities. A common approach is to use an attention mechanism, where the features of one modality are used as the query, the features of the other modality are used as the key and value, and then the attention weights are calculated. These weights represent the importance of different key modalities given the query modality.

[0081] 3. Feature fusion: Use the calculated attention weights to perform weighted fusion of features from different modalities. The fused features can contain better modality association information.

[0082] S104: Identify user behavior based on the fusion features.

[0083] The fused features can be used for user behavior classification. For example, the fused features are input into a behavior classification model to obtain user behavior information output by the behavior classification model.

[0084] For example, in the field of driver behavior recognition, accurate classification of the driver's driving behavior can be achieved through the fusion of features, and the driver's distracted driving behavior can be accurately identified.

[0085] The embodiment of the present application obtains user video data of multiple modes, and the user video data of multiple modes includes user video data of RGB mode, user video data of infrared mode and user video data of depth mode. Spatiotemporal features are extracted from the user video data of each mode, and the spatiotemporal features of the user video data of multiple modes are feature fused to obtain fused features, and user behavior is identified based on the fused features. The multimodal data obtained by the embodiment of the present application are all video data, and the required acquisition equipment is single, without the need to use multiple types of acquisition equipment, and the cost of the required acquisition equipment is low. The user video data of infrared mode and depth mode obtained can make up for the problem that the user video data of RGB mode is affected by lighting and lacks three-dimensional information, and can improve the information richness of multimodal fusion features. And by combining multimodal fusion with computer vision technology, accurate identification of user behavior can be achieved.

[0086] In one embodiment, after acquiring user video data in multiple modalities, the method further includes:

[0087] Divide the user video data of each modality into multiple video clips;

[0088] Extracting video frames from multiple video clips respectively and combining them into a video frame sequence;

[0089] Correspondingly, the step of extracting spatiotemporal features from user video data of each modality includes:

[0090] Spatiotemporal features are extracted from the video frame sequences of each modality.

[0091] Here, this embodiment adopts a segmented random sampling strategy to sample video frames of three visual modalities, RGB, IR and depth, to reduce redundant information of the video.

[0092] The segmented sampling method combines the advantages of random continuous sampling and fixed uniform sampling, aiming to simultaneously achieve simple long-term time series modeling and avoid video data redundancy and loss of key action information, thereby improving the validity of the data. For example, behaviors during driving can be divided into two categories: uniform regular behaviors and irregular behaviors. Most distracting behaviors are irregular behaviors, such as drinking water, adjusting the radio, etc. The constituent actions of these behaviors are not evenly distributed in the time dimension. For example, the behavior of drinking water includes two parts: picking up a water cup and drinking water. The action of picking up the water cup is relatively short, while the time of maintaining the drinking action is longer. Based on the above situation, this embodiment adopts a segmented random sampling method to perform video frame sampling.

[0093] Segmented random sampling is to evenly segment the entire video sequence, then randomly obtain a video frame in each segment, and finally combine them into the input video frame sequence of the subsequent recognition network. Figure 2 As shown, it is assumed that the entire video sequence includes video frames 1 to 9, which are divided into 3 segments. One video frame is randomly selected from each segment, for example, video frame 2, video frame 4 and video frame 9 are selected respectively, and combined into an input video frame sequence for the subsequent recognition network.

[0094] Specifically, consider a video V as a set of N video frames, namely {Frame1, Frame2…Frame N}, the video is evenly divided into L segments, denoted as {S1, S2, …S i ,…S L}, then from segment S i Randomly select a frame of image F i , forming L frames as the input data of the subsequent recognition network {F1, F2…F L}.

[0095] In one embodiment, after extracting video frames from the plurality of video clips and combining them into a video frame sequence, the method further includes:

[0096] Feature enhancement processing is performed on the infrared modality video frame sequence and the depth modality video frame sequence.

[0097] Infrared modality video data can be collected by a camera equipped with an infrared fill light. This data is insensitive to the lighting environment and can realize behavior recognition under low-light conditions. However, there are some problems with infrared modality data, such as loss of dark details and overexposure of highlight areas. Especially when driving at night, due to insufficient light, the video quality obtained by the infrared camera is easily affected by noise such as light spots. In order to solve these problems, the infrared modality video frames need to be enhanced.

[0098] The pixel value in the depth video frame reflects the distance between the point and the depth camera plane, so it can directly present the geometric shape of the human body. Compared with RGB images, depth video frames are not affected by factors such as ambient lighting and human wear, which makes them very robust. However, in the process of acquiring depth video, it is often interfered by background noise, which reduces the quality of the depth image frame. In order to solve this problem, feature enhancement processing is performed on the depth video frame.

[0099] In one embodiment, the step of performing feature enhancement processing on the infrared modality video frame sequence and the depth modality video frame sequence includes:

[0100] performing contrast enhancement, Gaussian filtering, and dilation and corrosion processing on the video frame sequence of the infrared modality in sequence;

[0101] The video frame sequence of the depth modality is sequentially subjected to power rate transformation, foreground enhancement, and dilation and erosion processing.

[0102] Figure 3 The figure is a schematic diagram of the processing flow of infrared video frame enhancement. First, the contrast of the IR video frame is enhanced to reduce information loss. Then, Gaussian filtering is used to reduce the white noise of the image. Finally, dilation and erosion operations are performed to enhance the features of the video frame.

[0103] Among them, enhancing the contrast of the video frame can enhance the local details of the video frame by increasing the image grayscale value proportionally. The specific calculation is shown in formula (2):

[0104] IR str (x,y)=IR ori (x,y)×a+b (2)

[0105] Where (x, y) represents the pixel position in the video frame, parameter a adjusts the contrast, and parameter b adjusts the brightness. str (x, y) is the gray value after contrast enhancement, IR ori (x, y) is the grayscale value of the original video frame.

[0106] Figure 4 is the pixel value histogram of the IR original image, Figure 5 It is the pixel value histogram after contrast enhancement of the original IR image. After enhancement, the pixel value distribution of the image has changed. Compared with the original image, the dark parts of the image frame are darker, the highlights are brighter, and the pixel values ​​at both ends increase.

[0107] After contrast enhancement, the video frames are filtered to reduce the noise and optimize image details. There is white noise in the video frames of the IR modality, which is similar to Gaussian white noise, so Gaussian filtering method is used to reduce noise.

[0108] For example, the calculation formula of the second-order Gaussian function distribution is shown in formula (3):

[0109]

[0110] Among them, σ is the Gaussian blur variance.

[0111] The erosion operation is used to remove noise from an image, segment connected areas, reduce the size of a target object, etc. The principle of the erosion operation is to traverse each pixel of the image under a given structural element (a structural element is a small grid that defines how to perform the dilation operation) and replace its value with the minimum value of the pixels in the neighborhood around the pixel. The structural element controls the range and shape of the eroded neighborhood. If any pixel in the neighborhood is black (0), the center pixel will also be set to black (0).

[0112] In the erosion operation, while eliminating noise, valuable information is also reduced. Therefore, if we want to increase this valuable information, we can use the expansion operation, which is the inverse operation of erosion.

[0113] Figure 6 It is a schematic diagram of the processing flow of enhancing the video frame of the depth modality, including power rate transformation, foreground enhancement, and corrosion expansion of the depth original image.

[0114] The depth map mainly provides distance information. In order to highlight the driver information, the power rate transformation is first used to stretch the image grayscale histogram. The calculation formula is shown in formula (4):

[0115] R=c×D μ (4)

[0116] Among them, c and μ are adjustment parameters, D is the original depth image showing the initial grayscale value, and R represents the grayscale image after the depth image is converted, with a value in the range of [0, 256]. When c = 1 and μ> 1, grayscale stretching occurs, image details are highlighted, and the image stretching area is mainly concentrated in the important grayscale range. When c = 1 and μ = 1.3, the brightness of the transformed image is significantly improved, the driver area shows a highlight part, and the pixel value distribution of the entire image is mainly concentrated at the high-brightness and high-dark ends. The effect is as follows Figure 7 and Figure 8 shown. Figure 7 is the original depth image, Figure 8 This is the depth image after power rate transformation. It can be seen that the brightness of the image after power rate transformation is significantly improved, and the driver area shows a highlight.

[0117] After the power rate conversion, the driver's outline is clearer, but there is some noise in the background. Foreground enhancement can highlight the foreground area and weaken the background, so the foreground enhancement method is used to suppress the background noise. This process is modeled by equation (5):

[0118]

[0119] Where ij is the pixel convolution kernel coordinate position, R represents the transferred grayscale image produced by the previous process, and m corresponds to the value of the brightest pixel of R. In fact, σ represents a depth frame where the foreground is enhanced and the brightness of the background is reduced. Black pixels given σ remain unchanged, while other pixels are kept updated. The brightness of the background area is significantly minimized, and the foreground area is mainly dark and is significantly highlighted corresponding to small R values.

[0120] In one embodiment, extracting spatiotemporal features from user video data of each modality includes:

[0121] The spatial and temporal features are extracted from the user video data of each modality via a residual convolutional neural network.

[0122] In order to alleviate the gradient vanishing problem that occurs during network deepening, the Inception module can be introduced. The Inception module can not only increase the network width, but also reduce the complexity of network gradient calculation and control the number of parameters in each layer within an appropriate range.

[0123] The detailed structure of InceptionV1 is as follows Fig. 9 As shown in the figure, multi-scale convolution is achieved through convolution layers of different sizes such as 1×1, 3×3, and 5×5, and then the output features of different scales are fused. This structure not only broadens the network width, but also increases the network's applicability to scale, which can effectively improve network performance.

[0124] The detailed structure of InceptionV2 is as follows Fig.10 As shown in the figure, based on InceptionV1, two 3×3 convolution kernels are used to replace the 5×5 convolution kernel to reduce the number of model parameters. The increase in convolution kernels will increase the number of channels in the output feature map, which will increase the number of parameters and calculation requirements of subsequent network layers. To alleviate this problem, a 1×1 convolution layer is added before the large convolution kernel calculation, and the number of channels for subsequent calculations is adjusted through this convolution layer, thereby effectively avoiding complex parameters and excessive calculations.

[0125] In order to solve the network degradation problem that occurs during model training, the present application embodiment introduces a residual structure. The residual structure consists of multiple groups of residual network structures, and the residual network structure consists of multiple residual blocks. The residual block is mainly composed of multiple convolutional layers and shortcut connections. The three-dimensional identity mapping residual block structure is as follows: Fig.11 As shown in the figure. The quick connection can cause the network to generate multiple branches, so that the gradient can be prevented from disappearing during the back propagation process. First, the input x is added to the input x after completing the nonlinear transformation of the convolution kernel of the 3D convolution layer and the Relu activation function, and then passes through the 3D convolution layer. The final output is obtained after the nonlinear transformation of the Relu activation function. The three-dimensional convolution residual block is similar, replacing the identity mapping branch with a 1×1×1 3D convolution operation.

[0126] When the number of network layers of 3DCNN is relatively shallow, the ability to model long time series of video sequences is insufficient. In addition, when the C3D network extracts the spatiotemporal features of video data, the network parameters are too many, resulting in complex calculations. On this basis, InceptionV1-3D was proposed, in which the main module is expanded from the InceptionV1 module to a 3D form, adding the time dimension, which can alleviate the problems of C3D to a certain extent. However, the existence of 3D convolution makes the network model not lightweight enough.

[0127] The embodiment of the present application combines the idea of ​​InceptionV2 module and convolution splitting to improve InceptionV2-3D, which can effectively reduce the number of model parameters. At the same time, on this basis, the residual structure is introduced to alleviate problems such as network degradation. The feature extraction network is called SR3D (Spatiotemporal-separable Residual3D convolutions). The basic framework of the SR3D network is as follows Fig.12 As shown in the figure, the Maxpool layer mainly implements data downsampling and reduces the dimension of features. The 1×1×1 convolution layer is mainly used to realize information interaction between different channels of the feature map and adjust the number of channels of the convolution output feature. The core module is the SR-Inc (Spatiotemporal-separable Residual Inception, SR-Inc) module. The structure of SR-Inc is as follows: Fig.13 As shown in the figure, it consists of Inception-V2+residual structure+convolution splitting, where the residual branch is convolution, and the network parameters are reduced by separating the three-dimensional convolution in time and space. Compared with the Inception-V1 module, the convolution kernel size is set to 3 to avoid the increase in computational requirements caused by large-size convolution.

[0128] In one embodiment, the step of fusing the spatiotemporal features of the user video data of the multiple modalities to obtain fused features includes:

[0129] For any two modal spatiotemporal features, the spatiotemporal features of one modality are used as queries, and the spatiotemporal features of the other modality are used as keys and values ​​to calculate the cross-modal attention weights.

[0130] Based on the cross-modal attention weights, the spatiotemporal features of the corresponding two modalities are weightedly fused to obtain features to be fused; any two modalities among the multiple modalities correspond to one feature to be fused respectively;

[0131] All the features to be fused are concatenated to obtain the fused features.

[0132] This embodiment implements multimodal fusion based on a multi-head attention mechanism. The attention function can be regarded as an operation that maps a set of queries and key-value pairs to outputs. The input is linearly transformed to obtain the query matrix Q, the key matrix K, and the value matrix V. The output is the weighted sum of the values. The scaled dot product attention used by the self-attention mechanism is faster and more space-saving in actual use, and alleviates the problem of gradient disappearance. The calculation process is as follows: Fig.14 As shown, the matmul function is a matrix multiplication function, Scale is a scaling layer, Msak(opt.) is a mask ratio score, and the SoftMax function is used to convert the correlation score into a probability distribution.

[0133] The calculation formula is shown in formula (6):

[0134]

[0135] Among them, d k is the number of columns of the query matrix and the key matrix, that is, the vector dimension. Attention(Q, K, V) represents the weighted output value of the query matrix Q, the key matrix K and the value matrix V. Softmax(*) is the processing function of the query matrix Q and the key matrix K. T is the transposed matrix symbol.

[0136] The multi-head attention mechanism can achieve attention to different internal information while keeping the total number of parameters unchanged, further improving the expressiveness of features. The multi-head attention mechanism is mainly implemented by splitting the query, key and value parameters multiple times. In order to better focus on different parts of the input, each group of split parameters is mapped to different subspaces of the high-dimensional space to calculate the corresponding attention weights. The calculation process of the multi-head attention mechanism is as follows: Fig.15 As shown in Figure 1, ScaledDot-ProductAttention is a scaled dot product attention, which is used to perform dot product operations on the query matrix Q and the key matrix K. Through multiple parallel calculations, the attention information in all subspaces is finally merged, as shown in equations (7) and (8).

[0137] MultiHead(Q,K,V)=Concat (h1,h2,...,h i )W 0 (7)

[0138] ht =Attention(QW t Q ,KW t K ,VW t V ) (8)

[0139] Where t represents the long index, W 0 , W t Q , W t K , W t V are the parameter matrices of Q, K, and V during linear mapping, Concat represents the concatenation operation, and h represents the number of heads in the multi-head attention, that is, the number of data splits. The attention calculated in different subspaces is different. Multi-head attention can fully explore the connection between inputs from different angles, so as to pay attention to the association and subtle differences between data.

[0140] Modal fusion is the core of multimodal methods, and the fusion effect determines the performance of behavior recognition. For example, when different modal visual features are extracted from the network, the modal features obtained are RGB, IR and Depth categories. The three modal features (spatiotemporal features) are used as the input of the multimodal fusion network. The structure of the multimodal fusion module based on multi-head attention is as follows: Fig.16 As shown in Figure 2, this module uses a two-layer multi-head attention mechanism to achieve feature enhancement within the modality and information interaction between different modalities. When the modal features are encoded through linear mapping, the input Z is obtained. i The eigenvector Z i The data is normalized through the LayerNorm layer, and then processed using the multi-head attention mechanism MSA to enhance the key features within the modality. At this time, feature y is obtained i The calculation formula is shown in formula (9):

[0141] y i =MSA(LN(Z i ))+Z i (9)

[0142] Where LN represents LayerNorm processing, MSA represents multi-head attention mechanism, and Z i Represents the modal characteristics after linear mapping of the input, y i Represents the processed output features, i∈[RGB, IR, Depth].

[0143] At this time, the characteristics y of three modes of self-enhancement can be obtained RGB ,y IR ,y Depth, and then perform interactive fusion between modalities. The information of adjacent modalities is fused in order by calculating cross-modal attention (CMA), where the query, key, and value are adjacent modal features, which are first linearly projected to the same dimension. The calculation formulas are shown in Equations (10) and (11).

[0144] F i =CMA(y i ,W proj y i+1 ) (10)

[0145]

[0146] Among them, W Q , W K , W V is the projection matrix of query, key and value during attention calculation, F represents the output feature after cross-modal fusion, and d k is the number of columns of the query matrix and the key matrix, CMA stands for cross-modal attention. i represents the feature, and i+1 represents another adjacent self-enhanced feature (y RGB ,y IR ,y Depth W porj Represents the projection matrix (or weight matrix), which is used to transform the modal y i+1 Projected to the dimensions required for attention calculation. x and y represent input sequences of different modalities, x is the query information, y contains the key and value information of the element, T represents the transposed matrix symbol, and the softmax function is used to convert the relevance score into a probability distribution.

[0147] After all the features are fused and processed by the LayerNorm layer and the MLP layer, the three feature vectors not only contain the unique information of another modality, but also enhance their own key features. Finally, the three processed features are directly Cancat operated to obtain the fused features for classification.

[0148] In one embodiment, identifying the user behavior based on the fusion feature includes:

[0149] The fusion feature is input into a classification model to obtain user behavior information output by the classification model, and the loss function of the classification model is a cross entropy loss function.

[0150] In response to the imbalance problem existing in samples, the embodiment of the present application introduces Focal Loss as the loss function during model training. Focal Loss is mainly used to address the problem of category imbalance that occurs when the accuracy of single-stage target detection is lower than that of two-stage target detection. Focal Loss is also applicable to classification problems. The Focal Loss function is improved on the basis of the standard cross entropy loss function. It adjusts the weight of the sample according to the classification difficulty of the sample, increases the impact of the difficult sample on the loss function, and makes the model pay more attention to the difficult samples.

[0151] The standard cross entropy formula is shown in formula (12).

[0152]

[0153] Among them, x i represents the i-th sample of the input, P represents the true probability distribution, Q represents the predicted probability distribution, and CE(,Q) represents the cross entropy. In the actual calculation of the cross entropy, only one probability value is used. Reducing the weight of samples that are easy to classify makes the model pay more attention to samples that are difficult to classify during training. The cross entropy loss function is modulated by adding a modulation factor (1-p t ) γ To achieve this, the adjustable focusing parameter γ ≥ 0. This loss function is the focal loss, and its calculation is shown in formula (13).

[0154] FL(p t )=-(1-p t ) γ log(p t ) (13)

[0155] Among them, FL(p t ) represents the loss function, p t Refers to the similarity between the output of the training sample and category y. The larger the similarity, the more accurate the classification. Category y refers to the recognition category, such as distraction behaviors such as making phone calls. The hyperparameter γ determines the degree of loss attenuation. Adding a modulation factor (1-p t ) γ To achieve the weights of easy-to-classify and difficult-to-classify samples. During driving, the driver's behavior in most cases complies with relevant laws and regulations, while the proportion of distracted driving behavior in the driving process is uneven, which will lead to the problem of unbalanced data samples. In the collected data samples, normal driving accounts for a high proportion, and there are fewer data on some abnormal behaviors. In response to the imbalance problem in the samples, the embodiment of the present application introduces Focal Loss as the loss function during model training to alleviate the problem of sample imbalance. Accurate classification of driver distracted driving behavior is achieved.

[0156] In one embodiment, the algorithm architecture of driver behavior recognition based on multimodal visual fusion is as follows: Fig.17 As shown in the figure, the input of the network is video sequences in three modes: RGB, IR and Depth. The main structure is divided into three parts: modal feature extraction, multimodal fusion and classifier.

[0157] In the modal feature extraction stage, we first use a segmented random sampling strategy to sample the video frames of the three visual modalities of RGB, IR and depth to reduce the redundant information of the video, and then enhance the video frames for the IR and Depth modalities to reduce image noise and improve data quality. In the modal feature extraction module, we design a lightweight 3D convolutional neural network as the backbone feature extraction network to extract spatiotemporal features from each modal video data. Depth , Feature RGB , Feature IR ). In the multimodal fusion module, a multi-head attention mechanism is introduced to enhance the internal features of the modality and focus on important features. At the same time, interactive fusion of features is achieved between different modalities to extract complementary information. The extracted features are concat-joined to obtain fused features. Classification and recognition are performed based on the fused features to ultimately achieve driving behavior recognition. Finally, in order to address the problem of unbalanced sample data, Focal Loss is used as the loss function of the classifier, which can better handle the imbalance of categories in the data, allowing the model to pay more attention to minority categories, thereby improving the performance of user behavior classification.

[0158] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0159] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0160] It should be noted that the technical solutions described in the embodiments of the present invention can be arbitrarily combined without conflict.

[0161] In addition, in the embodiments of the present invention, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0162] The present application also provides a user behavior identification device, which corresponds to the user behavior identification method of the receiving end, and each step in the above user behavior identification method embodiment is also fully applicable to the present device embodiment. The device includes:

[0163] An acquisition module, used to acquire user video data in multiple modes, wherein the user video data in multiple modes includes user video data in RGB mode, user video data in infrared mode, and user video data in depth mode;

[0164] An extraction module for extracting spatiotemporal features from user video data of each modality;

[0165] A fusion module, used for fusing the spatiotemporal features of the user video data of the multiple modes to obtain fusion features;

[0166] The recognition module is used to recognize user behavior based on the fusion feature.

[0167] In one embodiment, the device further comprises:

[0168] A division module, used for dividing the user video data of each modality into multiple video segments;

[0169] An extraction module, used for extracting video frames from multiple video clips respectively and combining them into a video frame sequence;

[0170] Correspondingly, the extraction module is used for:

[0171] Spatiotemporal features are extracted from the video frame sequences of each modality.

[0172] In one embodiment, the device further comprises:

[0173] The enhancement module is used to perform feature enhancement processing on the video frame sequence of the infrared modality and the video frame sequence of the depth modality.

[0174] In one embodiment, the enhancement module is specifically used for:

[0175] performing contrast enhancement, Gaussian filtering, and dilation and corrosion processing on the video frame sequence of the infrared modality in sequence;

[0176] The video frame sequence of the depth modality is sequentially subjected to power rate transformation, foreground enhancement, and dilation and erosion processing.

[0177] In one embodiment, the extraction module is specifically used for:

[0178] The spatial and temporal features are extracted from the user video data of each modality via a residual convolutional neural network.

[0179] In one embodiment, the fusion module is specifically used for:

[0180] For any two modal spatiotemporal features, the spatiotemporal features of one modality are used as queries, and the spatiotemporal features of the other modality are used as keys and values ​​to calculate the cross-modal attention weights.

[0181] Based on the cross-modal attention weights, the spatiotemporal features of the corresponding two modalities are weightedly fused to obtain features to be fused; any two modalities among the multiple modalities correspond to one feature to be fused respectively;

[0182] All the features to be fused are concatenated to obtain the fused features.

[0183] In one embodiment, the identification module is specifically used for:

[0184] The fusion feature is input into a classification model to obtain user behavior information output by the classification model, and the loss function of the classification model is a cross entropy loss function.

[0185] In actual application, the acquisition module, extraction module, fusion module and recognition module can be implemented by a processor in an electronic device, such as a central processing unit (CPU), a digital signal processor (DSP), a microcontroller unit (MCU) or a programmable gate array (FPGA).

[0186] It should be noted that: the user behavior identification device provided in the above embodiment only uses the division of the above modules as an example when performing user behavior identification. In actual applications, the above processing can be assigned to different modules as needed, that is, the internal structure of the device can be divided into different modules to complete all or part of the processing described above. In addition, the user behavior identification device provided in the above embodiment and the user behavior identification method embodiment belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0187] The user behavior recognition device can be in the form of an image file. After the image file is executed, it can be run in the form of a container or a virtual machine to implement the user behavior recognition method described in this application. Of course, it is not limited to the image file form. As long as some software forms that can implement the user behavior recognition method described in this application are within the scope of protection of this application.

[0188] Based on the hardware implementation of the above program modules and in order to implement the method of the embodiment of the present application, the embodiment of the present application also provides an electronic device. Fig.18 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown in FIG. Fig.18 As shown, the electronic device includes:

[0189] Communication interface 1801, capable of exchanging information with other devices such as network devices;

[0190] The processor 1802 is connected to the communication interface 1801 to implement information exchange with other devices and is used to execute the method provided by one or more technical solutions when running a computer program. The computer program is stored in the memory 1803.

[0191] Of course, in actual application, the various components in the electronic device are coupled together through the bus system 1804. It can be understood that the bus system 1804 is used to realize the connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig.18 Various buses are labeled as bus system 1804.

[0192] The memory 1803 in the embodiment of the present application is used to store various types of data to support the operation of the computer device. Examples of such data include: any computer program used to operate on the electronic device.

[0193] It can be understood that the memory 1803 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disk, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), direct memory bus random access memory (DRRAM). The memory described in the embodiments of the present application is intended to include but is not limited to these and any other suitable types of memory.

[0194] The method disclosed in the above embodiment of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the program in the memory and completes the steps of the above method in combination with its hardware.

[0195] Optionally, when the processor 1802 executes the program, it implements the corresponding processes implemented by the electronic device in each method of the embodiments of the present application, which will not be described again for the sake of brevity.

[0196] In an exemplary embodiment, the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a first memory storing a computer program, and the computer program can be executed by a processor of a computer device to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.

[0197] In the several embodiments provided in the present application, it should be understood that the disclosed devices, computer equipment and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0198] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0199] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0200] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, disks or optical disks.

[0201] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, electronic device, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0202] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program, which can be executed by the processor 1802 of the electronic device to complete the steps described in the user behavior identification method in the embodiment of the present application.

[0203] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0204] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.

[0205] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A behavior recognition method, characterized in that: The method comprises: Acquire user video data in multiple modalities, wherein the user video data in multiple modalities includes user video data in RGB modality, user video data in infrared modality, and user video data in depth modality; Extracting spatiotemporal features from user video data of each modality; Performing feature fusion on the spatiotemporal features of the user video data of the multiple modalities to obtain fused features; The user behavior is identified based on the fusion features.

2. The method according to claim 1, characterized in that After acquiring the user video data in multiple modes, the method further includes: Divide the user video data of each modality into multiple video clips; Extracting video frames from multiple video clips respectively and combining them into a video frame sequence; Correspondingly, the step of extracting spatiotemporal features from user video data of each modality includes: Spatiotemporal features are extracted from the video frame sequences of each modality.

3. The method according to claim 2, characterized in that After extracting video frames from the plurality of video clips and combining them into a video frame sequence, the method further comprises: Feature enhancement processing is performed on the infrared modality video frame sequence and the depth modality video frame sequence.

4. The method according to claim 3, characterized in that The feature enhancement processing of the infrared modality video frame sequence and the depth modality video frame sequence includes: performing contrast enhancement, Gaussian filtering, and dilation and corrosion processing on the video frame sequence of the infrared modality in sequence; The video frame sequence of the depth modality is sequentially subjected to power rate transformation, foreground enhancement, and dilation and erosion processing.

5. The method according to claim 1, characterized in that The step of extracting spatiotemporal features from user video data of each modality includes: The spatial and temporal features are extracted from the user video data of each modality via a residual convolutional neural network.

6. The method according to claim 1, characterized in that The step of fusing the spatiotemporal features of the user video data of the multiple modes to obtain fused features includes: For any two modal spatiotemporal features, the spatiotemporal features of one modality are used as queries, and the spatiotemporal features of the other modality are used as keys and values ​​to calculate the cross-modal attention weights. Based on the cross-modal attention weights, the spatiotemporal features of the corresponding two modalities are weightedly fused to obtain features to be fused; any two modalities among the multiple modalities correspond to one feature to be fused respectively; All the features to be fused are concatenated to obtain the fused features.

7. The method according to claim 1, characterized in that The identifying the user behavior based on the fusion feature includes: The fusion feature is input into a classification model to obtain user behavior information output by the classification model, and the loss function of the classification model is a cross entropy loss function.

8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the user behavior identification method according to any one of claims 1 to 7 are implemented.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the user behavior identification method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the steps of the user behavior identification method according to any one of claims 1 to 7.