A human behavior recognition method and device

By performing multi-sampling rate processing on the video, multiple sub-videos are obtained and their spatial feature maps are fused, and the global semantic features are combined, the problem of inaccurate recognition caused by user action rates in different age groups is solved, and the accuracy of human behavior recognition is improved.

CN114332693BActive Publication Date: 2025-07-18SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111541848.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-07-18
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The existing human behavior recognition model has different movement rates of users of different age groups, resulting in inaccurate recognition results, which is particularly difficult to capture the human behavior of the elderly.

Method used

The video to be processed is sampled through multiple different sampling rates, and multiple sub-videos are obtained. The number of video frames of each sub-video is different. The spatial feature map is extracted and fusion processed, and the human behavior type is determined based on global semantic features.

Benefits of technology

It improves the accuracy of human behavior type recognition and avoids identification errors caused by different movement rates of users of different age groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332693B_ABST
    Figure CN114332693B_ABST
Patent Text Reader

Abstract

The present disclosure provides a human behavior recognition method and apparatus. The method samples a video to be processed at multiple different sampling rates to obtain a plurality of sub-videos to be processed, wherein the number of video frames in each sub-video to be processed is different, that is, the action rate in each sub-video to be processed is different. In this way, in this embodiment, the human behavior type can be determined according to the spatial feature maps of the sub-videos to be processed with different action rates, that is, the factor of different action rates is considered in the process of human behavior category recognition, avoiding the problem of incorrect human behavior recognition caused by different action rates of users of different ages, thereby improving the accuracy of human behavior type recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method and apparatus for human behavior recognition. Background Art

[0002] Human behavior recognition based on video information is a hot issue in the field of computer vision. It mainly uses technologies such as image processing, image analysis, and computer vision to perform object detection, classification, and tracking on video sequences, and to understand and describe the behaviors in the video information.

[0003] Existing human behavior recognition models usually recognize human behaviors according to the key points (such as joint points) of the user's body parts. However, since the action rates of users of different ages are different, for example, the action rate of the elderly is slower compared to that of young people. If only recognizing human behaviors based on the key points of the user's body parts, it may lead to inaccurate human behavior recognition results, and thus it is difficult to effectively capture the human behaviors of the elderly. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a method, apparatus, computer device, and computer-readable storage medium for human behavior recognition to solve the problem of inaccurate human behavior recognition results in the prior art.

[0005] In a first aspect of embodiments of the present disclosure, a method for human behavior recognition is provided. The method includes:

[0006] Obtaining a plurality of to-be-processed sub-videos corresponding to a to-be-processed video and the global semantic feature of the to-be-processed video, where the number of video frames of each to-be-processed sub-video is different;

[0007] For each to-be-processed sub-video, extracting a spatial feature map of the to-be-processed sub-video;

[0008] Performing fusion processing on the spatial feature maps respectively corresponding to each to-be-processed sub-video to obtain a fusion feature;

[0009] Determining the human behavior type corresponding to the to-be-processed video according to the fusion feature and the global semantic feature.

[0010] In a second aspect of embodiments of the present disclosure, a human behavior recognition apparatus is provided. The apparatus includes:

[0011] An information acquisition module, configured to obtain a plurality of to-be-processed sub-videos corresponding to a to-be-processed video and the global semantic feature of the to-be-processed video, where the number of video frames of each to-be-processed sub-video is different;

[0012] A feature map extraction module, configured to extract the spatial feature map of each sub-video to be processed for each sub-video to be processed;

[0013] A feature fusion module, configured to perform fusion processing on the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fusion feature;

[0014] A type determination module, configured to determine the human behavior type corresponding to the video to be processed according to the fusion feature and the global semantic feature.

[0015] In a third aspect of the embodiments of the present disclosure, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0016] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0017] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: The embodiments of the present disclosure can first obtain a plurality of sub-videos to be processed corresponding to the video to be processed and the global semantic feature of the video to be processed; then, for each sub-video to be processed, the spatial feature map of the sub-video to be processed can be extracted; next, the spatial feature maps respectively corresponding to each sub-video to be processed can be subjected to fusion processing to obtain a fusion feature; finally, according to the fusion feature and the global semantic feature, the human behavior type corresponding to the video to be processed can be determined. Since in this embodiment, the video to be processed is sampled at multiple different sampling rates to obtain a plurality of sub-videos to be processed, where the number of video frames of each sub-video to be processed is different, that is, the action rate in each sub-video to be processed is different. In this way, this embodiment can determine the human behavior type according to the spatial feature maps of the sub-videos to be processed with different action rates, that is, the factor of different action rates is considered in the process of human behavior category recognition, avoiding the problem of incorrect human behavior recognition caused by the different action rates of users of different ages, thereby improving the accuracy of human behavior type recognition. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0019] Figure 1 It is a schematic diagram of the application scenario of the embodiments of the present disclosure;

[0020] Figure 2 It is a flowchart of the human behavior recognition method provided by the embodiments of the present disclosure;

[0021] Figure 3 It is a schematic diagram of the network architecture of the human behavior recognition model provided by the embodiments of the present disclosure;

[0022] Figure 4 It is a block diagram of the human behavior recognition device provided by the embodiments of the present disclosure;

[0023] Figure 5 It is a schematic diagram of the computer device provided by the embodiments of the present disclosure. Detailed implementation manners

[0024] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures and technologies are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art should understand that the present disclosure may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from obscuring the description of the present disclosure.

[0025] A human behavior recognition method and device according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0026] In the prior art, since existing human behavior recognition models usually recognize human behaviors based on key points (such as joint points) of the user's human body parts, but since the action rates of users of different age groups are different, for example, the action rate of the elderly is slower compared to that of young people. If only recognizing human behaviors based on the key points of the user's human body parts, it may lead to inaccurate human behavior recognition results, making it difficult to effectively capture the human behaviors of the elderly.

[0027] To solve the above problems, the present invention provides a human behavior recognition method. In this method, since the present embodiment samples the video to be processed at multiple different sampling rates to obtain multiple sub-videos to be processed, where the number of video frames in each sub-video to be processed is different, that is to say, the action rate in each sub-video to be processed is different. In this way, the present embodiment can determine the human behavior type according to the spatial feature maps of the sub-videos to be processed with different action rates, that is, the factors of different action rates are considered in the process of human behavior category recognition, avoiding the problem of incorrect human behavior recognition caused by the different action rates of users of different ages, thereby improving the accuracy of human behavior type recognition.

[0028] For example, the embodiments of the present invention can be applied to an application scenario as Figure 1 shown. In this scenario, it may include a terminal device 1 and a server 2.

[0029] The terminal device 1 can be hardware or software. When the terminal device 1 is hardware, it can be various electronic devices with functions of collecting images, storing images and supporting communication with the server 2, including but not limited to smart phones, tablet computers, laptop portable computers, digital cameras, monitors, video recorders and desktop computers, etc.; when the terminal device 1 is software, it can be installed in the above-mentioned electronic devices. The terminal device 1 can be implemented as multiple software or software modules, or can be implemented as a single software or software module, and the embodiments of the present disclosure do not limit this. Further, various applications can be installed on the terminal device 1, such as an image acquisition application, an image storage application, an instant messaging application, etc.

[0030] The server 2 can be a server that provides various services. For example, it can be a background server that receives requests sent by the terminal device 1 with which it establishes a communication connection. The background server can receive and analyze and process the requests sent by the terminal device 1 and generate a processing result. The server 2 can be a single server, or can be a server cluster composed of several servers, or can also be a cloud computing service center, and the embodiments of the present disclosure do not limit this.

[0031] It should be noted that the server 2 can be hardware or software. When the server 2 is hardware, it can be various electronic devices that provide various services for the terminal device 1. When the server 2 is software, it can be multiple software or software modules that provide various services for the terminal device 1, or can be a single software or software module that provides various services for the terminal device 1, and the embodiments of the present disclosure do not limit this.

[0032] The terminal device 1 and the server 2 can be communicatively connected via a network. The network can be a wired network connected by coaxial cables, twisted pairs, and optical fibers, or a wireless network that enables interconnection of various communication devices without wiring. For example, Bluetooth, Near Field Communication (NFC), Infrared, etc. The embodiments of the present disclosure do not limit this.

[0033] Specifically, the user can determine the video to be processed through the terminal device 1. For example, the terminal device 1 can collect the video to be processed, or select a video to be processed from multiple videos through the terminal device 1. Moreover, the terminal device 1 can send the video to be processed to the server 2. Then, after receiving the video to be processed, the server 2 can first obtain multiple sub-videos to be processed corresponding to the video to be processed and the global semantic features of the video to be processed. Next, for each sub-video to be processed, the server 2 can extract the spatial feature map of the sub-video to be processed. Subsequently, the server 2 can perform a fusion process on the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fusion feature. Finally, the server 2 can determine the human behavior type corresponding to the video to be processed based on the fusion feature and the global semantic feature, and return the human behavior type corresponding to the video to be processed to the terminal device 1. In this way, since the server 2 in this embodiment samples the video to be processed at multiple different sampling rates to obtain multiple sub-videos to be processed, where the number of video frames in each sub-video to be processed is different, that is, the action rate in each sub-video to be processed is different. In this way, the server 2 in this embodiment can determine the human behavior type based on the spatial feature maps of the sub-videos to be processed with different action rates. That is to say, the server 2 considers the factor of different action rates during the human behavior category recognition process, avoiding the problem of incorrect human behavior recognition caused by different action rates of users of different ages, and thus can improve the accuracy of human behavior type recognition.

[0034] It should be noted that the specific types, quantities, and combinations of the terminal device 1, the server 2, and the network can be adjusted according to the actual requirements of the application scenario. The embodiments of the present disclosure do not limit this.

[0035] It should be noted that the above application scenarios are only shown for the convenience of understanding the present disclosure, and the embodiments of the present disclosure are not limited in this regard. On the contrary, the embodiments of the present disclosure can be applied to any applicable scenario.

[0036] Figure 2 is a flowchart of a human behavior recognition method provided by an embodiment of the present disclosure. Figure 2 A human behavior recognition method can be performed byFigure 1 executed by the terminal device or server. As Figure 2 shown, the human behavior recognition method includes:

[0037] S201: Obtain a plurality of sub-videos to be processed corresponding to the video to be processed and the global semantic features of the video to be processed.

[0038] In this embodiment, the video to be processed can be understood as a video that needs to perform human behavior recognition. It can be understood that the video to be processed includes several video frames recording user actions. As an example, the video to be processed can be collected by a monitoring camera installed at a fixed position, can also be collected by a mobile terminal device, or can also be read from a storage device that pre-stores videos.

[0039] To avoid the problem of incorrect human behavior recognition caused by different action rates of different users, in this embodiment, after obtaining the video to be processed, the sub-videos to be processed at different action rates of the video to be processed can be obtained first, so that the factors of action rate change can be captured in the subsequent human behavior recognition process. It should be noted that the smaller the sampling rate of the sub-video to be processed, the fewer the number of video frames of the sub-video to be processed, and correspondingly, the faster the action rate in the sub-video to be processed. On the contrary, the larger the sampling rate of the sub-video to be processed, the more the number of video frames of the sub-video to be processed, and correspondingly, the slower the action rate in the sub-video to be processed. In one implementation, multiple preset sampling rates can be obtained, where each sampling rate is different; then, for each sampling rate, the video to be processed can be sampled to obtain the sub-video to be processed corresponding to the sampling rate. It can be understood that the number of video frames of each sub-video to be processed is different, that is, the action rate in each sub-video to be processed is different.

[0040] As an example, the video frame to be processed can be intercepted from the original video. For example, one video frame to be processed is intercepted from the original video every 10 seconds, and the video to be processed can be a video of 16 consecutive video frames. To consider the influence of the action speed, three sampling rates α / σ 2 、α / σ、α can be used to sample the video to be processed to obtain sub-videos to be processed with three time lengths; for example, the video to be processed can be a video of 16 consecutive video frames, and the sampling rate α can be the original sampling rate of the video frame to be processed, and the coefficient σ = 2. In this way, a sub-video to be processed with 16 frames, a sub-video to be processed with 8 frames, and a sub-video to be processed with 4 frames can be obtained.

[0041] In this embodiment, in order to be able to pay attention to the information features in video frames with relatively high importance, after obtaining the video to be processed in this embodiment, the global semantic features of the video to be processed can be extracted. Among them, the global semantic features of the video to be processed can be understood as the features that can reflect the semantic information of each video frame in the video to be processed. For example, they can reflect the appearance feature information of the user in the video. In this way, the information features of video frames with relatively high importance can be focused on during the subsequent human behavior recognition process by using the global semantic features of the video to be processed, thereby improving the recognition accuracy of human behavior categories.

[0042] S202: For each sub-video to be processed, extract the spatial feature map of the sub-video to be processed.

[0043] After obtaining multiple sub-videos to be processed of the video to be processed, for each sub-video to be processed, the semantic features and temporal-spatial features of each video frame in the sub-video to be processed can be extracted to obtain the spatial feature map of the sub-video to be processed. It can be understood that the spatial feature map of the sub-video to be processed can include the semantic features and temporal-spatial features of each video frame in the sub-video to be processed. Since the semantic features of each video frame can reflect the image content information in the video frame, for example, which key points of the user are included in the video frame and what the limb movements of the user are, and the temporal-spatial features of each video frame can reflect the sequential relationship between each video frame and the position information of each part in the video frame (such as the spatial position of the user's key points in the video frame and other spatial position information), in this way, the temporal information of the video frame can be enhanced by using the temporal-spatial features. Therefore, the spatial feature map of the sub-video to be processed can reflect the feature information of the human behavior type at the action rate corresponding to the sub-video to be processed. That is to say, the spatial feature map of the sub-video to be processed can be used to analyze the human behavior type of the video to be processed at the action rate corresponding to the sub-video to be processed.

[0044] S203: Perform a fusion process on the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fusion feature.

[0045] Since the action rates corresponding to each sub-video to be processed are different, and the spatial feature maps respectively corresponding to each sub-video to be processed can reflect the human behavior types of the video to be processed at the action rate corresponding to the sub-video to be processed. Therefore, in order to avoid the problem of incorrect human behavior recognition caused by the different action rates of users of different ages and exclude the interference of the movement rate on the recognition result of the human behavior type; in this embodiment, after obtaining the spatial feature maps respectively corresponding to each sub-video to be processed, the spatial feature maps respectively corresponding to each sub-video to be processed can be fused to obtain a fused feature; it can be understood that the fused feature includes the feature information of the human behavior types of the video to be processed at various movement rates.

[0046] Specifically, in one implementation manner, the method of fusing the spatial feature maps respectively corresponding to each sub-video to be processed can be one of the following methods: feature splicing processing method, feature superposition processing method, convolution processing method. Specifically, the splicing processing method can be: splicing the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a spliced spatial feature map, and using the spliced spatial feature map as the fused feature; the feature superposition processing method can be: superimposing the feature values of each pixel point in the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a spliced spatial feature map, and using the spliced spatial feature map as the fused feature; the convolution processing method can be using a convolution kernel to fuse the features of the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a spliced spatial feature map, and using the spliced spatial feature map as the fused feature.

[0047] S204: Determine the human behavior type corresponding to the video to be processed according to the fused feature and the global semantic feature.

[0048] Among them, the human behavior type can be understood as the type of behavior performed by the user. For example, the human behavior type can include walking, dancing, falling, eating, lying flat, etc.

[0049] After obtaining the fusion features and global semantic features of the video to be processed, since the global semantic features of the video to be processed can reflect the information features of video frames with relatively high importance, and can also reflect the appearance feature information of the user in the video and the semantic information of the overall picture content of some video frames; therefore, in the process of using the fusion features to identify the human behavior type corresponding to the video to be processed, the global semantic features of the video to be processed can also be combined to determine the human behavior type corresponding to the video to be processed. In this way, this embodiment can determine the human behavior type corresponding to the video to be processed according to the feature information of the human behavior type of the video to be processed at various motion speeds and the semantic information of the overall picture content of the video to be processed, so that the recognition result of the human behavior type corresponding to the video to be processed can be more accurate, that is, the recognition accuracy of the human behavior type corresponding to the video to be processed is improved.

[0050] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: The embodiments of the present disclosure can first obtain a plurality of sub-videos to be processed corresponding to the video to be processed and the global semantic features of the video to be processed; then, for each sub-video to be processed, a spatial feature map of the sub-video to be processed can be extracted; next, the spatial feature maps respectively corresponding to the respective sub-videos to be processed can be fused to obtain fusion features; finally, according to the fusion features and the global semantic features, the human behavior type corresponding to the video to be processed can be determined. Since this embodiment samples the video to be processed at multiple different sampling rates to obtain a plurality of sub-videos to be processed, where the number of video frames of each sub-video to be processed is different, that is to say, the action speed in each sub-video to be processed is different. In this way, this embodiment can determine the human behavior type according to the spatial feature maps of the sub-videos to be processed with different action speeds, that is, the factor of different action speeds is considered in the process of identifying human behavior categories, avoiding the problem of incorrect human behavior recognition caused by the different action speeds of users of different ages, thereby improving the accuracy of human behavior type recognition.

[0051] Next, an implementation manner of "obtaining the global semantic features of the video to be processed" in S201 will be introduced, that is, how to obtain the global semantic features of the video to be processed. In this embodiment, "obtaining the global semantic features of the video to be processed" in S201 may include the following steps:

[0052] S201a: Input the video to be processed into a high-resolution network to obtain multiple feature maps of different sizes.

[0053] Since the installation position of the camera for collecting video information is generally fixed under normal circumstances, while the distance between the user in the video information and the camera is not fixed, the sizes of the video image information of the user collected in this way are inconsistent, which will lead to inconsistent sizes of the features extracted based on the video information. And if pictures of the same scale are used, it will affect the recognition result. Therefore, in this embodiment, in order to ensure that the resolution of the feature map extracted for the video to be processed can be kept higher than the preset threshold when the distance between the user in the video information and the camera is not fixed, multiple feature maps of different sizes of the video to be processed can be extracted first by using a high-resolution network, that is, input the video to be processed into the high-resolution network to obtain multiple feature maps of different sizes.

[0054] In one implementation, as Figure 3 shown, the high-resolution network can be HighResolution Net (i.e., HRNet), and HRNet can maintain a high-resolution representation throughout the process. Starting with a high-resolution subnet as the first stage, high-to-low-resolution subnets are added one by one to form more stages, and multi-resolution subnets are connected in parallel. Information in the parallel multi-resolution subnets is repeatedly exchanged throughout the process for repeated multi-scale fusion. Specifically, the feature map of each picture can be obtained through the parallel multi-branch high-resolution network HRNet. The specific process is to downsample the high-resolution feature map to a low-resolution feature map and then restore it from the low-resolution feature map to a high-resolution feature map, so as to realize the process of multi-scale feature extraction. This process is repeated multiple times until the preset number of times is met. HRNet maintains the high resolution of the feature map throughout the process. In order to extract multi-scale features, the model gradually adds low-resolution feature map subnets in parallel to the main network of the high-resolution feature map, and different networks realize multi-scale fusion and feature extraction. In one implementation, the network can be constructed using 4 layers of features.

[0055] S201b: Input the multiple feature maps of different sizes into the global semantic feature extraction network to obtain the global semantic features of the video to be processed.

[0056] In order to focus on important picture frames, in this embodiment, a global semantic feature extraction network is set up, such as Figure 3 shown in the human behavior recognition model. The global semantic feature extraction network can include several convolutional layers. In this embodiment, multiple feature maps of different sizes can be input into the global semantic feature extraction network, so that the global semantic feature extraction network can extract global semantic information for each feature map and fuse the global semantic information of each feature map to obtain the global semantic features of the video to be processed.

[0057] Next, an implementation manner of "extracting the spatial feature map of the to-be-processed sub-video for each to-be-processed sub-video" in S202 will be introduced, that is, how to extract the spatial feature map of the to-be-processed sub-video. In this embodiment, "extracting the spatial feature map of the to-be-processed sub-video for each to-be-processed sub-video" in S202 may include the following steps:

[0058] For each to-be-processed sub-video, input the to-be-processed sub-video into a high-resolution network to obtain multiple feature maps of different sizes; and input the multiple feature maps of different sizes into a spatial feature extraction network to obtain the spatial feature map of the to-be-processed sub-video.

[0059] In this embodiment, as Figure 3 shown, a network branch for extracting the spatial feature map can be separately set for each to-be-processed sub-video, and the network architectures of the network branches for extracting the spatial feature maps of each to-be-processed sub-video are the same, and the difference lies only in that the network parameters of the extraction network branches are different. Specifically, each network branch for extracting the spatial feature map of a to-be-processed sub-video includes a high-resolution network (HRNet) and a spatial feature extraction network. In this embodiment, for each to-be-processed sub-video, the to-be-processed sub-video can be first input into the high-resolution network to obtain multiple feature maps of different sizes; and the multiple feature maps of different sizes are input into the spatial feature extraction network to obtain the spatial feature map of the to-be-processed sub-video. It should be noted that when the number of frames of a to-be-processed sub-video in the to-be-processed sub-videos is the same as the number of frames of the to-be-processed video, the high-resolution network in S201a and the high-resolution network in the network branch for extracting the spatial feature map of this to-be-processed sub-video can be the same high-resolution network.

[0060] As an example, as Figure 3 shown, the spatial feature extraction network is a bottom-up multi-feature map fusion network. Specifically, the spatial feature extraction network may include a semantic feature extraction model, a global feature pyramid network, a temporal attention feature extraction network, and a spatial feature pyramid network. The step of "inputting the multiple feature maps of different sizes into the spatial feature extraction network to obtain the spatial feature map of the to-be-processed sub-video" may include the following steps:

[0061] Step a: Input the multiple feature maps of different sizes into the semantic feature extraction model to obtain the semantic features of the to-be-processed sub-video.

[0062] In this embodiment, after obtaining multiple feature maps of different sizes, the multiple feature maps of different sizes can be input into the semantic feature extraction model so as to obtain the semantic features of the sub-video to be processed. For example, the semantic features of the sub-video to be processed may include low-level semantic features such as the contour, edge, color, texture, and shape features of the sub-video to be processed, and may also include high-level semantic features of specific image content.

[0063] In one implementation, as Figure 3 shown, the semantic feature extraction model may include four convolutional layers. First, the features can be passed from the bottom convolutional layer to the upper convolutional layer step by step so as to combine the low-level semantic information of the sub-video to be processed into the high-level semantic information to obtain the semantic features of the sub-video to be processed. Among them, the semantic features of the sub-video to be processed can be multi-layer temporal semantic feature maps.

[0064] Step b: Input the semantic features of the sub-video to be processed into the global feature pyramid network to obtain the global features of the sub-video to be processed.

[0065] After obtaining the semantic features of the sub-video to be processed, the semantic features of the sub-video to be processed can be input into the global feature pyramid network so as to use the global feature pyramid network to perform feature prediction on the multi-layer temporal semantic feature maps of the semantic features of the sub-video to be processed to obtain the global features of the sub-video to be processed. As an example, as Figure 3 shown, for the multi-layer temporal semantic feature maps in the semantic features of the sub-video to be processed, the global feature pyramid network can first take the average value of each pixel point of the multi-layer temporal semantic feature maps, and then superimpose the feature values corresponding to the pixel points in the temporal semantic feature maps on the top-down feature pyramid so as to pass down the high-level semantic information, thereby improving the richness of the detailed information of the global features of the sub-video to be processed extracted by the entire global feature pyramid network.

[0066] Step c: Input the global features of the sub-video to be processed into the temporal attention feature extraction network to obtain the temporal and spatial features of the sub-video to be processed.

[0067] In this embodiment, the temporal attention feature extraction network is used to enhance the spatial features of different time series. Specifically, the global feature of the sub-video to be processed can be input into the temporal attention feature extraction network to obtain the temporal-spatial feature of the sub-video to be processed. The temporal attention feature extraction network is specifically used to calculate the vector matching score between the spatial feature at each moment and the global feature of the sub-video to be processed. Among them, the higher the matching score, the greater the correlation between the feature output at this moment and the global feature of the sub-video to be processed. On the contrary, the lower the matching score, the smaller the correlation between the feature output at this moment and the global feature of the sub-video to be processed. Therefore, the spatial feature can be temporally processed according to the matching score to obtain the temporal-spatial feature of the sub-video to be processed.

[0068] Specifically, in order to calculate the weight coefficient for normalization, the temporal attention feature extraction network can use the softmax function to calculate the weight coefficient α of the similarity between the spatial feature at a moment corresponding to the global feature and the global feature. ij :

[0069]

[0070] where t j is the j-th spatial feature, t i is the i-th spatial feature, similary() is the function for calculating similarity, and T is the global feature of the sub-video to be processed.

[0071] Then, the weighted sum of the temporal-spatial features can be calculated, and then concatenated with the original features to obtain the concatenated spatial features:

[0072] t i ’ is the feature after weight integration, and T i ’ is the spatial feature obtained by directly concatenating all features.

[0073] At the same time, the feature maps of each frame are aggregated through the ReLU function:

[0074] where W is the convolution parameter, q h is the output feature map, and N is the number of spatial features.

[0075] The temporal weight matrix V is obtained by weighting with tanh, sigmoid and perceptron:

[0076] V = sigmoid(W u tanh(W s T ′ + W q q h + bs ) + b u );W s 、W q 、W u are all model systems, T’ is the spatial feature obtained by directly concatenating all features, and b s 、b u is the intercept of the model.

[0077] Then, for the temporal-spatial features of multiple frames, weight the temporal-spatial features and finally sum to output the temporal-spatial features v i is the weight coefficient.

[0078] Step d: Input the temporal-spatial features of the to-be-processed sub-video into the spatial feature pyramid network to obtain the spatial feature map of the to-be-processed sub-video.

[0079] In this embodiment, in order to capture more spatial location information, the temporal-spatial features of the to-be-processed sub-video can be input into the spatial feature pyramid network to obtain the spatial feature map of the to-be-processed sub-video. Among them, as Figure 3 shown, the spatial feature pyramid network is a bottom-up pyramid structure. The spatial feature pyramid network can horizontally connect the input temporal-spatial features of multiple frames from the bottom layer to the top layer to shorten the path and achieve the purpose of capturing both location and semantic information simultaneously, that is, the spatial feature map of the to-be-processed sub-video includes the semantic information and spatial information of the to-be-processed sub-video. Moreover, the spatial feature pyramid network can map the temporal-spatial features of multiple layers (such as four layers) together by element-wise addition (i.e., pixel-wise addition) and finally obtain the spatial feature map of the to-be-processed sub-video. It can be seen that the spatial feature pyramid can enhance the recognition feature information.

[0080] Next, an implementation manner of "determine the human behavior type corresponding to the to-be-processed video according to the fusion feature and the global semantic feature" in S204 will be introduced, that is, how to determine the human behavior type corresponding to the to-be-processed video. In this embodiment, "determine the human behavior type corresponding to the to-be-processed video according to the fusion feature and the global semantic feature" in S204 may include the following steps:

[0081] S204a: Perform pooling processing on the fusion feature to obtain the fusion pooling feature.

[0082] In this embodiment, as Figure 3As shown, the fused features can be first input into a pooling layer, and the fused features are pooled through the pooling layer to obtain fused pooled features. It can be understood that the average pooling process is performed on the fused features by using the pooling layer to obtain the average pooling fused features; then, the temporal average pooling can be used to aggregate the average pooling fused features between different frames, that is, the average features of the average pooling fused features are extracted to obtain the fused pooled features.

[0083] S204b: Pool the global semantic features to obtain global semantic pooled features.

[0084] In this embodiment, as Figure 3 shown, the global semantic features can be first input into a pooling layer, and the global semantic features are pooled through the pooling layer to obtain global semantic pooled features. It can be understood that the average pooling process is performed on the global semantic features by using the pooling layer to obtain the average pooling global semantic features; then, the temporal average pooling can be used to aggregate the average pooling global semantic features between different frames, that is, the average features of the average pooling global semantic features are extracted to obtain the global semantic pooled features.

[0085] S204c: Concatenate the fused pooled features and the global semantic pooled features to obtain concatenated features.

[0086] After obtaining the fused pooled features and the global semantic pooled features, the fused pooled features and the global semantic pooled features can be concatenated, that is, the feature vectors of the fused pooled features and the global semantic pooled features are connected to obtain concatenated features.

[0087] S204d: Input the concatenated features into a fully connected classifier to obtain the human behavior type corresponding to the video to be processed.

[0088] After obtaining the concatenated features, the concatenated features can be input into a fully connected classifier to obtain the human behavior type corresponding to the video to be processed. Among them, the fully connected classifier includes a batch normalization layer and a fully connected layer.

[0089] It should be noted that in one implementation manner of this embodiment, Figure 2 the method shown may further include:

[0090] S10: Obtain the time information and / or location information corresponding to the video to be processed.

[0091] In this embodiment, the location information corresponding to the video to be processed can be understood as the position corresponding to the image content in the video to be processed. As an example, the location information corresponding to the video to be processed can be determined by performing image recognition on the video frames in the video to be processed, or the location information corresponding to the video to be processed can be determined by obtaining the location information of the device that captures the video to be processed.

[0092] The time information corresponding to the video to be processed can be understood as the acquisition time of the video to be processed. As an example, the time information corresponding to the video to be processed can be obtained from the time attribute of the video to be processed, or the time recorded when the image capture device captures the video to be processed can be used as the time information corresponding to the video to be processed.

[0093] S20: Determine the event type corresponding to the video to be processed according to the time information and / or location information corresponding to the video to be processed, and the human behavior type.

[0094] The event type corresponding to the video to be processed can be understood as the event corresponding to the human behavior type in the video to be processed. In this embodiment, after obtaining the time information and / or location information corresponding to the video to be processed, and the human behavior type, the event type corresponding to the video to be processed can be determined according to the time information and / or location information corresponding to the video to be processed, and the human behavior type; for example, if the human behavior type is lying flat and the location information is a corridor, then the event type corresponding to the video to be processed can be determined as a fall in the corridor; another example is that if the human behavior type is standing, the time information is 12 o'clock at midnight, and the location information is a restaurant, then the event type corresponding to the video to be processed can be determined as the gathering of people in the restaurant at 12 o'clock at midnight.

[0095] S30: If the event type corresponding to the video to be processed belongs to an abnormal event, output a prompt message.

[0096] In this embodiment, an abnormal event database can be preset. A number of abnormal event types are preset in the abnormal event database. For example, abnormal event types such as lying down in the aisle and human body gathering in the restaurant at 12 o'clock midnight. After determining the event type corresponding to the video to be processed, it can first be determined whether the event type corresponding to the video to be processed is the same as at least one preset abnormal event type in the abnormal event database. If the event type corresponding to the processed video is the same as at least one preset abnormal event type in the abnormal event database, it can be determined that the event type corresponding to the video to be processed belongs to an abnormal event, and a prompt message can be output. Among them, the prompt message can be a sound prompt message, such as a siren sound, a voice broadcast reminder, etc.; the prompt message can also be a text prompt message. For example, a text prompt message is sent to a preset terminal device so that the text prompt message can be displayed to the user through the terminal device, so as to achieve the prompting effect; of course, the prompt message can also be other prompting methods, which will not be elaborated here.

[0097] In this way, this embodiment can monitor the abnormal situation of the user's behavior according to the time information and / or location information corresponding to the video to be processed, as well as the human body behavior type. If it is recognized that the user has an abnormal event, a prompt message can be output in real time, thus ensuring the user's personal safety and improving the user experience.

[0098] All the above optional technical solutions can be combined arbitrarily to form alternative embodiments of the present disclosure, which will not be elaborated one by one here.

[0099] The following is an embodiment of the device of the present disclosure, which can be used to execute the embodiment of the method of the present disclosure. For the details not disclosed in the embodiment of the device of the present disclosure, please refer to the embodiment of the method of the present disclosure.

[0100] Figure 4 is a schematic diagram of a human body behavior recognition device provided by an embodiment of the present disclosure. As Figure 4 shown, the human body behavior recognition device includes:

[0101] An information acquisition module 401, configured to acquire a plurality of sub-videos to be processed corresponding to the video to be processed and the global semantic feature of the video to be processed, wherein the number of video frames of each sub-video to be processed is different;

[0102] A feature map extraction module 402, configured to extract a spatial feature map of each sub-video to be processed for each sub-video to be processed;

[0103] A feature fusion module 403, configured to perform fusion processing on the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fusion feature;

[0104] A type determination module 404, configured to determine the human behavior type corresponding to the video to be processed according to the fusion feature and the global semantic feature.

[0105] In some embodiments, the information acquisition module 401 is configured to:

[0106] Obtain a plurality of preset sampling rates; wherein, each sampling rate is different;

[0107] For each sampling rate, sample the video to be processed to obtain a sub-video to be processed corresponding to the sampling rate.

[0108] In some embodiments, the information acquisition module 401 is configured to:

[0109] Input the video to be processed into a high-resolution network to obtain a plurality of feature maps of different sizes;

[0110] Input the plurality of feature maps of different sizes into a global semantic feature extraction network to obtain the global semantic feature of the video to be processed.

[0111] In some embodiments, the feature map extraction module 402 is configured to:

[0112] For each sub-video to be processed, input the sub-video to be processed into a high-resolution network to obtain a plurality of feature maps of different sizes; and input the plurality of feature maps of different sizes into a spatial feature extraction network to obtain the spatial feature map of the sub-video to be processed.

[0113] In some embodiments, the spatial feature extraction network includes a semantic feature extraction model, a global feature pyramid network, a temporal attention feature extraction network, and a spatial feature pyramid network; the feature map extraction module 402 is configured to:

[0114] Input the plurality of feature maps of different sizes into the semantic feature extraction model to obtain the semantic feature of the sub-video to be processed;

[0115] Input the semantic feature of the sub-video to be processed into the global feature pyramid network to obtain the global feature of the sub-video to be processed;

[0116] Input the global feature of the sub-video to be processed into the temporal attention feature extraction network to obtain the temporal and spatial feature of the sub-video to be processed;

[0117] Input the temporal and spatial feature of the sub-video to be processed into the spatial feature pyramid network to obtain the spatial feature map of the sub-video to be processed.

[0118] In some embodiments, the feature fusion module 403 is configured to:

[0119] Perform splicing processing, feature superposition processing, or convolution processing on the spatial feature maps respectively corresponding to each sub-video to be processed to obtain fused features.

[0120] In some embodiments, the type determination module 404 is configured to:

[0121] Perform pooling processing on the fused features to obtain fused pooling features;

[0122] Perform pooling processing on the global semantic features to obtain global semantic pooling features;

[0123] Splice the fused pooling features and the global semantic pooling features to obtain spliced features;

[0124] Input the spliced features into a fully connected classifier to obtain the human behavior type corresponding to the video to be processed.

[0125] In some embodiments, the device further includes a prompt module, configured to:

[0126] Obtain the time information and / or location information corresponding to the video to be processed;

[0127] Determine the event type corresponding to the video to be processed according to the time information and / or location information corresponding to the video to be processed, and the human behavior type;

[0128] If the event type corresponding to the video to be processed belongs to an abnormal event, output a prompt message.

[0129] According to the technical solution provided by the embodiments of the present disclosure, the human behavior recognition device includes: an information acquisition module, configured to acquire a plurality of sub-videos to be processed corresponding to the video to be processed and the global semantic feature of the video to be processed, wherein the number of video frames of each sub-video to be processed is different; a feature map extraction module, configured to extract the spatial feature map of each sub-video to be processed; a feature fusion module, configured to fuse the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fusion feature; and a type determination module, configured to determine the human behavior type corresponding to the video to be processed according to the fusion feature and the global semantic feature. Since in this embodiment, the video to be processed is sampled at multiple different sampling rates to obtain a plurality of sub-videos to be processed, wherein the number of video frames of each sub-video to be processed is different, that is to say, the action rate in each sub-video to be processed is different. In this way, this embodiment can determine the human behavior type according to the spatial feature maps of the sub-videos to be processed with different action rates, that is to say, the factor of different action rates is considered in the process of human behavior category recognition, avoiding the problem of incorrect human behavior recognition caused by different action rates of users of different ages, thereby improving the accuracy of human behavior type recognition.

[0130] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.

[0131] Figure 5 is a schematic diagram of the computer device 5 provided by the embodiments of the present disclosure. As Figure 5 shown, the computer device 5 of this embodiment includes: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501. When the processor 501 executes the computer program 503, the steps in the above method embodiments are implemented. Alternatively, when the processor 501 executes the computer program 503, the functions of each module / module in the above device embodiments are implemented.

[0132] Exemplarily, the computer program 503 can be divided into one or more modules / units. The one or more modules / units are stored in the memory 502 and executed by the processor 501 to complete the present disclosure. The one or more modules / units can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 503 in the computer device 5.

[0133] The computer device 5 can be a desktop computer, a notebook, a palm computer, a cloud server, or other computer devices. The computer device 5 can include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art can understand that Figure 5 merely examples of the computer device 5, which do not constitute a limitation on the computer device 5, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, a bus, etc.

[0134] The processor 501 can be a central processing module (Central Processing Unit, CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, DSP), application-specific integrated circuits (Application Specific Integrated Circuit, ASIC), field-programmable gate arrays (Field-Programmable Gate Array, FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0135] The memory 502 can be an internal storage module of the computer device 5. For example, the hard disk or memory of the computer device 5. The memory 502 can also be an external storage device of the computer device 5. For example, a plug-in hard disk, a smart media card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the computer device 5. Further, the memory 502 can also include both the internal storage module and the external storage device of the computer device 5. The memory 502 is used to store computer programs and other programs and data required by the computer device. The memory 502 can also be used to temporarily store data that has been output or will be output.

[0136] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned functional modules and module divisions are used as examples. In actual applications, the above-mentioned functions can be allocated to different functional modules or modules as needed, that is, the internal structure of the device can be divided into different functional modules or modules to complete all or part of the functions described above. Each functional module and module in the embodiments can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. In addition, the specific names of the functional modules and modules are only for the convenience of mutual distinction and do not limit the protection scope of the present disclosure. The specific working processes of the modules and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0137] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0138] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0139] In the embodiments provided by the present disclosure, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are only illustrative. For example, the division of modules or modules is only a logical function division. In actual implementation, there can be other division methods. Multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces. The indirect coupling or communication connection of the device or module can be in electrical, mechanical or other forms.

[0140] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0141] In addition, the functional modules in the various embodiments of the present disclosure may be integrated into one processing module, may exist separately as individual physical modules, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.

[0142] When the integrated module / module is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present disclosure, it may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented. The computer program may include computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0143] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included within the protection scope of the present disclosure.

Claims

1. A human behavior recognition method, characterized in that, The method includes: Obtaining a plurality of sub-videos to be processed corresponding to the video to be processed and the global semantic feature of the video to be processed, where the number of video frames of each sub-video to be processed is different; For each sub-video to be processed, extracting the spatial feature map of the sub-video to be processed; Fusing the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fused feature; Determining the human behavior type corresponding to the video to be processed according to the fused feature and the global semantic feature; The obtaining a plurality of sub-videos to be processed corresponding to the video to be processed includes: Obtaining a plurality of preset sampling rates; where each sampling rate is different; For each sampling rate, sampling the video to be processed to obtain the sub-video to be processed corresponding to the sampling rate; The extracting the spatial feature map of the sub-video to be processed for each sub-video to be processed includes: For each sub-video to be processed, inputting the sub-video to be processed into a high-resolution network to obtain a plurality of feature maps of different sizes; and inputting the plurality of feature maps of different sizes into a spatial feature extraction network to obtain the spatial feature map of the sub-video to be processed; The spatial feature extraction network includes a semantic feature extraction model, a global feature pyramid network, a temporal attention feature extraction network, and a spatial feature pyramid network; the inputting the plurality of feature maps of different sizes into the spatial feature extraction network to obtain the spatial feature map of the sub-video to be processed includes: Inputting the plurality of feature maps of different sizes into the semantic feature extraction model to obtain the semantic feature of the sub-video to be processed; Inputting the semantic feature of the sub-video to be processed into the global feature pyramid network to obtain the global feature of the sub-video to be processed; Inputting the global feature of the sub-video to be processed into the temporal attention feature extraction network to obtain the temporal-spatial feature of the sub-video to be processed; Inputting the temporal-spatial feature of the sub-video to be processed into the spatial feature pyramid network to obtain the spatial feature map of the sub-video to be processed.

2. The method according to claim 1, wherein The obtaining the global semantic feature of the video to be processed includes: Inputting the video to be processed into a high-resolution network to obtain a plurality of feature maps of different sizes; Inputting the plurality of feature maps of different sizes into a global semantic feature extraction network to obtain the global semantic feature of the video to be processed.

3. The method according to claim 1, wherein The fusing the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fused feature includes: Performing splicing processing, feature superposition processing, or convolution processing on the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fused feature.

4. The method according to claim 1, characterized in that The determining the human behavior type corresponding to the video to be processed according to the fused feature and the global semantic feature includes: Performing pooling processing on the fused feature to obtain a fused pooling feature; Performing pooling processing on the global semantic feature to obtain a global semantic pooling feature; Splicing the fused pooling feature and the global semantic pooling feature to obtain a spliced feature; Inputting the spliced feature into a fully connected classifier to obtain the human behavior type corresponding to the video to be processed.

5. According to the method described in any one of claims 1-4, characterized in that, The method further includes: Obtain the time information and / or location information corresponding to the video to be processed; Determine the event type corresponding to the video to be processed according to the time information and / or location information corresponding to the video to be processed, and the human behavior type; If the event type corresponding to the video to be processed belongs to an abnormal event, output a prompt message.

6. A human behavior recognition device, characterized in that, The device includes: An information acquisition module, configured to acquire a plurality of sub-videos to be processed corresponding to the video to be processed and the global semantic features of the video to be processed, wherein the number of video frames of each sub-video to be processed is different; A feature map extraction module, configured to extract the spatial feature map of each sub-video to be processed; A feature fusion module, configured to perform fusion processing on the spatial feature maps respectively corresponding to each sub-video to be processed to obtain a fusion feature; A type determination module, configured to determine the human behavior type corresponding to the video to be processed according to the fusion feature and the global semantic feature; The information acquisition module is specifically configured to: acquire a plurality of preset sampling rates; wherein each sampling rate is different; for each sampling rate, sample the video to be processed to obtain a sub-video to be processed corresponding to the sampling rate; The feature map extraction module is specifically configured to: for each sub-video to be processed, input the sub-video to be processed into a high-resolution network to obtain multiple feature maps of different sizes; and input the multiple feature maps of different sizes into a spatial feature extraction network to obtain the spatial feature map of the sub-video to be processed; Wherein, the spatial feature extraction network includes a semantic feature extraction model, a global feature pyramid network, a temporal attention feature extraction network, and a spatial feature pyramid network. The feature map extraction module is specifically configured to: input the multiple feature maps of different sizes into the semantic feature extraction model to obtain the semantic features of the sub-video to be processed; input the semantic features of the sub-video to be processed into the global feature pyramid network to obtain the global features of the sub-video to be processed; input the global features of the sub-video to be processed into the temporal attention feature extraction network to obtain the temporal-spatial features of the sub-video to be processed; input the temporal-spatial features of the sub-video to be processed into the spatial feature pyramid network to obtain the spatial feature map of the sub-video to be processed.

7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Video behavior recognition method based on multi-scale spatial-temporal feature aggregation

    CN112052795A

  • Behavior recognition method and device, electronic equipment and storage medium

    CN112597824A