Monkey behavior recognition method and device, computer equipment, readable storage medium and program product
Through the multi-level feature extraction and fusion based on the behavior recognition model based on the yolov8 network, the problem of low-cost and accurate identification of animal behavior in complex environments is solved, and efficient monkey behavior recognition is achieved.
Patent Information
- Application Number
- CN202510672953.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art is difficult to accurately identify monkey animal behavior in complex environments at low cost, especially in the absence of multimodal information.
The behavior recognition model built on the yolov8 network is adopted, and multi-level feature extraction is performed through the backbone network, and shallow and deep feature extraction is performed using the c2f module and the cross-stage feature processing module. The feature enhancement is performed by combining multi-scale pooling and self-attention modules. The neck network performs cross-scale feature fusion, and the head network performs behavior recognition through the attention mechanism.
It realizes low-cost, efficient and accurate identification of monkey behavior in complex contexts, improves recognition accuracy and reasoning speed, and enhances the ability to understand complex scenarios and handles targets of different sizes.
Smart Images

Figure CN120495826A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for monkey behavior recognition. Background Art
[0002] Monkeys possess complex behavioral patterns and social structures, and their behavioral research is of great significance in fields such as biology, ecology, and medicine. This is particularly true for wildlife conservation. Due to environmental degradation and climate change, monkeys face various challenges and require ex situ conservation. Analyzing and studying monkey behavior in captive environments is crucial for animal welfare and conservation. Computer-assisted automated recognition can complement manual observation, helping conservators promptly detect behavioral anomalies, identify diseases, or adjust group structure. With the development of information technology, image processing techniques have been increasingly applied to monkey behavior research, enabling efficient behavioral recognition of monkeys based on image or video data.
[0003] Among them, since the activity environment and behavioral movements of monkeys are highly complex and diverse, related technologies often need to combine multimodal information such as images and audio to perform behavior recognition, or rely on images with simple backgrounds for analysis, making it difficult to achieve low-cost and accurate recognition of the behavior of monkeys in complex environments. Summary of the Invention
[0004] Based on this, it is necessary to provide a monkey behavior recognition method, device, computer equipment, computer-readable storage medium and computer program product to address the above technical problems.
[0005] In a first aspect, the present application provides a method for identifying monkey behavior, comprising:
[0006] Obtain the image to be recognized;
[0007] A backbone network of an action recognition model built based on the yolov8 network is used to perform multi-level feature extraction on the image to be identified to obtain multiple levels of features; wherein the backbone network uses a c2f module for shallow feature extraction, a cross-stage feature processing module based on a generalized efficient layer aggregation network for deep feature extraction, and a multi-scale pooling module and a self-attention module for multi-scale feature enhancement;
[0008] The neck network of the behavior recognition model is used to perform cross-scale feature fusion on the hierarchical features from the backbone network to obtain fused features corresponding to multiple scales; wherein the neck network uses a fusion splicing module to perform feature splicing, a c2f module to perform shallow feature fusion, and the cross-stage feature processing module to perform deep feature fusion;
[0009] The head network of the behavior recognition model is used to process the fusion features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
[0010] In one embodiment, the hierarchical features include multi-scale features; the use of a multi-scale pooling module and a self-attention module to perform multi-scale feature enhancement includes: using the multi-scale pooling module to perform multi-scale pooling processing on the first feature output by the last cross-stage feature processing module in the backbone network to obtain a second feature; using the multi-head self-attention block, the first fully connected layer, the first activation layer, the second fully connected layer and the second activation layer in the self-attention module to process the second feature to obtain the multi-scale feature.
[0011] In one embodiment, the head network of the behavior recognition model is used to process each of the fused features through an attention mechanism to obtain a monkey behavior recognition result corresponding to the image to be recognized, including: inputting each of the fused features into the detection head of the corresponding scale in the head network; using each of the detection heads to process the fused features through a compression-excitation attention mechanism and a convolutional block attention mechanism to obtain a scale recognition result corresponding to each of the scales; and obtaining a monkey behavior recognition result corresponding to the image to be recognized based on the scale recognition results corresponding to each of the scales.
[0012] In one embodiment, the behavior recognition model is trained by the following steps: obtaining a monkey behavior image sample set, dividing the monkey behavior image sample set into a training set and a verification set; using the training set to train the behavior recognition model to be trained to obtain a candidate behavior recognition model; using the verification set to verify the candidate behavior recognition model to obtain the average precision mean of the candidate behavior recognition model; if the average precision mean does not converge, adjusting the hyperparameters of the candidate behavior recognition model, using the candidate behavior recognition model as the behavior recognition model to be trained, and executing the step of training the behavior recognition model to be trained using the training set until a target behavior recognition model that makes the average precision mean converge is obtained; and obtaining the behavior recognition model based on the target behavior recognition model.
[0013] In one embodiment, the method of using the training set to train the behavior recognition model to be trained to obtain a candidate behavior recognition model includes: using the behavior recognition model to be trained to perform monkey behavior recognition on each monkey behavior image sample in the training set to obtain a monkey behavior prediction result corresponding to each monkey behavior image sample; calculating the classification loss and regression loss of the behavior recognition model to be trained based on the difference between the monkey behavior prediction result corresponding to each monkey behavior image sample and the monkey behavior sample label, and obtaining the model loss of the behavior recognition model to be trained based on the classification loss and the regression loss; adjusting the model parameters of the behavior recognition model to be trained based on the model loss, and executing the step of using the behavior recognition model to be trained to perform monkey behavior recognition on each monkey behavior image sample in the training set, until a candidate behavior recognition model that converges the model loss is obtained.
[0014] In one embodiment, obtaining the monkey behavior image sample set includes: obtaining original image samples containing monkey objects; performing image transformation processing on at least part of the original image samples to obtain augmented image samples; and obtaining the monkey behavior image sample set based on the original image samples and the augmented image samples.
[0015] In a second aspect, the present application further provides a monkey behavior recognition device, comprising:
[0016] An image acquisition module, used to acquire an image to be identified;
[0017] A feature extraction module is used to perform multi-level feature extraction on the image to be identified using the backbone network of the behavior recognition model built based on the yolov8 network to obtain multiple levels of features; wherein the backbone network uses the c2f module for shallow feature extraction, uses the cross-stage feature processing module based on the generalized efficient layer aggregation network for deep feature extraction, and uses the multi-scale pooling module and the self-attention module for multi-scale feature enhancement;
[0018] A feature fusion module, configured to utilize the neck network of the behavior recognition model to perform cross-scale feature fusion on the hierarchical features from the backbone network to obtain fused features corresponding to multiple scales; wherein the neck network uses a fusion splicing module for feature splicing, a c2f module for shallow feature fusion, and the cross-stage feature processing module for deep feature fusion;
[0019] The result acquisition module is used to use the head network of the behavior recognition model to process each of the fusion features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
[0020] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0021] Obtain the image to be recognized;
[0022] A backbone network of an action recognition model built based on the yolov8 network is used to perform multi-level feature extraction on the image to be identified to obtain multiple levels of features; wherein the backbone network uses a c2f module for shallow feature extraction, a cross-stage feature processing module based on a generalized efficient layer aggregation network for deep feature extraction, and a multi-scale pooling module and a self-attention module for multi-scale feature enhancement;
[0023] The neck network of the behavior recognition model is used to perform cross-scale feature fusion on the hierarchical features from the backbone network to obtain fused features corresponding to multiple scales; wherein the neck network uses a fusion splicing module to perform feature splicing, a c2f module to perform shallow feature fusion, and the cross-stage feature processing module to perform deep feature fusion;
[0024] The head network of the behavior recognition model is used to process the fusion features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
[0025] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:
[0026] Obtain the image to be recognized;
[0027] A backbone network of an action recognition model built based on the yolov8 network is used to perform multi-level feature extraction on the image to be identified to obtain multiple levels of features; wherein the backbone network uses a c2f module for shallow feature extraction, a cross-stage feature processing module based on a generalized efficient layer aggregation network for deep feature extraction, and a multi-scale pooling module and a self-attention module for multi-scale feature enhancement;
[0028] The neck network of the behavior recognition model is used to perform cross-scale feature fusion on the hierarchical features from the backbone network to obtain fused features corresponding to multiple scales; wherein the neck network uses a fusion splicing module to perform feature splicing, a c2f module to perform shallow feature fusion, and the cross-stage feature processing module to perform deep feature fusion;
[0029] The head network of the behavior recognition model is used to process the fusion features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
[0030] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps:
[0031] Obtain the image to be recognized;
[0032] A backbone network of an action recognition model built based on the yolov8 network is used to perform multi-level feature extraction on the image to be identified to obtain multiple levels of features; wherein the backbone network uses a c2f module for shallow feature extraction, a cross-stage feature processing module based on a generalized efficient layer aggregation network for deep feature extraction, and a multi-scale pooling module and a self-attention module for multi-scale feature enhancement;
[0033] The neck network of the behavior recognition model is used to perform cross-scale feature fusion on the hierarchical features from the backbone network to obtain fused features corresponding to multiple scales; wherein the neck network uses a fusion splicing module to perform feature splicing, a c2f module to perform shallow feature fusion, and the cross-stage feature processing module to perform deep feature fusion;
[0034] The head network of the behavior recognition model is used to process the fusion features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
[0035] The above-mentioned monkey behavior recognition method, device, computer equipment, computer-readable storage medium and computer program product first obtain an image to be recognized, and then use the backbone network of the behavior recognition model constructed based on the yolov8 network to perform multi-level feature extraction on the image to be recognized to obtain multiple levels of features, wherein the backbone network uses the c2f module and the cross-stage feature processing module based on the generalized efficient layer aggregation network to perform shallow feature extraction and deep feature extraction, and uses the multi-scale pooling module and the self-attention module to perform multi-scale feature enhancement, and then uses the neck network of the behavior recognition model to perform cross-scale feature fusion on the features of each level from the backbone network to obtain fused features corresponding to multiple scales, wherein the neck network uses the fusion splicing module for feature splicing, uses the c2f module for shallow feature fusion, and uses the cross-stage feature processing module for deep feature fusion, and finally uses the head network of the behavior recognition model to process the fused features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized. This solution uses a cross-stage feature processing module based on a generalized efficient layer aggregation network for feature extraction or feature fusion in the backbone and neck networks of the behavior recognition model. This allows for a refined structure to replace the multi-layer nested structure of the traditional Yolov8 network, reducing computational complexity and improving inference speed while maintaining accuracy, thus facilitating low-cost and accurate recognition of monkey behaviors. Multi-scale feature enhancement using a multi-scale pooling module and a self-attention module in the backbone network enhances the expressiveness of features, thereby improving the accuracy of monkey behavior recognition. Feature concatenation using a fusion concatenation module in the neck network allows for richer details and more contextual information in the concatenated features, thereby improving the ability to understand complex scenes and process objects of varying sizes. Monkey behavior detection based on an attention mechanism in the head network reduces the computational load of the detection head and increases its sensitivity to foreground objects, further improving the accuracy of monkey behavior recognition. Consequently, this solution enables low-cost, efficient, and accurate detection of monkey behaviors in images to be recognized against complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 1 is a flow chart of a method for identifying monkey behavior according to an embodiment;
[0038] Figure 2 Schematic diagram of the structure of a behavior recognition model in one embodiment;
[0039] Figure 3 Schematic diagram of the structure of a cross-stage feature processing module in one embodiment;
[0040] Figure 4 A schematic diagram of the structure of a fusion and splicing module in one embodiment;
[0041] Figure 5 Schematic diagram of the structure of a self-attention module in one embodiment;
[0042] Figure 6 Schematic diagram of a process for obtaining a monkey behavior recognition result corresponding to an image to be recognized in one embodiment;
[0043] Figure 7 Schematic diagram of the structure of an attention module in one embodiment;
[0044] Figure 8 Schematic diagram of the training process of a behavior recognition model in one embodiment;
[0045] Figure 9 is a schematic diagram of an augmented image sample obtained by performing image transformation processing on an original image sample in one embodiment;
[0046] Figure 10 is a structural block diagram of a monkey behavior recognition device in one embodiment;
[0047] Figure 11 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0048] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0049] In one embodiment, Figure 1 As shown, a method for identifying monkey behavior is provided. This embodiment uses the method applied to a server as an example. It is understandable that the method can also be applied to a terminal, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0050] Step S101: obtaining an image to be recognized.
[0051] Step S102: Using the backbone network of the behavior recognition model constructed based on the yolov8 network, multi-level feature extraction is performed on the image to be recognized to obtain multiple levels of features.
[0052] In step S103 , the neck network of the behavior recognition model is used to perform cross-scale feature fusion on the features at each level from the backbone network to obtain fused features corresponding to multiple scales.
[0053] Step S104 , using the head network of the behavior recognition model, the fusion features are processed through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
[0054] The image to be identified may be an image containing one or more monkey objects. In step S101, the image to be identified may be obtained for subsequent behavior identification of each monkey object in the image.
[0055] The behavior recognition model used in this embodiment is a model built based on the yolov8 network, such as Figure 2 As shown, the model can include a backbone network, a neck network, and a head network.
[0056] In step S102, by inputting the image to be identified into the behavior recognition model, the backbone network of the model can be used to perform multi-level feature extraction on the image to be identified to obtain multiple levels of features, and then input them into the neck network. In the shallow layer of the backbone network, the behavior recognition model uses the c2f module to perform shallow feature extraction, and uses a cross-stage feature processing module based on the Generalized Efficient Layer Aggregation Network (GELAN) to perform deep feature extraction in the deep layer of the backbone network. The multi-scale pooling module and the self-attention module are used to perform multi-scale feature enhancement on the features output by the last cross-stage feature processing module. The multiple levels of features extracted by the backbone network may include shallow features extracted by the c2f module, deep features extracted by the cross-stage feature processing module, and multi-scale features processed by the multi-scale pooling module and the self-attention module.
[0057] For example, Figure 2As shown, the backbone network of the action recognition model can sequentially use two c2f modules to extract shallow features, with the features output by the second c2f module used as shallow features for the neck network. Based on the shallow features, the backbone network can then sequentially use two cross-stage feature processing modules to extract deep features, with the features output by the first cross-stage feature processing module used as deep features for the neck network. Subsequently, the backbone network can sequentially use a multi-scale pooling module and a self-attention module to perform multi-scale feature enhancement on the features output by the last cross-stage feature processing module, with the features output by the self-attention module used as multi-scale features for the neck network. Two 3×3 convolutional layers can be used before the first c2f module to process features, and a 3×3 convolutional layer can be used between each c2f module and the cross-stage feature processing module to process features. The multi-scale pooling module used in the backbone network can be a Spatial Pyramid Pooling Enhanced with Efficient Layer Aggregation Network (SPPELAN), and the self-attention module can be a module that uses the self-attention mechanism for feature processing.
[0058] The cross-stage feature processing module used in the behavior recognition model can be a module built based on the Generalized Efficient Layer Aggregation Network (GELAN). For example, the structure of the cross-stage feature processing module can be as follows: Figure 3 As shown in part (a), the module can first use a 1×1 convolutional layer to process the input features and perform a split operation on the processed features to obtain the first sub-feature and the second sub-feature. The second sub-feature is then processed using the first basic operation block and the second basic operation block in sequence. The features output by the first sub-feature, the second sub-feature, the first basic operation block, and the second basic operation block can then be concatenated to obtain the final output features of the cross-stage feature processing module.
[0059] For example, Figure 3 As shown in part (a), the first basic operation block and the second basic operation block in the cross-stage feature processing module can respectively include a Light Cross Stage Partial Network (LightCSP) and a 1×1 convolutional layer. For example, Figure 3As shown in part (b), the lightweight cross-stage local network can process the input features through two branches, where the first branch can use a 1×1 convolution layer to process the features, and the second branch can use a 1×1 convolution layer, a 3×3 depthwise convolution layer (DWConv), a combined convolution block, and an efficient channel attention layer (ECA) to process the features. The features processed by the first and second branches are then concatenated and convolved through a 1×1 convolution layer to obtain the features output by the lightweight cross-stage local network. For example, Figure 3 As shown in part (c), the combined convolution block can process the input features using a 1×1 convolution layer and a 3×3 depth convolution layer in sequence, and then add the input features to the features output by the 3×3 depth convolution layer element by element (represented in the figure as ), and obtain the features of the combined convolutional block output.
[0060] In step S103, the neck network of the behavior recognition model can be used to perform cross-scale feature fusion on the features at each level from the backbone network. Figure 2 , the behavior recognition model can use the fusion splicing module (FuseConcat) in the neck network for feature splicing, and use the c2f module in the shallow layer of the neck network for shallow feature fusion and the cross-stage feature processing module in the deep layer of the neck network for deep feature fusion. For example, the structure of the fusion splicing module used in the neck network can be as follows: Figure 4 As shown in the figure, it can first concatenate the two input features to obtain the concatenated features, then use two 1×1 convolutional layers to process the concatenated features, and add the features output by the two 1×1 convolutional layers element by element (represented in the figure as ), and obtain the features output by the fusion splicing module. The structure of the cross-stage feature processing module used in the neck network is similar to that in the backbone network, and will not be repeated here. Figure 2 As shown, the neck network of the action recognition model can output multiple fused features corresponding to multiple scales to the head network.
[0061] Among them, in step S104, the head network of the behavior recognition model can be used to process each fused feature to obtain the monkey behavior recognition result corresponding to the image to be recognized. Specifically, the head network of the behavior recognition model can include detection heads corresponding to different scales, and each detection head can process the fused features of the corresponding scale through the attention mechanism to obtain the scale recognition result corresponding to the scale. Then, multiple scale recognition results can be fused to obtain the monkey behavior recognition result corresponding to the image to be recognized. Exemplarily, the monkey behavior recognition result corresponding to the image to be recognized may include the position information of each monkey object in the image to be recognized and the behavior classification label. Exemplarily, the position information of the monkey object may include the prediction box coordinates, and the behavior classification label of the monkey object may indicate its corresponding monkey behavior type, and the monkey behavior type may include but is not limited to walking, sitting, jumping, climbing, eating, hanging, etc.
[0062] In the above-mentioned monkey behavior recognition method, by using a cross-stage feature processing module based on a generalized efficient layer aggregation network for feature extraction or feature fusion in the backbone network and deep layers of the neck network of the behavior recognition model, the multi-layer nested structure in the traditional yolov8 network can be replaced with a refined structure, while maintaining accuracy while reducing computational complexity and improving inference speed, which is conducive to achieving low-cost and accurate recognition of monkey behaviors. By using a multi-scale pooling module and a self-attention module for multi-scale feature enhancement in the backbone network, the expressive power of features can be enhanced, which is conducive to improving the recognition accuracy of monkey behaviors. By using a fusion splicing module for feature splicing in the neck network, the spliced features can have richer details and more contextual information, which is conducive to improving the ability to understand complex scenes and the ability to process objects of different sizes. By using an attention mechanism to detect monkey behaviors in the head network, the computational load of the detection head can be reduced and the sensitivity of the detection head to foreground objects can be improved, which is conducive to further improving the recognition accuracy of monkey behaviors. Therefore, the method can achieve low-cost, efficient and accurate detection of the behaviors of various monkey objects in images to be recognized with complex backgrounds.
[0063] In an exemplary embodiment, the hierarchical features include multi-scale features; multi-scale feature enhancement is performed using a multi-scale pooling module and a self-attention module, including: using the multi-scale pooling module to perform multi-scale pooling processing on the first feature output by the last cross-stage feature processing module in the backbone network to obtain a second feature; using the multi-head self-attention block, the first fully connected layer, the first activation layer, the second fully connected layer and the second activation layer in the self-attention module to process the second feature to obtain a multi-scale feature.
[0064] Specifically, the features input from the backbone network to the neck network in the behavior recognition model may include multi-scale features output from the self-attention module. Figure 2As shown in the figure, the multi-scale pooling module in the backbone network can receive the first feature output by the last cross-stage feature processing module in the backbone network, perform multi-scale pooling on it to obtain the second feature, and then input it into the self-attention module. The self-attention module can process the second feature to obtain a multi-scale feature and input it into the neck network.
[0065] For example, Figure 5 As shown, the self-attention module can first use the 1×1 convolution layer and the multi-headed self-attention block (MHSA) to process the second feature in sequence, and perform residual connection on the features output by the 1×1 convolution layer and the multi-headed self-attention block to obtain the first intermediate feature. Then, the self-attention module can use the first fully connected layer, the first activation layer, the second fully connected layer, and the second activation layer to process the first intermediate feature in sequence to obtain the second intermediate feature. Subsequently, the first intermediate feature and the second intermediate feature are residually connected to obtain the third intermediate feature, and the third intermediate feature is processed by the 1×1 convolution layer to obtain the multi-scale feature output by the self-attention module. Among them, as shown in Figure 5 As shown, the first activation layer can use ReLU as the activation function, and the second activation layer can use SiLU as the activation function.
[0066] In this embodiment, by using a multi-scale pooling module and a self-attention module in the backbone network to enhance multi-scale features, the multi-scale pooling and multi-head self-attention mechanisms can effectively capture richer deep information in the image to be recognized. By introducing a two-layer activation function structure to increase network depth and nonlinear expression capabilities, the accuracy and generalization performance of the behavior recognition model are effectively improved. As a result, this behavior recognition model can achieve more accurate monkey behavior recognition in the image to be recognized.
[0067] In an exemplary embodiment, Figure 6 As shown in the figure, the head network of the behavior recognition model is used to process the fused features through the attention mechanism to obtain the monkey behavior recognition results corresponding to the image to be recognized, which may include:
[0068] Step S601: Input each fusion feature into the detection head of the corresponding scale in the head network.
[0069] Among them, such as Figure 2 As shown, the head network of the behavior recognition model can include multiple detection heads corresponding to different scales. In this step, the fusion features corresponding to different scales output by the neck network can be input into the detection heads of corresponding scales in the head network respectively.
[0070] In step S602 , each detection head is used to process the fused features through the compression-excitation attention mechanism and the convolution block attention mechanism to obtain scale recognition results corresponding to each scale.
[0071] Please continue to refer to Figure 2 Each detection head in the head network of the behavior recognition model can include a classification branch and a regression branch. The classification branch can be used to classify the monkey's behavior, and the regression branch can be used to detect the monkey's position. Thus, by processing the fused features of the corresponding scale using each detection head, a scale recognition result corresponding to each scale can be obtained. The scale recognition result can include the scale position prediction information of each monkey object, the scale behavior classification label, and the corresponding confidence level.
[0072] Among them, Figure 2 As shown in , each branch of the detection head can use the attention module and the 3×3 convolution layer in turn to process the fusion features to obtain the corresponding scale recognition results. Figure 7 As shown in the figure, the attention module can sequentially process the fused features using Squeeze-and-Excitation Networks (SE), Convolutional Block Attention Module (CBAM), and a 1×1 convolutional layer to obtain processed fused features. Subsequently, the processed fused features can be processed using a 3×3 convolutional layer to obtain the corresponding scale position prediction information and scale prediction confidence, or the scale classification confidence corresponding to each monkey behavior type.
[0073] Step S603 , obtaining a monkey behavior recognition result corresponding to the image to be recognized based on the scale recognition results corresponding to each scale.
[0074] Among them, the head network of the behavior recognition model can obtain the monkey behavior recognition results corresponding to the image to be identified based on the scale recognition results corresponding to each scale through confidence filtering, non-maximum suppression and other processing.
[0075] In this embodiment, by introducing the compressed excitation attention mechanism and the convolutional block attention mechanism into the head network of the behavior recognition model, the sensitivity of the detection head to foreground objects can be enhanced, the representation ability of the detection head can be improved, and the computational complexity of the detection head can be reduced. Therefore, the behavior recognition model can be used to realize efficient and accurate monkey behavior recognition of the image to be recognized.
[0076] In an exemplary embodiment, Figure 8 As shown, the behavior recognition model in this application can be trained through the following steps:
[0077] Step S801: Acquire a monkey behavior image sample set, and divide the monkey behavior image sample set into a training set and a validation set.
[0078] The monkey behavior image sample set may include multiple monkey behavior image samples. Each monkey behavior image sample may include one or more monkey objects. The monkey behavior image sample set may be divided into a training set and a validation set, each of which may include multiple monkey behavior image samples. For example, the number of monkey behavior image samples included in the training set may be greater than the number of monkey behavior image samples included in the validation set.
[0079] Step S802: Use the training set to train the behavior recognition model to be trained to obtain a candidate behavior recognition model.
[0080] The behavior recognition model to be trained may be a pre-built model, and the structure of the model may be consistent with the structure of the behavior recognition model. In this step, the training set may be used to iteratively train the behavior recognition model to be trained until a candidate behavior recognition model that meets the preset conditions is obtained. Exemplarily, when training the behavior recognition model to be trained, a loss function corresponding to the model may be constructed and the model loss of the behavior recognition model to be trained may be calculated using the function. When the model loss of the model converges, or when the number of iterations reaches a preset threshold, it may be determined that the model obtained from the training meets the preset conditions, thereby obtaining a candidate behavior recognition model.
[0081] Step S803: Verify the candidate behavior recognition model using the verification set to obtain the average precision of the candidate behavior recognition model.
[0082] The candidate behavior recognition model can be validated using the validation set to obtain the monkey behavior prediction results output by the candidate behavior recognition model for each monkey behavior image sample in the validation set. The mean average precision (mAP) of the candidate behavior recognition model can then be calculated based on the difference between the monkey behavior prediction results corresponding to each monkey behavior image sample and the monkey behavior sample label.
[0083] After obtaining the mean average precision of the candidate behavior recognition model, it can be determined whether it has converged. If the mean average precision has not converged, step S804 can be executed. If the mean average precision has converged, the candidate behavior recognition model can be used as the target behavior recognition model and step S805 can be executed. For example, the convergence of the mean average precision can mean that the mean average precision is greater than a preset value, or that the mean average precision no longer increases.
[0084] Step S804: If the mean of the average precision has not converged, adjust the hyperparameters of the candidate behavior recognition model, use the candidate behavior recognition model as the behavior recognition model to be trained, and execute the step of training the behavior recognition model to be trained using the training set until the target behavior recognition model that makes the mean of the average precision converge is obtained.
[0085] If the mean average precision of the candidate behavior recognition model does not converge, the hyperparameters of the candidate behavior recognition model may be adjusted. For example, the hyperparameters of the candidate behavior recognition model may include batch size, learning rate, etc. Subsequently, the candidate behavior recognition model may be used as the behavior recognition model to be trained, and steps S802 and S803 may be performed again based on the adjusted hyperparameters until a target behavior recognition model is obtained that can converge the mean average precision.
[0086] Step S805: Obtain a behavior recognition model according to the target behavior recognition model.
[0087] After obtaining the target behavior recognition model that converges the mean precision, the model can be used as a behavior recognition model for performing monkey behavior recognition on the image to be recognized.
[0088] In this embodiment, by dividing the monkey behavior image sample set into a training set and a validation set, the model is first trained with the training set, and then the average precision of the model on the validation set is calculated. When the average precision has not converged, the hyperparameters of the model are updated and trained again. This can improve the accuracy of the model while avoiding overfitting, effectively improve the adaptability and robustness of the model in the real environment, and ensure that the behavior recognition model can have high generalization performance and stable recognition accuracy on different images to be recognized.
[0089] In an exemplary embodiment, training a behavior recognition model to be trained using a training set to obtain a candidate behavior recognition model can include: using the behavior recognition model to be trained to perform monkey behavior recognition on each monkey behavior image sample in the training set to obtain a monkey behavior prediction result corresponding to each monkey behavior image sample; calculating the classification loss and regression loss of the behavior recognition model to be trained based on the difference between the monkey behavior prediction result corresponding to each monkey behavior image sample and the monkey behavior sample label, and obtaining the model loss of the behavior recognition model to be trained based on the classification loss and regression loss; adjusting the model parameters of the behavior recognition model to be trained based on the model loss, and executing the step of performing monkey behavior recognition on each monkey behavior image sample in the training set using the behavior recognition model to be trained, until a candidate behavior recognition model that converges with the model loss is obtained.
[0090] When using the training set to train the behavior recognition model to be trained, the training set can be used to iteratively train the behavior recognition model to be trained, and during each training session, the model loss is calculated based on the difference between the monkey behavior prediction results corresponding to each monkey behavior image sample and the monkey behavior sample label, until a candidate behavior recognition model that converges with the model loss is obtained. The monkey behavior prediction results corresponding to the monkey behavior image samples may include the predicted position information and predicted behavior classification information of each monkey object, and the monkey behavior sample label may include the sample position information and sample behavior category of each monkey object. The model loss of the behavior recognition model to be trained can be calculated based on the classification loss and regression loss of the model by weighted summation or other methods.
[0091] For example, the classification loss of the action recognition model to be trained can be calculated using the binary cross-entropy loss shown in the following formula:
[0092]
[0093] Where, is the binary cross entropy loss, is the sample size, is the sample behavior category of the i-th monkey behavior image sample, The monkey behavior prediction result corresponding to the i-th monkey behavior image sample corresponds to The predicted probability of .
[0094] For example, the distribution focal loss (Distribution Focal Loss) of the monkey behavior image sample can be calculated by the following formula, and then the regression loss is calculated based on the distribution focal loss of each monkey behavior image sample in the training set:
[0095]
[0096] Where, is the distribution focus loss, is the vertex coordinate value of the sample target box corresponding to the monkey behavior image sample, and for The adjacent integer coordinate values of and They are the outputs of the monkey behavior model corresponding to and Among them, according to the distribution focus loss of each monkey behavior image sample, the regression loss can be obtained by summing up.
[0097] After obtaining the model loss, the model parameters of the behavior recognition model to be trained can be adjusted based on the model loss, and training can be performed again. When the model loss no longer decreases, it can be determined that the model loss has converged, and the model obtained from this training can be used as a candidate behavior recognition model.
[0098] In this embodiment, by introducing distributed focal loss and binary cross entropy loss to calculate the regression loss and classification loss during the model training process, the regression branch and classification branch in the behavior recognition model to be trained can be effectively constrained respectively, so that the model parameters of the behavior recognition model to be trained can be adjusted based on the model loss and the training is repeated, so that a candidate behavior recognition model with higher classification and positioning accuracy can be obtained.
[0099] In an exemplary embodiment, obtaining a set of monkey behavior image samples may include: obtaining original image samples containing monkey objects; performing image transformation processing on at least part of the original image samples to obtain augmented image samples; and obtaining a set of monkey behavior image samples based on the original image samples and the augmented image samples.
[0100] The original image sample may be an image collected at a monkey activity site, and may include one or more monkey objects. In this embodiment, a plurality of original image samples may be first obtained, and then image transformation processing may be performed on one or more of the original image samples to obtain augmented image samples. For example, Figure 9 As shown, for the original image sample ( Figure 9 The image transformation process performed in part (a) can include random cropping ( Figure 9 Part (b), random rotation ( Figure 9 Part (c), random scaling ( Figure 9 Part (d), horizontal flip ( Figure 9 Middle (e) part), vertical flip ( Figure 9 Middle (f) part), color jitter ( Figure 9 (g) part), random mixing ( Figure 9 (h) in the figure) to obtain multiple augmented image samples corresponding to the original image sample. It is understandable that one or more image transformation processes can be performed on each original image sample to obtain one or more augmented image samples, wherein each augmented image sample can be obtained by performing one image transformation process on the original image sample, or can be obtained by performing multiple image transformation processes on the original image sample (for example, random scaling and color dithering can be performed on the original image sample).
[0101] Among them, using the original image samples and augmented image samples, a monkey behavior image sample set for model training can be constructed.
[0102] In this embodiment, by transforming the original image samples, the diversity and amount of model training data can be effectively improved, so that the model can be trained using the monkey behavior image sample set, which can improve the accuracy and generalization ability of the behavior recognition model.
[0103] In an exemplary embodiment, a method for identifying monkey behavior is provided.
[0104] Among them, the monkey behavior recognition method in this embodiment may include: obtaining an image to be recognized; using the backbone network of the behavior recognition model constructed based on the yolov8 network to perform multi-level feature extraction on the image to be recognized to obtain multiple levels of features; using the neck network of the behavior recognition model to perform cross-scale feature fusion on the features of each level from the backbone network to obtain fusion features corresponding to multiple scales; using the head network of the behavior recognition model to process each fusion feature through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
[0105] Specifically, the behavior recognition model in this embodiment can be a model built based on the yolov8 network, and its structure can be as follows: Figure 2 shown.
[0106] Among them, the behavior recognition model uses the c2f module to extract shallow features in the shallow layer of the backbone network and uses the deep layer of the backbone network. Figure 3 The cross-stage feature processing module shown in the figure performs deep feature extraction. The cross-stage feature processing module can reduce redundant calculations by splitting and merging feature maps, and uses simple convolution to combine basic operation blocks, thereby forming a refined backbone network and reducing the multi-layer loop calculation caused by the multi-layer nested structure of the C2F module used in the traditional YOLOv8 network.
[0107] At the same time, the behavior recognition model adds the following after the multi-scale pooling module of the backbone network: Figure 5 The self-attention module shown in the figure utilizes a multi-head self-attention mechanism to process the semantically rich features output by the multi-scale pooling module. This parallel processing of multiple independent attention heads captures richer information within the features, improving the model's expressiveness and generalization performance. Furthermore, the self-attention module utilizes the first fully connected layer, the first activation layer, the second fully connected layer, and the second activation layer to process features, further reducing model complexity and improving generalization.
[0108] Among them, the neck network of the behavior recognition model can use the Feature Pyramid Networks (FPN) and Path Aggregation Network (PAN) FPN+PAN structure to fuse the feature maps of different scales extracted by the backbone network. Figure 4 The fusion splicing module shown in the figure performs feature splicing, and uses the c2f module in the shallow layer of the neck network to perform shallow feature fusion and the deep layer of the neck network to use the Figure 3 The cross-stage feature processing module shown in the figure performs deep feature fusion. The fusion and splicing module fuses input features into new features with richer texture and edge details, as well as more contextual information. This improves the model's understanding of complex scenes and its ability to handle objects of varying sizes in images. Using the cross-stage feature processing module in the neck network simplifies the neck network and improves the efficiency of feature fusion.
[0109] The head network of the behavior recognition model can include detection heads corresponding to different scales. Each detection head can include two branches, namely, a classification branch and a regression branch, which can output the predicted box and behavior classification label of the monkey object in the image to be recognized. Each branch of the detection head can include the following: Figure 7 The attention module and 3×3 convolution layer are shown. By improving the detection head of the traditional Yolov8 network and replacing the original ordinary convolution with a lightweight attention module, the computational complexity of the detection head can be reduced and the detection speed can be improved.
[0110] The training process of the behavior recognition model in this embodiment may include:
[0111] First, original image samples containing monkey objects are obtained, and image transformation processing is performed on at least a portion of the original image samples to obtain augmented image samples. Then, a set of monkey behavior image samples can be obtained based on the original image samples and the augmented image samples. The set of monkey behavior image samples can be divided into a training set and a validation set.
[0112] Then, the behavior recognition model to be trained is used to perform monkey behavior recognition on each monkey behavior image sample in the training set, and the monkey behavior prediction results corresponding to each monkey behavior image sample are obtained. Based on the difference between the monkey behavior prediction results corresponding to each monkey behavior image sample and the monkey behavior sample label, the classification loss and regression loss of the behavior recognition model to be trained are calculated, and the model loss of the behavior recognition model to be trained is obtained based on the classification loss and regression loss. Then, the model parameters of the behavior recognition model to be trained are adjusted based on the model loss, and the step of performing monkey behavior recognition on each monkey behavior image sample in the training set using the behavior recognition model to be trained is performed until a candidate behavior recognition model that converges with the model loss is obtained. The classification loss of the behavior recognition model to be trained can be obtained by calculating its binary cross entropy loss, and the regression loss can be obtained based on the distribution focal loss statistics of the model.
[0113] After obtaining the candidate behavior recognition model, the validation set can be used to validate the candidate behavior recognition model and obtain the mean average precision of the candidate behavior recognition model. If the mean average precision does not converge, the batch size, learning rate, and other hyperparameters of the candidate behavior recognition model can be adjusted. The candidate behavior recognition model can then be used as the behavior recognition model to be trained, and the steps of training the behavior recognition model to be trained using the training set can be repeated until the target behavior recognition model is obtained for which the mean average precision converges. Finally, the behavior recognition model can be derived based on the target behavior recognition model.
[0114] To verify the monkey behavior recognition method using the above-mentioned behavior recognition model in this embodiment, the behavior recognition model in this embodiment (hereinafter referred to as the "model of this application") and the Yolov5, Yolov6, Yolov7, Yolov8, Yolov9, and Yolov10 models were verified using the same set of images to be recognized. The corresponding parameters, billion floating-point operations per second, mean average precision at an intersection of Union (IoU) threshold of 0.5 (hereinafter referred to as "mAP50"), and mean average precision with IoU thresholds ranging from 0.5 to 0.95 (hereinafter referred to as "mAP50-95") were calculated for each model. The indicators of each model are shown in Table 1. As can be seen from Table 1, the model of this application achieved a higher recognition accuracy rate.
[0115] Table 1
[0116]
[0117] Among them, this embodiment further conducts ablation experiments on the various improvements of the above-mentioned behavior recognition model relative to the YOLOv8 network. The corresponding designs in the traditional YOLOv8 network are replaced with single improvements, and Model 1 (replacing the C2F module in the backbone network and the neck network with a cross-stage feature processing module), Model 2 (adding a self-attention module), Model 3 (replacing the splicing operation with a fusion splicing module), and Model 4 (replacing the detection head of the YOLOv8 network with a detection head based on the attention mechanism) are constructed. The indicators of each model are shown in Table 2. Combining Tables 1 and 2, it can be seen that the combination of the various improvements can synergistically improve the recognition accuracy of monkey behaviors.
[0118] Table 2
[0119]
[0120] At the same time, this embodiment also targets the compression incentive attention mechanism (hereinafter referred to as "SE"), the convolutional block attention mechanism (hereinafter referred to as "CAMB"), and Figure 5 The self-attention module (hereinafter referred to as "AB") shown in Figure 7 Experiments comparing the application of the attention module (hereafter referred to as "AttentionBlock") at different model locations were conducted. By replacing the convolutional layers in the cross-stage feature processing module with SE, CAMB, and AB, respectively, and applying the three different cross-stage feature processing modules to the deep layers of the backbone and neck networks of the Yolov8 network, Models 5, 6, and 7 were constructed. By adding SE, CAMB, and AB modules after the multi-scale pooling module of the Yolov8 network, Models 8, 9, and 10 were constructed. By replacing the convolutional modules in the detection head of the Yolov8 network with SE, CAMB, AB, and AttentionBlock, Models 11, 12, 13, and 14 were constructed, respectively. The performance metrics of each model are shown in Table 3.
[0121] Table 3
[0122]
[0123] In this embodiment, in order to solve the problem of difficulty in capturing monkey behavior features, a behavior recognition model is constructed from the perspective of improving feature extraction capabilities, and the model is used to perform monkey behavior recognition on the images to be identified. Among them, based on the traditional Yolov8 network, the C2F module is replaced with a cross-stage feature processing module in the deep layer of the backbone network and the neck network, which can optimize the network structure to reduce the amount of calculation and improve the network's expression ability; by introducing a self-attention module after the multi-scale pooling module, it is possible to use multi-head self-attention to capture richer information in the features, improving the model's expression ability and generalization performance; by using a fusion splicing module in the neck network, it is possible to enhance the expression ability at different scales, improve the model's ability to understand complex scenes and its ability to handle targets of different sizes; by adding a channel attention mechanism to the head network, it is possible to enhance the sensitivity of the detection head to foreground objects, improve the representation ability of the detection head and reduce the amount of calculation of the detection head. At the same time, in this embodiment, by transforming the monkey behavior images, a monkey behavior image sample set with rich diversity and large data volume is constructed, and then the monkey behavior image sample set is used for model training, which can further improve the accuracy and generalization ability of the model. Therefore, the behavior recognition model used in this embodiment can achieve low-cost, efficient and accurate recognition of monkey behavior in complex environments.
[0124] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0125] Based on the same inventive concept, embodiments of the present application also provide a monkey behavior recognition device for implementing the aforementioned monkey behavior recognition method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more monkey behavior recognition device embodiments provided below can be found in the aforementioned limitations of the monkey behavior recognition method and will not be further elaborated here.
[0126] In an exemplary embodiment, Figure 10 As shown, a monkey behavior recognition device 1000 is provided, comprising:
[0127] Image acquisition module 1001, used to acquire the image to be identified;
[0128] A feature extraction module 1002 is configured to perform multi-level feature extraction on the image to be identified using a backbone network of a behavior recognition model constructed based on a YOLOv8 network to obtain multiple levels of features; wherein the backbone network uses a C2F module for shallow feature extraction, a cross-stage feature processing module based on a generalized efficient layer aggregation network for deep feature extraction, and a multi-scale pooling module and a self-attention module for multi-scale feature enhancement;
[0129] A feature fusion module 1003 is configured to use the neck network of the behavior recognition model to perform cross-scale feature fusion on the hierarchical features from the backbone network to obtain fused features corresponding to multiple scales; wherein the neck network uses a fusion splicing module to perform feature splicing, a c2f module to perform shallow feature fusion, and the cross-stage feature processing module to perform deep feature fusion;
[0130] The result acquisition module 1004 is used to use the head network of the behavior recognition model to process each of the fusion features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
[0131] In an exemplary embodiment, the hierarchical features include multi-scale features; the feature extraction module 1002 is used to: use the multi-scale pooling module to perform multi-scale pooling processing on the first feature output by the last cross-stage feature processing module in the backbone network to obtain a second feature; use the multi-head self-attention block, the first fully connected layer, the first activation layer, the second fully connected layer and the second activation layer in the self-attention module to process the second feature to obtain the multi-scale feature.
[0132] In an exemplary embodiment, the result acquisition module 1004 is used to: input each of the fused features into the detection head of the corresponding scale in the head network; use each of the detection heads to process the fused features through the compression-excitation attention mechanism and the convolution block attention mechanism to obtain the scale recognition results corresponding to each of the scales; and obtain the monkey behavior recognition result corresponding to the image to be identified based on the scale recognition results corresponding to each of the scales.
[0133] In an exemplary embodiment, the behavior recognition model is trained by the following steps: obtaining a sample set of monkey behavior images, and dividing the sample set of monkey behavior images into a training set and a verification set; using the training set to train the behavior recognition model to be trained to obtain a candidate behavior recognition model; using the verification set to verify the candidate behavior recognition model to obtain the average precision mean of the candidate behavior recognition model; if the average precision mean does not converge, adjusting the hyperparameters of the candidate behavior recognition model, using the candidate behavior recognition model as the behavior recognition model to be trained, and executing the step of training the behavior recognition model to be trained using the training set until a target behavior recognition model that makes the average precision mean converge is obtained; and obtaining the behavior recognition model based on the target behavior recognition model.
[0134] In an exemplary embodiment, the method of using the training set to train the behavior recognition model to be trained to obtain a candidate behavior recognition model includes: using the behavior recognition model to be trained to perform monkey behavior recognition on each monkey behavior image sample in the training set to obtain a monkey behavior prediction result corresponding to each monkey behavior image sample; calculating the classification loss and regression loss of the behavior recognition model to be trained based on the difference between the monkey behavior prediction result corresponding to each monkey behavior image sample and the monkey behavior sample label, and obtaining the model loss of the behavior recognition model to be trained based on the classification loss and the regression loss; adjusting the model parameters of the behavior recognition model to be trained based on the model loss, and executing the step of using the behavior recognition model to be trained to perform monkey behavior recognition on each monkey behavior image sample in the training set until a candidate behavior recognition model that converges the model loss is obtained.
[0135] In an exemplary embodiment, obtaining a set of monkey behavior image samples includes: obtaining original image samples containing monkey objects; performing image transformation processing on at least part of the original image samples to obtain augmented image samples; and obtaining the set of monkey behavior image samples based on the original image samples and the augmented image samples.
[0136] Each module in the monkey behavior recognition device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0137] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 11As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data such as images to be identified, relevant parameters of the behavior recognition model, and monkey behavior image sample sets. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a monkey behavior recognition method is implemented.
[0138] Those skilled in the art will understand that Figure 11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0139] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0140] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0141] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0142] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0143] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0144] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0145] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for identifying monkey behavior, characterized in that: The method comprises: Obtain the image to be recognized; A backbone network of an action recognition model built based on the yolov8 network is used to perform multi-level feature extraction on the image to be identified to obtain multiple levels of features; wherein the backbone network uses a c2f module for shallow feature extraction, a cross-stage feature processing module based on a generalized efficient layer aggregation network for deep feature extraction, and a multi-scale pooling module and a self-attention module for multi-scale feature enhancement; The neck network of the behavior recognition model is used to perform cross-scale feature fusion on the hierarchical features from the backbone network to obtain fused features corresponding to multiple scales; wherein the neck network uses a fusion splicing module to perform feature splicing, a c2f module to perform shallow feature fusion, and the cross-stage feature processing module to perform deep feature fusion; The head network of the behavior recognition model is used to process the fusion features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
2. The method according to claim 1, characterized in that The hierarchical features include multi-scale features; The multi-scale feature enhancement using the multi-scale pooling module and the self-attention module includes: Using the multi-scale pooling module, the first feature output by the last cross-stage feature processing module in the backbone network is subjected to multi-scale pooling processing to obtain a second feature; The multi-scale feature is obtained by processing the second feature using the multi-head self-attention block, the first fully connected layer, the first activation layer, the second fully connected layer, and the second activation layer in the self-attention module.
3. The method according to claim 1, characterized in that The head network of the behavior recognition model is used to process the fusion features through an attention mechanism to obtain a monkey behavior recognition result corresponding to the image to be recognized, including: Inputting each of the fused features into the detection head of the corresponding scale in the head network; Utilizing each of the detection heads, the fused features are processed through a compression-excitation attention mechanism and a convolutional block attention mechanism to obtain scale recognition results corresponding to each of the scales; According to the scale recognition results corresponding to each of the scales, a monkey behavior recognition result corresponding to the image to be recognized is obtained.
4. The method according to any one of claims 1 to 3, characterized in that The behavior recognition model is trained by the following steps: Obtaining a monkey behavior image sample set, and dividing the monkey behavior image sample set into a training set and a validation set; Using the training set to train the behavior recognition model to be trained to obtain a candidate behavior recognition model; Verifying the candidate behavior recognition model using the verification set to obtain an average precision of the candidate behavior recognition model; If the mean average precision does not converge, adjust the hyperparameters of the candidate behavior recognition model, use the candidate behavior recognition model as the behavior recognition model to be trained, and perform the step of training the behavior recognition model to be trained using the training set until a target behavior recognition model is obtained that makes the mean average precision converge; The behavior recognition model is obtained according to the target behavior recognition model.
5. The method according to claim 4, characterized in that The method of training the behavior recognition model to be trained using the training set to obtain a candidate behavior recognition model includes: Performing monkey behavior recognition on each monkey behavior image sample in the training set using the behavior recognition model to be trained to obtain a monkey behavior prediction result corresponding to each monkey behavior image sample; Calculating the classification loss and regression loss of the behavior recognition model to be trained based on the difference between the monkey behavior prediction results corresponding to each of the monkey behavior image samples and the monkey behavior sample labels, and obtaining the model loss of the behavior recognition model to be trained based on the classification loss and the regression loss; The model parameters of the behavior recognition model to be trained are adjusted according to the model loss, and the step of using the behavior recognition model to be trained to perform monkey behavior recognition on each of the monkey behavior image samples in the training set is performed until a candidate behavior recognition model that converges the model loss is obtained.
6. The method according to claim 4, characterized in that The step of obtaining a sample set of monkey behavior images comprises: Obtaining original image samples containing monkey objects; Performing image transformation processing on at least part of the original image samples to obtain augmented image samples; The monkey behavior image sample set is obtained according to the original image samples and the augmented image samples.
7. A monkey behavior recognition device, characterized in that: The device comprises: An image acquisition module, used to acquire an image to be identified; A feature extraction module is used to perform multi-level feature extraction on the image to be identified using the backbone network of the behavior recognition model built based on the yolov8 network to obtain multiple levels of features; wherein the backbone network uses the c2f module for shallow feature extraction, uses the cross-stage feature processing module based on the generalized efficient layer aggregation network for deep feature extraction, and uses the multi-scale pooling module and the self-attention module for multi-scale feature enhancement; A feature fusion module, configured to utilize the neck network of the behavior recognition model to perform cross-scale feature fusion on the hierarchical features from the backbone network to obtain fused features corresponding to multiple scales; wherein the neck network uses a fusion splicing module for feature splicing, a c2f module for shallow feature fusion, and the cross-stage feature processing module for deep feature fusion; The result acquisition module is used to use the head network of the behavior recognition model to process each of the fusion features through the attention mechanism to obtain the monkey behavior recognition result corresponding to the image to be recognized.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.