Animal behavior recognition method and device, animal behavior model training method and device and computer storage medium

By using multi-branch neural networks and training audio data in animal behavior recognition, the problem that traditional methods are difficult to conduct comprehensive monitoring and long-term data collection in wild environments is solved, and animal behavior recognition is achieved that is suitable for both wild and captive environments.

CN120145202AInactive Publication Date: 2025-06-13SICHUAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510622319.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional animal behavior recognition methods rely on manual analysis of video data, making it difficult to conduct comprehensive monitoring and long-term data acquisition in wild environments, limiting its application scenarios.

Method used

By using multi-branched neural networks, the model is trained using training audio data. The recording device is worn on the target animal and can collect audio data in a wild environment and train the model to identify animal behavior.

Benefits of technology

This method enables the animal behavior recognition model to be applicable in both the wild and captive environments, improves the model's expression ability and multi-task processing ability, and expands the application scenarios of animal behavior recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145202A_ABST
    Figure CN120145202A_ABST
Patent Text Reader

Abstract

The invention provides an animal behavior recognition method and device, an animal behavior recognition model training method and device and a computer storage medium. The animal behavior recognition model training method comprises the steps that training data features of training audio data are extracted; the training audio data are multiple pieces of audio data, used for animal behavior recognition model training, of the target animal; inputting the training data features into a multi-branch neural network to train a plurality of branches in the multi-branch neural network; under the condition that the multiple branches meet the training conditions, animal behavior recognition model training is stopped; the animal behavior recognition model is configured to recognize a behavior of a target animal. According to the embodiment of the invention, the multi-branch neural network is trained by adopting the training audio data, and the recording equipment for collecting the audio of the target animal can be worn on the target animal, is not limited by the environment, and is suitable for both field living animals and captive animals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of behavior recognition, and more particularly, to an animal behavior recognition and its model training method, device, and computer storage medium. Background Art

[0002] Traditional animal behavior recognition mainly relies on research experts to manually analyze and label video data to judge animal behaviors, and then obtain information such as their activity rhythms and physical states. Researchers need to combine video materials and have the ability to accurately judge behavior types through specialized training.

[0003] However, the above method is more suitable for captive animals. For the wild environment, it is difficult to deploy all-round monitoring for long-term monitoring. Therefore, data collection is extremely difficult, which limits the use of this animal behavior recognition method. Summary of the Invention

[0004] In view of this, the purpose of the embodiments of this application is to provide an animal behavior recognition and its model training method, device, and computer storage medium, which can increase the application scenarios of the animal recognition model for animal behavior recognition.

[0005] In a first aspect, the embodiments of this application provide an animal behavior recognition model training method, including: extracting training data features of training audio data; where the training audio data is multiple audio data of a target animal for training an animal behavior recognition model; inputting the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network; stopping the training of the animal behavior recognition model when the multiple branches reach the training conditions; where the animal behavior recognition model is configured to recognize the behaviors of the target animal.

[0006] In the above implementation process, by using training audio data to train a multi-branch neural network, since only the audio of the target animal needs to be collected for training audio, the recording device for collecting the audio of the target animal can be worn on the target animal, without being restricted by the environment, and is applicable to both wild animals and captive animals, which can increase the application scenarios of the animal recognition model for animal behavior recognition. In addition, by training multiple branches in the multi-branch neural network, the multiple branches can compete and learn, which can improve the expression ability and multi-task processing ability of the model.

[0007] In one embodiment, the multi-branch neural network includes an overlapping activation penalty; the step of inputting the training data features into the multi-branch neural network to train multiple branches in the multi-branch neural network includes: inputting the training data features into the multi-branch neural network, and adjusting the weight coefficients of the overlapping regions in the multiple branches through the overlapping activation penalty; and training the multiple branches with adjusted weight coefficients of the overlapping regions through the training data features.

[0008] In the above implementation process, by introducing an overlapping activation penalty into the animal behavior recognition model, the overlapping degree between the activation maps of different branches can be calculated, and this overlap can be penalized, forcing the model to learn more diverse and complementary feature representations. Furthermore, the quality and diversity of the features can be improved, and the generation of redundant or circular content can be avoided.

[0009] In one embodiment, the multi-branch neural network includes an imbalance loss penalty; the step of inputting the training data features into the multi-branch neural network to train multiple branches in the multi-branch neural network includes: inputting the training data features into the multi-branch neural network, and adjusting the outputs of the multiple branches through the imbalance loss penalty; and training the multiple branches with adjusted outputs through the training data features.

[0010] In the above implementation process, by introducing an imbalance loss penalty into the animal behavior recognition model, an adjustment or weighting of the loss function can be performed so that the model can better focus on the minority classes, thereby improving the performance of the model on the minority classes, avoiding one branch becoming overly dominant during the competition process, and improving the output accuracy of the model.

[0011] In one embodiment, the multi-branch neural network includes an overlapping activation penalty and an imbalance loss penalty; the step of inputting the training data features into the multi-branch neural network to train multiple branches in the multi-branch neural network includes: inputting the training data features into the multi-branch neural network, and adjusting the weight coefficients of the overlapping regions in the multiple branches through the overlapping activation penalty; and adjusting the outputs of the multiple branches through the imbalance loss penalty; and training the multiple branches with adjusted weight coefficients of the overlapping regions and outputs of the multiple branches through the training data features.

[0012] In the above implementation process, by introducing overlapping activation penalty and imbalance loss penalty in the animal behavior recognition model, not only can the overlap degree between activation maps of different branches be calculated and this overlap be penalized, forcing the model to learn more diverse and complementary feature representations, thereby improving the quality and diversity of features and avoiding the generation of redundant or cyclic content. It is also possible to make an adjustment or weighting of the loss function so that the model can better focus on the minority classes, thus improving the model's performance on the minority classes, avoiding one branch becoming overly dominant during the competition process, and improving the accuracy of the model output.

[0013] In one embodiment, before extracting the training data features of the training audio data, the method further includes: converting the training audio data into a preset format; cropping the training audio data converted into the preset format to obtain one or more training audio segments; screening the training audio segments that meet the annotation requirements among the training audio segments, and marking the training audio segments that meet the annotation requirements; the extracting of the training data features of the training audio data includes: extracting the training data features of the marked training audio segments.

[0014] In the above implementation process, by converting the training audio data into a preset format and performing cropping, the training audio data can be simplified and the training efficiency can be improved. Additionally, by screening out the training audio segments that meet the annotation requirements among the training audio segments, some irrelevant training audio segments and difficult-to-mark audio segments can be filtered out, reducing the invalid audio segments and improving the accuracy of the audio segments.

[0015] In one embodiment, before extracting the training data features of the training audio data, the method further includes: removing the noise in the training audio data; extracting the Mel spectrogram feature map of the training audio data after removing the noise; the extracting of the training data features of the training audio data includes: extracting the training data features in the Mel spectrogram feature map.

[0016] In the above implementation process, by removing the noise in the training audio data before extracting the training data features of the training audio data, the clarity of the training audio data can be improved and the accuracy of model training can be enhanced. Additionally, by extracting the Mel spectrogram feature map of the training audio data after removing the noise, the audio features can be converted into a visual representation, improving the visualization ability of audio processing and enhancing the accuracy of animal action recognition.

[0017] In a second aspect, an embodiment of the present application further provides an animal behavior recognition method, including: extracting data features of audio data of a target animal; identifying the data features through any one of the multiple branches in the multi-branch neural network trained by the method in the first aspect or any possible implementation manner of the first aspect to obtain an identification result of the audio data; and determining the behavior of the target animal according to the classification probability of the identification result.

[0018] In a third aspect, an embodiment of the present application further provides an animal behavior recognition model training device, including: a first extraction module configured to extract training data features of training audio data, where the training audio data is multiple audio data of a target animal for training an animal behavior recognition model; a training module configured to input the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network, and the training module is further configured to stop training of the animal behavior recognition model when the multiple branches meet the training conditions, where the animal behavior recognition model is configured to recognize the behavior of the target animal.

[0019] In a fourth aspect, an embodiment of the present application further provides an animal behavior recognition device, including: a second extraction module configured to extract data features of audio data of a target animal; an identification module configured to identify the data features through any one of the multiple branches in the multi-branch neural network trained by the method in the first aspect or any possible implementation manner of the first aspect to obtain an identification result of the audio data; and a determination module configured to determine the behavior of the target animal according to the classification probability of the identification result.

[0020] In a fifth aspect, an embodiment of the present application further provides an electronic device, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and when the electronic device runs, the machine-readable instructions, when executed by the processor, execute the steps of the method in the first aspect or any possible implementation manner of the first aspect, or the second aspect or any possible implementation manner of the second aspect.

[0021] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is run by a processor, it executes the steps in the first aspect or any possible implementation manner of the first aspect, or the second aspect or any possible implementation manner of the second aspect.

[0022] To make the above objects, features, and advantages of the present application more obvious and understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following accompanying drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related accompanying drawings can also be obtained based on these drawings.

[0024] Figure 1 It is a block diagram of an electronic device provided by an embodiment of the present application; Figure 2 It is a flowchart of a method for training an animal behavior recognition model provided by an embodiment of the application; Figure 3 It is a flowchart of an animal behavior recognition method provided by an embodiment of the present application; Figure 4(a) is a schematic diagram of the visualization results of the accuracy of the baseline and after adopting competitive fusion learning, taking the audio data of 5 behaviors of 5 giant pandas as an example provided by an embodiment of the present application; Figure 4(b) is a schematic diagram of the visualization results of the F1-score of the baseline and after adopting competitive fusion learning, taking the audio data of 5 behaviors of 5 giant pandas as an example provided by an embodiment of the present application; Figure 5 It is a schematic diagram of the functional modules of an animal behavior recognition model training device provided by an embodiment of the present application; Figure 6 It is a schematic diagram of the functional modules of an animal behavior recognition device provided by an embodiment of the present application. Detailed implementation manners

[0025] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0026] It should be noted that similar reference numerals and letters denote similar items in the following accompanying drawings. Therefore, once an item is defined in one accompanying drawing, it does not need to be further defined and explained in subsequent accompanying drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0027] With the progress of captive breeding technology, the return of captive endangered species to the wild has become an increasing focus of biologists. For example, the giant panda. Since 2003, the China Conservation and Research Center for the Giant Panda has carried out a systematic wild training project aimed at preparing captive giant pandas for release into the wild. Selecting suitable individuals for release is crucial for improving their survival rate in the natural environment. In recent years, researchers have invested a great deal of effort in wild training experiments to screen suitable release candidates, and one of the key criteria for selection is to analyze the proportion of different behaviors exhibited by giant pandas within a set monitoring time.

[0028] Traditional giant panda behavior recognition mainly relies on research experts to manually analyze and label videos to judge the behaviors of giant pandas, and then obtain information such as their activity rhythms and physical states. Researchers need to combine video materials and go through specialized training to have the ability to accurately judge behavior types.

[0029] Through long-term research, the present invention has found that the current methods have the following disadvantages: 1. Traditional manual video annotation requires a large number of professional researchers, a great deal of time and financial resources, and the annotation effect is highly subjective. Moreover, long-term listening in a noisy environment has a great impact on the hearing and mental state of researchers.

[0030] 2. The method of automatically recognizing giant panda behaviors using videos is not applicable in the wild. It is difficult to deploy all-round monitoring in the wild for long-term monitoring. Therefore, data collection is extremely difficult, and this method is only applicable to captive environments.

[0031] In view of this, the present application proposes a method for training an animal behavior recognition model. By using training audio data to train a multi-branch neural network, since only the audio of the target animal needs to be collected for training the audio, the recording device for collecting the audio of the target animal can be worn on the target animal, without being restricted by the environment, and is applicable to both wild animals and captive animals, which can increase the application scenarios of the animal recognition model for animal behavior recognition. In addition, by training multiple branches in the multi-branch neural network, multiple branches can compete in learning, which can improve the expression ability and multi-task processing ability of the model.

[0032] To facilitate the understanding of this embodiment, first, an electronic device for implementing a method for training an animal behavior recognition model and / or an animal behavior recognition method disclosed in the present application will be introduced in detail.

[0033] As Figure 1 shown, it is a block diagram of the electronic device. The electronic device 100 may include a memory 111 and a processor 113. Those of ordinary skill in the art can understand that Figure 1 the structure shown Figure 1 is only schematic and does not limit the structure of the electronic device 100. For example, the electronic device 100 may also include more or fewer components than Figure 1 shown, or have a different configuration from

[0034] The above-mentioned memory 111 and processor 113 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The above-mentioned processor 113 is used to execute the executable module stored in the memory.

[0035] Among them, the memory 111 can be, but is not limited to, random access memory (Random Access Memory, referred to as RAM), read-only memory (Read Only Memory, referred to as ROM), programmable read-only memory (Programmable Read-Only Memory, referred to as PROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, referred to as EPROM), electrically erasable programmable read-only memory (Electric Erasable Programmable Read-Only Memory, referred to as EEPROM), etc. Among them, the memory 111 is used to store a program, and after receiving an execution instruction, the processor 113 executes the program. The method executed by the electronic device 100 defined by the process disclosed in any embodiment of the embodiments of the present application can be applied to the processor 113 or implemented by the processor 113.

[0036] The above-mentioned processor 113 may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor 113 can be a general-purpose processor, including a central processing unit (Central Processing Unit, referred to as CPU), a network processor (Network Processor, referred to as NP), etc.; it can also be a digital signal processor (digital signal processor, referred to as DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, referred to as ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0037] Optionally, the animal behavior recognition model training method and the animal behavior recognition method can use the same electronic device or multiple different electronic devices. When the animal behavior recognition model training method and the animal behavior recognition method use different electronic devices, the structures of these different electronic devices may be the same or different. The selection of the electronic devices for the animal behavior recognition model training method and the animal behavior recognition method can be made according to the actual situation.

[0038] The electronic device 100 in this embodiment can be used to execute each step in the various methods provided in the embodiments of the present application. The implementation process of the animal behavior recognition model training method will be described in detail through several embodiments below.

[0039] Please refer to Figure 2 , which is a flowchart of the animal behavior recognition model training method provided in the embodiments of the present application. The specific process shown below will be elaborated in detail. Figure 2 shown will be elaborated in detail.

[0040] Step S201, extract the training data features of the training audio data.

[0041] Among them, the training audio data is multiple audio data for training the animal behavior recognition model of the target animal.

[0042] The target animal here can be a panda, dog, wild boar, monkey, elk, etc., and the target animal can be selected according to the actual situation.

[0043] The above-mentioned training audio data can be obtained by wearing a recording device on the body of the target animal. For example, a recording pen, microphone, sound card, etc. These recording devices can be worn on the neck, head, or limbs of the target animal through ropes or chains, and the specific type and wearing position of the recording device can be adjusted according to the actual situation.

[0044] It can be understood that since the audio data of the target animal is collected and the recording device can be worn on the body of the target animal, the corresponding audio data of the target animal in the wild can also be well collected, which can better adapt to the wild environment and realize the action recognition of wild animals.

[0045] In one embodiment, the training audio data can be audio data continuously collected from multiple target animals for multiple days.

[0046] The above-mentioned training data features can be extracted by a neural network. For example, the ResNet50 backbone neural network.

[0047] Optionally, the training data features can include Mel spectrogram, pitch, timbre, etc., and the training data features can be selected according to the actual situation.

[0048] Step S202: Input the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network.

[0049] The multi-branch neural network here refers to a neural network structure including multiple branch structures. These multiple branch structures can be the same or different. The multiple branch structures in the multi-branch neural network can be selected according to the actual situation.

[0050] Among them, during the model training process, multiple branches in the multi-branch neural network compete in learning. After the model training is completed, multiple branches in the multi-branch neural network are almost the same.

[0051] Optionally, the multi-branch neural network includes one or more of overlapping activation penalty and imbalance loss penalty, and the penalties included in the multi-branch neural network can be selected according to the actual situation.

[0052] Step S203: Stop the training of the animal behavior recognition model when multiple branches meet the training conditions.

[0053] The training conditions here can be that all training data features are trained, or the number of training times is reached, or the recognition accuracy of the trained model reaches the preset requirements, etc. The training conditions can be selected according to the actual situation.

[0054] Among them, the animal behavior recognition model is configured to recognize the behavior of the target animal.

[0055] The above multi-branch neural network is set in the animal behavior recognition model. When recognizing the behavior of the target animal, it is recognized through any one of the branches in the animal behavior recognition model.

[0056] The outputs of multiple branches in the animal behavior recognition model here are the same.

[0057] Optionally, the behaviors of the target animal can include eating, resting, moving, nursing, and drinking water, etc., and the behaviors of the target animal can be selected according to the actual situation.

[0058] In the above implementation process, by using training audio data to train the multi-branch neural network, since only the audio of the target animal needs to be collected for training audio, the recording device for collecting the audio of the target animal can be worn on the target animal, without being restricted by the environment, and it is applicable to both wild animals and captive animals, which can increase the application scenarios of the animal recognition model for animal behavior recognition. In addition, by training multiple branches in the multi-branch neural network, multiple branches can compete in learning, which can improve the expression ability and multi-task processing ability of the model.

[0059] In a possible implementation, step S202 includes: inputting the training data features into a multi-branch neural network, adjusting the weight coefficients of the overlapping regions in multiple branches through overlapping activation penalty; and training the multiple branches with the adjusted weight coefficients of the overlapping regions using the training data features.

[0060] Here, the overlapping activation penalty refers to a penalty mechanism imposed on the generated vocabulary or structures in order to suppress the appearance of repeated vocabulary or phrases when the model generates text or sequences. Its core purpose is to improve the quality and diversity of the generated features and avoid generating redundant or cyclic content.

[0061] In a neural network, different branches or layers may produce highly similar activation responses to certain regions of the input image, that is, there is an overlap in the activation regions. The overlapping activation penalty forces the model to learn more diverse and complementary feature representations by calculating the degree of overlap between the activation maps of different branches and penalizing this overlap. For example, in an animal behavior recognition task, if multiple branches all overly focus on a certain local region of the animal, such as the pitch feature, while ignoring other important features, then the features extracted by these branches may have high similarity, which is not conducive to accurately distinguishing different animal individuals. Through the overlapping activation penalty, different branches can be more inclined to focus on different features of the animal.

[0062] It can be understood that since the goal of the animal behavior recognition model in the embodiments of this application is to encourage multiple branches to compete with each other and extract richer class-related features. Therefore, the goal of this animal behavior recognition model is to minimize the overlap between high-response regions, which forces multiple branches to pay more attention to low-response regions.

[0063] In one embodiment, the overlapping activation penalty can be expressed as: ; where ⊙ represents element-wise multiplication, C represents the number of categories, is the overlapping activation penalty, represents the variable generated by the first branch, represents the variable generated by the second branch.

[0064] In the above implementation process, by introducing the overlapping activation penalty in the animal behavior recognition model, the model can be forced to learn more diverse and complementary feature representations by calculating the degree of overlap between the activation maps of different branches and penalizing this overlap. Furthermore, the quality and diversity of the features can be improved, and redundant or cyclic content can be avoided.

[0065] In a possible implementation, step S202 includes: inputting the training data features into a multi-branch neural network, adjusting the outputs of multiple branches through an imbalance loss penalty; and training the multiple branches after adjusting the outputs of the multiple branches with the training data features.

[0066] Here, the imbalance loss penalty refers to an adjustment or weighting of the loss function during training to address the class imbalance problem, so that the model can better focus on the minority classes (usually positive samples or rare classes), thereby improving the performance of the model on the minority classes. Its core purpose is to balance the model's attention to different classes and prevent the model from being overly biased towards the majority classes (usually negative samples or common classes).

[0067] It can be understood that since the goal of the animal behavior recognition model in the embodiments of the present application is to prevent one branch from becoming overly dominant during the competition while the other branch remains completely inactive or only slightly active, the embodiments of the present application constrain the final outputs of multiple branches by adding an imbalance loss penalty to the model.

[0068] In one embodiment, the imbalance loss penalty can be expressed as: ; where C represents the number of classes, is the overlapping activation penalty, represents the variable generated by the first branch, represents the variable generated by the second branch.

[0069] In the above implementation process, by introducing an imbalance loss penalty into the animal behavior recognition model, an adjustment or weighting of the loss function can be performed so that the model can better focus on the minority classes, thereby improving the performance of the model on the minority classes, preventing one branch from becoming overly dominant during the competition, and improving the accuracy of the model output.

[0070] In a possible implementation, step S202 includes: inputting the training data features into a multi-branch neural network, adjusting the weight coefficients of the overlapping regions in multiple branches through an overlapping activation penalty; and adjusting the outputs of multiple branches through an imbalance loss penalty; training the multiple branches after adjusting the weight coefficients of the overlapping regions and the outputs of multiple branches with the training data features.

[0071] It can be understood that both an overlapping activation penalty and an imbalance loss penalty can be added to the animal behavior recognition model, so that redundant or cyclic content is avoided in the animal behavior recognition model, and at the same time, one branch is prevented from becoming overly dominant during the competition, increasing the accuracy of the model output.

[0072] The overlapping activation penalty and the imbalance loss penalty here can jointly train a neural network model under the combined action of the cross-entropy loss function. Its expression can be shown as follows: ; Among them, is to guide the model to learn regions related to categories, is the overlapping activation penalty, 𝞴 is the balance weight (gradually decreasing from 1 to 0 with the number of training rounds, and gradually decreasing from 1 to 0 with the number of training rounds, so as to achieve strong competition in the early stage and reduced competition in the late stage), is the imbalance loss penalty.

[0073] It should be understood that through the training of the above embodiments, in the test stage of the animal behavior recognition model obtained, since the outputs of each branch are almost the same, that is, only the result of one branch needs to be used as the total output result, and finally the probability of each category is output through the activation function for classification, without adding additional computational consumption.

[0074] Each branch structure in the multi-branch neural network here generates its own class activation map and class score.

[0075] In the above implementation process, by introducing the overlapping activation penalty and the imbalance loss penalty in the animal behavior recognition model, not only can the overlapping degree between the activation maps of different branches be calculated and this overlap be punished, forcing the model to learn more diverse and complementary feature representations, thereby improving the quality and diversity of features and avoiding generating redundant or circular content. It can also be a kind of adjustment or weighting of the loss function so that the model can better focus on the minority classes, thereby improving the performance of the model on the minority classes, avoiding one branch becoming overly dominant in the competition process, and improving the output accuracy of the model.

[0076] In a possible implementation manner, before step S201, the method further includes: converting the training audio data into a preset format; cropping the training audio data converted into the preset format to obtain one or more training audio segments; screening the training audio segments that meet the annotation requirements among the training audio segments, and marking the training audio segments that meet the annotation requirements.

[0077] The preset format here can include: preset channels, preset audio, etc. For example, the preset format is a mono-channel, 44100Hz audio format.

[0078] Optionally, converting the training audio data into the preset format can be achieved through audio editing software, audio conversion tools, etc., and the way of converting the training audio data into the preset format can be selected according to the actual situation.

[0079] The above-mentioned training audio segments can be audio segments of equal length. For example, segments of 1 minute in length each.

[0080] Of course, the training audio segments can also be audio segments of different lengths, and the training audio segments can be selected according to the actual situation.

[0081] It can be understood that after the training audio data is cropped into one or more training audio segments, some irrelevant or difficult-to-label segments can be screened out. Then, the remaining training audio data is labeled. For example, labels such as eating, resting, moving, feeding, and drinking can be performed, and then multiple audio samples are obtained to form a target animal behavior audio data set.

[0082] In one embodiment, step S201 includes: extracting the training data features of the labeled training audio segments.

[0083] In the above implementation process, by converting the training audio data into a preset format and performing cropping, the training audio data can be simplified and the training efficiency can be improved. In addition, by screening out the training audio segments that meet the labeling requirements in the training audio segments, some irrelevant training audio segments and difficult-to-label audio segments can be screened out, reducing the invalid audio segments and improving the accuracy of the audio segments.

[0084] In a possible implementation manner, before step S201, the method further includes: removing the noise in the training audio data; extracting the Mel spectrogram feature map of the training audio data after removing the noise.

[0085] Optionally, the noise in the training audio data can be removed by methods such as spectral subtraction, Wiener filtering, low-pass filtering, and high-pass filtering, and the method for removing the noise in the audio data can be selected according to the actual situation.

[0086] In one embodiment, a non-stationary noise reduction algorithm can be used to remove the noise in the training audio data.

[0087] The Mel spectrogram feature map here is a visual representation based on Mel frequency cepstral coefficients. It decomposes the audio signal on the Mel frequency scale to obtain the energy distribution of each Mel frequency band, thereby more intuitively displaying the frequency characteristics of the audio signal.

[0088] Optionally, the Mel spectrogram feature map can be extracted by methods such as Fourier transform method, deep learning method, and method based on open source tools, and the extraction method of the Mel spectrogram feature map can be selected according to the actual situation.

[0089] In one embodiment, multiple important parameters can be set when generating the Mel spectrogram feature map. For example, "n_fft", "hop_length", and "n_mel". Among them, 'n_fft' represents the length of the FFT window, which is 0.16 s. "hop_length" refers to the frame shift, which is 0.6 s. And "n_mel" refers to the number of Mel filters, which is set to 64. In this way, a 1-minute long audio sample is converted into a tensor of dimension (1, 64, 1001).

[0090] It should be understood that in this method for training an animal behavior recognition model, a leave-one-out cross-validation strategy is adopted. Each time, the data of one individual is selected as the test set, and the experiment is repeated until all individuals have been tested. For example, if there are a total of 5 individuals, then five cross-validations are performed. The batch size can be set to 16, the learning rate is initialized to 10^−4, and the learning rate optimization adopts a cosine annealing strategy. Each time the model is trained, it can be for 40 epochs (which refers to the number of times all data samples are iterated through in one training process).

[0091] In one embodiment, step S201 includes: extracting the training data features from the Mel spectrogram feature map.

[0092] In the above implementation process, before extracting the training data features from the training audio data, removing the noise in the training audio data can improve the clarity of the training audio data and the accuracy of model training. Additionally, by extracting the Mel spectrogram feature map of the noise-removed training audio data, the audio features can be converted into a visual representation, improving the visualization ability of audio processing and the accuracy of animal action recognition.

[0093] Please refer to Figure 3 , which is the flowchart of the animal behavior recognition method provided by the embodiments of this application. The following will elaborate in detail on Figure 3 the specific process shown.

[0094] Step S301: Extract the data features of the audio data of the target animal.

[0095] The data features here can be extracted through a neural network. For example, a ResNet50 backbone neural network.

[0096] Optionally, the data features can include Mel spectrograms, pitch, timbre, etc., and the data features can be selected according to the actual situation.

[0097] Step S302: Use any one of the multiple branches in the multi-branch neural network trained by the method in the above embodiment to recognize the data features and obtain the recognition result of the audio data.

[0098] The outputs of multiple branches in the multi-branch neural network here are basically the same.

[0099] Step S303: Determine the behavior of the target animal according to the classification probability of the recognition result.

[0100] To facilitate the understanding of the embodiments of the present application, the following takes giant pandas with 5 individuals and 5 behaviors as an example to demonstrate the advantages of animal action recognition in the embodiments of the present application: Among them, the training data includes audio data of 5 giant pandas with 5 behaviors. As shown in Table 1, the method in the embodiments of the present application achieves an average accuracy rate of 92.61%, and its standard deviation is only 2.73%. Compared with other methods, it has a higher accuracy rate and a lower standard deviation. Therefore, it can be shown that the recognition performance of the animal behavior recognition method in the embodiments of the present application is better and has little fluctuation among different individuals.

[0101]

[0102] In addition, by comparing the baseline and the visualization results after adopting competitive fusion learning, as shown in FIGS. 4(a) and 4(b), it can be clearly seen from FIGS. 4(a) and 4(b) that the model trained by the animal behavior recognition model training method in the embodiments of the present application can obtain richer features.

[0103] In the above implementation process, by using audio data analysis to target animal behavior, the recording device for collecting the audio of the target animal can be worn on the target animal, without being restricted by the environment, and is applicable to both wild animals and captive animals, which can increase the application scenarios of the animal recognition model for animal behavior recognition. In addition, by recognizing animal behavior in the trained multi-branch neural network, the accuracy of animal behavior recognition can be improved.

[0104] Based on the same inventive concept, an animal behavior recognition model training device corresponding to the animal behavior recognition model training method is also provided in the embodiments of the present application. Since the principle of solving problems by the device in the embodiments of the present application is similar to that of the foregoing animal behavior recognition model training method embodiments, the implementation of the device in this embodiment can refer to the description in the embodiments of the above method, and the repeated parts will not be elaborated.

[0105] Please refer to Figure 5 , which is a schematic diagram of the functional modules of the animal behavior recognition model training device provided by the embodiments of the present application. Each module in the animal behavior recognition model training device in this embodiment is used to execute each step in the above method embodiments. The animal behavior recognition model training device includes a first extraction module 401 and a training module 402; among them, The first extraction module 401 is used to extract the training data features of the training audio data; wherein, the training audio data is a plurality of audio data for training an animal behavior recognition model of a target animal.

[0106] The training module 402 is used to input the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network; wherein, the training module is further used to stop the training of the animal behavior recognition model when the multiple branches reach the training conditions; wherein, the animal behavior recognition model is configured to recognize the behavior of the target animal.

[0107] In a possible implementation manner, the training module 402 is further used to: input the training data features into a multi-branch neural network, and adjust the weight coefficients of the overlapping regions in the multiple branches through the overlapping activation penalty; and train the multiple branches with the adjusted weight coefficients of the overlapping regions through the training data features.

[0108] In a possible implementation manner, the training module 402 is further used to: input the training data features into a multi-branch neural network, and adjust the outputs of the multiple branches through the unbalanced loss penalty; and train the multiple branches with the adjusted outputs of the multiple branches through the training data features.

[0109] In a possible implementation manner, the training module 402 is further used to: input the training data features into a multi-branch neural network, and adjust the weight coefficients of the overlapping regions in the multiple branches through the overlapping activation penalty; and adjust the outputs of the multiple branches through the unbalanced loss penalty; and train the multiple branches with the adjusted weight coefficients of the overlapping regions and the outputs of the multiple branches through the training data features.

[0110] In a possible implementation manner, the animal behavior recognition model training device further includes a first preprocessing module, configured to convert the training audio data into a preset format; crop the training audio data converted into the preset format to obtain one or more training audio segments; screen the training audio segments that meet the annotation requirements among the training audio segments, and mark the training audio segments that meet the annotation requirements. In a possible implementation manner, the first extraction module 401 is specifically used to: extract the training data features of the marked training audio segments.

[0111] In a possible implementation manner, the animal behavior recognition model training device further includes a second preprocessing module, configured to remove the noise in the training audio data; and extract the Mel spectrogram feature map of the training audio data after removing the noise.

[0112] In a possible implementation manner, the first extraction module 401 is specifically configured to: extract the training data features from the Mel spectrogram feature map.

[0113] Based on the same inventive concept, an animal behavior recognition device corresponding to the animal behavior recognition method is further provided in the embodiments of the present application. Since the principle of solving problems by the device in the embodiments of the present application is similar to that of the foregoing embodiments of the animal behavior recognition method, the implementation of the device in this embodiment can refer to the description in the embodiments of the above method, and repeated parts will not be described again.

[0114] Please refer to Figure 6 , which is a schematic diagram of the functional modules of the animal behavior recognition device provided in the embodiments of the present application. Each module in the animal behavior recognition device in this embodiment is used to execute each step in the above method embodiments. The animal behavior recognition device includes a second extraction module 501, a recognition module 502, and a determination module 503; wherein, The second extraction module 501 is used to extract the data features of the audio data of the target animal.

[0115] The recognition module 502 is used to recognize the data features through any one of the multiple branches in the multi-branch neural network trained by the method in the above embodiments, and obtain the recognition result of the audio data.

[0116] The determination module 503 is used to determine the behavior of the target animal according to the classification probability of the recognition result.

[0117] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the animal behavior recognition model training method and / or the animal behavior recognition method described in the above method embodiments.

[0118] The computer program product of the animal behavior recognition model training method and / or the animal behavior recognition method provided in the embodiments of the present application includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the steps of the animal behavior recognition model training method and / or the animal behavior recognition method described in the above method embodiments. Specifically, it can refer to the above method embodiments, which will not be described again here.

[0119] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0120] In addition, the functional modules in each embodiment of this application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0121] When the above-mentioned functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "comprising..." do not exclude the presence of additional identical elements in the process, method, article or device comprising the said elements. The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0122] As described above, this is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or replacements, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A method for training an animal behavior recognition model, characterized in that: include: Extracting training data features of training audio data; wherein the training audio data is a plurality of audio data of a target animal used for training an animal behavior recognition model; Inputting the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network; When the plurality of branches meet the training conditions, stopping the animal behavior recognition model training; Wherein, the animal behavior recognition model is configured to recognize the behavior of the target animal.

2. The method according to claim 1, characterized in that in, The multi-branch neural network includes an overlapping activation penalty; The step of inputting the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network includes: Inputting the training data features into a multi-branch neural network, and adjusting the weight coefficients of the overlapping regions in the multiple branches by the overlapping activation penalty; The plurality of branches after adjusting the weight coefficients of the overlapping areas are trained by using the training data features.

3. The method according to claim 1, characterized in that in, Introducing an imbalance loss penalty into the multi-branch neural network; The step of inputting the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network includes: Inputting the training data features into a multi-branch neural network, and adjusting the outputs of the multiple branches by the imbalance loss penalty; The multiple branches after adjusting the outputs of the multiple branches are trained using the training data features.

4. The method according to claim 1, characterized in that in, The multi-branch neural network includes an overlapping activation penalty and an imbalance loss penalty; The step of inputting the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network includes: Inputting the training data features into a multi-branch neural network, and adjusting the weight coefficients of the overlapping regions in the multiple branches by the overlapping activation penalty; and Adjusting the outputs of the plurality of branches by the imbalance loss penalty; The plurality of branches after adjusting the weight coefficients of the overlapping areas and the outputs of the plurality of branches are trained through the training data features.

5. The method according to any one of claims 1 to 4, characterized in that: Before extracting the training data features of the training audio data, the method further includes: Converting the training audio data into a preset format; Cut and convert the training audio data into a preset format to obtain one or more training audio segments; Screening the training audio segments that meet the labeling requirements in the training audio segments, and marking the training audio segments that meet the labeling requirements; The step of extracting training data features of training audio data comprises: Extract training data features of the labeled training audio segments.

6. The method according to any one of claims 1 to 4, characterized in that: Before extracting the training data features of the training audio data, the method further includes: Removing noise from the training audio data; Extract the Mel-spectrogram feature map of the training audio data after removing the noise; The step of extracting training data features of training audio data comprises: Extract training data features from the mel-spectrogram feature map.

7. A method for identifying animal behavior, characterized in that: extracting data features of audio data of a target animal; Recognize the data feature by any one of the multiple branches in the multi-branch neural network trained by the method according to any one of claims 1 to 6, and obtain the recognition result of the audio data; The behavior of the target animal is determined according to the classification probability of the recognition result.

8. An animal behavior recognition model training device, characterized in that: include: A first extraction module is used to extract training data features of training audio data; wherein the training audio data is a plurality of audio data of a target animal used for training an animal behavior recognition model; A training module, used for inputting the training data features into a multi-branch neural network to train multiple branches in the multi-branch neural network; Wherein, the training module is also used to stop the animal behavior recognition model training when the multiple branches meet the training conditions; wherein, the animal behavior recognition model is configured to recognize the behavior of the target animal.

9. An animal behavior recognition device, characterized in that: include: A second extraction module, for extracting data features of the audio data of the target animal; A recognition module, configured to recognize the data feature by any one of the multiple branches in the multi-branch neural network trained by the method according to any one of claims 1 to 6, and obtain a recognition result of the audio data; A determination module is used to determine the behavior of the target animal according to the classification probability of the recognition result.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the method according to any one of claims 1 to 6 and / or the steps of the method according to claim 7.

Citation Information

Patent Citations

  • Model training method, animal behavior recognition method, device and equipment

    CN114299551A

  • Industrial image defect detection method based on multi-head unbalanced semi-supervised network

    CN116630696A