Unbiased animal behavior recognition method and system based on multi-scale feature fusion

Through the multi-scale feature fusion network and the classifier calibration module driven by neural collapse, the problem of suboptimal sampling frequency and unbalanced categories in animal behavior recognition is solved, and the accurate identification of various behaviors is achieved, especially improving the accuracy of a few categories.

CN120260142AActive Publication Date: 2025-07-04HANGZHOU DIANZI UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510748275.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

Existing deep learning models have problems of suboptimal sampling frequency and category imbalance in animal behavior recognition, resulting in low accuracy of specific behavior recognition, especially insufficient accuracy of a few categories.

Method used

A multi-scale feature fusion network combined with a classifier calibration module driven by neural collapse is adopted to construct an unbiased animal behavior recognition system by adaptively fusion of different sampling frequency data and dynamic calibration of classifier weights.

Benefits of technology

It improves the accuracy of identification of various behaviors, especially the recognition performance of a few categories, and improves the classification effect of the model under the condition of category imbalance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260142A_ABST
    Figure CN120260142A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of animal behavior detection, and particularly relates to an unbiased animal behavior recognition method and system based on multi-scale feature fusion. The method comprises the following steps: S1, adaptively fusing data of different sampling frequencies through a multi-scale feature fusion module based on a hybrid expert system so as to capture specific characterization of various unbiased animal behaviors; s2, maximizing an included angle between weight vectors of different categories of classifiers through a classifier calibration module driven by neural collapse, and completing construction of an unbiased animal behavior classification model for realizing unbiased animal behavior classification under category imbalance; and S3, optimizing the unbiased animal behavior classification model by using class equilibrium focus loss as a loss function. According to the method, the multi-scale sampling feature fusion network can be constructed, and a classifier dynamic calibration mechanism is combined, so that various animal behaviors can be accurately distinguished.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of animal behavior detection, and particularly relates to an unbiased animal behavior recognition method and system based on multi-scale feature fusion. Background Art

[0002] Deep learning is driving the rapid development of refined animal activity recognition (AAR) technology based on wearable sensors. This technology can achieve real-time monitoring of animal behavior and early detection of diseases, thereby improving livestock management and animal welfare. Research shows that deep learning models perform well in distinguishing various animal behaviors. For example, some researchers developed a multiplayer perceptron (MLP) based on sensor data to identify behaviors such as grazing, ruminating, and resting of cows, with an accuracy rate of over 90%. Currently, convolutional neural networks (CNNs) are the most widely used models in AAR tasks, and their accuracy rate in identifying various animal behaviors also exceeds 90%, such as identifying the standing and walking of horses, the ruminating and salt supplementation of cows, the grazing and activity level of sheep, and the eating and nursing behaviors of pigs. In addition, due to the advantage of recurrent neural networks (RNNs) in processing time series information such as sensor data, their applications are becoming increasingly widespread. Most studies show that by combining RNNs and CNNs to construct a hybrid model, the behavior recognition performance can be effectively improved, and the effect is significantly better than that of a single network structure.

[0003] Although significant progress has been made in AAR tasks using deep learning, there are still problems with low classification accuracy for certain specific behavior categories in practical applications. This problem mainly stems from the classification confusion between different animal behaviors, and its root causes can be attributed to the following two aspects: One is the suboptimal sampling frequency strategy. Existing methods usually adopt a unified sampling frequency to optimize the overall classification performance. This "one-size-fits-all" strategy ignores the significant differences in time scales among different behavior patterns and cannot meet the specific sampling frequency requirements of various behaviors. For example, some researchers compared the sheep activity recognition performance based on triaxial accelerometer and gyroscope data at different sampling frequencies (8 Hz, 16 Hz, and 32 Hz) and found that the 32 Hz sampling frequency could obtain the optimal classification effect, with an overall accuracy rate of 95%. Other researchers evaluated the influence of different sampling frequencies (100 Hz, 50 Hz, 25 Hz, and 12.5 Hz) on horse behavior classification using triaxial acceleration and angular velocity data, and the results showed that the 25 Hz sampling frequency brought the best overall performance.

[0004] Second, there is the problem of class imbalance. Classifiers trained on class-imbalanced datasets tend to be biased towards the majority class, resulting in low accuracy for minority behavior classes. For example, some researchers used an MLP network to identify five different daily behaviors of cows, with an overall accuracy of 97.75%. However, the accuracies of the "walking" and "drinking" behaviors were only 67.69% and 48.08% respectively. This is mainly because the sample sizes of these two behaviors only account for 7.61% and 4.97% of the overall. Similarly, some researchers also previously observed a similar phenomenon in the study of horse behavior recognition, where the overall accuracy was 93.37%, but the accuracy of the "natural walking" behavior, which is a minority class, was only 24.75%. For this class imbalance problem, existing research either alleviates it by balancing the class distribution through resampling techniques or by adjusting the loss function to increase the penalty for minority classes. However, these methods focus more on improving the training effect of the classifier by optimizing external factors, while ignoring the internal structural characteristics that a well-trained classifier should possess under class-balanced conditions.

[0005] To solve the above problems, the present invention develops an unbiased animal behavior recognition method based on multi-scale feature fusion, which enhances the model's recognition ability for each specific behavior by simultaneously customizing features and calibrating the classifier. Summary of the Invention

[0006] The present invention aims to overcome the problems in the prior art, such as insufficient feature extraction (suboptimal sampling frequency) and unbalanced classification performance in specific behavior recognition. It provides an unbiased animal behavior recognition method and system based on multi-scale feature fusion that can accurately distinguish various behaviors by constructing a multi-scale sampling feature fusion network and combining a classifier dynamic calibration mechanism.

[0007] To achieve the above invention objective, the present invention adopts the following technical solutions: An unbiased animal behavior recognition method based on multi-scale feature fusion, comprising the following steps: S1, through a multi-scale feature fusion module based on a mixture of experts system, adaptively fuse data with different sampling frequencies to capture the specific characteristics of various behaviors; S2, through a classifier calibration module driven by neural collapse, maximize the angle between the weight vectors of different class classifiers to complete the construction of an unbiased animal behavior classification model for achieving unbiased animal behavior classification under class imbalance; S3, use class-balanced focal loss as the loss function to optimize the unbiased animal behavior classification model.

[0008] Preferably, step S1 includes the following steps: S11, collect corresponding data respectively at , , ; , , ; Among them, setting the values of the three sampling frequencies follows ; S12, input , , into their respective feature extractors to capture features at different scales , , ; Among them and respectively represent the number of feature channels and the dimension of the coordinate axis, , , refers to the time dimension corresponding to the features obtained on each branch, and ; S13, perform global average pooling operations on the features extracted from the three branches respectively to generate three feature vectors, namely , and ; S14, input the three generated feature vectors into a router-based soft weighted fusion layer for adaptive fusion.

[0009] Preferably, step S14 includes the following steps: S141, generate a combined feature vector by calculating the average value of the three feature vectors; S142, process the generated combined feature vector through a two-layer multi-layer perceptron MLP to obtain a logical value ; (1); S143, perform SoftMax operation on the obtained logical value to obtain the contribution rates corresponding to the features of the three different branches, denoted as : (2); Among them, the parameter represents the participation degree of features at different sampling frequencies; S144, respectively adopt three projection layers, namely , to align the feature vectors into a unified embedding space and obtain their respective projected features ; S145, set the final fusion feature Expressed as a weighted sum of each projection feature, i.e.: (3); where the weight is the corresponding contribution rate .

[0010] Preferably, in step S2, the classifier calibration module driven by neural collapse includes an isometric tight frame classifier with fixed parameters and a fully connected layer classifier with learnable parameters; the isometric tight frame classifier with fixed parameters specifically includes the following process: S211, randomly synthesize classifier vectors at the start of training and ensure that the classifier vectors conform to the simplex isometric tight frame property, specifically as follows: Set a set containing vectors , where represents the feature dimension, represents the number of classes, and satisfies the condition ; if the following conditions are met, this vector set is regarded as a simplex isometric tight frame: (4); where allows rotation and satisfies ; is the identity matrix, represents a vector of length and all values are 1; all vectors within the simplex isometric tight frame have equal norms and the same pairwise angles, i.e.: (5); (6); where the pairwise angles of represent the maximum isometric separation between S212, considering that the normalized class feature prototype represents the classifier weight vectors of each class, construct the normalized vector as the fixed weight matrix of the isometric tight frame classifier and keep it unchanged throughout the training process; S213, project the fusion feature obtained from equation (3) through the projection layer to obtain the feature ; set for dimension adjustment to reach the neural collapse state in the classifier calibration module later; S214, the feature Normalize to : (7); S215, multiply with the classifier weights predefined in step S212 to obtain a logical value : (8); wherein, represents a learnable temperature coefficient for scaling the feature product , and both and have a finite range. represents the category corresponding to the obtained logical value, wherein .

[0011] Preferably, the parameter-learnable fully connected layer classifier specifically includes the following process: S221, randomly initialize the parameters of the fully connected layer classifier and allow these parameters to be normally trained during the training process; S222, represent the weights of each category classifier as for processing the fused feature to generate a logical value , that is: (9); Preferably, step S2 further includes the following steps: S23, linearly combine the logical value output by the isometric tight frame classifier with fixed parameters and the logical value output by the parameter-learnable fully connected layer classifier to obtain a combined logical value : (10); wherein, is used to regulate the fixed degree of the parameters in the classifier during the classification stage.

[0012] Preferably, step S3 includes the following steps: S31, set to represent the finally output logical value, then the class balance focal loss is expressed as: (11); (12); wherein, represents the category The number of samples; is the sample size adjustment factor, used to control the growth rate of the effective sample size as increases; is the weight smoothing factor, used to adjust the weight decay rate of easily classified samples.

[0013] The present invention also provides an unbiased animal behavior recognition system based on multi-scale feature fusion, including; A feature extraction module, used to adaptively fuse data with different sampling frequencies through a multi-scale feature fusion module based on a mixture of experts system, so as to capture the specific characteristics of various behaviors; A behavior classification module, used to maximize the angle between the weight vectors of different category classifiers through a classifier calibration module driven by neural collapse, complete the construction of an unbiased animal behavior classification model, so as to achieve unbiased animal behavior classification under class imbalance; A model optimization module, used to optimize the unbiased animal behavior classification model by using class-balanced focal loss as the loss function.

[0014] Compared with the prior art, the beneficial effects of the present invention are: (1) The present invention proposes an unbiased animal behavior recognition method based on multi-scale feature fusion, and for the first time turns the research focus to the recognition accuracy of a single specific behavior. By simultaneously optimizing the dual mechanisms of feature customization and classifier calibration, it breaks through the limitation of traditional methods that only focus on the overall behavior classification accuracy; (2) Aiming at the different sampling frequency requirements of different behaviors, the present invention designs a multi-scale feature fusion module based on a mixture of experts system; this module realizes the adaptive selection and fusion of data with different sampling frequencies by constructing a multi-branch feature extraction network; different from the prior art, the method of the present invention does not need to rely on prior behavior labels to determine the optimal sampling frequency combination, and effectively captures the specific characteristics of various behaviors through the optimization criterion of maximizing the feature expression ability, and solves the technical bottleneck that it is difficult for traditional single sampling frequency to adapt to the extraction of diverse behavior characteristics; (3) To address the classifier bias problem caused by class imbalance, the present invention proposes a classifier calibration module driven by neural collapse; this module introduces an equiangular tight frame constraint during the classifier training process to maximize the angle between the weight vectors of different category classifiers, significantly improving the inter-class separation; this design draws on the ideal internal structure characteristics of the classifier under balanced conditions, alleviates the bias phenomenon of the classifier under class imbalance, and thus improves the recognition accuracy of minority behavior categories. Brief Description of the Drawings

[0015] Figure 1 is a flowchart of an unbiased animal behavior recognition method based on multi-scale feature fusion in the present invention; Figure 2Schematic diagram of an architecture of the soft weighted fusion layer based on a router in the present invention; Figure 3 Schematic diagram of the distribution of activity categories of the sheep dataset (a) and the cattle dataset (b) provided by an embodiment of the present invention; Figure 4 Comparison diagram of confusion matrices corresponding to three different variants (unbiased animal behavior recognition method based on multi-scale feature fusion (c), variant without the multi-scale feature fusion module based on a mixture of experts system (b), and variant without both the multi-scale feature fusion module based on a mixture of experts system and the classifier calibration module driven by neural collapse (a)); Figure 5 t-SNE visualization diagram of single-sampling-frequency features and multi-sampling-frequency fusion features; Figure 6 Comparison diagram of the influence of the classifier calibration module driven by neural collapse on the cosine similarity distribution between weight vectors of various categories. Detailed implementation manners

[0016] To more clearly illustrate the embodiments of the present invention, the specific implementation manners of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts, and other implementation manners can also be obtained.

[0017] As Figure 1 and Figure 2 shown, the present invention provides an unbiased animal behavior recognition method based on multi-scale feature fusion, which specifically includes the following steps: Step 1, feature extraction stage: adaptively fuse data with different sampling frequencies through a multi-scale feature fusion module based on a mixture of experts system to capture the specific representations of various behaviors.

[0018] Step 2, behavior classification stage: through a classifier calibration module driven by neural collapse, maximize the angle between the weight vectors of different category classifiers, and complete the construction of an unbiased animal behavior classification model to achieve unbiased animal behavior classification under class imbalance.

[0019] Step 3, model optimization stage: use class-balanced focal loss as the loss function to optimize the unbiased animal behavior classification model.

[0020] For step 1, it specifically includes the following steps: 1-1, at sampling frequencies , , collect corresponding data , , ; Among them, setting the values of three sampling frequencies follows the relationship; 1 - 2, input , , into their respective feature extractors to capture features at different scales , , ; Among them and represent the number of feature channels and the dimension of the coordinate axis respectively, , , refers to the time dimension corresponding to the features obtained on each branch, and ; 1 - 3, perform global average pooling operations on the features extracted from the three branches respectively to generate three feature vectors, namely , and ; 1 - 4, input the three generated feature vectors into a router - based soft - weighted fusion layer for adaptive fusion.

[0021] For another example, as shown in Figure 2 , step 1 - 4 includes the following steps: 1 - 4 - 1, generate a combined feature vector by calculating the average value of the three feature vectors; 1 - 4 - 2, process the generated combined feature vector through a two - layer multi - layer perceptron MLP to obtain a logical value ; (1); 1 - 4 - 3, perform SoftMax operation on the obtained logical value to obtain the contribution rates corresponding to the features of the three different branches, denoted as : (2); Among them, the parameter represents the participation degree of features at different sampling frequencies; 1 - 4 - 4, respectively adopt three projection layers, namely , to align the feature vectors into a unified embedding space and obtain their respective projected features ; 1 - 4 - 5, set the final fused feature to be expressed as the weighted sum of each projected feature, that is: (3); Among them, the weight is the corresponding contribution rate .

[0022] For step 2, the classifier calibration module driven by neural collapse proposed in this embodiment includes an isometric tight frame classifier with fixed parameters and a fully connected layer classifier with learnable parameters.

[0023] Among them, the isometric tight frame classifier with fixed parameters specifically includes the following process: 2-1-1. At the beginning of training, randomly synthesize classifier vectors and ensure that the classifier vectors conform to the simplex isometric tight frame property, specifically as follows: Set a set containing vectors , where represents the feature dimension, represents the number of categories, and satisfies the condition ; If the following conditions are met, this vector set is regarded as a simplex isometric tight frame: (4); Among them, is allowed to rotate and satisfy ; is the identity matrix, represents a vector with a length of and all values being 1; All vectors within the simplex isometric tight frame have equal norms and the same pairwise angles, that is: (5); (6); Among them, the pairwise angles of represent the maximum isometric separation between 2-1-2. Considering that the normalized class feature prototype represents the classifier weight vectors of each category, construct the normalized vector as the fixed weight matrix of the isometric tight frame classifier and keep it unchanged throughout the training process; 2-1-3. Project the fused feature obtained from equation (3) through the projection layer to obtain the feature ; Set for dimension adjustment, which is used to reach the neural collapse state in the classifier calibration module later; 2-1-4. Normalize the feature to : (7); 2-1-5, multiply with the classifier weights predefined in step 2-1-2 to obtain the logical value : (8); where, represents the learnable temperature coefficient for scaling the feature product and both and have a finite range, represents the category corresponding to the obtained logical value, where ;

[0024] Additionally, the fully connected layer classifier with learnable parameters specifically includes the following process: 2-2-1, randomly initialize the fully connected layer classifier parameters and allow these parameters to be normally trained during the training process; 2-2-2, represent the weights of each category classifier as for processing the fused feature to generate the logical value i.e.: (9); 2-3, finally, linearly combine the logical value output by the isometric tight frame classifier with fixed parameters and the logical value output by the fully connected layer classifier with learnable parameters to obtain the combined logical value : (10); where, is used to regulate the degree of fixation of the parameters in the classifier during the classification stage.

[0025] For step 3, it specifically includes the following steps: Set to represent the finally output logical value, then the class balance focal loss is expressed as: (11); (12); where, represents the number of samples of category ; is the sample size adjustment factor for controlling the effective sample size with respect to Increased growth rate; is the weight smoothing factor, which is used to adjust the weight decay rate of easily classified samples. Based on experimental verification, the present invention respectively sets and to 0.9999 and 0.5.

[0026] The following further illustrates the significant technical effects of the present invention in combination with experiments.

[0027] 1. Experimental data and settings 1-1. Experimental data The present invention uses two open-source animal behavior recognition datasets to verify the proposed method, which are respectively collected from the behaviors of sheep and cattle.

[0028] Sheep dataset: The sheep dataset is collected from five sheep in two farms, including three small domestic dwarf sheep in one farm and two larger wild sheep in another farm. Six triaxial accelerometers and triaxial gyroscopes in different directions are installed on the collars of each sheep, and the sampling frequency is 100 Hz. The present invention has screened out a total of 42,943 data samples, and each sample corresponds to a 2-second behavior segment, expressed as a tensor with a dimension of 1×36×200. Figure 3 (a) in shows the class distribution of these samples, covering five main activities, including standing (43.15%), galloping (0.87%), eating (35.35%), trotting (0.44%), and walking (20.19%). Obviously, there is a serious class imbalance problem in this dataset, and the imbalance ratio is 98.05.

[0029] Cattle dataset: The cattle dataset is collected from six free-ranging cows, and a triaxial accelerometer is installed on the collar of each cow, and the sampling frequency is 25 Hz. The present invention uses a 2-second random sliding window to segment the original acceleration time series data, and a total of 10,429 data samples are generated. Each sample is expressed as a tensor with a dimension of 1×3×50. Figure 3 (b) in shows the class distribution of the samples, covering five main activities, including eating (6.10%), moving (16.29%), resting (54.25%), ruminating (19.32%), and salt supplementation (4.04%). Obviously, there is also a class imbalance problem in this dataset, and the imbalance ratio is 13.44.

[0030] 1.2. Experimental settings The present invention uses precision, recall, F1-score, and accuracy as indicators to evaluate the overall performance of the classification network. To verify the generalization ability of the present invention, the "Leave-one-out Cross-validation" method is adopted on the sheep dataset, where "one" represents the data of a single animal individual. Since not all individuals in the cattle dataset cover all behaviors, another commonly used verification method, the Stratified Five-fold Cross-validation method, is adopted. In each round, the samples of 3 folds, 1 fold, and 1 fold are used as the training set, validation set, and test set respectively, and this is repeated 5 rounds while ensuring that the validation set and test set are different in each round.

[0031] To prevent model overfitting, the present invention adds an L2 regularization term to the loss function, where the weight decay coefficients of the sheep dataset and the cattle dataset are set to 1×10 -4 and 6×10 -2 . During the optimization process, the Adam optimizer is used, and the initial learning rate is set to 1×10 -4 on the sheep dataset and 5×10 -4 on the cattle dataset, and a stepped learning rate decay strategy is adopted, with the learning rate decaying to 0.1 times the original every 20 training epochs. In the present invention, the model training is carried out for 100 epochs in total, and the batch size is set to 256. During the training process, the performance of the validation set is continuously monitored and the model parameters with the highest accuracy are saved. Finally, an independent test set is used to evaluate the performance of the saved best model. Based on the conclusions of previous studies, 12.5 Hz is used as the sampling frequency of the baseline model for the sheep dataset, and this frequency has been proven to be able to obtain the optimal recognition performance. For the cattle dataset, through preliminary experimental comparative analysis, 25 Hz is determined as the best sampling frequency, so it is used as the sampling frequency of the baseline model.

[0032] To verify the effectiveness of the proposed method, the present invention designed a comprehensive comparative experiment. During the training process, based on the preliminary experimental results, the sheep dataset adopted a combination of three sampling frequencies: 50Hz, 25Hz, and 12.5Hz, and the cattle dataset adopted a combination of three sampling frequencies: 25Hz, 12.5Hz, and 5Hz. This is the first time to introduce a multi-sampling frequency fusion strategy in animal behavior recognition. The comparative experiment includes two aspects: First, compare with the baseline method that only uses a single optimal sampling frequency; Second, compare with a variety of advanced class imbalance processing methods, including resampling methods, including K-Means Synthetic Minority Over-sampling Technique (KMeansSMOTE, KMS) and Random UnderSampler (RUS), and reweighting methods, including Cost-sensitive Cross-entropy Loss (CS_CE), Class-balanced Focal Loss (CB_FL), and Adaptive Class Suppression Loss (ACSL). All tests were carried out on the NVIDIA GeForce RTX 3090 graphics processing unit using the PyTorch platform.

[0033] 2. Performance Comparison between the Present Method and Existing Methods Tables 1 and 2 respectively show the performance comparison results between the method proposed in the present invention and existing methods on the sheep and cattle datasets. Among them, the hyperparameters and are respectively set as follows: for the sheep dataset ( =0.4, =0.3), for the cattle dataset ( =0.8, (=0.1). The experimental results show that the proposed method achieves the optimal performance on both datasets: it obtains an accuracy of 93.17%, an F1-score of 87.17%, and a recall rate of 91.71% on the sheep dataset; and reaches an accuracy of 91.42%, an F1-score of 90.62%, a precision of 88.06%, and a recall rate of 93.64% on the cattle dataset. In particular, compared with the baseline method, the proposed method has a significant improvement in the recall rate indicator, that is, it is improved by 16.98% and 15.78% on the sheep and cattle datasets respectively. These improvements are in line with the original intention of this invention to achieve a high recall rate and maximize the classification accuracy of various behaviors. In addition, the performance improvement of the KMS method is limited, which may be due to its oversampling strategy leading to the overrepresentation of minority class samples, thus causing the problem of model overfitting. In contrast, the RUS method significantly reduces the performance due to undersampling on the dataset with limited sample size, resulting in the loss of key information. It is worth noting that CB_FL has a better improvement in the recall rate than other methods, and this result verifies the rationality of this invention in choosing the class-balanced focal loss as the optimization target.

[0034] Table 1 Performance comparison data table of the method proposed in this invention and existing methods on the sheep dataset

[0035] Table 2 Performance comparison data table of the method proposed in this invention and existing methods on the cattle dataset

[0036] 3. Ablation experiments 3.1. Evaluation of the multi-scale feature fusion module based on the hybrid expert system and the classifier calibration module driven by neural collapse: To verify the contributions of the newly proposed multi-scale feature fusion module based on the hybrid expert system and the classifier calibration module driven by neural collapse in this invention, this invention conducted experiments on the sheep dataset for animal behavior recognition methods with and without these two modules respectively. The experimental results are presented through various evaluation indicators in Table 3, and the recall confusion matrix is presented in Figure 4 . Table 3 shows that after adding the classifier calibration module, the F1-score, precision, and recall rate are slightly improved by 0.51%, 0.29%, and 0.99% respectively. Correspondingly, the classification accuracy of most behaviors has been improved to varying degrees (such as (b) in Figure 4 and Figure 4in (a) comparison), including the minority classes "galloping" and "trotting" increased by 3.48% and 1.06% respectively. This indicates that the classifier calibration module proposed in the present invention helps to improve the classification performance of minority classes. After further adding the multi-scale feature fusion module based on the mixture of experts system, the overall performance is significantly improved, that is, the accuracy, F1-score, precision and recall rate are increased by 1.81%, 4.78%, 5.33% and 3.37% respectively. Combining Figure 4 as shown in (c) therein, ideal recall rates are obtained for all behaviors, and the accuracy of most behaviors exceeds 90%. This is attributed to the fact that the feature fusion module can capture the optimal motion patterns of each behavior category by adaptively fusing features of multiple sampling frequencies, thus ensuring the classification accuracy of various behaviors.

[0037] Table 3 Comparison data table of behavior recognition performance with and without the multi-scale feature fusion module based on the mixture of experts system and the classifier calibration module driven by neural collapse

[0038] 3.2. Analysis of the multi-scale feature fusion module based on the mixture of experts system Analysis of the advantages of the soft weighted fusion mechanism: The multi-scale feature fusion module proposed in the present invention innovatively adopts the architecture of the mixture of experts system. Its core advantage lies in that the integrated soft routing mechanism can dynamically allocate contribution degrees for the features extracted at different sampling frequencies, and use these contribution degrees as weights for soft feature fusion. To evaluate the effectiveness of this method, the present invention conducts a comparative experiment with traditional fusion methods (including addition, mean, multiplication and splicing), and the results are shown in Table 4. The experimental data show that the soft weighted fusion method proposed in the present invention is superior to other methods in all evaluation indexes. This advantage mainly stems from the following two aspects: First, the dynamic routing mechanism of the mixture of experts system can automatically identify the optimal sampling frequency of specific behavior patterns; Second, the soft weighted strategy can adaptively strengthen the feature representation at the optimal sampling frequency, so as to accurately capture the customized feature patterns of different behavior categories.

[0039] Table 4 Comparison data table of the performance of different feature fusion methods

[0040] Analysis of the advantages of multi-scale feature fusion: Since a single sampling frequency is difficult to meet the feature extraction requirements of each behavior category, the present invention proposes a multi-scale feature fusion module based on a mixture of experts system to solve this problem by integrating data of multiple sampling frequencies. To verify its effectiveness, the present invention conducted a comparative experiment by comparing it with a method that only uses a single sampling frequency (50Hz, 25Hz, 12.5Hz). As shown in Table 5, multi-scale feature fusion shows significant improvement in all evaluation metrics, verifying its robustness and generalization ability. In addition, the present invention also visualized the feature distribution before and after fusion through t-distributed Stochastic Neighbor Embedding (t-SNE). As Figure 5 shown, compared with the independent features extracted by each expert network before fusion, the fused features show higher clustering compactness on samples of the same behavior category, and at the same time, the separability between different categories is significantly enhanced. This result indicates that multi-scale feature fusion can not only adaptively integrate discriminative information at different frequencies, but also effectively enhance the intra-class consistency and inter-class distinguishability of features, thereby improving the accuracy and robustness of behavior recognition.

[0041] Table 5 Performance comparison data table of the model under multi-sampling frequency fusion and single sampling frequency

[0042] 3.3. Analysis of the classifier calibration module driven by neural collapse Analysis of the influence of hyperparameters : The hyperparameter in formula (10) is used to control the constraint strength of the classifier weights during training, and its value directly affects the optimization dynamics and final performance of the model. To explore its influence law, the present invention fixed the parameter at 0.4 and systematically evaluated the performance of the model when . As shown in Table 6. When , the method reaches the peak performance, and the accuracy, F1 score, precision, and recall rate are 93.17%, 87.17%, 83.80%, and 91.71% respectively. When exceeds 0.3, the model performance gradually decreases; especially when , the classifier weights are completely fixed in the initial equiangular tight frame state, resulting in the model losing its optimization ability, verifying the necessity of dynamic adjustment of classifier parameters.

[0043] Table 6 Data table of the influence of different values on the model performance

[0044] Visualization of the cosine similarity of classifier vectors: The cosine similarity between classifier weight vectors directly determines the geometric characteristics of the decision boundary and is a key factor affecting the performance of class discrimination. To verify the effectiveness of the classifier calibration module driven by neural collapse, the present invention systematically compares and analyzes the changes in the cosine similarity distribution of classifier weight vectors before and after the introduction of the module through experiments, as Figure 6 shown. The experimental results show that without using the classifier calibration module, the cosine similarity distribution range is relatively wide (-0.5 to +0.5), and a significant proportion of the sample similarities exceed 0.4, indicating that there is a small angle between classifier vectors, which will lead to a blurred decision boundary and increase the risk of misclassification. In contrast, after introducing the classifier calibration module driven by neural collapse, the cosine similarity distribution range is significantly narrowed to -0.2 to +0.2, proving that this module can effectively promote the classifier vectors to form an approximately equiangular distribution in the feature space. This optimization of geometric characteristics significantly improves the discrimination performance of the classifier under the condition of class imbalance and effectively alleviates the problem of the model's bias towards the majority class.

[0045] 3.4. Model Robustness Evaluation under the Condition of Class Imbalance The unbiased animal behavior recognition method based on multi-scale feature fusion proposed by the present invention effectively maximizes the angle between different classifier weight vectors by introducing an equiangular tight frame classifier as a regularization constraint. To systematically evaluate the robustness of the method under different degrees of class imbalance, the present invention constructs a progressive imbalance experimental scenario by downsampling the minority class samples of the sheep dataset: the minority class samples are respectively reduced to 1 / 2 and 1 / 5 of the original number, so that the class imbalance rate gradually increases from the baseline of 98.05 to 196.1 and 490.25. As shown in Table 7, under different degrees of imbalance, the proposed method maintains a significant advantage in key indicators such as accuracy, F1 score, and precision. It is worth noting that as the degree of imbalance increases (the imbalance rate increases from 196.1 to 490.25), the performance advantage of the proposed method compared with the baseline method becomes more prominent, which fully proves the strong robustness of the proposed equiangular tight frame regularization mechanism in extreme imbalance scenarios.

[0046] Table 7 Performance Comparison Data Table of the Method of the Present Invention and the Baseline Method under Different Imbalance Rates

[0047] The present invention develops an unbiased animal behavior recognition method based on multi-scale feature fusion, which enhances the recognition ability of the unbiased animal behavior classification model proposed by this method for each specific behavior by customizing features and calibrating the classifier simultaneously. Specifically, considering that different behaviors require different sampling frequencies to achieve the best classification effect for their respective categories, the present invention designs a multi-scale feature fusion module based on a mixture of experts system. This module can adaptively fuse data with multiple different sampling frequencies, thereby extracting personalized features of specific behavior categories. At the same time, in order to alleviate the problem of classifier bias towards the majority class caused by class imbalance, the present invention proposes a neural collapse-driven classifier calibration module. This module introduces a fixed isometric tight frame classifier in the classification stage, and effectively alleviates the classifier bias phenomenon by maximizing the angle between different classifier weight vectors, thereby improving the classification performance for minority classes.

[0048] The above description only elaborates on the preferred embodiments and principles of the present invention in detail. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.

Claims

1. An unbiased animal behavior recognition method based on multi-scale feature fusion, characterized in that, It includes the following steps: S1. Through a multi-scale feature fusion module based on a mixture of experts system, adaptively fuse data with different sampling frequencies to capture the specific characteristics of various behaviors; S2. Through a classifier calibration module driven by neural collapse, maximize the angle between the weight vectors of different classifiers to complete the construction of an unbiased animal behavior classification model for unbiased animal behavior classification under class imbalance; S3. Use class-balanced focal loss as the loss function to optimize the unbiased animal behavior classification model.

2. The unbiased animal behavior recognition method based on multi-scale feature fusion according to claim 1, wherein Step S1 includes the following steps: S11, collect corresponding data respectively at sampling frequencies , , ; , , ; Among them, setting the values of three sampling frequencies follows the relationship; S12, input , , into their respective feature extractors to capture features at different scales , , ; wherein and represent the number of feature channels and the dimension of the coordinate axis respectively, , , refers to the time dimension corresponding to the features obtained on each branch, and ; S13, perform global average pooling operations on the features extracted from the three branches respectively to generate three feature vectors, namely , and ; S14. Input the three generated feature vectors into a router-based soft weighted fusion layer for adaptive fusion.

3. The unbiased animal behavior recognition method based on multi-scale feature fusion according to claim 2, wherein Step S14 includes the following steps: S141. Generate a combined feature vector by calculating the average of the three feature vectors; S142, process the generated combined feature vector through a two-layer multi-layer perceptron (MLP) to obtain a logical value ; (1); S143. Perform a SoftMax operation on the obtained logical values to obtain the contribution rates corresponding to the features of the three different branches, denoted as :[[]]END]] (2); Among them, the parameter represents the participation degree of features at different sampling frequencies; S144, respectively adopt three projection layers, namely , which are used to align the feature vectors into a unified embedding space to obtain their respective projected features ; S145, set the final fusion feature Expressed as a weighted sum of each projection feature, i.e.: (3); Among them, the weight is the corresponding contribution rate .

4. The unbiased animal behavior recognition method based on multi-scale feature fusion according to claim 3, characterized in that, In step S2, the classifier calibration module driven by neural collapse includes an isometric tight frame classifier with fixed parameters and a fully connected layer classifier with learnable parameters; the isometric tight frame classifier with fixed parameters specifically includes the following process: S211. Randomly synthesize classifier vectors at the beginning of training and ensure that the classifier vectors conform to the simplex isometric tight frame property, specifically as follows: A set is set up to contain a set of vectors , where represents the feature dimension, represents the number of categories, and satisfies the condition ; If the following conditions are met, the vector set is regarded as a simplex isometric tight frame: (4); Among them, rotation is allowed and satisfies ; is the identity matrix, represents a vector of length and all elements are 1; all vectors within a simplex equiangular tight frame have equal norms and the same pairwise angles; S212. Construct the normalized vector as the fixed weight matrix of the equiangular tight frame classifier and keep it unchanged throughout the training process; S213, through the projection layer project the fused features obtained from Equation (3) to obtain features ; set for dimensionality adjustment to achieve the neural collapse state in the subsequent classifier calibration module; S214, normalize the feature to ; S215. Multiply by the classifier weights predefined in step S212 to obtain a logical value : (5); Among them, represents a learnable temperature coefficient for scaling the feature product , and both have a finite range; represents the category corresponding to the obtained logical value, where .

5. The unbiased animal behavior recognition method based on multi-scale feature fusion according to claim 4, characterized in that The fully connected layer classifier with learnable parameters specifically includes the following process: S221. Randomly initialize the parameters of the fully connected layer classifier and allow the parameters to be normally trained during the training process; S222, represent the weights of each category classifier as , which is used to process the fused feature to generate a logical value , that is: (6)。 6. The unbiased animal behavior recognition method based on multi-scale feature fusion according to claim 5, wherein Step S2 also includes the following steps: S23, linearly combine the logical value output by the equiangular tight frame classifier with fixed parameters and the logical value output by the fully connected layer classifier with learnable parameters to obtain a combined logical value : (7); Among them, It is used to adjust the degree of fixation of the parameters in the classifier during the classification stage.

7. The unbiased animal behavior recognition method based on multi-scale feature fusion according to claim 6, characterized in that, Step S3 includes the following steps: S31, Set represents the logical value of the final output, then the class-balanced focal loss is expressed as: (8); (9); Among them, represents the number of samples of the category; is the sample size adjustment factor, used to control the growth rate of the effective sample size as increases; is the weight smoothing factor, used to adjust the weight decay rate of easily classified samples.

8. An unbiased animal behavior recognition system based on multi-scale feature fusion, which is used to implement the unbiased animal behavior recognition method based on multi-scale feature fusion according to any one of claims 1-7, characterized in that, The unbiased animal behavior recognition system based on multi-scale feature fusion includes; A feature extraction module for adaptively fusing data with different sampling frequencies through a multi-scale feature fusion module based on a mixture of experts system to capture the specific characteristics of various behaviors; A behavior classification module for maximizing the angle between the weight vectors of different classifiers through a classifier calibration module driven by neural collapse to complete the construction of an unbiased animal behavior classification model for unbiased animal behavior classification under class imbalance; A model optimization module for optimizing the unbiased animal behavior classification model using class-balanced focal loss as the loss function.

Citation Information

Patent Citations

  • Multi-modal feature fusion image classification method and application in humanoid robot

    CN118628802A

  • Long-tail multi-label image classification method based on nerve collapse

    CN119580003A

  • Dynamic model calibration method and device based on simplex equiangular tight frame classifier

    CN119762859A

  • Animal anomaly detection method and system based on hybrid expert model

    CN119832341A

  • Analogical prompting-based incremental learning method, apparatus and device, and storage medium

    WO2025020418A1