A pet food preference intelligent evaluation method
Patent Information
- Application Number
- CN202610661930.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-09-22
AI Technical Summary
然而,现有的计算机视觉方法主要针对人类行为识别任务设计,直接应用于宠物进食行为分析存在以下技术挑战:首先,宠物面部解剖结构与人类差异显著,猫科动物具有43块面部肌肉而人类有42块,但肌肉分布和运动模式完全不同,犬科动物的面部肌肉数量约为26块,表情表达方式与人类差异更大,现有的人脸表情识别模型(如FER2013、AffectNet等预训练模型)难以直接迁移应用,迁移准确率通常低于40%;其次,宠物进食行为具有快速、连续、周期性的特点,猫的舔食频率通常在3-5Hz,狗的舔食频率在4-8Hz,咀嚼频率在1-3Hz,传统的单帧图像分类方法采样率通常为1-2fps,难以捕捉这种高频时序动态特征;再次,不同宠物种类(猫、狗)之间存在显著的物种特异性行为差异,猫科动物倾向于通过面部表情(耳朵位置、瞳孔大小)表达情绪,而犬科动物更多通过身体语言(尾巴摆动、身体姿势)表达情绪,同一物种的不同个体之间也存在明显的个体差异,研究表明个体间进食速度的变异系数可达35%-50%,缺乏有效的跨物种迁移学习机制和个体自适应校准方法;最后,现有方法通常只能输出单一的分类标签或回归值,缺乏对嗜好性进行多维度、细粒度评估的能力,而嗜好性实际上是一个多维度的复合概念,涉及进食速度、积极性、情绪、姿态等多个方面
第一,本发明实现了非接触式多维度自动评估,仅需普通RGB摄像头即可对宠物进食嗜好性进行客观、量化的智能评估,无需佩戴任何传感器或接触式设备,不会对宠物造成应激反应,保证了进食行为的自然性和评估结果的真实性,同时从进食速度、进食积极性、情绪状态、身体姿态和进食完整度五个维度进行综合评估,相比传统的单一指标评估方法(如仅测量进食量)能够更全面、准确地反映宠物对食品的嗜好程度,解决了现有方法评估维度单一、主观性强、效率低下的技术问题。
Smart Images

Figure CN122799151A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of pet food evaluation technology, and in particular to an intelligent assessment method for pet eating preferences. Background Technology
[0002] With the rapid development of the pet economy, the pet food industry has become one of the world's most promising consumer markets. Pet food preferences are a core factor influencing consumer purchasing decisions, and accurately assessing pets' preferences for different foods has significant theoretical and practical value for pet food research and development, quality control, and personalized feeding.
[0003] Currently, pet food preference assessment mainly relies on the following methods: The first is manual observation, where professionals with a background in pet behavior observe and record the pet's eating process, judging the degree of preference based on experience. This method is highly subjective, easily influenced by the observer's personal experience, fatigue level, and subjective judgment. The inter-annotator consistency coefficient (Kappa value) is usually below 0.60, and it is inefficient, typically requiring more than 30 minutes per assessment, making it difficult to apply on a large scale in industrial production environments. The second method is the two-bowl selection method, where two test foods are placed in two identical bowls, and relative preference is assessed by statistically analyzing the pet's first-choice rate, the proportion of food consumed, and the eating time. This method can only perform pairwise comparisons, failing to capture behavioral details and emotional state information during the pet's eating process. The assessment dimension is singular, and it is easily influenced by bowl position preferences. Studies show that approximately 3 0% of pets exhibit significant location preferences; the third method is sensor monitoring, which monitors the amount and time of food intake by installing strain gauge weight sensors at the bottom of the food bowl or using radio frequency identification technology. This method requires additional hardware, with the cost of a single set of equipment typically ranging from 500 to 2000 yuan. It can only obtain limited indirect indicators such as the amount of food consumed and the time spent eating, and cannot capture key behavioral characteristics that directly reflect preferences, such as facial expressions and body postures; the fourth method is physiological indicator monitoring, which indirectly assesses preferences by monitoring physiological indicators such as heart rate variability, salivary amylase activity, and skin conductivity before and after eating. This method requires contact biosensors, and the process of wearing the sensors may cause stress to the pet, increasing cortisol levels by 15%-30%, affecting the naturalness of the eating behavior and the objectivity of the assessment results. Moreover, the equipment is expensive, with the cost of a single test typically exceeding 200 yuan.
[0004] With the rapid development of deep learning technology, computer vision-based methods for analyzing animal behavior have gradually attracted attention from academia and industry. Convolutional neural networks have made breakthrough progress in tasks such as image classification and object detection, achieving a Top-5 accuracy of over 98% on the ImageNet dataset. Vision Transformer (ViT) achieves effective modeling of global image features through a self-attention mechanism, demonstrating superior performance to CNNs in multiple vision tasks. However, existing computer vision methods are primarily designed for human behavior recognition tasks, and directly applying them to pet eating behavior analysis presents the following technical challenges: First, the facial anatomy of pets differs significantly from that of humans. Felines have 43 facial muscles, while humans have 42, but the muscle distribution and movement patterns are completely different. Canines have approximately 26 facial muscles, and their facial expressions differ even more from humans. Existing facial expression recognition models (such as pre-trained models like FER2013 and AffectNet) are difficult to directly transfer and apply, with transfer accuracy typically below 40%. Second, pet eating behavior is characterized by its rapid, continuous, and periodic nature. Cats typically lick at a frequency of 3-5 Hz, dogs at 4-8 Hz, and chew at 1-3 Hz. Traditional single-frame image classification methods, with their high sampling rates, cannot meet these requirements. The frame rate is typically 1-2 fps, making it difficult to capture such high-frequency temporal dynamic features. Furthermore, there are significant species-specific behavioral differences between different pet species (cats and dogs). Felines tend to express emotions through facial expressions (ear position, pupil size), while canines express emotions more through body language (tail wagging, body posture). There are also significant individual differences among individuals of the same species; studies show that the coefficient of variation in eating speed between individuals can reach 35%-50%, lacking effective cross-species transfer learning mechanisms and individual adaptive calibration methods. Finally, existing methods typically only output a single classification label or regression value, lacking the ability to conduct multi-dimensional, fine-grained assessments of cravings. Cravings are actually a multi-dimensional, composite concept involving eating speed, enthusiasm, emotion, posture, and other aspects.
[0005] In summary, there is an urgent need for a non-contact, high-precision, and multi-dimensional intelligent assessment method for pet eating preferences. This method should be able to automatically extract multimodal behavioral characteristics of pets during the eating process using only a regular RGB camera, and comprehensively assess their food preferences, providing objective and quantitative technical support for pet food development and personalized feeding. Summary of the Invention
[0006] To address the problems existing in the prior art, this invention provides an intelligent assessment method for pet eating preferences, comprising the following steps: S1. Pre-build and train the MSTF-Net model, including the construction of a standardized dataset of pet eating behavior, dataset partitioning and preprocessing, MSTF-Net model construction and training, and MSTF-Net model performance evaluation. The MSTF-Net model consists of six functional modules: video preprocessing module, spatial feature extraction branch, temporal feature extraction branch, cross-modal attention fusion module, species-aware dynamic weight allocation module, and multi-task classification and regression head. S2. Collect video data of the pet's eating process using camera equipment, and preprocess the collected video, including video decoding and framing, adaptive background modeling and foreground separation; S3. Using a pre-trained MSTF-Net model, inference is performed on the pre-processed pet eating process video data to obtain the pet eating preference assessment results, wherein the pet eating preference assessment results include eating speed score, eating enthusiasm score, emotional state score, body posture score and comprehensive preference score.
[0007] Optionally, in step S1, the construction of a standardized dataset of pet eating behavior specifically includes: Recruit n experimental pets as data collection subjects, and ensure that the experimental pets fully cover different breeds, age groups and body sizes; Prepare k different types of pet food as test samples, and ensure that the pet food fully covers different categories and flavors. The amount of each food should be calculated according to 2%-3% of the pet's body weight. Establish a standardized experimental environment: control the environmental parameters as follows: light intensity 300-500 lux, color temperature 4000-5000K, ambient temperature 20-25℃, relative humidity 40%-60%, background noise below 45dB, use an RGB camera with a resolution of not less than 1920×1080 pixels, a frame rate of not less than 30fps, and a color depth of 8bit, install the camera directly in front of the pet food bowl at a 45°±5° downward angle, with the lens optical axis at a 45° angle to the ground, the horizontal distance between the camera and the center of the food bowl is 60-80cm, the vertical height is 50-70cm, and the field of view covers the food bowl and an area with a radius of 50cm around it; Each pet was tested three times independently for each food, with an interval of at least 24 hours between two consecutive tests to eliminate the effects of satiety and memory. The pets were fasting for 4-6 hours before each test. The complete process of the pet from entering the camera's field of vision to leaving the food bowl area was recorded for each test, and videos within a specific time range were selected as valid videos. Each video was independently annotated by at least five annotators with backgrounds in pet behavior or animal science. The annotation dimensions included five dimensions: eating speed rating, eating enthusiasm rating, emotional state rating, body posture rating, and overall preference rating. Each dimension was rated using a 1-5 level Likert scale, where level 1 represents the lowest preference, level 2 represents the lowest preference, level 3 represents the highest preference, level 4 represents the highest preference, and level 5 represents the highest preference. The median of the five annotators' ratings for each video was used as the final label value to reduce the influence of extreme values. The Fleiss' Kappa coefficient was also calculated, and annotated data with a Fleiss' Kappa value of not less than 0.70 were selected as valid annotated data to ensure annotation quality.
[0008] Optionally, in step S1, the dataset partitioning and preprocessing specifically includes: The entire dataset was randomly divided into training, validation and test sets according to a ratio of 70%, 15% and 15%, respectively. During the division, it was strictly ensured that all video samples of the same pet belonged to the same dataset to avoid data leakage. Stratified sampling is used to ensure that the distribution of preference levels in each dataset is consistent with the overall distribution. The video frames are standardized by normalizing pixel values to the [0,1] interval, subtracting the mean of the ImageNet dataset [0.485, 0.456, 0.406] and dividing by the standard deviation [0.229, 0.224, 0.225], and finally the pixel value distribution approximates the standard normal distribution N(0,1).
[0009] Optionally, step S1, the construction and training of the MSTF-Net model, specifically includes: The video preprocessing module uses a pet target detection network based on the YOLOv8-L architecture to detect pets in each frame of the input video; it employs the DeepSORT multi-target tracking algorithm to perform cross-frame correlation and continuous tracking of the detected pet targets; and it uses a pet keypoint detection network based on the HRNet-W48 architecture to locate 17 body keypoints of the pet, including: point 1 is the tip of the nose, point 2 is the center of the left eye, point 3 is the center of the right eye, point 4 is the base of the left ear, point 5 is the base of the right ear, point 6 is the tip of the left ear, point 7 is the tip of the right ear, point 8 is the shoulder joint of the left forelimb, point 9 is the shoulder joint of the right forelimb, point 10 is the paw tip of the left forelimb, point 11 is the paw tip of the right forelimb, and point 12 is the spine in the back. Points 13, 14, 15, 16, and 17 are the tail root, left hind leg hip joint, right hind leg hip joint, left hind leg paw tip, and right hind leg paw tip, respectively. The original video frames are cropped based on the pet detection bounding boxes, with the cropped area being 20% larger than the bounding box to preserve contextual information. The cropped images are then uniformly adjusted to 224×224 pixels using bilinear interpolation, generating a sequence of full-body pet images. Facial region bounding boxes are calculated based on five facial key points: nose tip, eyes, and ear roots. The bounding box is the smallest bounding rectangle containing these five points, expanded by 30%. The facial region is cropped and uniformly adjusted to 112×112 pixels using bilinear interpolation, generating a sequence of pet facial region images. The spatial feature extraction branch adopts a three-stream parallel architecture, including three parallel sub-networks: global appearance feature stream, local facial feature stream, and skeletal pose feature stream. This three-stream architecture can capture feature information at different spatial scales and semantic levels: the global appearance feature stream uses Swin Transformer V2-Base as the backbone network, with input being a 224×224×3 pixel image of the pet's full-body region. V2-Base employs a hierarchical shift-window attention mechanism, comprising four stages. The feature map resolutions for each stage are 56×56, 28×28, 14×14, and 7×7, respectively. The embedding dimensions for each stage are 128, 256, 512, and 1024, respectively. The number of Transformer blocks in each stage is 2, 2, 18, and 2, totaling 24 blocks. Each block contains alternating stacks of multi-head self-attention for windows and multi-head self-attention for shift windows. The window size is 7×7, and the number of attention heads is 4, 8, 16, and 32, respectively. Log-CPB relative position encoding is used, and the weights are initialized using pre-trained weights from the ImageNet-22K dataset. The output of the final stage is passed through a global average pooling layer to obtain a 1024-dimensional global appearance feature vector F. globalThe local facial feature flow uses EfficientNet-B4 as the backbone network. The input is a 112×112×3 pixel pet facial region image. EfficientNet-B4 employs a compound scaling strategy with a depth coefficient α=1.4, a width coefficient β=1.2, and a resolution coefficient γ=1.3. It contains 7 MBConv module stages, each stage using MBConv1 and MBConv6×6 stages respectively. The number of layers in each stage is 2, 4, 4, 6, 6, 8, and 2, for a total of 32 layers. The number of output channels is 24, 32, 56, 112, 160, 272, and 448 respectively. A Squeeze-and-Excitation attention mechanism is used, with an SE reduction ratio of 4. Weights are initialized using pre-trained weights from the ImageNet-1K dataset. The output of the last stage passes through a global average pooling layer and a 1×1 convolutional layer to obtain a 512-dimensional facial feature vector F. face The skeletal pose feature stream uses a spatiotemporal graph convolutional network to encode the keypoint sequence. Seventeen keypoints are constructed as a graph structure. Node features consist of three dimensions: the normalized (x, y) coordinates and confidence score of the keypoint. The total number of nodes is V=17. There are 16 edges following the anatomical topology of a pet skeleton. The adjacency matrix A is a 17×17 symmetric matrix. The graph convolutional network contains four spatiotemporal graph convolutional blocks. Each block contains a spatial graph convolutional layer, a temporal convolutional layer, and residual connections. The output channels of the four blocks are 64, 128, 256, and 256 respectively. Each block is followed by a BatchNorm layer and a ReLU activation function. Finally, global average pooling is used to obtain a 256-dimensional skeletal pose feature vector F. pose ; the global appearance feature vector F global Facial feature vector F face and skeletal pose feature vector F pose A 1792-dimensional feature vector is obtained by concatenating along the channel dimension. This vector is then fused using a multilayer perceptron containing two fully connected layers. The first fully connected layer maps the 1792 dimensions to 1024 dimensions and applies GELU activation and Dropout regularization with a Dropout probability p=0.1. The second fully connected layer maps the 1024 dimensions to 512 dimensions, resulting in a 512-dimensional comprehensive spatial feature vector F. spatial ; The temporal feature extraction branch employs a Transformer-based temporal encoder design to capture the dynamic changes and temporal dependencies in pet eating behavior: it extracts the spatial feature vector F from T consecutive video frames. spatialThe input sequence is composed of elements, and the input tensor has a shape of (B, T, D), where B is the batch size, T is the sequence length, and D is the feature dimension. First, the input feature sequence is encoded using a learnable temporal position encoding method. This method employs a hybrid encoding approach, combining sine and cosine position encoding with learnable position embeddings. Sine and cosine encoding provides absolute positional information, while the learnable embeddings capture task-related positional patterns. The position encoding dimension is the same as the feature dimension, both being 512-dimensional. The position encoding matrix PE∈R {T×D} The temporal encoder consists of N stacked Transformer encoder layers. Each encoder layer contains a multi-head self-attention sub-layer, a feedforward neural network sub-layer, two layer normalization operations, and two residual connections, employing a Pre-LN structure. The multi-head self-attention sub-layer uses H parallel attention heads, each head having a dimension d. k =d v =D / H, the total computational complexity of attention is O=T²D, and the attention calculation uses the scaled dot product attention mechanism: Attention(Q,K,V)=softmax(QK). T / √d k In the multi-head self-attention mechanism, a temporal convolutional attention bias module (TCAB) is introduced. The TCAB extracts local temporal features from the key matrix K using one-dimensional depthwise separable convolutions. The kernel size of the one-dimensional convolution is K. size =5, stride=1, padding=2 to maintain sequence length, depthwise separable convolution parameters are only 1 / D of standard convolution, the convolution output is added to the original attention score as an attention bias term, the calculation formula of the temporal convolutional attention bias module TCAB is expressed as Attention TCAB(Q,K,V) =softmax(QK T / √d k +Conv1D dw(K) Let Q, K, and V be the query matrix, key matrix, and value matrix, respectively, all with shapes (B, H, T, d). k ), d k Conv1D represents the dimension of the key vector. dw This represents a one-dimensional depthwise separable convolution operation; the feedforward neural network sublayer contains two linear transformation layers and a GELU activation function. The first linear layer expands the 512-dimensional array to 2048-dimensional arrays, and the second linear layer compresses the 2048-dimensional arrays back to 512-dimensional arrays. Dropout regularization is applied between the two linear layers with a dropout probability p=0.1; the temporal encoder outputs a sequence of feature vectors over T time steps, with an output tensor shape of (B,T,D). These feature vectors are adaptively aggregated into a single feature vector by a temporal attention pooling layer. The temporal attention pooling layer first uses a learnable query vector q∈R... DCalculate the attention score a for the feature vector at each time step. t =softmax(q·h t / √D), and then use the normalized attention score as the weight to perform a weighted sum F on the feature vectors at each time step. temporal =Σ(a t ·h t Finally, a 512-dimensional temporal feature vector F is obtained. temporal ; The cross-modal attention fusion module employs a design combining a bidirectional cross-attention mechanism and a gated fusion mechanism to achieve deep semantic alignment and information fusion of spatial and temporal features: the spatial-to-temporal cross-attention calculation process involves converting the spatial feature vector F... spatial ∈R D Through the linear projection matrix W Qs ∈R {D×D} Projection yields the query vector Q s =W Qs ·F spatial ∈R D The time series feature vector F temporal ∈R D Through the linear projection matrix W Kt ∈R {D×D} and W Vt ∈R {D×D} The key vector K is obtained by projecting it separately. t =W Kt ·F temporal ∈R D Sum vector V t =W Vt ·F temporal ∈R D The attention level of spatial features to temporal features is calculated by scaling the dot product attention, with the attention weight being a scalar α. s2t =softmax(Q s ·K t T / √D The spatial feature vector F enhanced with temporal information is obtained. s2t =α s2t ·V t ∈R D Where D is the vector dimension equal to 512; the calculation process of temporal to spatial cross-attention is symmetrical to the above, and the temporal feature vector F is... temporal Through linear projection W Qt Obtain the query vector Q t , the spatial feature vector F spatial Through linear projection W Ks and W Vs Obtain the key vector K s Sum vector V sThe spatial information-enhanced temporal feature vector F is calculated. s2t =α s2t ·V t ∈R D The gating fusion mechanism will convert the original spatial features F spatial Original time series features F temporal Cross-enhancement feature F s2t and F t2s Four 512-dimensional feature vectors are concatenated along the channel dimension to obtain a 2048-dimensional vector F. concat =[F spatial ;F temporal ;F s2t ;F t2s ]∈R {4D} The adaptive fusion weights are calculated using a gated network, which consists of a linear layer and a sigmoid activation function σ(·), outputting a 4D=2048 dimensional gated vector G=σ(W). g ·F concat +b g )∈R {4D}, Where b g ∈R {4D} As the bias vector, F is obtained by fusing the feature vectors through element-wise multiplication. gated =G⊙F concat ∈R {4D} Finally, the 4D=2048 dimensions are compressed to D'=1024 dimensions using a mapping network containing a linear layer, resulting in the final fused feature vector F. fused =W proj ·F gated ∈R {D'} Where D'=1024; The species perception dynamic weight allocation module is used to adaptively adjust the weight coefficients of the five evaluation dimensions based on the pet species category and individual behavioral characteristics. First, it fuses the feature F using the species classification head. fused For classification, the species classification head consists of two fully connected layers. The first fully connected layer maps the 1024 dimensions to 256 dimensions and applies the ReLU activation function. The second fully connected layer maps the 256 dimensions to 2 dimensions and applies the Softmax activation function to output the class probability distribution P. species =[p cat ,p dog ]∈R 2 Based on the species classification results, retrieve the corresponding basic weight vector W from the pre-set species weight database. base ∈R 5 Simultaneously, the fusion feature F fusedThe input is an individual feature encoder, which consists of three fully connected layers with dimensionality transformations of 1024, 512, 256, and 128 respectively. ReLU activation and Dropout regularization are applied between layers, with a Dropout probability p=0.2. The output is a 128-dimensional individual feature vector F. individual ∈R {128} This vector encodes the current individualized behavioral pattern of the pet; the individual feature vector F individual The input weight generation network consists of two fully connected layers with dimensionality transformations of 128→64→5. The middle layers use the ReLU activation function, and the final layer uses the Softmax activation function to ensure that the sum of the five output weight values is 1, resulting in the individual weight adjustment vector W. adjust ∈R 5 The final dynamic weight vector W dynamic ∈R 5 The weighted combination of the base weight and the individual adjustment weight is calculated using the formula W. dynamic =α·W base +(1-α)·W adjust Where α is a learnable balance coefficient, through α raw After transformation, α = 0.5 + 0.4·σ(α) raw The constraint values are within the range of [0.5, 0.9], ensuring that the basic weights always dominate while allowing for individual adjustments; The multi-task classification and regression head comprises six output branches: five parallel sub-dimensional rating heads and one comprehensive preference rating head. It employs a hard-parameter-shared multi-task learning architecture. The five sub-dimensional rating heads correspond to eating speed, eating enthusiasm, emotional state, body posture, and eating completeness ratings, respectively. Each sub-dimensional rating head uses the same network structure, consisting of three fully connected layers with dimensionality transformations of 1024, 256, 64, and 5, respectively. The first two layers use ReLU activation and Dropout regularization, with a Dropout probability p=0.3 to enhance generalization. The last layer uses a Softmax activation function to output the probability distribution P of the five preference levels. i ∈R 5 (i=1,2,3,4,5); The eating speed scoring head introduces an attention gating mechanism at the input end to affect the fused features F. fused Mid- and temporal features F temporal The relevant components are weighted and enhanced, with a focus on dynamic features such as licking frequency and chewing frequency; the feeding enthusiasm scoring head focuses on behavioral indicators such as the pet's movement speed when approaching food, the delay time of first contact with food, and the number of times the pet leaves the food bowl; the emotional state scoring head introduces an attention gating mechanism at the input end to affect the fused features F. fused Central and facial features Fface The relevant components are weighted and enhanced, with a focus on facial expression features such as ear orientation angle, eye opening degree, pupil size changes, and beard extension degree; the body posture scoring head introduces an attention gating mechanism at the input end to enhance the fused features F. fused In terms of skeletal posture features F pose The relevant components are weighted and enhanced, with a focus on body language characteristics such as tail wagging frequency and amplitude, back arching, and standing posture. The feeding integrity scoring head scores the feeding integrity by analyzing the effective feeding time percentage and the trend of food intake changes throughout the feeding process. The comprehensive preference scoring head assigns a dynamic weight vector W to the probability distribution of the scores for the five sub-dimensions. dynamic To perform weighted fusion, first, the 5-dimensional probability distribution P of each sub-dimension is... i The corresponding rank value vector L=[1,2,3,4,5] T The expected score is obtained by performing a dot product calculation. i =P i ·L=Σ(P i [j]·j) (j=1,2,3,4,5), and then the five expected scores are dynamically weighted and summed to obtain the comprehensive score Score. total =Σ(W dynamic [i]·Score i (i=1,2,3,4,5), and finally the comprehensive score is mapped to an integer level of 1-5 through four thresholds as the final comprehensive preference level output; The MSTF-Net model training strategy employs a multi-task joint learning framework, simultaneously optimizing five sub-dimensional scoring tasks, one comprehensive scoring task, and one auxiliary species classification task. The sub-dimensional scoring tasks utilize the label-smoothed cross-entropy loss function, calculated using the formula L. CEsmooth =-Σ((1-ε)·y true [c]+ε / K)·log(y pred [c]), where ε is the label smoothing coefficient, c is the category index, K is the number of categories equal to 5, and y true For the one-hot encoding of the real label, y pred The predicted probability distribution is the output of Softmax; the comprehensive scoring task uses a weighted combination of the mean squared error loss function and the ordinal regression loss function. The mean squared error loss is used to calculate the squared difference L between the predicted comprehensive score and the true comprehensive score. MSE =(Score pred -Score true )², the ordinal regression loss uses the CORAL framework to transform the 5-level classification problem into 4 binary classification sub-problems to maintain the orderliness of the rating levels. ordinal =-Σ(y k·log(σ(f k ))+(1-y k )·log(1-σ(f k (k=1,2,3,4), where y k f is the true label for the k-th binary sub-classification problem. k The logit is the model output; the standard cross-entropy loss L is used for the auxiliary species classification task. species =-Σy sp ·log(P species The total loss function is defined as L. total =Σ(λ i ·L CEsmoothi )+λ6·L MSE +λ7·L ordinal +λ8·L species L CEsmoothi Let λ be the label smoothing cross-entropy loss for the i-th sub-dimension (i=1,2,3,4,5). i These are the weight coefficients for each loss term; the optimizer uses the AdamW algorithm with an initial learning rate of lr. init Set to 1×10 -4 Weight decay coefficient decay Set β1 to 0.05, β2 to 0.9, and ε to 1×10⁻⁶. -8 The learning rate scheduling strategy uses cosine annealing, with the learning rate gradually decreasing from its initial value to lr according to a cosine curve. min =1×10 -6 The attenuation formula is lr t =lr min +0.5·(lr init -lr min )·(1+cos(π·t / T max Total number of training rounds T max The training run is set to 200 rounds, with each round iterating through the entire training set once. Data augmentation strategies include random horizontal flipping, random affine transformation, random color jittering, random erasure, temporal random sampling, and Mixup augmentation. During training, a validation set is used to evaluate model performance. An early stopping mechanism is triggered when the validation set loss does not decrease for 15 consecutive rounds to prevent overfitting.
[0010] Optionally, in step S1, the MSTF-Net model performance evaluation specifically includes: The MSTF-Net model was evaluated on the test set, and evaluation metrics for each scoring dimension were calculated, including classification accuracy, weighted F1 score, mean absolute error (MAE), root mean square error (RMSE), Spearman rank correlation coefficient (Spearman-ρ), and Cohen Kappa coefficient (Cohen-κ).
[0011] After adopting the above technical solution, the beneficial effects of the present invention are as follows: First, this invention achieves non-contact, multi-dimensional automatic assessment. It only requires a regular RGB camera to objectively and quantitatively assess a pet's eating preferences without the need for any sensors or contact devices. This avoids causing stress to the pet, ensuring the naturalness of the eating behavior and the authenticity of the assessment results. It comprehensively assesses five dimensions: eating speed, eating enthusiasm, emotional state, body posture, and completeness of eating. Compared with traditional single-indicator assessment methods (such as measuring only the amount of food consumed), it can more comprehensively and accurately reflect the pet's food preferences, solving the technical problems of existing methods such as single assessment dimensions, strong subjectivity, and low efficiency.
[0012] Second, the MSTF-Net model proposed in this invention innovatively designs a temporal convolutional attention bias mechanism (TCAB) and a bidirectional cross-modal attention fusion mechanism. The temporal convolutional attention bias mechanism effectively enhances the model's ability to model the temporal sequence of continuous periodic eating actions such as licking and chewing by introducing a one-dimensional depthwise separable convolutional bias term into the standard Transformer self-attention. Compared with the standard self-attention mechanism, it can improve the accuracy of local temporal pattern recognition by 12%-15%. The bidirectional cross-modal attention fusion mechanism achieves deep semantic alignment of static appearance features and dynamic behavioral features through spatial-temporal bidirectional cross-attention and gating fusion.
[0013] Third, the species-aware dynamic weight allocation module designed in this invention can automatically load the preset basic weight configuration according to the pet species category, and further learn the individualized weight adjustment vector through the individual feature encoder (1024→128 dimensions) and the weight generation network (128→5 dimensions). The adaptive fusion of basic weight and individual weight is achieved through the learnable balance coefficient α∈[0.5,0.9]. This effectively handles the interspecies behavioral differences between felines with rich facial expressions and canines with obvious body language, as well as the differences in eating habits between different individuals of the same species, thereby improving the model's cross-species generalization ability and individual adaptability. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is an overall flowchart of the intelligent assessment method for pet eating preferences in this embodiment of the invention; Figure 2 This is a network structure diagram of the MSTF-Net model in an embodiment of the present invention; Figure 3 This is a schematic diagram of the temporal convolutional attention bias mechanism in an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the definition of key points on a pet's body and the topological connections of its skeleton in an embodiment of the present invention. Figure 5 This is a diagram showing the loss curve and learning rate curve of the MSTF-Net model training process in an embodiment of the present invention; Figure 6 This is a comparison chart of the accuracy of different deep learning methods on the comprehensive liking assessment task in the embodiments of the present invention; Figure 7 This is a normalized confusion matrix diagram of the MSTF-Net model on a comprehensive preference five-class classification task in an embodiment of the present invention; Figure 8 The figures show the ROC curves and AUC values of the MSTF-Net model for each category on the comprehensive preference five-class classification task in this embodiment of the invention. Figure 9 Box plots showing the distribution of cat and dog preferences for eight test foods in embodiments of the present invention; Figure 10 This is an attention weighting analysis diagram showing the contribution of five types of behavioral characteristics to the assessment of addiction in an embodiment of the present invention; Figure 11 This is a comparison diagram of the temporal characteristics of high and low craving samples during the eating process in an embodiment of the present invention. Figure 12 This is a comparison chart of the evaluation errors of different test individuals before and after applying the individual calibration mechanism in an embodiment of the present invention; Figure 13 This is a comparison chart of ablation experiment results for each functional module of the MSTF-Net model in this embodiment of the invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] like Figure 1 As shown in the figure, this embodiment of the invention provides a method for intelligently assessing a pet's eating habits, including the following steps: S1. Pre-build and train the MSTF-Net model
[0018] This step includes constructing a standardized dataset of pet eating behavior, splitting and preprocessing the dataset, building and training the MSTF-Net model, and evaluating the performance of the MSTF-Net model.
[0019] (1) Construction of a standardized dataset of pet eating behavior
[0020] In this embodiment, 120 experimental pets were recruited as data collection subjects, including 60 domestic cats and 60 domestic dogs. The domestic cat sample should include British Shorthair cats (12), American Shorthair cats (10), Ragdoll cats (8), Siamese cats (6), and Chinese domestic cats (24), with an age range of 1-8 years (mean 3.2±1.8 years) and a weight range of 2.8-7.2 kg (mean 4.5±1.2 kg). The domestic dog sample includes Golden Retrievers (10), Labrador Retrievers (10), Corgis (10), Poodles (8), Bichon Frises (6), and Chinese domestic dogs (16), with an age range of 1-10 years (mean 4.1±2.3 years) and a weight range of 3.5-28.6 kg (mean 12.8±7.5 kg). The sample distribution should meet the Shannon-Wiener diversity index. H'>1.5, all experimental pets participated in the experiment after their owners signed informed consent forms, and the experimental process complied with animal ethics standards.
[0021] Prepare 8 different types of pet food as test samples, including 3 types of dry food (chicken-flavored dry food A, fish-flavored dry food B, and beef-flavored dry food C), 3 types of wet food (canned chicken wet food A, canned tuna wet food B, and canned beef wet food C), and 2 types of treats (freeze-dried chicken treat A and nutritional paste treat B). The amount of each food should be calculated based on 2.5% of the pet's body weight.
[0022] A standardized experimental environment was established: light intensity was controlled at 380±20 lux (measured using a TES-1335 illuminance meter), color temperature at 4500K, ambient temperature at 23±1℃, relative humidity at 50±5%, and background noise below 40dB. A Logitech C920 Pro HD camera was used for video capture, with an output resolution of 1920×1080 pixels, a frame rate of 30fps, and H.264 encoding. The camera was mounted directly in front of the food bowl at a 45° downward angle, 70cm horizontally from the center of the bowl, and 60cm vertically, covering the food bowl and an area with a radius of 50cm around it.
[0023] Each pet underwent three independent replicate experiments for each food, with an interval of at least 24 hours between each experiment to eliminate the effects of satiety and memory effects. Pets were fasting for 4-6 hours before each experiment (6 hours for cats and 4 hours for dogs). The entire process of the pet entering the camera's field of vision and leaving the food bowl area was recorded for each experiment. The effective video duration ranged from 30 seconds to 10 minutes, and a total of 2,880 effective feeding videos were collected (120 pets × 8 types × 3 experiments), with a total duration of approximately 286 hours and an average video duration of 5.9 ± 2.3 minutes.
[0024] Five master's students with backgrounds in pet behavior independently annotated each video, assigning ratings for five dimensions of preference: eating speed, eating enthusiasm, emotional state, body posture, and overall preference. Each dimension was rated using a 1-5 level Likert scale, where level 1 represents the lowest preference (extremely dislike / very poor), level 2 represents low preference (somewhat dislike / poor), level 3 represents moderate preference (average), level 4 represents high preference (somewhat like / good), and level 5 represents the highest preference (very like / excellent). The detailed criteria for the five sub-dimensional preference ratings are defined as follows: For the eating speed rating, level 1 indicates extremely slow eating or complete refusal to eat with a licking frequency of less than 1 time / second (<1Hz), and level 2 indicates slow eating with a licking frequency of 1-2 times. The licking frequency is 2-4 times / second (1-2Hz), level 3 indicates normal eating speed and licking frequency of 2-4 times / second (2-4Hz), level 4 indicates relatively fast eating and licking frequency of 4-6 times / second (4-6Hz), and level 5 indicates very fast and eager eating and licking frequency of more than 6 times / second (>6Hz). The evaluation criteria for eating enthusiasm are as follows: level 1 indicates no interest in food and an approach delay time of more than 60 seconds (>60s), level 2 indicates low interest and an approach delay time of 30-60 seconds (30-60s), and level 3 indicates moderate interest and an approach delay time of 1 second (1-2Hz). 0-30 seconds (10-30s), Level 4 indicates moderate interest with an approximate delay of 3-10 seconds (3-10s), and Level 5 indicates strong interest with an approximate delay of less than 3 seconds (<3s); the criteria for judging emotional state are as follows: Level 1 indicates negative emotions, tension, and fear, with an ear pressure angle >30° and pupil dilation ratio >1.3; Level 2 indicates slightly negative emotions, alertness, and unease, with ears turned to the side and body stiff; Level 3 indicates neutral emotions, calmness, with ears in a natural position and a peaceful expression; and Level 4 indicates slightly positive emotions, relaxation. Furthermore, the ears naturally tilt forward at an angle of 10-20°, the eyes are relaxed, and a level 5 indicates a positive emotional state of pleasure and satisfaction, with half-closed eyes (eye opening <50%), whiskers pointing forward at an angle >45°, possibly accompanied by purring (cats) or tail wagging (dogs); the evaluation criteria for body posture are as follows: level 1 indicates a tense and stiff posture with the body curled up and the tail tucked between the hind legs; level 2 indicates a slightly tense posture with the body somewhat stiff and the tail drooping; level 3 indicates a normal and natural standing or sitting posture with the tail naturally drooping or horizontal; and level 4 indicates a relaxed posture with the body stretched out and the tail wagging naturally at a frequency of 0.5-2Hz, Level 5 indicates a very relaxed and stretched posture with pleasant body swaying and tail held high and wagging (dogs) or tail held upright and trembling slightly (cats); the criteria for judging the completeness of eating are as follows: Level 1 indicates almost no food intake and less than 10% (<10%) of the supply; Level 2 indicates eating a small amount and 10%-30% (10-30%) of the supply; Level 3 indicates eating about half and 30%-60% (30-60%) of the supply; Level 4 indicates eating most of the food and 60%-90% (60-90%) of the supply; Level 5 indicates... The food was completely consumed, exceeding 90% of the allotted amount (>90%). The overall preference level was calculated based on a weighted score of five sub-dimensions: Level 1 indicates extremely dislike with a weighted overall score below 1.8 (<1.8); Level 2 indicates moderate dislike with a weighted overall score of 1.8-2.6 ([1.8, 2.6]); Level 3 indicates neutral with a weighted overall score of 2.6-3.4 ([2.6, 3.4]); Level 4 indicates somewhat liking with a weighted overall score of 3.4-4.2 ([3.4, 4.2]); and Level 5 indicates very liking with a weighted overall score above 4.2 (≥4.2).
[0025] For each video segment, the median score from five annotators was used as the final label value to reduce the impact of extreme values. Meanwhile, the Fleiss' Kappa coefficient for inter-annotator consistency was calculated to be 0.78 (95% confidence interval: 0.74-0.82), which is considered to be at a "good consistency" level.
[0026] (2) Dataset partitioning and preprocessing
[0027] In this embodiment, the entire dataset is randomly divided into 2016 training sets, 432 validation sets, and 432 test sets according to a ratio of 70%, 15%, and 15%, respectively. During the division, it is strictly ensured that all video samples of the same pet belong to the same dataset to avoid data leakage. Stratified sampling is used to ensure that the distribution of preference levels in each dataset is consistent with the overall distribution. The video frames are standardized, including normalizing the pixel values to the [0,1] interval, subtracting the mean of the ImageNet dataset [0.485, 0.456, 0.406] and dividing by the standard deviation [0.229, 0.224, 0.225], so that the final pixel value distribution approximates the standard normal distribution N(0,1).
[0028] (3) MSTF-Net model construction and training
[0029] In this embodiment, the overall architecture of the MSTF-Net model is as follows: Figure 2As shown, the model consists of six functional modules: video preprocessing module, spatial feature extraction branch, temporal feature extraction branch, cross-modal attention fusion module, species-aware dynamic weight allocation module, and multi-task classification and regression head. The model adopts an end-to-end differentiable design and supports backpropagation gradient flow. Its implementation is based on the PyTorch 2.0 deep learning framework in the Python 3.9 environment. The model training is performed on a workstation equipped with an NVIDIA RTX 4090 graphics card (24GB GDDR6X video memory), an Intel Core i9-13900K CPU, 64GB DDR5 memory, Ubuntu 22.04 LTS operating system, CUDA version 12.1, and cuDNN version 8.9.
[0030] The video preprocessing module uses a pet detection network based on the YOLOv8-L architecture to detect pets in each frame of the input video. The backbone of the YOLOv8-L network is CSPDarknet53, which contains 53 convolutional layers. The feature pyramid uses a PANet (Path Aggregation Network) structure to achieve multi-scale feature fusion. The detection head output includes bounding box coordinates (x,y,w,h) (normalized to [0,1]), target confidence score (0-1), and class probability (cat / dog binary classification). The detection input resolution is 640×640 pixels, the confidence threshold is set to 0.5, the IoU threshold for non-maximum suppression (NMS) is set to 0.45, and the single-frame detection latency is approximately 8-12ms (on NVIDIA RTX). (4090 on) The DeepSORT multi-target tracking algorithm is used to perform cross-frame association and continuous tracking of detected pet targets. The DeepSORT algorithm uses a Kalman filter for target state prediction. The state vector contains 8 components [x,y,a,h,vx,vy,va,vh] (center coordinates, aspect ratio, height and velocity). A 128-dimensional ReID feature vector is used for appearance similarity matching. The ReID features are extracted by a pre-trained OSNet network. The matching process uses the Hungarian Algorithm to solve for the optimal allocation. The maximum number of frames lost (Max Age) is set to 30 frames, and the minimum number of frames confirmed (Min Age) is set to 30 frames. The Hits were set to 3 frames, with a single-frame tracking latency of approximately 2-3ms. A pet keypoint detection network based on the HRNet-W48 architecture was used to locate 17 body keypoints of the pet. HRNet-W48 employs a multi-resolution parallel convolutional structure with four parallel branches, containing resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively, and the number of channels in each branch being 48, 96, 192, and 384 respectively. The output feature map resolution is 1 / 4 of the input image. Heatmap regression was used to predict the keypoint locations, with a heatmap resolution of 64×64. Figure 4As shown, the 17 key points are defined as follows: Point 1 is the nose tip; Point 2 is the center of the left eye; Point 3 is the center of the right eye; Point 4 is the left ear base; Point 5 is the right ear base; Point 6 is the left ear tip; Point 7 is the right ear tip; Point 8 is the left front shoulder; Point 9 is the right front shoulder; Point 10 is the left front paw; Point 11 is the right front paw; Point 12 is the midpoint of the spine; Point 13 is the tail base; Point 14 is the left hip; Point 15 is the right hip; and Point 16 is the left hind paw. The 17th point is the right hind paw. Each keypoint output includes (x,y) normalized coordinates and a confidence score. The single-frame keypoint detection latency is approximately 15-20ms. The original video frame is cropped based on the pet detection bounding box, with the cropped area being 20% outside the bounding box to preserve contextual information. The calculation formula is x... new =x-0.1*w、y new =y-0.1*h、w new =1.2*w、h new =1.2*h, the cropped image is uniformly adjusted to 224×224 pixels using bilinear interpolation to generate a sequence of pet full-body images; the facial region bounding box is calculated based on 5 facial key points: nose tip, eyes, and ear roots. The bounding box is the smallest bounding rectangle containing these 5 points, expanded by 30%. The facial region is cropped and uniformly adjusted to 112×112 pixels using bilinear interpolation to generate a sequence of pet facial region images.
[0031] The spatial feature extraction branch adopts a three-stream parallel architecture, including three parallel sub-networks: global appearance feature stream, local facial feature stream, and skeletal pose feature stream. This three-stream architecture can capture feature information at different spatial scales and semantic levels. The global appearance feature stream uses Swin Transformer V2-Base as the backbone network, with an input of a 224×224×3 pixel image of the pet's full-body region. Swin Transformer V2-Base employs a hierarchical shifted window attention mechanism, comprising four stages. The feature map resolutions for each stage are 56×56, 28×28, 14×14, and 7×7, respectively. The embedding dimension of each stage... The dimensions are 128, 256, 512, and 1024 respectively. The number of Transformer blocks in each stage is 2, 2, 18, and 2 respectively, for a total of 24 blocks. Each block contains alternating stacks of window multi-head self-attention (W-MSA) and shift window multi-head self-attention (SW-MSA). The window size is 7×7, and the number of attention heads is 4, 8, 16, and 32 respectively. Log-CPB (Log-spaced Continuous Position Bias) relative position encoding is used, and the weights are initialized using pre-trained weights from the ImageNet-22K dataset. The output of the last stage is passed through a global average pooling layer to obtain a 1024-dimensional global appearance feature vector, denoted as F. global The number of parameters in this branch is approximately 88M; the local facial feature flow uses EfficientNet-B4 as the backbone network, with an input of a 112×112×3 pixel pet facial region image. EfficientNet-B4 employs a compound scaling strategy. The scaling mechanism (with a depth coefficient α = 1.4, a width coefficient β = 1.2, and a resolution coefficient γ = 1.3) comprises seven MBConv module stages. Each stage sequentially uses MBConv1 (expansion ratio 1) and MBConv6 (expansion ratio 6) × 6 stages, with the number of layers in each stage being 2, 4, 4, 6, 6, 8, and 2, totaling 32 layers. The number of output channels is 24, 32, 56, 112, 160, 272, and 448, respectively. It employs a Squeeze-and-Excitation (SE) attention mechanism with an SE reduction ratio of 4. Weights are initialized using pre-trained weights from the ImageNet-1K dataset. The output of the final stage is processed by a global average pooling layer and a 1×1 convolutional layer (reducing the dimensionality to 512) to obtain a 512-dimensional facial feature vector, denoted as F. faceThe branch has approximately 19M parameters. The skeletal pose feature stream uses a spatiotemporal graph convolutional network (ST-GCN) to encode the keypoint sequence, constructing a graph structure from 17 keypoints. The node features are the (x,y) normalized coordinates and confidence scores of the keypoints, totaling 3 dimensions. The total number of nodes is V=17. The edge connections follow the pet skeletal anatomical topology, with a total of 16 edges (6 for the head, 4 for the forelimbs, 3 for the trunk, and 4 for the hindlimbs). The adjacency matrix A is a 17×17 symmetric matrix. The graph convolutional network contains 4 spatiotemporal graph convolutions. The algorithm is a block, each containing a spatial graph convolutional layer (using a partitioning strategy to divide neighboring nodes into three subsets: the node itself, nearest neighbors with a distance of 1, and distant neighbors with a distance greater than 1), a temporal convolutional layer (kernel size 9, dilation rate 1), and residual connections. The output channels of the four blocks are 64, 128, 256, and 256 respectively. Each block is followed by a BatchNorm layer and a ReLU activation function. Finally, global average pooling (in both the temporal and node dimensions) yields a 256-dimensional skeletal pose feature vector, denoted as F. pose The number of parameters in this branch is approximately 1.5M; the global appearance feature vector F global (1024-dimensional) facial feature vector F face (512-dimensional) and skeletal pose feature vector F pose (256-dimensional) concatenation along the channel dimension yields a 1792-dimensional feature vector. This vector is then fused using a multilayer perceptron (MLP) containing two fully connected layers. The first fully connected layer maps the 1792-dimensional vector to a 1024-dimensional vector (weight matrix W1∈R^). {1024×1792} It employs the GELU activation function (Gaussian Error Linear Unit) and Dropout regularization, with a Dropout probability p=0.1. The second fully connected layer maps the 1024 dimensions to 512 dimensions (weight matrix W2∈R). {512×1024} Finally, a 512-dimensional comprehensive spatial feature vector is obtained, denoted as F. spatial .
[0032] The temporal feature extraction branch employs a Transformer-based temporal encoder design to capture the dynamic changes and temporal dependencies in pet eating behavior: it extracts the spatial feature vector F from T consecutive video frames. spatialThe input sequence is a tensor of shape (B, T, D), where B is the batch size, T is the sequence length, and D is the feature dimension (512). The default value for T is 64, corresponding to a video segment of approximately 2.13 seconds at 30fps. This time window is sufficient to cover multiple complete licking-chewing cycles. First, the input feature sequence is encoded using learnable temporal position encoding. This is a hybrid encoding method combining sine and cosine position encoding with learnable position embeddings. Sine and cosine encoding provides absolute positional information, while the learnable embeddings capture task-related positional patterns. The positional encoding dimension is the same as the feature dimension, both being 512. The positional encoding matrix is PE∈R. {T×D} The temporal encoder consists of N stacked Transformer encoder layers, where N is 6 by default. Each encoder layer contains a multi-head self-attention sublayer, a feed-forward neural network sublayer, two layer normalization operations (normalization dimension is the last dimension), and two residual connections. It adopts a pre-LN structure (normalization before attention / FFN). The multi-head self-attention sublayer uses H parallel attention heads, each head having a dimension d. k =d v =D / H, the total computational complexity of attention is O=T²D, and the attention calculation uses the scaled dot product attention mechanism: Attention(Q,K,V)=softmax(QK). T / √d k In the multi-head self-attention mechanism, a temporal convolutional attention bias module (TCAB) is introduced, such as... Figure 3 As shown, the Temporal Convolutional Attention Bias Module (TCAB) extracts local temporal features from the key matrix K using one-dimensional depthwise separable convolution. The kernel size K of the one-dimensional convolution is... size =5, stride=1, padding=2 to maintain sequence length, depthwise separable convolution parameters are only 1 / D (approximately 0.2%) of standard convolution. The convolution output is added to the original attention score as an attention bias term. The calculation formula for the temporal convolutional attention bias module TCAB is expressed as Attention. TCAB(Q,K,V) =softmax(QK T / √d k +Conv1D dw(K) Let Q, K, and V be the query matrix, key matrix, and value matrix, respectively, all with shapes (B, H, T, d). k ), d kFor a key vector with a dimension of 64, Conv1D dw This represents a one-dimensional depthwise separable convolution operation (kernel size 5, input and output channels are both d). k =64), this mechanism can enhance the model's ability to model local temporal patterns of continuous eating actions (such as continuous licking, periodic chewing), and can improve the recognition accuracy of local patterns within 5 frames by 12%-15% compared with the standard attention mechanism; the feedforward neural network sublayer contains two linear transformation layers and a GELU activation function. The first linear layer expands 512 dimensions to 2048 dimensions (expansion ratio r=4), and the weight matrix W_ff1∈R {2048 ×512} The second linear layer compresses the 2048-dimensional structure back to 512-dimensional structure, with the weight matrix W... ff2 ∈R {512×2048} Dropout regularization is applied between the two linear layers, with a dropout probability p=0.1. The temporal encoder outputs a sequence of feature vectors over T time steps, with an output tensor shape of (B,T,D). This tensor is adaptively aggregated into a single feature vector through a temporal attention pooling layer. The temporal attention pooling layer first uses a learnable query vector q∈R. D (Initialized to a standard normal distribution) Calculate the attention score a for the feature vector at each time step. t =softmax(q·h t / √D), and then use the normalized attention score as the weight to perform a weighted sum F on the feature vectors at each time step. temporal =Σ(a t ·h t This pooling method, compared to simple average pooling, can adaptively focus on key time steps (such as starting to eat, pausing, accelerating, etc.), ultimately yielding a 512-dimensional temporal feature vector, denoted as F. temporal .
[0033] The cross-modal attention fusion module employs a design combining bidirectional cross-attention and gated fusion to achieve deep semantic alignment and information fusion of spatial and temporal features. The spatial-to-temporal cross-attention calculation process involves converting the spatial feature vector F... spatial ∈R D Through the linear projection matrix W Qs ∈R {D×D} Projection yields the query vector Q s =W Qs ·F spatial ∈R D The time series feature vector Ftemporal ∈R D Through the linear projection matrix W Kt ∈R {D×D} and W Vt ∈R {D×D} The key vector K is obtained by projecting it separately. t =W Kt ·F temporal ∈R D Sum vector V t =W Vt ·F temporal ∈R D The attention level of spatial features to temporal features is calculated by scaling the dot product attention, with the attention weight being a scalar α. s2t =softmax(Q s ·K t T / √D The spatial feature vector F enhanced with temporal information is obtained. s2t =α s2t ·V t ∈R D Where D is the vector dimension equal to 512; the calculation process of temporal to spatial cross-attention is symmetrical to the above, and the temporal feature vector F is... temporal Through linear projection W Qt Obtain the query vector Q t , the spatial feature vector F spatial Through linear projection W Ks and W Vs Obtain the key vector K s Sum vector V s The spatial information-enhanced temporal feature vector F is calculated. s2t =α s2t ·V t ∈R D The gating fusion mechanism will convert the original spatial features F spatial Original time series features F temporal Cross-enhancement feature F s2t and F t2s Four 512-dimensional feature vectors are concatenated along the channel dimension to obtain a 2048-dimensional vector F. concat =[F spatial ;F temporal ;F s2t ;F t2s ]∈R {4D} Adaptive fusion weights are calculated using a gated network, which contains a linear layer (weight matrix W). g ∈R {4D×4D} The sigmoid activation function σ(·) outputs a 4D=2048-dimensional gated vector G=σ(W). g ·F concat +bg )∈R {4D}, Where b g ∈R {4D} As the bias vector, the fused feature vector is calculated using element-wise multiplication (Hadamard Product) to obtain F. gated =G⊙F concat ∈R {4D} Finally, a mapping network containing a linear layer is used to compress the 4D=2048 dimension to D'=1024 dimension, with the weight matrix W. proj ∈R {D'×4D} The final fused feature vector F is obtained. fused =W proj ·F gated ∈R {D'} , where D'=1024.
[0034] The species perception dynamic weight allocation module adaptively adjusts the weight coefficients of five evaluation dimensions based on pet species category and individual behavioral characteristics. This module addresses the significant differences in behavioral expression between cats and dogs: firstly, it fuses features F through a species classification head pair. fused The classification process determines whether the current pet belongs to the Felidae or Canidae family. The species classification head consists of two fully connected layers. The first fully connected layer maps the 1024 dimensions to 256 dimensions (weight matrix W). sp1 ∈R {256×1024} And it uses the ReLU activation function (Rectified Linear Unit), and the second fully connected layer maps 256 dimensions to 2 dimensions (weight matrix W). sp2 ∈R {2×256} The softmax activation function is used to output the class probability distribution P. species =[p cat ,p dog ]∈R 2 The species classification accuracy reached 99.2% on the validation set; based on the species classification results (argmax(P)... species Retrieve the corresponding basic weight vector W from the preset species weight library. base ∈R 5 The basic weight vector for felines is set to W. basecat =[0.25,0.20,0.22,0.18,0.15] corresponds to five dimensions: eating speed, eating enthusiasm, emotional state, body posture, and eating completeness (summing to 1.0). The basic weight vector for canines is set to W. basedog=[0.22,0.23,0.18,0.22,0.15], this weighting configuration is based on the conclusions of animal behavior research: felines have richer facial expressions (more precise facial muscle control), therefore emotional state has a higher weight (0.22 vs 0.18), while canines have more obvious body language (larger tail swing and more changes in body posture), therefore body posture has a higher weight (0.22 vs 0.18); at the same time, the fused feature F fused The input is an individual feature encoder, which consists of three fully connected layers with dimensionality transformations of 1024, 512, 256, and 128 respectively. ReLU activation and Dropout regularization are applied between layers, with a Dropout probability p=0.2. The output is a 128-dimensional individual feature vector F. individual ∈R {128} This vector encodes the current individualized behavioral pattern of the pet; the individual feature vector F individual The input weight generation network consists of two fully connected layers with dimensionality transformations of 128→64→5. The middle layers use the ReLU activation function, and the final layer uses the Softmax activation function to ensure that the sum of the five output weight values is 1, resulting in the individual weight adjustment vector W. adjust ∈R 5 The final dynamic weight vector W dynamic ∈R 5 The weighted combination of the base weight and the individual adjustment weight is calculated using the formula W. dynamic =α·W base +(1-α)·W adjust Where α is a learnable balance coefficient (scalar), initialized to 0.7, and through α raw After transformation, α = 0.5 + 0.4·σ(α) raw The constraint values are within the range of [0.5, 0.9], ensuring that the basic weights always dominate while allowing for individual adjustments.
[0035] The multi-task classification and regression head comprises six output branches: five parallel sub-dimensional rating heads and one comprehensive preference rating head. It employs a hard parameter sharing multi-task learning architecture. The five sub-dimensional rating heads correspond to the Eating Speed Score, Eating Eagerness Score, Emotional State Score, Body Posture Score, and Eating Completeness Score, respectively. Each sub-dimensional rating head uses the same network structure, consisting of three fully connected layers with dimensionality transformations of 1024, 256, 64, and 5, respectively. The first two layers use ReLU activation and Dropout regularization, with a Dropout probability p=0.3 to enhance generalization. The last layer uses a Softmax activation function to output the probability distribution P of the five preference levels. i ∈R 5 (i=1,2,3,4,5), each branch has approximately 0.28M parameters; the eating speed scoring head introduces an attention gating mechanism at the input end to affect the fused features F. fused Mid- and temporal features F temporal The relevant components are weighted and enhanced, with a focus on dynamic features such as lick rate (times / second) and chewing rate (times / second); the feeding enthusiasm scoring head focuses on behavioral indicators such as the pet's approach speed (cm / second), first contact delay (seconds), and leaving the food bowl (number of times); the emotional state scoring head introduces an attention gating mechanism at the input end to enhance the fused features F. fused Central and facial features F face The relevant components are weighted and enhanced, with a focus on facial expression features such as ear angle (forward / backward / lateral, in degrees), eye opening degree (Eye Aperture Ratio, 0-1), pupil dilation ratio, and whisker spread angle. The body posture scoring head introduces an attention gating mechanism at the input end to enhance the fused features F. fused In terms of skeletal posture features F poseThe relevant components are weighted and enhanced, with a focus on body language features such as tail wag frequency (Hz), tail wag amplitude (degrees), back arch ratio, and stance width (ratio of forelimb to hindlimb). The feeding integrity scoring head scores the feeding integrity by analyzing the effective eating time ratio (effective eating time / total time) and the trend of food intake changes (estimated by inferring changes in pixels in the food bowl area). The comprehensive preference scoring head assigns the probability distribution of scores for the five sub-dimensions to a dynamic weight vector W. dynamic To perform weighted fusion, first, the 5-dimensional probability distribution P of each sub-dimension is... i The corresponding rank value vector L=[1,2,3,4,5] T The expected score is obtained by performing a dot product calculation. i =P i ·L=Σ(P i [j]·j) (j=1,2,3,4,5), and then the five expected scores are dynamically weighted and summed to obtain the comprehensive score Score. total =Σ(W dynamic [i]·Score i (i=1,2,3,4,5), and finally, the comprehensive score is mapped to an integer level of 1-5 through four thresholds (1.8,2.6,3.4,4.2) as the final comprehensive preference level output: Score total <1.8 is mapped to level 1, 1.8≤Score total <2.6 is mapped to level 2, 2.6≤Score total <3.4 is mapped to level 3, 3.4≤Score total <4.2 is mapped to level 4, Score total ≥4.2 is mapped to level 5.
[0036] The MSTF-Net model training strategy employs a multi-task learning framework, simultaneously optimizing five sub-dimensional scoring tasks, one comprehensive scoring task, and one auxiliary species classification task. The sub-dimensional scoring tasks utilize the label smoothing cross-entropy loss function, calculated using the formula L. CEsmooth =-Σ((1-ε)·y true [c]+ε / K)·log(y pred [c]), where ε is the label smoothing coefficient, c is the category index, K is the number of categories equal to 5, and ytrue One-hot encoding of the real label ∈ R K y pred The predicted probability distribution ∈ R of the Softmax output K Label smoothing can prevent the model from becoming overconfident and improve generalization ability. The comprehensive scoring task uses a weighted combination of mean squared error loss (MSE Loss) and ordinal regression loss. The MSE Loss is used to calculate the squared difference L between the predicted comprehensive score and the true comprehensive score. MSE =(Score pred -Score true )², the ordinal regression loss uses the CORAL (Consistent Rank Logits) framework to transform the 5-level classification problem into 4 binary sub-problems (whether it is ≥2 levels, whether it is ≥3 levels, whether it is ≥4 levels, whether it is ≥5 levels) to maintain the orderliness of the rating levels. The ordinal regression loss L ordinal =-Σ(y k ·log(σ(f k ))+(1-y k )·log(1-σ(f k (k=1,2,3,4), where y k f is the true label for the k-th binary sub-classification problem. k The logit is the model output; the standard cross-entropy loss L is used for the auxiliary species classification task. species =-Σy sp ·log(P species The total loss function is defined as L. total =Σ(λ i ·L CEsmoothi )+λ6·L MSE +λ7·L ordinal +λ8·L species L CEsmoothi Let λ be the label smoothing cross-entropy loss for the i-th sub-dimension (i=1,2,3,4,5). i The weight coefficients for each loss term are determined through hyperparameter search and are set to the default values: λ1=λ2=λ3=λ4=λ5=1.0, λ6=0.5, λ7=0.3, λ8=0.2. The optimizer uses the AdamW algorithm (Adam with Decoupled Weight Decay), with an initial learning rate lr. init Set to 1×10 -4 Weight decay coefficient decaySet to 0.05 (only for non-biased and non-normalized layer parameters), β1 to 0.9, β2 to 0.999, and ε to 1×10. -8 The learning rate scheduling strategy employs cosine annealing, where the learning rate gradually decreases from its initial value according to a cosine curve to lr. min =1×10 -6 The attenuation formula is lr t =lr min +0.5·(lr init -lr min )·(1+cos(π·t / T max Total number of training rounds T max The training run is set to 200 rounds, with each round iterating through the entire training set once. Data augmentation strategies include random horizontal flipping (probability 0.5), random affine transformation (rotation angle ±10°, translation ratio ±5%, scaling ratio 0.9-1.1), random color dithering (brightness ±0.2, contrast ±0.2, saturation ±0.2, hue ±0.05), random erasing (probability 0.1, area ratio 0.02-0.2, aspect ratio 0.3-3.3), temporal random sampling (randomly selecting T=64 frames from the original frame sequence, allowing repeated sampling), and Mixup enhancement (mixing coefficient α~Beta(0.2,0.2), performing linear interpolation mixing on the input image and labels simultaneously). During training, a validation set is used to evaluate model performance. When the validation set loss does not decrease for 15 consecutive rounds, an early stopping mechanism is triggered to prevent overfitting.
[0037] like Figure 5 As shown, the model training loss decreased rapidly in the first 50 epochs, then gradually converged and stabilized. The training loss decreased from an initial 2.45 to 0.18 (a decrease of 92.7%), and the validation loss decreased from 2.52 to 0.21 (a decrease of 91.7%). The model reached its optimal validation performance (87.5% overall accuracy on the validation set) in the 156th epoch, triggering an early stopping mechanism. The learning rate was adjusted from 1×10⁻⁶ using a cosine annealing strategy. -4 Smooth decay to 1×10 -6 Each training session lasts approximately 45 minutes, with a total training time of approximately 117 hours.
[0038] (4) Performance evaluation of the MSTF-Net model
[0039] In this embodiment, the MSTF-Net model is evaluated on the test set, and evaluation metrics for each scoring dimension are calculated: Classification accuracy = number of correctly predicted samples / total number of samples, which is the proportion of samples whose predicted class is exactly the same as the true class.
[0040] Calculate the weighted F1 score:
[0041] Weighted-F1=Σ(n c / N)·F1 c
[0042] F1 c =2·P c ·R c / (P c +R c )
[0043] Among them, the accuracy P of each level is taken into account. c (Precision) and Recall R c (Recall) and according to the number of samples n c The weighted average is the F1 score for class c; Mean Absolute Error (MAE) = Σ|y pred -y true | / N, calculate the mean of the absolute differences between the predicted score and the actual score; Root mean square error RMSE = √(Σ(y) pred -y true )² / N), calculates the square root of the mean of the squared differences between the predicted score and the actual score, and is more sensitive to larger errors; Spearman rank correlation coefficient Spearman-ρ=1-6Σd i² / (N(N²-1)), evaluates the consistency between predicted scores and actual scores in ranking, with a value range of [-1, 1]. The closer the value is to 1, the better the consistency in ranking. Cohen-κ coefficient (p) o -p e ) / (1-p e To assess the true consistency of classification results after excluding random consistency, p o To observe the consistency rate, p e κ represents the expected consistency rate, where κ > 0.8 indicates almost perfect consistency.
[0044] The detailed indicators for each rating dimension are as follows: The accuracy rate for the eating speed rating was 87.3% (95% CI: 84.1-90.5%), with a weighted F1 score of 0.864, MAE of 0.312, RMSE of 0.428, and Spearman correlation coefficient ρ of 0.923; the accuracy rate for the eating enthusiasm rating was 85.6%, with a weighted F1 score of 0.849, MAE of 0.347, and Spearman-ρ of 0.908; the accuracy rate for the emotional state rating was 82.4%, with a weighted F1 score of 0.817, MAE of 0.398, and Spearman-ρ of 0.8. 76; The accuracy rate of the body posture score was 84.1%, the weighted F1 score was 0.833, the MAE was 0.372, and the Spearman-ρ was 0.891; the accuracy rate of the food integrity score was 89.5%, the weighted F1 score was 0.891, the MAE was 0.268, and the Spearman-ρ was 0.945; the accuracy rate of the comprehensive craving score was 86.8% (95% CI: 83.6-90.0%), the weighted F1 score was 0.862, the MAE was 0.324, the RMSE was 0.445, the Spearman-ρ was 0.917, and the Cohen-κ was 0.834.
[0045] like Figure 6 As shown, the performance of the MSTF-Net model of this invention is compared with other deep learning methods: the pure CNN method based on ResNet-50 has an overall preference accuracy of 78.2% ± 1.8%; the pure Transformer method based on ViT-B / 16 has an accuracy of 81.5% ± 1.5%; the temporal modeling method based on ResNet-50+LSTM has an accuracy of 82.9% ± 1.6%; the video understanding method based on TimeSformer has an accuracy of 84.1% ± 1.4%; and the MSTF-Net method of this invention has an accuracy of 86.8% ± 1.2%. One-way ANOVA showed significant differences among the methods (F=18.7, p<0.001), and the post-hoc Tukey HSD test showed that the differences between MSTF-Net and all other compared methods were statistically significant (p<0.05).
[0046] like Figure 7 As shown, the normalized confusion matrix of the MSTF-Net model exhibits high values for the main diagonal elements, with classification accuracies of 85.0%, 84.0%, 86.0%, 88.0%, and 90.0% for levels 1 to 5, respectively. There is a certain degree of confusion between adjacent levels (e.g., level 2 being misclassified as level 1 or 3), consistent with the ordinal characteristics of preference scoring, and the misclassification rate across two or more levels is less than 3%.
[0047] like Figure 8As shown, the ROC curves of the MSTF-Net model on the five-class classification task indicate that each class has good classification performance. The AUC values for levels 1 to 5 are 0.94, 0.92, 0.93, 0.95, and 0.96, respectively. The macro-AUC is 0.94 and the micro-AUC is 0.95.
[0048] like Figure 9 As shown, the distribution of cat and dog preferences for eight tested food items was analyzed. Cats showed the highest preference for wet food B (canned tuna), with an average composite score of 4.52±0.48 (N=180), followed by treat A (freeze-dried chicken) at 4.38±0.55. Their preference for dry food C (beef flavor) was relatively low, with an average score of 2.87±0.82. One-way ANOVA showed a significant difference in cats' preferences for different foods (F=42.3, p<0.001). Dogs showed the highest preference for wet food A (canned chicken), with an average composite score of 4.41±0.52 (N=180), followed by dry food C (beef flavor) at 4.23±0.58. Their preference for treat B (nutritional paste) was relatively low, with an average score of 3.12±0.91. This result is consistent with the findings in pet behavior studies where cats prefer fish protein and dogs prefer meat protein, thus validating the effectiveness of this model.
[0049] like Figure 10 As shown, the contribution of various behavioral features to the assessment of cravings was analyzed by visualizing the attention weights of the MSTF-Net model. The average contribution of eating action features (licking / chewing frequency) was the highest at 28.5% ± 2.1%, followed by approach behavior features (approach speed / delay time) at 22.3% ± 1.8%, facial expression features (ears / eyes / mouth) at 18.7% ± 1.5%, body posture features (trunk / limbs) at 16.2% ± 1.4%, and tail movements at 14.3% ± 1.2%. The differences in the contribution of the five features were statistically significant (χ²=89.4, p<0.001). These results indicate that eating actions and approach behaviors are the most important indicators reflecting cravings, consistent with the conclusions of behavioral studies.
[0050] like Figure 11As shown, the temporal characteristics of the eating process were compared and analyzed between high-interest samples (comprehensive score 4-5, N=156) and low-interest samples (comprehensive score 1-2, N=98). The eating speed of the high-interest samples rapidly increased to a peak within 5 seconds of starting eating (normalized value approximately 0.85±0.08) and remained stable until the end of eating, with eating enthusiasm consistently maintained at a high level (normalized value 0.75-0.85). The eating speed of the low-interest samples increased slowly and had a lower peak (normalized value approximately 0.45±0.12), with obvious pauses and fluctuations during eating (average number of pauses 2.3±1.1 times), and eating enthusiasm gradually decreased over time (slope -0.003 / second). The differences between the two groups in peak eating speed and number of pauses were statistically significant (independent samples t-test p<0.001).
[0051] like Figure 12 As shown, the individual calibration effect of the species perception dynamic weight allocation module was verified. Ten test pets were selected (5 cats, numbered P1-P5, and 5 dogs, numbered P6-P10), and the assessment error (MAE) before and after applying individual calibration was calculated. Without individual calibration (using only fixed species base weights), the average MAE of the 10 pets was 0.52±0.18, with large differences in error between individuals (range 0.31-0.78, coefficient of variation CV=34.6%). After applying individual calibration (using dynamic weights W_dynamic), the average MAE decreased to 0.31±0.08, and the differences in error between individuals significantly decreased (range 0.22-0.42, CV=25.8%). Paired-samples t-test showed that the difference before and after calibration was highly statistically significant (t=8.72, df=9, p<0.001). The individual calibration mechanism reduced the mean of cross-individual assessment error by 40.4% and the standard deviation by 55.6%.
[0052] like Figure 13As shown, the effectiveness of each functional module of the MSTF-Net model was verified through ablation experiments. The overall preference accuracy of the complete model was 86.8% ± 1.2%. After removing the temporal feature extraction branch, the accuracy decreased to 81.2% ± 1.8%, a decrease of 5.6 percentage points (relative decrease of 6.5%), indicating that temporal features are crucial for capturing dynamic changes in feeding behavior. After removing the cross-modal attention fusion module, the accuracy decreased to 83.5% ± 1.5%, a decrease of 3.3 percentage points, indicating that deep fusion of spatial and temporal features significantly improved model performance. After removing the dynamic weight allocation module, the accuracy decreased to 84.6% ± 1.4%, a decrease of 2.2 percentage points, indicating that species perception and individual adaptation mechanisms effectively improved the model's generalization ability. After removing the TCAB mechanism, the accuracy decreased to 85.1% ± 1.3%, a decrease of 1.7 percentage points, indicating that temporal convolutional attention bias effectively enhanced the modeling ability of local temporal patterns. The differences between each ablation variant and the complete model were statistically significant (paired t-test p < 0.05).
[0053] S2. Collect video data of the pet's eating process and preprocess it.
[0054] In this embodiment, the technical requirements for collecting video data of the pet's eating process using a camera device are as follows: the camera device is an RGB digital camera with the following specifications: the image sensor type is CMOS (Complementary Metal-Oxide-Semiconductor); the sensor size is not less than 1 / 2.8 inches (diagonal length ≥ 6.35 mm); the effective pixel count is not less than 2 million pixels (5 million pixels or more recommended); and the output resolution is not less than 1920×1080 pixels (Full resolution). HD resolution, with higher resolutions of 2560×1440 or 3840×2160 preferred. Video frame rate should be no less than 30fps (frames per second), with 60fps recommended for capturing finer motion details. Horizontal field of view (HFOV) should be 90°-120°, with 100°-110° recommended to cover the entire feeding area at a suitable distance. Automatic exposure (AE) and automatic white balance (AWB) are supported to adapt to different lighting conditions. Automatic gain control (AGC) is supported in low light environments with a minimum illumination of 1 lux (F2.0 aperture). Lens focal length should be adjustable from 2.8-12mm (zoom lens) or 3.6mm (fixed focal length lens), with an aperture of F1.8-F2.8. H.264 (AVC) or H.265 (HEVC) video encoding formats are supported; H.265 can reduce the bitrate by approximately 50% while maintaining the same image quality. Recommended models include [list of models]. Suitable camera brands include C920 / C922 / BRIO, Hikvision DS-2CD series, and Dahua IPC-HDW series. The camera should be installed directly in front of the pet's food bowl at a 45°±5° downward angle (i.e., the angle between the lens optical axis and the horizontal plane is 45°), with the lens optical axis pointing towards the center of the food bowl. The horizontal distance between the camera and the center of the food bowl should be 60-80cm (70cm recommended), and the vertical height should be 50-70cm (60cm recommended). Ensure the field of view completely covers the pet's face, torso, and limbs while eating, and includes at least 80% of the pet's body area. Video recording is automatically triggered when a pet enters a 1-meter radius around the food bowl (this can be achieved using an infrared sensor or motion detection algorithm). Recording ends when the pet leaves the food bowl area (the center of the bounding box is more than 50cm from the center of the food bowl) for more than 30 seconds or when the recording duration reaches the preset maximum of 10 minutes.
[0055] The detailed technical solution for preprocessing the acquired video data includes: using an adaptive background modeling algorithm based on a Gaussian Mixture Model (GMM) for foreground separation. Each pixel position is modeled as a mixture distribution of K Gaussian components. The number of Gaussian components K in the GMM is set to 5, the learning rate α is set to 0.01 (to control the background update speed), the background ratio threshold T is set to 0.7 (components with accumulated weights exceeding T are considered background), the initial variance is set to 36 (grayscale value), and the variance threshold is set to 16 (to determine whether a pixel matches a certain Gaussian component). Bilateral filtering is used to perform edge-preserving denoising on the separated foreground region. Bilateral filtering considers both spatial distance and grayscale similarity, with the spatial domain standard deviation σ... s Set to 9 (pixels), grayscale standard deviation σ r Set to 75 (grayscale value), filter window size is For low-light environments (average brightness < 50 / 255), contrast-limited adaptive histogram equalization (CLAHE) is used to enhance the brightness of video frames. The clip limit parameter is set to 2.0 (limiting the contrast amplification factor), and the tileGridSize is set to 8×8 pixels. Each tile is independently histogram equalized and then bilinearly interpolated and fused. The video frame sampling strategy adopts adaptive keyframe extraction. The default sampling rate is 8 frames per second (1 frame is taken every 3.75 frames). When the target detection module detects that the distance between the pet's nose tip key point and the edge of the food bowl is less than 5cm, it is determined to be in the feeding state (Active Eating State) and the sampling rate is automatically increased to 16 frames per second (1 frame is taken every 1.875 frames) to capture more detailed feeding actions. The position of the food bowl edge can be obtained through initial calibration or automatic detection.
[0056] S3. Use a pre-trained MSTF-Net model to perform inference on the pre-processed pet feeding process video data.
[0057] In this embodiment, a pet eating preference intelligent evaluation system is constructed based on a trained MSTF-Net model.
[0058] System hardware configuration: The main control unit uses an NVIDIA Jetson Orin Nano edge computing module, equipped with a 6-core Arm Cortex-A78AE CPU and a 1024-core Ampere GPU, 8GB LPDDR5 memory (102GB / s bandwidth), supports 40TOPS INT8 inference computing power, and has an adjustable TDP of 7-15W; the image acquisition unit uses a Sony IMX219 CMOS image sensor, 1 / 4-inch sensor size, 8 megapixels (3280×2464), supports 1920×1080@30fps and 1280×720@60fps video acquisition, F2.0 aperture, and a 77° field of view; the storage unit includes 64GB eMMC 5.1 internal storage (300MB / s read speed) and 128GB microSD expansion storage (UHS-I U3); the display unit uses a 3.5-inch IPS LCD touchscreen with a resolution of 480×320 pixels and a brightness of 350 nits, used for local display of evaluation results; the communication unit supports WiFi. It features 802.11ax (802.11ax, maximum speed 1.2Gbps) and Bluetooth 5.0 (BLE Low Energy mode) for data synchronization with cloud servers and user mobile apps.
[0059] System software architecture: The operating system uses Ubuntu 20.04 LTS for Jetson (JetPack 5.1.1); the deep learning inference framework uses NVIDIA TensorRT 8.5, and the MSTF-Net model trained by PyTorch is exported and converted into a TensorRT engine (.engine file) via ONNX format, using FP16 mixed precision inference to balance accuracy and speed; the application layer uses Python 3.8+Flask 2.3 framework to build a local RESTful API service, providing interfaces such as / predict, / history, and / config; the front end uses WeChat Mini Program (base library version 2.30+) to implement a cross-platform user interface, supporting iOS and Android systems.
[0060] System Workflow: When the built-in pyroelectric infrared sensor (PIR, detection distance 3 meters, response time <0.5 seconds) detects a pet approaching the feeder, video acquisition is triggered, and the camera begins recording the feeding process. The recorded video data is encoded with H.264 and transmitted in real time to Jetson Orin Nano for local inference. The MSTF-Net TensorRT engine outputs the preference score results. The evaluation results include 5 sub-dimension scores and 1 comprehensive level, stored in JSON format in a local SQLite 3 database (approximately 2KB per record), and simultaneously uploaded to Alibaba Cloud OSS storage service via HTTPS encrypted transmission. Users can view the evaluation results, historical preference trend curves (supporting filtering by food type / time period), and personalized feeding suggestions (food recommendations based on preference ranking) in real time through a WeChat mini program.
[0061] Performance metrics: The MSTF-Net TensorRT engine achieves a single complete inference time of approximately 1.2 seconds on Jetson Orin Nano (64 frames of input, FP16 precision), including approximately 0.3 seconds for preprocessing (object detection + tracking + keypoints), approximately 0.6 seconds for feature extraction, and approximately 0.3 seconds for fusion and classification, meeting near real-time evaluation requirements (latency <2 seconds); the system's standby power consumption is approximately 2W (only PIR sensor active), and its operating power consumption is approximately 8W (full load inference); the storage space for a single evaluation record is approximately 50KB (including 2 keyframe thumbnails 320×240 JPEG and scoring data JSON).
[0062] The developed intelligent pet eating preference assessment system has undergone a 3-month beta test in 50 households (25 cat owners and 25 dog owners). During the test, the system completed 12,680 valid preference assessments (6,340 for cats and 6,340 for dogs). The system's stability indicators were: crash rate <0.1% and inference success rate >99.5%. User satisfaction survey results (5-point Likert scale) showed that 92% of users believed the system's output preference scores were largely consistent with their observations (score ≥4); 88% of users considered the personalized feeding suggestions generated by the system to be valuable (score ≥4); and 95% of users indicated they were willing to continue using the system (NPS: 72).
[0063] Although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the claims.
Claims
1. A method for intelligently assessing a pet's eating habits, characterized in that, Includes the following steps: S1. Pre-build and train the MSTF-Net model, including the construction of a standardized dataset of pet eating behavior, dataset partitioning and preprocessing, MSTF-Net model construction and training, and MSTF-Net model performance evaluation. The MSTF-Net model consists of six functional modules: video preprocessing module, spatial feature extraction branch, temporal feature extraction branch, cross-modal attention fusion module, species-aware dynamic weight allocation module, and multi-task classification and regression head. S2. Collect video data of the pet eating process to be evaluated using camera equipment, and preprocess the collected video, including video decoding and framing, adaptive background modeling and foreground separation; S3. Using a pre-trained MSTF-Net model, inference is performed on the pre-processed pet eating process video data to obtain the pet eating preference assessment results, wherein the pet eating preference assessment results include eating speed score, eating enthusiasm score, emotional state score, body posture score and comprehensive preference score.
2. The intelligent assessment method for pet eating preferences according to claim 1, characterized in that, In step S1, the construction of the standardized dataset for pet eating behavior specifically includes: Recruit n experimental pets as data collection subjects, and ensure that the experimental pets fully cover different breeds, age groups and body sizes; Prepare k different types of pet food as test samples, and ensure that the pet food fully covers different categories and flavors. The amount of each food should be calculated according to 2%-3% of the pet's body weight. Establish a standardized experimental environment: control the environmental parameters as follows: light intensity 300-500 lux, color temperature 4000-5000K, ambient temperature 20-25℃, relative humidity 40%-60%, background noise below 45dB, use an RGB camera with a resolution of not less than 1920×1080 pixels, a frame rate of not less than 30fps, and a color depth of 8bit, install the camera directly in front of the pet food bowl at a 45°±5° downward angle, with the lens optical axis at a 45° angle to the ground, the horizontal distance between the camera and the center of the food bowl is 60-80cm, the vertical height is 50-70cm, and the field of view covers the food bowl and an area with a radius of 50cm around it; Each pet was tested three times independently for each food, with an interval of at least 24 hours between two consecutive tests to eliminate the effects of satiety and memory. The pets were fasting for 4-6 hours before each test. The complete process of the pet from entering the camera's field of vision to leaving the food bowl area was recorded for each test, and videos within a specific time range were selected as valid videos. Each video was independently annotated by at least five annotators with backgrounds in pet behavior or animal science. The annotation dimensions included five dimensions: eating speed rating, eating enthusiasm rating, emotional state rating, body posture rating, and overall preference rating. Each dimension was rated using a 1-5 level Likert scale, where level 1 represents the lowest preference, level 2 represents the lowest preference, level 3 represents the highest preference, level 4 represents the highest preference, and level 5 represents the highest preference. The median of the five annotators' ratings for each video was used as the final label value to reduce the influence of extreme values. The Fleiss' Kappa coefficient was also calculated, and annotated data with a Fleiss' Kappa value of not less than 0.70 were selected as valid annotated data to ensure annotation quality.
3. The intelligent assessment method for pet eating preferences according to claim 1, characterized in that, In step S1, the dataset partitioning and preprocessing specifically include: The entire dataset was randomly divided into training, validation and test sets according to a ratio of 70%, 15% and 15%, respectively. During the division, it was strictly ensured that all video samples of the same pet belonged to the same dataset to avoid data leakage. Stratified sampling is used to ensure that the distribution of preference levels in each dataset is consistent with the overall distribution. The video frames are standardized by normalizing pixel values to the [0,1] interval, subtracting the mean of the ImageNet dataset [0.485, 0.456, 0.406] and dividing by the standard deviation [0.229, 0.224, 0.225], and finally the pixel value distribution approximates the standard normal distribution N(0,1).
4. The intelligent assessment method for pet eating preferences according to claim 1, characterized in that, In step S1, the MSTF-Net model construction and training specifically include: The video preprocessing module uses a pet target detection network based on the YOLOv8-L architecture to detect pets in each frame of the input video; it employs the DeepSORT multi-target tracking algorithm to perform cross-frame correlation and continuous tracking of the detected pet targets; and it uses a pet keypoint detection network based on the HRNet-W48 architecture to locate 17 body keypoints of the pet, including: point 1 is the tip of the nose, point 2 is the center of the left eye, point 3 is the center of the right eye, point 4 is the base of the left ear, point 5 is the base of the right ear, point 6 is the tip of the left ear, point 7 is the tip of the right ear, point 8 is the shoulder joint of the left forelimb, point 9 is the shoulder joint of the right forelimb, point 10 is the paw tip of the left forelimb, point 11 is the paw tip of the right forelimb, and point 12 is the spine in the back. Points 13, 14, 15, 16, and 17 are the tail root, left hind leg hip joint, right hind leg hip joint, left hind leg paw tip, and right hind leg paw tip, respectively. The original video frames are cropped based on the pet detection bounding boxes, with the cropped area being 20% larger than the bounding box to preserve contextual information. The cropped images are then uniformly adjusted to 224×224 pixels using bilinear interpolation, generating a sequence of full-body pet images. Facial region bounding boxes are calculated based on five facial key points: nose tip, eyes, and ear roots. The bounding box is the smallest bounding rectangle containing these five points, expanded by 30%. The facial region is cropped and uniformly adjusted to 112×112 pixels using bilinear interpolation, generating a sequence of pet facial region images. The spatial feature extraction branch adopts a three-stream parallel architecture, including three parallel sub-networks: global appearance feature stream, local facial feature stream, and skeletal pose feature stream. This three-stream architecture can capture feature information at different spatial scales and semantic levels: the global appearance feature stream uses Swin Transformer V2-Base as the backbone network, with input being a 224×224×3 pixel image of the pet's full-body region. V2-Base employs a hierarchical shift-window attention mechanism, comprising four stages. The feature map resolutions for each stage are 56×56, 28×28, 14×14, and 7×7, respectively. The embedding dimensions for each stage are 128, 256, 512, and 1024, respectively. The number of Transformer blocks in each stage is 2, 2, 18, and 2, totaling 24 blocks. Each block contains alternating stacks of multi-head self-attention for windows and multi-head self-attention for shift windows. The window size is 7×7, and the number of attention heads is 4, 8, 16, and 32, respectively. Log-CPB relative position encoding is used, and the weights are initialized using pre-trained weights from the ImageNet-22K dataset. The output of the final stage is passed through a global average pooling layer to obtain a 1024-dimensional global appearance feature vector F. global The local facial feature flow uses EfficientNet-B4 as the backbone network. The input is a 112×112×3 pixel pet facial region image. EfficientNet-B4 employs a compound scaling strategy with a depth coefficient α=1.4, a width coefficient β=1.2, and a resolution coefficient γ=1.
3. It contains 7 MBConv module stages, each stage using MBConv1 and MBConv6×6 stages respectively. The number of layers in each stage is 2, 4, 4, 6, 6, 8, and 2, for a total of 32 layers. The number of output channels is 24, 32, 56, 112, 160, 272, and 448 respectively. A Squeeze-and-Excitation attention mechanism is used, with an SE reduction ratio of 4. Weights are initialized using pre-trained weights from the ImageNet-1K dataset. The output of the last stage passes through a global average pooling layer and a 1×1 convolutional layer to obtain a 512-dimensional facial feature vector F. face The skeletal pose feature stream uses a spatiotemporal graph convolutional network to encode the keypoint sequence. Seventeen keypoints are constructed as a graph structure. Node features consist of three dimensions: the normalized (x, y) coordinates and confidence score of the keypoint. The total number of nodes is V=17. There are 16 edges following the anatomical topology of a pet skeleton. The adjacency matrix A is a 17×17 symmetric matrix. The graph convolutional network contains four spatiotemporal graph convolutional blocks. Each block contains a spatial graph convolutional layer, a temporal convolutional layer, and residual connections. The output channels of the four blocks are 64, 128, 256, and 256 respectively. Each block is followed by a BatchNorm layer and a ReLU activation function. Finally, global average pooling is used to obtain a 256-dimensional skeletal pose feature vector F. pose ; the global appearance feature vector F global Facial feature vector F face and skeletal pose feature vector F pose A 1792-dimensional feature vector is obtained by concatenating along the channel dimension. This vector is then fused using a multilayer perceptron containing two fully connected layers. The first fully connected layer maps the 1792 dimensions to 1024 dimensions and applies GELU activation and Dropout regularization with a Dropout probability p=0.
1. The second fully connected layer maps the 1024 dimensions to 512 dimensions, resulting in a 512-dimensional comprehensive spatial feature vector F. spatial ; The temporal feature extraction branch employs a Transformer-based temporal encoder design to capture the dynamic changes and temporal dependencies in pet eating behavior: it extracts the spatial feature vector F from T consecutive video frames. spatial The input sequence is composed of elements, and the input tensor has a shape of (B, T, D), where B is the batch size, T is the sequence length, and D is the feature dimension. First, the input feature sequence is encoded using a learnable temporal position encoding method. This method employs a hybrid encoding approach, combining sine and cosine position encoding with learnable position embeddings. Sine and cosine encoding provides absolute positional information, while the learnable embeddings capture task-related positional patterns. The position encoding dimension is the same as the feature dimension, both being 512-dimensional. The position encoding matrix PE∈R {T×D} The temporal encoder consists of N stacked Transformer encoder layers. Each encoder layer contains a multi-head self-attention sub-layer, a feedforward neural network sub-layer, two layer normalization operations, and two residual connections, employing a Pre-LN structure. The multi-head self-attention sub-layer uses H parallel attention heads, each head having a dimension d. k =d v =D / H, the total computational complexity of attention is O=T²D, and the attention calculation uses the scaled dot product attention mechanism: Attention(Q,K,V)=softmax(QK). T / √d k In the multi-head self-attention mechanism, a temporal convolutional attention bias module (TCAB) is introduced. The TCAB extracts local temporal features from the key matrix K using one-dimensional depthwise separable convolutions. The kernel size of the one-dimensional convolution is K. size =5, stride=1, padding=2 to maintain sequence length, depthwise separable convolution parameters are only 1 / D of standard convolution, the convolution output is added to the original attention score as an attention bias term, the calculation formula of the temporal convolutional attention bias module TCAB is expressed as Attention TCAB(Q,K,V) =softmax(QK T / √d k +Conv1D dw(K) Let Q, K, and V be the query matrix, key matrix, and value matrix, respectively, all with shapes (B, H, T, d). k ), d k Conv1D represents the dimension of the key vector. dw This represents a one-dimensional depthwise separable convolution operation; the feedforward neural network sublayer contains two linear transformation layers and a GELU activation function. The first linear layer expands the 512-dimensional array to 2048-dimensional arrays, and the second linear layer compresses the 2048-dimensional arrays back to 512-dimensional arrays. Dropout regularization is applied between the two linear layers with a dropout probability p=0.1; the temporal encoder outputs a sequence of feature vectors over T time steps, with an output tensor shape of (B,T,D). These feature vectors are adaptively aggregated into a single feature vector by a temporal attention pooling layer. The temporal attention pooling layer first uses a learnable query vector q∈R... D Calculate the attention score a for the feature vector at each time step. t =softmax(q·h t / √D), and then use the normalized attention score as the weight to perform a weighted sum F on the feature vectors at each time step. temporal =Σ(a t ·h t Finally, a 512-dimensional temporal feature vector F is obtained. temporal ; The cross-modal attention fusion module employs a design combining a bidirectional cross-attention mechanism and a gated fusion mechanism to achieve deep semantic alignment and information fusion of spatial and temporal features: the spatial-to-temporal cross-attention calculation process involves converting the spatial feature vector F... spatial ∈R D Through the linear projection matrix W Qs ∈R {D×D} Projection yields the query vector Q s =W Qs ·F spatial ∈R D The time series feature vector F temporal ∈R D Through the linear projection matrix W Kt ∈R {D×D} and W Vt ∈R {D×D} The key vector K is obtained by projecting it separately. t =W Kt ·F temporal ∈R D Sum vector V t =W Vt ·F temporal ∈R D The attention level of spatial features to temporal features is calculated by scaling the dot product attention, with the attention weight being a scalar α. s2t =softmax(Q s ·K t T / √D The spatial feature vector F enhanced with temporal information is obtained. s2t =α s2t ·V t ∈R D Where D is the vector dimension equal to 512; the calculation process of temporal to spatial cross-attention is symmetrical to the above, and the temporal feature vector F is... temporal Through linear projection W Qt Obtain the query vector Q t , the spatial feature vector F spatial Through linear projection W Ks and W Vs Obtain the key vector K s Sum vector V s The spatial information-enhanced temporal feature vector F is calculated. s2t =α s2t ·V t ∈R D The gating fusion mechanism will convert the original spatial features F spatial Original time series features F temporal Cross-enhancement feature F s2t and F t2s Four 512-dimensional feature vectors are concatenated along the channel dimension to obtain a 2048-dimensional vector F. concat =[F spatial ;F temporal ;F s2t ;F t2s ]∈R {4D} The adaptive fusion weights are calculated using a gated network, which consists of a linear layer and a sigmoid activation function σ(·), outputting a 4D=2048 dimensional gated vector G=σ(W). g ·F concat +b g )∈R {4D}, Where b g ∈R {4D} As the bias vector, F is obtained by fusing the feature vectors through element-wise multiplication. gated =G⊙F concat ∈R {4D} Finally, the 4D=2048 dimensions are compressed to D'=1024 dimensions using a mapping network containing a linear layer, resulting in the final fused feature vector F. fused =W proj ·F gated ∈R {D'} Where D'=1024; The species perception dynamic weight allocation module is used to adaptively adjust the weight coefficients of the five evaluation dimensions based on the pet species category and individual behavioral characteristics. First, it fuses the feature F using the species classification head. fused For classification, the species classification head consists of two fully connected layers. The first fully connected layer maps the 1024 dimensions to 256 dimensions and applies the ReLU activation function. The second fully connected layer maps the 256 dimensions to 2 dimensions and applies the Softmax activation function to output the class probability distribution P. species =[p cat ,p dog ]∈R 2 Based on the species classification results, retrieve the corresponding basic weight vector W from the pre-set species weight database. base ∈R 5 Simultaneously, the fusion feature F fused The input is an individual feature encoder, which consists of three fully connected layers with dimensionality transformations of 1024, 512, 256, and 128 respectively. ReLU activation and Dropout regularization are applied between layers, with a Dropout probability p=0.
2. The output is a 128-dimensional individual feature vector F. individual ∈R {128} This vector encodes the current individualized behavioral pattern of the pet; the individual feature vector F individual The input weight generation network consists of two fully connected layers with dimensionality transformations of 128→64→5. The middle layers use the ReLU activation function, and the final layer uses the Softmax activation function to ensure that the sum of the five output weight values is 1, resulting in the individual weight adjustment vector W. adjust ∈R 5 The final dynamic weight vector W dynamic ∈R 5 The weighted combination of the base weight and the individual adjustment weight is calculated using the formula W. dynamic =α·W base +(1-α)·W adjust Where α is a learnable balance coefficient, through α raw After transformation, α = 0.5 + 0.4·σ(α) raw The constraint values are within the range of [0.5, 0.9], ensuring that the basic weights always dominate while allowing for individual adjustments; The multi-task classification and regression head comprises six output branches: five parallel sub-dimensional rating heads and one comprehensive preference rating head. It employs a hard-parameter-shared multi-task learning architecture. The five sub-dimensional rating heads correspond to eating speed, eating enthusiasm, emotional state, body posture, and eating completeness ratings, respectively. Each sub-dimensional rating head uses the same network structure, consisting of three fully connected layers with dimensionality transformations of 1024, 256, 64, and 5, respectively. The first two layers use ReLU activation and Dropout regularization, with a Dropout probability p=0.3 to enhance generalization. The last layer uses Softmax activation to output the probability distribution P of the five preference levels. i ∈R 5 (i=1,2,3,4,5); The eating speed scoring head introduces an attention gating mechanism at the input end to affect the fused features F. fused Mid- and temporal features F temporal The relevant components are weighted and enhanced, with a focus on dynamic features such as licking frequency and chewing frequency; the feeding enthusiasm scoring head focuses on behavioral indicators such as the pet's movement speed when approaching food, the delay time of first contact with food, and the number of times the pet leaves the food bowl; the emotional state scoring head introduces an attention gating mechanism at the input end to affect the fused features F. fused Central and facial features F face The relevant components are weighted and enhanced, with a focus on facial expression features such as ear orientation angle, eye opening degree, pupil size changes, and beard extension degree; the body posture scoring head introduces an attention gating mechanism at the input end to enhance the fused features F. fused In terms of skeletal pose features F pose The relevant components are weighted and enhanced, with a focus on body language characteristics such as tail wagging frequency and amplitude, back arching, and standing posture. The feeding integrity scoring head scores the feeding integrity by analyzing the effective feeding time percentage and the trend of food intake changes throughout the feeding process. The comprehensive preference scoring head assigns a dynamic weight vector W to the probability distribution of the scores for the five sub-dimensions. dynamic To perform weighted fusion, first, the 5-dimensional probability distribution P of each sub-dimension is... i The corresponding rank value vector L=[1,2,3,4,5] T The expected score is obtained by performing a dot product calculation. i =P i ·L=Σ(P i [j]·j) (j=1,2,3,4,5), and then the five expected scores are dynamically weighted and summed to obtain the comprehensive score Score. total =Σ(W dynamic [i]·Score i (i=1,2,3,4,5), and finally the comprehensive score is mapped to an integer level of 1-5 through four thresholds as the final comprehensive preference level output; The MSTF-Net model training strategy employs a multi-task joint learning framework, simultaneously optimizing five sub-dimensional scoring tasks, one comprehensive scoring task, and one auxiliary species classification task. The sub-dimensional scoring tasks utilize the label-smoothed cross-entropy loss function, calculated using the formula L. CEsmooth =-Σ((1-ε)·y true [c]+ε / K)·log(y pred [c]), where ε is the label smoothing coefficient, c is the category index, K is the number of categories equal to 5, and y true For the one-hot encoding of the real label, y pred The predicted probability distribution is the output of Softmax; the comprehensive scoring task uses a weighted combination of the mean squared error loss function and the ordinal regression loss function. The mean squared error loss is used to calculate the squared difference L between the predicted comprehensive score and the true comprehensive score. MSE =(Score pred -Score true )², the ordinal regression loss uses the CORAL framework to transform the 5-level classification problem into 4 binary classification sub-problems to maintain the orderliness of the rating levels. ordinal =-Σ(y k ·log(σ(f k ))+(1-y k )·log(1-σ(f k (k=1,2,3,4), where y k f is the true label for the k-th binary sub-classification problem. k The logit is the model output; the standard cross-entropy loss L is used for the auxiliary species classification task. species =-Σy sp ·log(P species The total loss function is defined as L. total =Σ(λ i ·L CEsmoothi )+λ6·L MSE +λ7·L ordinal +λ8·L species L CEsmoothi Let λ be the label smoothing cross-entropy loss for the i-th sub-dimension (i=1,2,3,4,5). i These are the weight coefficients for each loss term; the optimizer uses the AdamW algorithm with an initial learning rate of lr. init Set to 1×10 -4 Weight decay coefficient decay Set β1 to 0.05, β2 to 0.9, and ε to 1×10⁻⁶. -8 The learning rate scheduling strategy uses cosine annealing, with the learning rate gradually decreasing from its initial value to lr according to a cosine curve. min =1×10 -6 The attenuation formula is lr t =lr min +0.5·(lr init -lr min )·(1+cos(π·t / T max Total number of training rounds T max The training run is set to 200 rounds, with each round iterating through the entire training set once. Data augmentation strategies include random horizontal flipping, random affine transformation, random color jittering, random erasure, temporal random sampling, and Mixup augmentation. During training, a validation set is used to evaluate model performance. An early stopping mechanism is triggered when the validation set loss does not decrease for 15 consecutive rounds to prevent overfitting.
5. The intelligent assessment method for pet eating preferences according to claim 4, characterized in that, In step S1, the MSTF-Net model performance evaluation specifically includes: The MSTF-Net model was evaluated on the test set, and evaluation metrics for each scoring dimension were calculated, including classification accuracy, weighted F1 score, mean absolute error (MAE), root mean square error (RMSE), Spearman rank correlation coefficient (Spearman-ρ), and Cohen Kappa coefficient (Cohen-κ).