Pet health monitoring method, device, equipment, medium and product

By collecting multimodal data and analyzing large models, combined with veterinary knowledge graphs, the problem of single-dimensional pet health monitoring has been solved, enabling comprehensive monitoring and early warning of pet health, and improving the accuracy of identifying potential health problems.

CN121709236APending Publication Date: 2026-03-20GREE ELECTRIC APPLIANCE INC OF ZHUHAI +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511683842.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Current pet health monitoring technologies are limited to a single dimension, failing to meet the needs for comprehensive monitoring and early warning, and making it difficult to accurately identify potential health problems.

Method used

By acquiring multi-modal data (cameras, microphone arrays, smart litter boxes, etc.) from various sources, such as pet gait, eating speed, vocalizations, frequency and weight of excrement, we can conduct in-depth analysis by combining large models and veterinary knowledge graphs to achieve multi-modal data fusion and health risk assessment.

Benefits of technology

It enables comprehensive monitoring and early warning of pet health, improves the accuracy of identifying potential health problems, reduces false alarms and missed alarms, and meets pet owners' needs for accurate monitoring and timely intervention of pet health.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121709236A_ABST
    Figure CN121709236A_ABST
Patent Text Reader

Abstract

The invention provides a pet health monitoring method, device and equipment, a medium and a product, and is applied to the technical field of pet health management.The method comprises the steps that pet monitoring data of multiple modes is acquired; fusing the pet monitoring data of the multiple modes to obtain fused monitoring data, and analyzing the fused monitoring data by adopting a large model to obtain a health analysis result corresponding to each mode; according to a health analysis result, determining a health anomaly probability corresponding to each mode, and performing weighted summation on the health anomaly probabilities by adopting weight information corresponding to each mode to obtain a health risk value; according to the health risk value, pet health early warning is carried out, pet health monitoring is carried out according to the data of multiple modes, the requirements for comprehensive monitoring and early warning of pet health are met, the pet monitoring data of multiple modes are fused and analyzed through a large model, and the requirement for accurate recognition of potential health problems of pets is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pet health management technology, and in particular to a method, apparatus, equipment, medium and product for pet health monitoring. Background Technology

[0002] As people's living standards improve and the demand for pet companionship grows, pet health monitoring has gradually become a core concern for pet-owning families. However, current pet health monitoring technologies are limited in scope, only capable of monitoring basic feeding or single physiological indicators. This makes it difficult to meet the needs for comprehensive pet health monitoring and early warning. Furthermore, relying solely on simple data comparisons (such as comparing temperature data with preset thresholds) to determine abnormalities is insufficient for accurately identifying potential health problems in pets. Summary of the Invention

[0003] In view of the above problems, a method, apparatus, equipment, medium, and product for pet health monitoring are proposed to overcome or at least partially solve the above problems, including: A method for monitoring pet health, the method comprising: Acquire pet monitoring data in multiple modalities; The pet monitoring data from the various modalities are fused to obtain fused monitoring data. A large model is then used to analyze the fused monitoring data to obtain health analysis results for each modality. Based on the health analysis results, the probability of health abnormality corresponding to each modality is determined, and the weight information corresponding to each modality is used to perform a weighted summation of the probabilities of health abnormality to obtain the health risk value; Based on the stated health risk values, a health warning for the pet is issued.

[0004] Optionally, a large model is used to analyze the fused monitoring data to obtain health analysis results corresponding to each modality, including: A large model was used to conduct a preliminary analysis of the pet monitoring data for each modality in the fused monitoring data, and then the preliminary analysis results were combined with the pet monitoring data for other modalities for further analysis. Based on the results of the reanalysis, the health analysis results corresponding to each modality are determined.

[0005] Optionally, the multi-modal pet monitoring data includes pet video data collected by a camera module, and the preliminary analysis results of the pet video data include gait analysis results and feeding analysis results; a preliminary analysis of the pet monitoring data of each modality in the fused monitoring data is performed, including: Feature points are extracted from the pet video data in the fused monitoring data to obtain the current gait features, and the current gait features are compared with a preset gait feature template to obtain gait analysis results; The pet video data in the fused monitoring data is segmented into food area and container background area. The feeding speed is determined based on the amount of pixel reduction in the food area, and the feeding analysis result is determined based on the fluctuation range of the feeding speed.

[0006] Optionally, the multimodal pet monitoring data includes pet audio data collected by a microphone array module, and the preliminary analysis results of the pet audio data include emotion analysis results; a preliminary analysis of the pet monitoring data of each modality in the fused monitoring data is performed, including: Voiceprint features are extracted from the pet audio data in the fused monitoring data to obtain the current voiceprint features, and the emotion analysis results are determined based on the current voiceprint features.

[0007] Optionally, the multi-modal pet monitoring data includes pet excrement data collected by the pet excrement processing module, and the preliminary analysis results of the pet excrement data include excrement analysis results; preliminary analysis is performed on the pet monitoring data of each modality in the fused monitoring data, including: The pet excrement data in the fused monitoring data is subjected to multi-dimensional excrement analysis to obtain excrement analysis results; wherein, the pet excrement data includes excrement weight data, excrement biochemical data, and excrement image data, the excrement multi-dimensional analysis includes excrement weight analysis, excrement biochemical analysis, and excrement image analysis, and the excrement analysis results include excrement weight analysis results, excrement biochemical analysis results, and excrement image analysis results.

[0008] Optionally, the pet monitoring data from the multiple modalities are fused to obtain fused monitoring data, including: The pet monitoring data from the various modalities were normalized in dimensions, and an attention mechanism was used to fuse the normalized pet monitoring data to obtain fused monitoring data.

[0009] Optionally, before performing dimensionality normalization on the pet monitoring data of the various modalities, the method further includes: Spatiotemporal alignment was performed on the pet monitoring data of the various modalities.

[0010] A device for monitoring pet health, the device comprising: The pet monitoring data acquisition module is used to acquire pet monitoring data in multiple modalities. The health analysis module is used to fuse the pet monitoring data of the multiple modalities to obtain fused monitoring data, and to analyze the fused monitoring data using a large model to obtain the health analysis results corresponding to each modality; The health risk value determination module is used to determine the probability of health abnormality corresponding to each modality based on the health analysis results, and to perform a weighted summation of the probabilities of health abnormality using the weight information corresponding to each modality to obtain the health risk value. The pet health early warning module is used to issue pet health early warnings based on the stated health risk values.

[0011] An electronic device includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the method described above.

[0012] A computer-readable storage medium on which a computer program is stored, which, when executed by a processor, implements the method described above.

[0013] A computer program product includes a computer program that, when executed by a processor, implements the method described above.

[0014] The embodiments of the present invention have the following advantages: In this embodiment of the invention, pet monitoring data of multiple modalities is acquired; the pet monitoring data of multiple modalities is fused to obtain fused monitoring data, and a large model is used to analyze the fused monitoring data to obtain health analysis results corresponding to each modality; based on the health analysis results, the probability of health abnormality corresponding to each modality is determined, and the weight information corresponding to each modality is used to perform weighted summation of the probability of health abnormality to obtain a health risk value; based on the health risk value, pet health early warning is performed, realizing pet health monitoring based on multiple modal data, meeting the needs of comprehensive pet health monitoring and early warning, and the fusion of pet monitoring data of multiple modalities and analysis using a large model meet the needs of accurate identification of potential health problems in pets. Attached Figure Description

[0015] To more clearly illustrate the technical solution of the present invention, the accompanying drawings used in the description of the present invention will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a flowchart of a pet health monitoring method provided in some embodiments of the present invention; Figure 2 This is a flowchart of the steps of a method for monitoring pet health provided in some embodiments of the present invention; Figure 3 This is a flowchart of the steps of a second method for pet health monitoring provided in some embodiments of the present invention; Figure 4 This is a structural block diagram of a pet health monitoring device provided in some embodiments of the present invention. Detailed Implementation

[0017] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0018] In related technologies, the monitoring dimensions are relatively singular, such as monitoring based solely on basic feeding conditions or a single physiological indicator. Specifically, a single sensor (such as the weight sensor of a smart feeder) is used to monitor the pet's food intake; or simple data comparison (comparing the pet's body temperature data with a preset threshold) is used to determine whether the pet is exhibiting abnormalities.

[0019] Because it relies on only a single data point, it cannot comprehensively reflect the pet's health status. It lacks in-depth analysis of the pet's health condition and cross-modal correlation analysis, resulting in poor accuracy in early warnings. Furthermore, the lack of integrated analysis of multi-dimensional data and support from professional knowledge leads to low accuracy in early warnings, with a high probability of false alarms and missed alarms. This makes it difficult to accurately identify potential health problems in pets, and consequently, to discover various potential health risks, failing to meet pet owners' needs for comprehensive monitoring and early warning of their pets' health.

[0020] Based on this, embodiments of the present invention utilize various devices such as cameras, microphone arrays, and smart litter boxes to collect multi-dimensional data on pets, including gait, eating speed, vocalizations, frequency and weight of excrement.

[0021] Breaking through the limitations of single devices, it enables comprehensive monitoring of pet behavior and physiological characteristics, and can collect a wide range of information reflecting the pet's health status, providing a rich data foundation for subsequent analysis and early warning, and meeting the needs of pet owners for comprehensive monitoring of their pets' health.

[0022] Secondly, this invention integrates the collected multimodal data and conducts in-depth analysis using a large-scale model and veterinary knowledge graph to achieve early warning of potential health problems in pets. By leveraging multimodal data fusion technology and the powerful learning capabilities of the large-scale model, potential correlations between data points are uncovered, and professional veterinary knowledge is combined to improve the accuracy of warnings. This enhances the accuracy of identifying potential health problems in pets, reduces false alarms and missed alarms, achieves early warning, and meets pet owners' needs for precise monitoring and timely intervention of their pets' health.

[0023] In some examples, embodiments of the present invention may employ the following system, specifically including: 1. Data Acquisition Layer: Camera Module: High-definition cameras with wide-angle shooting capabilities are installed in areas where pets typically move around (such as the living room, bedroom, and near the pet's bed) to comprehensively capture the pet's activities. Using computer vision technology, skeletal motion recognition and high-speed tracking algorithms are employed to analyze the pet's gait and determine if there are any abnormalities such as limping. Simultaneously, the system monitors the pet's eating process, calculates the eating speed, and analyzes whether its eating habits have changed.

[0024] Among them, eating habits refer to the pattern characteristics of a pet's eating behavior, including eating speed (food consumption per unit time g / min), eating frequency (number of times to eat in 24 hours), eating duration (duration of a single meal), and eating method (biting / licking / pausing).

[0025] Microphone Array Module: This module uses a microphone array to collect pet vocalizations. Utilizing voice recognition technology and voiceprint analysis algorithms, it identifies the type of pet vocalizations and emotional changes, such as whether the pet is in pain or anxiety. Combined with a deep learning emotion recognition engine, it can accurately distinguish various emotional states of pets and promptly detect health problems reflected in abnormal vocalizations.

[0026] Pet waste management module: A built-in pet waste management module (including a weight sensor and detection module) is installed in the litter box to monitor the frequency and weight of waste in real time. By analyzing the waste data and combining it with a veterinary knowledge graph, it can determine if the pet has potential health problems such as urinary tract diseases or digestive issues. For example, sudden changes in waste weight or abnormal frequency may be signs of health problems.

[0027] 2. Data transmission layer: like Figure 1The system employs wireless transmission technologies (such as Wi-Fi and Bluetooth) to transmit collected video, audio, and sensor data to the home smart gateway in real time. The smart gateway performs initial data integration and preprocessing, and then uploads the data to a cloud server via the internet, ensuring stable data transmission and efficient processing. The cloud server processes the pet monitoring data from various modalities, including fusion, analysis, early warning feedback, and risk assessment.

[0028] The initial integration of smart gateways includes format standardization, timestamp synchronization, data compression, and invalid data filtering for video, audio, and sensor data. For example, video (30fps), audio (44.1kHz), and sensor data (10Hz) with different sampling rates are unified into a 1Hz timeline, and missing values ​​are supplemented through linear interpolation to ensure that the data structure uploaded to the cloud is standardized and minimized in size.

[0029] In some examples, during data transmission, encryption protocols such as Secure Sockets Layer (SSL) / Transport Layer Security (TLS) can be used to encrypt video, audio, and sensor data to prevent theft or tampering. In the data storage stage, pet health data stored on cloud servers and local devices is encrypted using encryption algorithms such as Advanced Encryption Standard (AES) to ensure data security.

[0030] 3. Data Processing and Analysis Layer: Multimodal data fusion: Employing spatiotemporal alignment algorithms, this method eliminates temporal discrepancies between different data types (video, audio, sensor data) to construct a unified dynamic biometric database. It deeply integrates behavioral data captured by cameras, sound data acquired by microphone arrays, and sensor data from the smart litter box, providing a comprehensive and accurate data foundation for subsequent analysis.

[0031] Large-scale model analysis: Utilizing a large-scale model, combined with massive amounts of pet health data and a veterinary knowledge graph, deep learning and analysis are performed on the fused multimodal data. Through comprehensive analysis of multi-dimensional data such as gait, eating speed, vocalizations, and excrement frequency and weight, early warnings of potential health problems in pets are achieved. For example, when video analysis shows a pet's limping gait, audio recognition identifies painful cries, and sensors detect a sudden decrease in excrement weight, the large-scale model, combined with the veterinary knowledge graph (e.g., abnormal gait + changes in excrement = urinary system disease), comprehensively calculates a health risk value, identifies a high-probability health problem, and triggers an alert. Through comprehensive analysis of this multi-dimensional data, potential health problems in pets can be identified more accurately.

[0032] 4. Early warning and feedback layer: When the large-scale model analysis detects potential health problems in a pet, the system immediately sends an alert to the pet owner via a mobile application (APP), detailing the possible health issues and suggesting appropriate measures. Simultaneously, pet owners can use the APP to view their pet's health data and behavior analysis reports in real time, allowing for convenient and timely monitoring of their pet's health. Furthermore, the system can integrate with veterinary hospital information systems to provide pet owners with online consultations and appointment scheduling services.

[0033] In some examples, user authentication and access control mechanisms can be established, allowing only authorized pet owners and relevant staff to access pet health data and system functions. Multi-factor authentication methods, such as passwords, fingerprint recognition, and facial recognition, are employed to enhance user authentication security. User actions are also logged and audited to promptly detect and address any abnormal activity.

[0034] As examples, besides home use, this system can also be applied to veterinary hospitals and pet boarding facilities. In veterinary hospitals, veterinarians can use the system to remotely monitor and diagnose pets' health, improving efficiency and accuracy. In pet boarding facilities, staff can use the system to monitor the health of boarded pets in real time, promptly identify and address health issues, and provide better boarding services for pet owners. Furthermore, the system's monitoring data can be used as a basis for pet health assessments, providing data support for pet insurance pricing and claims.

[0035] The present invention will be further described below with reference to the accompanying drawings: Reference Figure 2 The diagram illustrates a flowchart of a pet health monitoring method according to some embodiments of the present invention, which may specifically include the following steps: Step 201: Obtain pet monitoring data in multiple modalities.

[0036] As some examples, multimodal pet monitoring data may include pet video data collected by a camera module, pet audio data collected by a microphone array module, and pet excrement data collected by a pet waste processing module.

[0037] The camera module collects pet video data, which can be used to capture pets' daily activities, gait characteristics, and eating behavior; the microphone array module collects pet audio data, which can be used to analyze pets' vocalizations and emotional changes, and identify abnormal states such as pain and anxiety; and the pet excrement processing module collects pet excrement data, including excrement weight, biochemical indicators, and image data, which can be used to monitor pets' excretion frequency, weight changes, and the health status of excrement.

[0038] like Figure 1 Camera modules and microphone array modules can be deployed in the pet's daily activity area, and pet excrement processing modules can be placed in the smart litter box to collect data and obtain pet monitoring data for multiple modalities.

[0039] Step 202: The pet monitoring data of the multiple modalities are fused to obtain fused monitoring data, and a large model is used to analyze the fused monitoring data to obtain the health analysis results corresponding to each modality.

[0040] After acquiring pet monitoring data from multiple modalities, pet video data, pet audio data, and pet excrement data can be fused to obtain fused monitoring data. Then, a large-scale model is used to comprehensively analyze the fused monitoring data to obtain health analysis results corresponding to each modality. Specifically, the fused monitoring data can include pet video data, pet audio data, and pet excrement data; the health analysis results can include pet video analysis results, pet audio analysis results, and pet excrement analysis results.

[0041] For example, analyzing pet video data can help determine if there are any abnormalities in gait or motor dysfunction such as lameness; extracting voiceprint features and analyzing emotions from pet audio data can help identify whether the vocalizations contain abnormal emotional signals such as pain or anxiety; and conducting multi-dimensional analysis of pet excrement data, including whether there are sudden changes in excrement weight, whether biochemical indicators are abnormal, and whether image features show abnormalities (such as changes in color or texture), can comprehensively determine whether the pet has potential health risks such as urinary system diseases or digestive problems.

[0042] In some examples, when fusing pet monitoring data from multiple modalities, spatiotemporal alignment algorithms can be used to eliminate time discrepancies between different data sources, ensuring that video, audio, and sensor data are synchronized on the timeline. Then, dimensionality normalization is used to map data of different dimensions (such as gait characteristics, voiceprint parameters, excrement weight, etc.) to a uniform numerical range, avoiding analytical biases caused by differences in magnitude.

[0043] Then, an attention mechanism is used to dynamically allocate the weights of each modality of data. For example, when limping is detected in a pet, the analysis priority of gait data is increased, thereby generating fused monitoring data.

[0044] After the fused monitoring data is processed by a large model, independent health analysis results for modalities such as gait, vocalization, and excrement can be output. For example, abnormal gait may indicate joint disease, vocalization anxiety may be associated with environmental stress, and a sudden decrease in excrement weight may indicate digestive system problems.

[0045] Finally, the results of the multimodal analysis were cross-validated by combining veterinary knowledge graphs. For example, when gait abnormalities and decreased excrement weight occur simultaneously, the risk of urinary system diseases is given priority, forming a comprehensive health assessment covering behavioral, physiological and emotional dimensions.

[0046] As some examples, a large model can be a pre-trained artificial intelligence large model (AI large model). Through deep learning algorithms, combined with massive amounts of pet health data and veterinary knowledge graphs, large models have the ability to comprehensively analyze multimodal data.

[0047] In some examples, the training process for pre-trained large AI models is as follows: 1. Data collection and annotation: We collected a large amount of pet monitoring data and corresponding health record information from pets of different breeds, ages, and health conditions across multiple modalities to create a dataset. A professional veterinary team and data annotation personnel then meticulously labeled the pet behavior, vocalizations, excrement, and other information in the dataset, clearly identifying normal and abnormal conditions and noting potential health problems.

[0048] 2. Model Training: The large model is trained using a labeled dataset, and its parameters are tuned to enable it to accurately identify and analyze multimodal data of pets and predict potential health problems.

[0049] By employing techniques such as transfer learning and incremental learning, the performance of large models is continuously optimized, improving their adaptability to new data and situations. For example, based on a large amount of existing pet health data for training, incremental learning is used to enable large models to quickly learn and adapt to new features, particularly emerging pet disease types or behavioral patterns.

[0050] 3. Model Evaluation and Optimization: The trained large-scale model is evaluated using metrics such as cross-validation, accuracy, recall, and F1 score to verify its performance and accuracy. Based on the evaluation results, the large-scale model is optimized and adjusted, such as by adjusting the model structure, increasing the amount of training data, and improving the training algorithm, to continuously improve the prediction accuracy and reliability of the large-scale model.

[0051] In some embodiments of the present invention, the pet monitoring data of the multiple modalities are fused to obtain fused monitoring data, including: normalizing the dimensions of the pet monitoring data of the multiple modalities, and using an attention mechanism to fuse the dimension-normalized pet monitoring data to obtain fused monitoring data.

[0052] As some examples, dimensionality normalization refers to scaling pet monitoring data of different modalities according to a uniform standard, so that various types of data are mapped to a similar numerical range, thereby eliminating analytical bias caused by differences in data units.

[0053] When fusing pet monitoring data from multiple modalities, features can be extracted from the pet monitoring data from different modalities to obtain features such as gait features, voiceprint features, excrement weight features, and excrement images. These features are then normalized in dimension, and a self-attention mechanism is used to fuse the dimension-normalized pet monitoring data to obtain fused monitoring data.

[0054] As examples, pet monitoring data can be fused using a feature fusion model to obtain fused monitoring data. Feature fusion models can include: 1. Modal Embedding Layer: This layer extracts features from pet monitoring data across different modalities and normalizes the dimensions of these features. For example, video gait features (18-dimensional dynamic vectors) are mapped to 64 dimensions through a fully connected layer; audio voiceprint features (32-dimensional Mel-Frequency Cepstral Coefficients (MFCCs)) are upscaled to 64 dimensions through a Convolutional Neural Network (CNN) (1 convolutional layer + pooling); and pet excrement data (excrement weight features + excrement image features, totaling 12 dimensions) are compressed to 64 dimensions through an autoencoder, ensuring consistent dimensionality across all modal features.

[0055] 2. Cross-Attention Module: Employs a multimodal Transformer (self-attention mechanism) architecture, comprising two intramodal attention layers (processing temporal features within the same modality) and one cross-modal attention layer (calculating mutual information between features from different modalities). The number of attention heads is set to 8, and weights are calculated using a scaled dot-product mechanism. Stable training is achieved through residual connections and layer normalization.

[0056] In some embodiments of the present invention, before dimensional normalization of the pet monitoring data of the multiple modalities, the method further includes: spatiotemporal alignment of the pet monitoring data of the multiple modalities.

[0057] In practical applications, spatiotemporal alignment refers to eliminating the time-axis deviation of pet monitoring data of various modalities (such as video, audio, and sensor data) through spatiotemporal alignment algorithms.

[0058] In some examples, the Network Time Protocol (NTP) can be used to unify data collected by devices such as cameras, microphone arrays, and smart litter boxes to the same time base. For instance, the clocks of each device can be calibrated using the NTP to ensure that the timestamp error between video frames, audio sampling points, and sensor data is less than 10ms.

[0059] As some examples, spatiotemporal alignment techniques can be used to perform spatiotemporal alignment on pet monitoring data of multiple modalities. These techniques may include: 1. Time Synchronization: The camera module, microphone array module, and pet waste processing module all carry high-precision timestamps (calibrated based on the NTP protocol, with a time error of <10ms). Data at different sampling rates (video 30fps, audio 44.1kHz, sensor 10Hz) are unified to a 1Hz time axis using linear interpolation. Missing data is supplemented through forward padding (missing duration <3 seconds) or LSTM prediction (missing duration 3-10 seconds).

[0060] 2. Spatial Calibration: Since the camera coordinate system and the microphone array coordinate system are two different standards, spatial calibration can be used to link and fuse them during alignment. A spatial transformation matrix is ​​established by calibrating the camera coordinate system and the microphone array coordinate system (using Zhang Zhengyou's calibration method, with a reprojection error of <0.5 pixels). The pet's position information (2D coordinates detected by the camera) in the pet video data is converted into array-relative coordinates, achieving spatial correlation between sound and image, with a positioning error ≤10cm.

[0061] For example, the Zhang Zhengyou calibration method is first used to calculate the relative positional relationship between the camera coordinate system and the microphone array coordinate system through the calibration board image, generating a spatial transformation matrix; then the 2D coordinates of the pet detected by the camera are converted into the relative coordinates of the microphone array through this matrix (such as mapping the camera position to the physical coordinate system of the microphone array), achieving accurate association between the sound source and the image position, with a positioning error ≤10cm.

[0062] In some embodiments of the present invention, a large model is used to analyze the fused monitoring data to obtain health analysis results corresponding to each modality, including: Sub-step 11: Using a large model, perform a preliminary analysis of the pet monitoring data for each modality in the fused monitoring data, and then perform a further analysis by combining the preliminary analysis results with the pet monitoring data for other modalities.

[0063] As examples, preliminary analysis refers to using a large model to perform preliminary feature extraction and analysis on pet monitoring data of each modality in the fused monitoring data, while reanalysis refers to combining the results of the preliminary analysis with pet monitoring data of other modalities to make a comprehensive judgment.

[0064] For example, in the initial analysis, the large model identified gait abnormalities in pet video data, suggesting a possible motor dysfunction in the pets. Simultaneously, the initial analysis of pet audio data identified distress signals in the vocalizations. In the subsequent analysis phase, these preliminary findings can be comprehensively considered in conjunction with pet excrement data (such as whether there were sudden changes in excrement weight or abnormal biochemical indicators).

[0065] In some examples, for pet video data, convolutional neural networks can be used to extract key features such as gait and posture to determine whether there are abnormalities such as lameness or slow movement; for pet audio data, recurrent neural networks are used to analyze voiceprint features to identify emotional signals such as pain and anxiety in the cries; for pet excrement data, multilayer perceptrons are used to analyze weight, biochemical indicators, etc. to determine excretion frequency and health status.

[0066] Following initial analysis, the large-scale model cross-validates the results from different modal analyses. For example, when video analysis reveals abnormal gait in a pet, it combines this data with audio recordings of painful cries and a sudden decrease in fecal weight to comprehensively determine if the pet has joint disease or digestive system problems. Through initial analysis and cross-validation, the large-scale model can more accurately identify potential health problems in pets, avoiding misjudgments or omissions caused by single-modal data, and ultimately outputting comprehensive health analysis results covering behavioral, physiological, and emotional dimensions.

[0067] In some embodiments of the present invention, the multi-modal pet monitoring data includes pet video data collected by a camera module, and the preliminary analysis results of the pet video data include gait analysis results and feeding analysis results; preliminary analysis of the pet monitoring data of each modality in the fused monitoring data includes: Sub-step 111: Extract feature points from the pet video data in the fused monitoring data to obtain the current gait features, and compare the current gait features with the preset gait feature template to obtain the gait analysis results.

[0068] As examples, features can be extracted from pet video data based on keypoint detection models (such as PetNet-KeyPoint) to determine the coordinates of feature points and obtain the current gait features.

[0069] For example, a deep learning-based keypoint detection model is used to detect 18 key nodes in real time on the pet's torso (neck, back, waist) and limbs (shoulder, elbow, wrist, hip, knee, ankle), with a detection frame rate of no less than 30fps and a localization accuracy error of ≤2 pixels. A heatmap regression algorithm is used to generate the probability distribution of each keypoint, and non-maximum suppression (NMS) is used to select the optimal feature point coordinates.

[0070] Among them, the core calculation for selecting the optimal feature point coordinates for gait analysis is to calculate the displacement vector and joint angle change curves through the precise pixel coordinates of 18 key points (such as shoulder and knee joints) to generate a dynamic gait feature template, ensuring the accuracy of subsequent Euclidean distance comparison.

[0071] As examples, a pre-set gait feature template can be a set of typical gait features extracted from a large amount of healthy pet video data. The gait feature template can include the gait features of pets of different breeds, ages and sizes in a normal state.

[0072] In some examples, gait data from at least 500 healthy individuals of different pet breeds (e.g., felines and canines categorized by size as small / medium / large) can be collected to pre-establish a breed-specific benchmark database. Dynamic gait feature templates are generated by calculating the displacement vectors of key nodes and the joint angle change curves (e.g., the range of knee flexion and extension angles) during the walking cycle.

[0073] For example, the walking cycle of a pet is determined based on pet video data. Then, the normalized coordinates of key nodes within the cycle are extracted, the total displacement vector and Euclidean distance of each node in the cycle are calculated, the angles of key joints in each frame are calculated through key points and change curves are generated, and finally, the node displacement set and joint angle curves of 500 healthy individuals are integrated to generate gait feature templates according to breed and body type.

[0074] As examples, gait analysis results can be obtained by comparing the Euclidean distance between the current gait features and the gait feature template, as well as the distance value of the walking cycle and the symmetry analysis of left and right limb movements, such as limping in pets.

[0075] For example, the Euclidean distance between the current gait and the baseline template is calculated in real time based on real-time pet video data. When the distance value exceeds a threshold for three consecutive walking cycles (set based on the 3σ principle, with threshold differences between different breeds ≤5%), a lameness warning is triggered. At the same time, the symmetry analysis of left and right limb movements (such as stride length difference >15% and stride frequency difference >10%) is combined to distinguish between single-limb and multi-limb abnormalities, which are used as the gait analysis results.

[0076] Here, Euclidean distance is the distance between the real-time collected current gait data (such as normalized coordinates or displacement vectors of key nodes) and the baseline template (a breed-specific gait feature template generated based on data from 500 healthy pets).

[0077] In some examples, the coordinates of key points within a pet's walking cycle can be obtained in real time from pet video data, its displacement vector can be calculated, and then compared with the Euclidean distance of the displacement vector of the corresponding healthy pet in the gait feature template. When the distance exceeds a threshold (e.g., 5.0) for three consecutive cycles, an alert is triggered.

[0078] Specifically, exceeding the threshold for three consecutive walking cycles means that the Euclidean distance between the pet's gait characteristics and the gait characteristic template characteristics calculated in real time exceeds a preset threshold (e.g., threshold 5.0) in three complete steps. For example, the distance is 4.9 in the first cycle (not exceeded), 5.2 in the second cycle (exceeded), and 5.5 in the third cycle (exceeded). Exceeding the threshold three times consecutively triggers a limp warning.

[0079] Sub-step 112 involves segmenting the pet video data in the fused monitoring data to obtain the food area and the container background area. The feeding speed is determined based on the amount of pixel reduction in the food area, and the feeding analysis result is determined based on the fluctuation range of the feeding speed.

[0080] In practical applications, image processing techniques can be used to segment the pet feeding bowl area in pet video data at the pixel level, separating the food from the container background. The feeding speed can be determined based on the amount of pixel reduction in the food area and the feeding behavior. The feeding analysis results (such as eating too fast, eating too slow, or abnormal eating) can be determined based on the fluctuation range of the feeding speed.

[0081] As some examples, the results of food intake analysis can be determined in the following ways: 1. Food region segmentation: A mask-based region-based convolutional neural network (Mask R-CNN) is used to segment the feeding bowl region at the pixel level, and the food and container background are separated by the HSV color space thresholding method.

[0082] 2. Eating Action Recognition: An eating behavior classification model is trained based on a Temporal Convolutional Network (TCN) to identify three states: biting, licking, and pausing. A sliding window (window size of 1 second, step size of 0.5 seconds) is used to label the behavior sequence, and the recognition accuracy is ≥92%.

[0083] Among these factors, feeding actions (biting / licking / pausing) affect eating speed and duration. Biting is fast (10g / min) and short; licking is slow (3g / min) and long; pausing interrupts eating speed and prolongs eating time. By identifying and labeling feeding actions, eating speed can be accurately calculated (distinguishing between action consumption rate) and abnormal patterns can be detected (such as frequent pauses which may indicate pain).

[0084] 3. Calculation of eating speed: The amount of food pixels reduced per unit time is calculated by the frame difference method and converted into actual weight consumption (the mapping relationship between food pixels and food weight needs to be established in advance through calibration experiments, with an error ≤3g).

[0085] For example, first establish a linear mapping relationship between weight and pixels, then filter the feeding frame sequence to segment the food region to obtain the number of pixels in each frame, calculate the total pixel reduction and corresponding duration, and finally use the mapping relationship to convert the pixel reduction into total weight consumption, and further calculate the feeding speed.

[0086] Among them, the eating speed can be represented as the average weight consumption rate (g / min) during the continuous eating phase. When the speed fluctuation range is greater than 30% within 24 hours or the duration of a single feeding is greater than twice the baseline value (the judgment condition can be determined based on the statistical distribution of healthy pet data), it can be judged as an abnormal eating condition.

[0087] The feeding duration can be determined by using the pause state of the feeding action recognition as the start and end point, from the time interval from the detection of the start of feeding (the pet approaches the food bowl) to the end of feeding (the food pixels do not decrease for 10 consecutive seconds).

[0088] As examples, adaptive exposure control algorithms can be used to improve the visibility of feature points in low-light environments by combining image enhancement techniques for illumination variations (10-10000 lux); and Gaussian Mixture Models (GMMs) can be established using dynamic backgrounds to eliminate the interference of furniture movement and lighting changes on detection.

[0089] Among them, the Gaussian mixture model can update the background image in real time. By dividing the scene into multiple Gaussian distributions, it distinguishes between static backgrounds (such as furniture) and dynamic objects (pets). When furniture moves or the lighting changes, the background model is automatically adjusted, treating the interference as background updates and identifying the movement of the pet as a foreground object, thereby eliminating false detections.

[0090] In some embodiments of the present invention, the multimodal pet monitoring data includes pet audio data collected by a microphone array module, and the preliminary analysis results of the pet audio data include emotion analysis results; preliminary analysis of the pet monitoring data of each modality in the fused monitoring data includes: Sub-step 113: Extract voiceprint features from the pet audio data in the fused monitoring data to obtain the current voiceprint features, and determine the emotion analysis result based on the current voiceprint features.

[0091] As examples, pet audio data can be preprocessed, and then voiceprint analysis can be used to extract the current voiceprint features. Based on the current voiceprint features, the corresponding emotion label can be determined and used as the emotion analysis result.

[0092] For example, preprocessing is performed on pet audio data acquired by a microphone array module (8 channels, sampling rate 44.1kHz, 16-bit quantization), including pre-emphasis (filter coefficient 0.97), framing (frame length 20ms, frame shift 10ms), and applying a Hanning window. A 50-dimensional feature vector is extracted, consisting of 24-dimensional Mel-Frequency Cepstral Coefficients (MFCC), 13-dimensional first-order differences, and 13-dimensional second-order differences. This feature vector is then reduced to 32-dimensional voiceprint features using Principal Component Analysis (PCA).

[0093] Then, a three-layer bidirectional Long Short-Term Memory Network (LSTM) is used, with 128, 64, and 32 hidden units in each layer, respectively. A dropout layer (dropout rate 0.3) is added to prevent overfitting. The input is a 100-frame feature sequence (corresponding to 1 second of audio), and the output layer uses a normalized exponential function (softmax) to output five emotion labels: pain, hunger, excitement, fear, and normal.

[0094] When preprocessing pet audio data, the Minimum Controlled Recursive Averaging (MCRA) algorithm can be used to estimate the background noise power spectrum in real time, and a noise update threshold can be set (updates are paused when the signal-to-noise ratio (SNR) is <5dB) to ensure that noise is updated during the pet's silent period (2 consecutive seconds without effective voiceprints) in order to suppress environmental noise and extract the effective features of the pet's vocalizations more clearly.

[0095] When preprocessing pet audio data, an improved Normalized Least Mean Square (NLMS) algorithm can be used to denoise each channel signal of the microphone array module, with a filter order of 128 and a convergence factor of 0.01. Combined with beamforming technology (delay-summation beamforming, main lobe width < 30°), the target sound source (the pet's location is provided by camera positioning) is enhanced, and sidelobe noise is suppressed (SNR improvement ≥ 15dB after noise reduction).

[0096] As examples, a unique voiceprint model can be built for each pet, and the threshold for classifying emotion labels can be optimized through adaptive incremental learning (the model is updated every 30 minutes of new audio data), thereby improving the individual recognition accuracy to over 95%.

[0097] In some embodiments of the present invention, the multi-modal pet monitoring data includes pet excrement data collected by the pet excrement processing module, and the preliminary analysis results of the pet excrement data include excrement analysis results; preliminary analysis of the pet monitoring data of each modality in the fused monitoring data includes: Sub-step 114 involves performing multi-dimensional analysis of pet excrement data from the fused monitoring data to obtain excrement analysis results. The pet excrement data includes excrement weight data, excrement biochemical data, and excrement image data. The multi-dimensional analysis includes excrement weight analysis, excrement biochemical analysis, and excrement image analysis. The excrement analysis results include excrement weight analysis results, excrement biochemical analysis results, and excrement image analysis results.

[0098] In practical applications, a pet waste processing module can be configured in a smart litter box to collect pet waste data, such as waste weight, biochemical data, and image data. The results of the waste analysis can be determined by analyzing this data. This pet waste processing module can include a weight sensor, a biosensor array, and an embedded camera.

[0099] In some examples, fecal weight analysis results can indicate whether the fecal weight is within the normal range and the trend of fecal weight changes by analyzing whether the amount of feces is gradually increasing (which may indicate digestive problems) or decreasing (which may indicate insufficient food intake).

[0100] Excrement biochemical analysis results can reveal key biochemical indicators in excrement, such as pH value, protein content, and occult blood. A biosensor array (including pH electrodes, protein detection membranes, and occult blood test strips) is used to rapidly detect excrement biochemical data and compare it with baseline values ​​for healthy pet excrement (e.g., cat urine pH 6.0-6.5) in a veterinary knowledge graph. When a persistently abnormal pH value (greater than 7.0, indicating a urinary tract infection), excessive protein content (greater than 0.3 g / dL, indicating kidney problems), or a positive occult blood test is detected, a corresponding biochemical abnormality warning is generated.

[0101] The results of excrement image analysis can include features such as the shape, color, and texture of the excrement. Excrement image data is acquired using an embedded camera, and image processing algorithms (such as a morphological classification model based on convolutional neural networks) are used to identify the excrement's shape (e.g., normal strip-shaped, loose, or hard stool). Color space conversion (RGB to HSV) is used to analyze the excrement's color (e.g., yellow is normal, black may indicate gastrointestinal bleeding, and white may indicate parasites), and texture analysis is used to determine the excrement's texture (e.g., the presence of undigested food particles may indicate indigestion). When abnormal shapes (e.g., persistent loose stools), abnormal colors (e.g., red), or abnormal textures (e.g., containing large amounts of mucus) are detected, corresponding image anomaly warnings can be generated.

[0102] As examples, the weight sensor can be a high-precision strain gauge weight sensor (such as the HX711+FAS-D10) that monitors changes in excrement weight in real time, providing key physiological data for health assessment. It has a measurement range of 0-5 kg, a resolution of 0.1 g, and a sampling frequency of 10 Hz. A temperature compensation algorithm (error ≤ 0.5 g within the range of -10℃ to 40℃) eliminates the influence of ambient temperature, and a stable weight value is recorded every 30 seconds (a fluctuation of < 0.3 g over 5 consecutive sampling points is considered stable).

[0103] The biosensor array integrates three detection electrodes (pH: 0-14, accuracy ±0.1; occult blood: sensitivity ≥5ng / mL; urobilinogen: detection range 0-16μmol / L), uses an electrochemical workstation (scanning rate 50mV / s) for signal acquisition, has a detection response time of <30 seconds, and automatically starts the cleaning module (UV disinfection + distilled water rinsing) after each detection.

[0104] The embedded camera (2 megapixels, LED fill light) can capture images of pet excrement. It uses the YOLOv8-nano model (a lightweight model version in the YOLOv8 algorithm series) for target detection, identifying solid, liquid, and semi-solid types (accuracy ≥90%). Combined with a veterinary knowledge graph, it uses color histogram analysis (comparison of R / G / B channel mean values) to determine whether there are abnormalities such as hematuria (R channel percentage > 30%) and jaundice (Y channel brightness > 20% of the baseline value).

[0105] Sub-step 12: Based on the results of the re-analysis, determine the health analysis results corresponding to each modality.

[0106] Specifically, the reanalysis results can be combined with a veterinary knowledge graph to determine the health analysis results corresponding to each modality, thereby more comprehensively and accurately determining the health analysis results corresponding to each modality.

[0107] For example, if gait analysis results show that the pet is lame, but the feeding analysis results and excrement analysis results are not abnormal, and the veterinary knowledge graph shows that the pet has no health problems, then it may be a behavior that the pet is imitating or deliberately performing, rather than a health problem. In this case, the health analysis results corresponding to each modality can be marked as no abnormality.

[0108] If the gait analysis results are normal, but the feeding analysis and excrement analysis results are abnormal, such as eating too slowly and abnormal excrement biochemical indicators, it may indicate that the pet has a relatively serious health problem and requires further in-depth examination.

[0109] In some examples, health analysis results for each modality can be recorded, and a database of health analysis results can be established for subsequent data traceability and comparative analysis. Simultaneously, these health analysis results can be presented to pet owners in an intuitive and easy-to-understand manner, such as by generating visual charts via a mobile app. These charts include detailed explanations of the health analysis results for different modalities, annotations of abnormal conditions, and corresponding recommended measures.

[0110] For more complex health analysis results, an online veterinary consultation portal can be provided, allowing pet owners to easily obtain professional medical advice. Furthermore, tiered alerts are implemented based on the severity of the health analysis results. For minor abnormalities, notifications are sent via the app; for serious health risks, in addition to app alerts, pet owners can be notified urgently via SMS or phone to ensure their pets receive timely treatment.

[0111] Step 203: Based on the health analysis results, determine the health abnormality probability corresponding to each modality, and use the weight information corresponding to each modality to perform a weighted summation of the health abnormality probabilities to obtain the health risk value.

[0112] As examples, the probability of health anomalies can be determined by quantifying the results of health analysis, representing the likelihood that health anomalies were detected in the data acquired for that modality. For instance, corresponding anomaly scoring criteria can be set for the data collected in each modality, and different scores can be assigned to each modality based on the degree of anomaly in the preliminary and subsequent analysis results, thereby calculating the probability of health anomalies for each modality.

[0113] After obtaining the probability of health abnormalities for each modality, a weighted sum can be performed using the weight information corresponding to each modality. The weight information can be determined based on the degree of influence of different modalities on the pet's overall health. For example, some modalities may have a more critical impact on the pet's health and should therefore be given higher weights. By performing a weighted sum, a comprehensive health risk value can be obtained, which can more comprehensively reflect the pet's health status and potential risks.

[0114] In practical applications, appropriate weighting information can be determined based on a large amount of pet health data and veterinary expertise to ensure the accuracy and reliability of health risk values.

[0115] For example, the health anomaly probabilities corresponding to each modality (0.8 for pet video data + 0.7 for pet audio data + 0.9 for pet excrement data) are weighted and summed (the weights are optimized by grid search on the validation set). The health risk value is calculated as Σ (health anomaly probability × weight). An alert is triggered when the health risk value is greater than 0.75.

[0116] Among them, the health anomaly probability (e.g., 0.8 for pet video data) can represent the probability value (between 0 and 1) of the data acquired by this modality detecting health anomalies, and is used to quantify the degree of anomalies of each data source. Setting the health anomaly probability can avoid false alarms from a single modality (e.g., anomalies only in pet video data may be a misjudgment). The health risk value is calculated by weighted fusion (e.g., 0.4 for pet video data, 0.3 for pet audio data, and 0.3 for pet excrement data), ensuring that the early warning is more reliable.

[0117] Step 204: Based on the health risk value, issue a health warning for the pet.

[0118] In practical applications, when the system detects a potential health problem in a pet, it can send a push notification to the pet owner via the app. The notification includes the type of warning, a description of the possible health problem, and suggested measures.

[0119] The push notifications utilize multiple methods, such as pop-up alerts, sound alerts, and vibration alerts, to ensure pet owners receive timely warning information. Additionally, a historical warning record page is included in the app, allowing pet owners to easily review past warning messages and their handling.

[0120] In this embodiment of the invention, pet monitoring data of multiple modalities is acquired; the pet monitoring data of multiple modalities is fused to obtain fused monitoring data, and a large model is used to analyze the fused monitoring data to obtain health analysis results corresponding to each modality; based on the health analysis results, the probability of health abnormality corresponding to each modality is determined, and the weight information corresponding to each modality is used to perform weighted summation of the probability of health abnormality to obtain a health risk value; based on the health risk value, pet health early warning is performed, realizing pet health monitoring based on multiple modal data, meeting the needs of comprehensive pet health monitoring and early warning, and the fusion of pet monitoring data of multiple modalities and analysis using a large model meet the needs of accurate identification of potential health problems in pets.

[0121] Reference Figure 3The diagram illustrates a flowchart of another method for pet health monitoring provided by some embodiments of the present invention, which may specifically include the following steps: Step 301: Obtain pet monitoring data in multiple modalities.

[0122] Step 302: The pet monitoring data of the various modalities are normalized in dimensions, and an attention mechanism is used to fuse the pet monitoring data after the dimensions are normalized to obtain fused monitoring data.

[0123] Step 303: Using a large model, perform a preliminary analysis of the pet monitoring data for each modality in the fused monitoring data, and then perform a further analysis by combining the preliminary analysis results with the pet monitoring data for other modalities.

[0124] Step 304: Based on the results of the reanalysis, determine the health analysis results corresponding to each modality.

[0125] Step 305: Based on the health analysis results, determine the probability of health abnormality corresponding to each modality, and use the weight information corresponding to each modality to perform a weighted summation of the probabilities of health abnormality to obtain the health risk value.

[0126] Step 306: Based on the health risk value, issue a pet health warning.

[0127] In this embodiment of the invention, pet monitoring data of multiple modalities is acquired; the pet monitoring data of multiple modalities is fused to obtain fused monitoring data, and a large model is used to analyze the fused monitoring data to obtain health analysis results corresponding to each modality; based on the health analysis results, the probability of health abnormality corresponding to each modality is determined, and the weight information corresponding to each modality is used to perform weighted summation of the probability of health abnormality to obtain a health risk value; based on the health risk value, pet health early warning is performed, realizing pet health monitoring based on multiple modal data, meeting the needs of comprehensive pet health monitoring and early warning, and the fusion of pet monitoring data of multiple modalities and analysis using a large model meet the needs of accurate identification of potential health problems in pets.

[0128] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0129] Reference Figure 4The diagram shows a structural schematic of a pet health monitoring device according to some embodiments of the present invention, which may specifically include the following modules: The pet monitoring data acquisition module 401 is used to acquire pet monitoring data in multiple modalities; The health analysis module 402 is used to fuse the pet monitoring data of the multiple modalities to obtain fused monitoring data, and to analyze the fused monitoring data using a large model to obtain the health analysis results corresponding to each modality. The health risk value determination module 403 is used to determine the probability of health abnormality corresponding to each modality based on the health analysis results, and to use the weight information corresponding to each modality to perform a weighted summation of the probability of health abnormality to obtain the health risk value; The pet health early warning module 404 is used to issue pet health early warnings based on the health risk values.

[0130] In some embodiments of the present invention, the health analysis module 402 includes: The first analysis submodule is used to perform preliminary analysis on pet monitoring data of each modality in the fused monitoring data using a large model, and to perform further analysis by combining the preliminary analysis results with pet monitoring data of other modalities. The second analysis submodule is used to determine the health analysis result corresponding to each modality based on the results of the re-analysis.

[0131] In some embodiments of the present invention, the multimodal pet monitoring data includes pet video data collected by a camera module, and the preliminary analysis results of the pet video data include gait analysis results and feeding analysis results; the first analysis submodule includes: The gait analysis unit is used to extract feature points from the pet video data in the fused monitoring data to obtain the current gait features, and compare the current gait features with a preset gait feature template to obtain the gait analysis results; The feeding analysis unit is used to segment the pet video data in the fused monitoring data to obtain the food area and the container background area. Based on the amount of pixel reduction in the food area, the feeding speed is determined, and based on the fluctuation range of the feeding speed, the feeding analysis result is determined.

[0132] In some embodiments of the present invention, the multimodal pet monitoring data includes pet audio data collected by a microphone array module, and the preliminary analysis results of the pet audio data include emotion analysis results; the first analysis submodule includes: The emotion analysis unit is used to extract voiceprint features from the pet audio data in the fused monitoring data, obtain the current voiceprint features, and determine the emotion analysis result based on the current voiceprint features.

[0133] In some embodiments of the present invention, the multimodal pet monitoring data includes pet excrement data collected by the pet excrement processing module, and the preliminary analysis results of the pet excrement data include excrement analysis results; the first analysis submodule includes: The excrement analysis unit is used to perform multi-dimensional analysis of pet excrement data in the fused monitoring data to obtain excrement analysis results; wherein, the pet excrement data includes excrement weight data, excrement biochemical data, and excrement image data, the excrement multi-dimensional analysis includes excrement weight analysis, excrement biochemical analysis, and excrement image analysis, and the excrement analysis results include excrement weight analysis results, excrement biochemical analysis results, and excrement image analysis results.

[0134] In some embodiments of the present invention, the health analysis module 402 includes: The pet monitoring data fusion submodule is used to normalize the dimensions of the pet monitoring data of the various modalities, and to fuse the normalized pet monitoring data using an attention mechanism to obtain fused monitoring data.

[0135] In some embodiments of the present invention, the apparatus further includes: The pet monitoring data alignment module is used to perform spatiotemporal alignment of the pet monitoring data from the various modalities.

[0136] In this embodiment of the invention, pet monitoring data of multiple modalities is acquired; the pet monitoring data of multiple modalities is fused to obtain fused monitoring data, and a large model is used to analyze the fused monitoring data to obtain health analysis results corresponding to each modality; based on the health analysis results, the probability of health abnormality corresponding to each modality is determined, and the weight information corresponding to each modality is used to perform weighted summation of the probability of health abnormality to obtain a health risk value; based on the health risk value, pet health early warning is performed, realizing pet health monitoring based on multiple modal data, meeting the needs of comprehensive pet health monitoring and early warning, and the fusion of pet monitoring data of multiple modalities and analysis using a large model meet the needs of accurate identification of potential health problems in pets.

[0137] Some embodiments of the present invention also provide an electronic device, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the method described above.

[0138] Some embodiments of the present invention also provide a computer-readable storage medium on which a computer program is stored, and which, when executed by a processor, implements the method described above.

[0139] Some embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the method described above.

[0140] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0141] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0142] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0143] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0144] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0145] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0147] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0148] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the aforementioned element.

[0149] The above provides a detailed description of the method, apparatus, equipment, medium, and product for pet health monitoring. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for monitoring pet health, characterized in that, The method includes: Acquire pet monitoring data in multiple modalities; The pet monitoring data from the various modalities are fused to obtain fused monitoring data. A large model is then used to analyze the fused monitoring data to obtain health analysis results for each modality. Based on the health analysis results, the probability of health abnormality corresponding to each modality is determined, and the weight information corresponding to each modality is used to perform a weighted summation of the probabilities of health abnormality to obtain the health risk value; Based on the stated health risk values, a health warning for the pet is issued.

2. The method according to claim 1, characterized in that, A large model is used to analyze the fused monitoring data to obtain health analysis results for each modality, including: A large model was used to conduct a preliminary analysis of the pet monitoring data for each modality in the fused monitoring data, and then the preliminary analysis results were combined with the pet monitoring data for other modalities for further analysis. Based on the results of the reanalysis, the health analysis results corresponding to each modality are determined.

3. The method according to claim 2, characterized in that, The multimodal pet monitoring data includes pet video data collected by the camera module, and the preliminary analysis results of the pet video data include gait analysis results and feeding analysis results. A preliminary analysis of the pet monitoring data for each modality in the fused monitoring data was performed, including: Feature points are extracted from the pet video data in the fused monitoring data to obtain the current gait features, and the current gait features are compared with a preset gait feature template to obtain gait analysis results; The pet video data in the fused monitoring data is segmented into food area and container background area. The feeding speed is determined based on the amount of pixel reduction in the food area, and the feeding analysis result is determined based on the fluctuation range of the feeding speed.

4. The method according to claim 2, characterized in that, The multimodal pet monitoring data includes pet audio data collected by a microphone array module. Preliminary analysis results of the pet audio data include emotion analysis results. A preliminary analysis of the pet monitoring data for each modality in the fused monitoring data is performed, including: Voiceprint features are extracted from the pet audio data in the fused monitoring data to obtain the current voiceprint features, and the emotion analysis results are determined based on the current voiceprint features.

5. The method according to claim 2, characterized in that, The multi-modal pet monitoring data includes pet excrement data collected by the pet excrement processing module, and the preliminary analysis results of the pet excrement data include excrement analysis results; a preliminary analysis of the pet monitoring data of each modality in the fused monitoring data is performed, including: The pet excrement data in the fused monitoring data is subjected to multi-dimensional excrement analysis to obtain excrement analysis results; wherein, the pet excrement data includes excrement weight data, excrement biochemical data, and excrement image data, the excrement multi-dimensional analysis includes excrement weight analysis, excrement biochemical analysis, and excrement image analysis, and the excrement analysis results include excrement weight analysis results, excrement biochemical analysis results, and excrement image analysis results.

6. The method according to any one of claims 1-5, characterized in that, The pet monitoring data from the various modalities are fused to obtain fused monitoring data, including: The pet monitoring data from the various modalities were normalized in dimensions, and an attention mechanism was used to fuse the normalized pet monitoring data to obtain fused monitoring data.

7. The method according to claim 6, characterized in that, Before performing dimensionality normalization on the pet monitoring data from the various modalities, the following steps are also included: Spatiotemporal alignment was performed on the pet monitoring data of the various modalities.

8. A device for monitoring pet health, characterized in that, The device includes: The pet monitoring data acquisition module is used to acquire pet monitoring data in multiple modalities. The health analysis module is used to fuse the pet monitoring data of the multiple modalities to obtain fused monitoring data, and to analyze the fused monitoring data using a large model to obtain the health analysis results corresponding to each modality; The health risk value determination module is used to determine the probability of health abnormality corresponding to each modality based on the health analysis results, and to perform a weighted summation of the probabilities of health abnormality using the weight information corresponding to each modality to obtain the health risk value. The pet health early warning module is used to issue pet health early warnings based on the stated health risk values.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Method for determining risk score, storage medium, electronic device and product

    CN122221039A

  • A multi-pet cross-device pet identity attribution and health warning method and system

    CN122245826A

  • Multi-modal based pet living area behavior detection and analysis method and system

    CN122245828A