A multi-modal fusion cat health and emotion intelligent monitoring method and system
By integrating multimodal fusion algorithms that combine cat sounds, facial expressions, fur condition, and body temperature data, the problem of superficial data fusion and environmental interference in pet health and emotion monitoring has been solved. This enables accurate and real-time monitoring of cats' physiological health and emotional state, and is suitable for intelligent monitoring of family pets and auxiliary diagnosis of pet medical care.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENTRAL SOUTH UNIVERSITY OF FORESTRY AND TECHNOLOGY
- Filing Date
- 2026-05-21
- Publication Date
- 2026-06-19
AI Technical Summary
Existing technologies in the field of pet health and emotion monitoring suffer from shallow data fusion, non-standardized feature extraction, lack of anti-interference design, and technical positioning bias, resulting in unstable recognition results, high false alarm rates, and difficulty in working accurately in complex home environments.
A multimodal fusion-based intelligent monitoring method for cat health and emotions is adopted, integrating cat voice, facial expression, fur condition, and body temperature data. A two-level fusion algorithm with feature layer and decision layer is designed, and dimensionality reduction is performed by combining principal component analysis and t-distribution random neighborhood embedding. A random forest classifier with dynamic weight allocation is used for decision-making.
It enables accurate and real-time monitoring of cats' physiological health and emotional state, and is suitable for intelligent monitoring of family pets and auxiliary diagnosis of pet medical care. It improves the objectivity and accuracy of monitoring results and enhances the stability and adaptability of the algorithm in complex environments.
Smart Images

Figure CN122245788A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence multimodal monitoring algorithm technology, and in particular to a multimodal fusion intelligent monitoring method and system for cat health and emotion. Background Technology
[0002] With the rapid development of the pet economy and the popularization of smart home devices, pet health and emotion monitoring has become a research hotspot for the application of artificial intelligence in vertical fields.
[0003] At the research level, multimodal fusion research in pet health and emotion recognition is showing rapid growth, with a rich variety of related papers published. Domestically, some studies have constructed two-way interactive design models for emotional communication between pets and their owners, while others emphasize improving user experience through intelligent multimodal interaction. Internationally, some studies have achieved refined analysis of cat behavior and emotions based on specific fusion architectures; for dogs, some studies have achieved high emotion recognition results by fusing inertial sensors and physiological data, while others have integrated multiple modal data and combined them with specific algorithms to achieve high classification accuracy.
[0004] In terms of applications, domestic companies have already launched smart collars based on accelerometers to monitor pet activity levels and sleep quality; foreign companies have developed smart collars integrating body temperature and heart rate sensors, enabling early warning of basic diseases. Regarding emotion recognition, some patents attempt to analyze animal intentions using acoustic signals.
[0005] Despite some progress in the field of pet health and mood monitoring, existing technologies still have many shortcomings and cannot fully meet practical needs. These shortcomings are as follows: 1. Superficial Data Fusion: Current research largely focuses on human-pet interaction or single-dimensional emotion recognition, with insufficient exploration of deep integration between emotional states and physiological health monitoring. Multimodal data fusion often simply splices together raw data, lacking deep correlation at the feature and decision levels. For example, when faced with complex pet physical conditions, it is difficult to distinguish between stress-induced fever and pathological fever. This results in poor performance of monitoring systems in complex home environments, with high false alarm rates and unstable recognition results. 2. Non-standardized feature extraction: In the feature extraction process, algorithms for humans or general animals are directly applied without fully considering the specificities of pet species such as dogs and cats. Factors such as the fact that a dog's or cat's coat affects heat conduction, and the significant differences in facial structure among different breeds, can lead to data distortion and misjudgment. For example, when constructing a canine emotion dataset containing multiple breeds and a large number of images, it was found that the facial structure and body posture of different breeds differ significantly. This directly affects the generalization ability of emotion recognition, making it difficult to accurately apply the recognition results to different breeds of dogs. 3. Lack of Anti-interference Design: Existing technologies have not been specifically optimized for common interference sources in the home environment, such as changes in lighting and background noise. In real-world scenarios, such as poor lighting or interactions between multiple pets, the recognition rate drops significantly. For example, some research-developed pet emotion recognition systems perform significantly worse in complex environments, failing to meet practical needs and struggling to work stably and accurately in various home environments. 4. Technological Misalignment: Current technology focuses on entertaining "AI translation" functions while neglecting accurate health monitoring. It lacks specific judgment logic and warning thresholds for health issues such as pet skin diseases and respiratory abnormalities. This makes it difficult for the system to accurately distinguish between a pet's stress response and pathological state, and it cannot deeply analyze the complex motivations behind a pet's emotions, thus failing to provide pet owners with comprehensive and accurate health and emotional information. Summary of the Invention
[0006] The purpose of this invention is to provide a multimodal fusion intelligent monitoring method and system for cat health and emotions, which integrates four types of data: cat voice, facial expression, fur condition, and body temperature. Through a two-level fusion algorithm design of feature layer and decision layer, it achieves accurate and real-time monitoring of cat physiological health and emotional state, thereby solving at least one of the aforementioned problems in the prior art.
[0007] In a first aspect, the present invention provides a multimodal fusion method for intelligent monitoring of feline health and emotions, the method specifically comprising: Acoustic signals, facial images, full-body images, and body temperature values of cats were collected to construct multimodal raw data; Data cleaning and feature extraction were performed on the multimodal raw data to obtain acoustic features, facial features, hair features, and body temperature features; Acoustic features, facial features, hair features, and body temperature features are normalized and then dimensionality is reduced by combining principal component analysis and t-distribution random neighborhood embedding to extract cross-modal core fusion feature sets. Based on a cross-modal core fusion feature set, a random forest classifier with dynamic weight allocation is used to make decisions, identify the cat's health status and emotional state, and output monitoring results including health status, emotional category and their confidence probability.
[0008] Secondly, the present invention provides a multimodal fusion intelligent monitoring system for cat health and emotions, the system specifically comprising: The data acquisition module is used to collect acoustic signals, facial images, full-body images, and body temperature values of cats to construct multimodal raw data; The feature extraction module is used to clean and extract features from multimodal raw data to obtain acoustic features, facial features, hair features, and body temperature features. The feature fusion module is used to normalize acoustic features, facial features, hair features and body temperature features, and combine principal component analysis and t-distribution random neighborhood embedding to perform dimensionality reduction and extract cross-modal core fusion feature set. The classification decision module is used to make decisions based on cross-modal core fusion feature sets and a random forest classifier with dynamic weight allocation to identify the cat's health status and emotional state, and output monitoring results including health status, emotional category and their confidence probability.
[0009] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, and a computer program stored in the memory, wherein when the computer program is executed on the processor, it implements a multimodal fusion intelligent monitoring method for cat health and emotion as described in any of the above methods.
[0010] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a multimodal fusion intelligent monitoring method for cat health and emotion as described in any of the above methods.
[0011] Compared with the prior art, the present invention has at least one of the following technical effects: 1. This invention integrates four types of data: cat sounds, facial expressions, fur condition, and body temperature. Through a two-level fusion algorithm design of feature layer and decision layer, it achieves accurate and real-time monitoring of cat's physiological health and emotional state. It is suitable for application in scenarios such as intelligent monitoring of family pets and auxiliary diagnosis of pet medical care. 2. This invention uses a two-level multimodal fusion algorithm consisting of a feature layer and a decision layer to mine the correlation features between multimodal data, thus solving the problem of one-sided analysis by single-modal algorithms; 3. This invention designs a standardized quantitative feature extraction algorithm for cats' face, fur, voice, and body temperature, transforming qualitative state judgments into quantitative numerical analysis, making the monitoring results of cats' health and emotional state more objective and accurate. 4. This invention solves the problem of environmental interference in acoustic, visual, and body temperature data by using preprocessing algorithms such as wavelet denoising + spectral subtraction, image enhancement, and ambient temperature compensation, allowing the algorithm to maintain stable output in complex home environments and significantly improving environmental adaptability. 5. Each algorithm module of this invention adopts a lightweight model and efficient feature engineering design, without complex computing power requirements. It can be flexibly deployed in various software systems and terminal programs for intelligent pet monitoring, and is suitable for various application scenarios such as home pet monitoring and pet medical auxiliary diagnosis. It has strong practicality and applicability. 6. This invention performs targeted processing on multimodal raw data to obtain various features, laying the foundation for subsequent accurate fusion and classification recognition, and improving the accuracy of health and emotion monitoring; 7. This invention uses a combination of denoising algorithms and multiple feature extraction methods to process acoustic signals, effectively filtering out noise and comprehensively characterizing the acoustic state of cats, thereby improving the quality and usability of acoustic features. 8. This invention, by constructing a key point detection model and calculating facial features based on key points, can accurately obtain information about the emotional state of a cat's face, thus enhancing the accuracy and stability of facial feature extraction. 9. This invention utilizes a hair segmentation model and texture analysis algorithm to obtain hair features, effectively eliminating interference and accurately characterizing the health status of cat hair, thereby improving the accuracy of hair feature extraction. 10. This invention establishes a linear regression model to compensate for environmental temperature values, eliminates environmental interference, obtains accurate body temperature characteristics, and improves the reliability of body temperature data in health monitoring. 11. This invention normalizes multimodal features and combines principal component analysis with t-distributed random neighborhood embedding for dimensionality reduction, extracting cross-modal core fusion feature sets, reducing feature dimensionality while retaining key information, and improving classification efficiency and accuracy; 12. This invention performs principal component analysis on the normalized feature set, selects principal components for dimensionality reduction by calculating the covariance matrix, etc., reduces the amount of data while retaining the main information, reduces computational complexity and improves the efficiency of subsequent processing. 13. This invention inputs the PCA dimensionality reduction feature set into the t-distributed random neighborhood embedding algorithm, and extracts more discriminative cross-modal core fusion features by constructing a probability distribution and minimizing KL divergence optimization, thereby improving the classification and recognition effect. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart illustrating a multimodal fusion-based intelligent monitoring method for cat health and emotions provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the algorithm architecture of a multimodal fusion intelligent monitoring method for cat health and emotions provided in an embodiment of the present invention; Figure 3 This is a flowchart illustrating a multimodal fusion algorithm provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a multimodal fusion intelligent monitoring system for cat health and emotions provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0014] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0015] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0016] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0017] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0018] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0019] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0020] The following is an explanation of algorithm-related terminology: MFCC: Mel Frequency Cepstral Coefficient, is a commonly used feature extraction metric in acoustic signal processing that can effectively characterize the spectral features of sound; CNN: Convolutional Neural Network, is a deep learning algorithm that excels at image feature extraction and keypoint detection tasks; PCA: Principal Component Analysis is a feature dimensionality reduction algorithm that can remove redundant features while preserving core information. t-SNE (t-Distributed Stochastic Neighbor Embedding) is a nonlinear dimensionality reduction algorithm suitable for the fusion and dimensionality reduction of high-dimensional cross-modal features. STFT: Short-Time Fourier Transform, is a core time-frequency analysis method used to analyze the changes in signal frequency components over time; MFCC: Mel-Frequency Cepstral Coefficients (MFCC) is a feature extraction technique that simulates the human auditory system to analyze and identify sounds.
[0021] In this application embodiment, the entity executing the process includes a terminal device. This terminal device includes, but is not limited to, devices capable of executing the methods disclosed in this application, such as servers, computers, smartphones, and tablets. Figure 1 A flowchart illustrating a multimodal fusion-based intelligent monitoring method for feline health and emotions, as disclosed in an embodiment of the present invention, is shown below in detail: S101 collects acoustic signals, facial images, full-body images, and body temperature values of cats to construct multimodal raw data; S102, perform data cleaning and feature extraction on the multimodal raw data to obtain acoustic features, facial features, hair features and body temperature features; S103 normalizes acoustic features, facial features, hair features, and body temperature features, and combines principal component analysis and t-distribution random neighborhood embedding for dimensionality reduction to extract cross-modal core fusion feature sets. S104, based on a cross-modal core fusion feature set, makes decisions through a random forest classifier with dynamic weight allocation, identifies the cat's health and emotional state, and outputs monitoring results including health status, emotional category and their confidence probability.
[0022] In this embodiment, a high-sensitivity microphone is used to collect the cat's acoustic signals, capturing subtle differences in the cat's various vocalizations, including pitch, frequency, and volume. These acoustic signals contain important information about the cat's mood and health status. Simultaneously, a high-definition camera is used to capture facial and full-body images of the cat from different angles. Facial images clearly show the cat's facial expressions, such as the degree of eye opening and closing, ear orientation, and mouth shape, while full-body images help observe the cat's posture and fur condition. Additionally, a precise temperature measurement instrument, such as an infrared thermometer or an implantable temperature sensor (under animal welfare and safety regulations), is used to collect the cat's body temperature. The collected acoustic signals, facial images, full-body images, and body temperature values are integrated to construct a multimodal raw dataset, providing the foundational data for subsequent analysis and processing.
[0023] The collected multimodal raw data was cleaned to remove noise and outliers. For acoustic signals, filtering algorithms were used to remove environmental noise interference, ensuring that only the cat's own vocalizations were retained. For facial and full-body images, image enhancement techniques, such as adjusting brightness, contrast, and sharpness, were employed to improve image quality while removing blurred areas and irrelevant backgrounds.
[0024] After data cleaning, feature extraction is performed. For acoustic signals, signal processing techniques are used to extract acoustic features, such as Mel-frequency cepstral coefficients (MFCCs), which effectively describe the spectral characteristics of sound and reflect the unique patterns of a cat's meows. For facial images, computer vision algorithms are used to extract facial features, including the position, shape, and texture information of key areas such as the eyes, ears, and mouth. These features can visually reflect changes in the cat's facial expressions. For full-body images, the focus is on extracting fur features, observing the luster, fluffiness, and presence of abnormalities such as hair loss and dandruff. Body posture features are also analyzed, such as whether the cat is standing, sitting, lying down, or curled up, as different postures may indicate different health and emotional states. Body temperature values are directly used as a feature in subsequent analysis. Through this step, acoustic features, facial features, fur features, and body temperature features are obtained.
[0025] Because the data range and distribution of different features may vary significantly, normalization is required for acoustic features, facial features, hair features, and body temperature features to ensure more accurate and stable subsequent analysis. Normalization unifies the data of each feature to a specific range, such as the [0,1] interval, eliminating the influence of different dimensions. After normalization, dimensionality reduction is performed using Principal Component Analysis (PCA) and t-distributed random neighborhood embedding (t-SNE). PCA projects high-dimensional feature data into a low-dimensional space through linear transformation, preserving the main feature information and reducing the dimensionality and computational cost. t-distributed random neighborhood embedding is a non-linear dimensionality reduction method that maintains the local structure of data in the high-dimensional space within the low-dimensional space, further uncovering the intrinsic relationships between data. By combining these two methods, a cross-modal core fusion feature set is extracted. This feature set integrates key information from different modalities, providing a more comprehensive and accurate reflection of the cat's health and emotional state.
[0026] Based on the extracted cross-modal core fusion feature set, a dynamically weighted random forest classifier is used for decision-making. The random forest classifier consists of multiple decision trees, each independently judging the cat's health and emotional state. During classification, weights are dynamically assigned to each feature according to its importance in monitoring the cat's health and emotion, giving higher weights to features that have a greater impact on the classification result, thereby improving classification accuracy. By combining the judgments of multiple decision trees, the cat's health status is determined, such as healthy or ill (further subdivided into disease types, such as skin diseases, respiratory diseases, etc.), and its emotional state, such as happy, angry, fearful, calm, etc. Simultaneously, the confidence probability of each classification result is calculated, reflecting the reliability of the classification result. Finally, the monitoring results, including health status, emotional category, and their confidence probabilities, are output to the pet owner, providing comprehensive and accurate information about the cat's health and emotional state so that appropriate measures can be taken in a timely manner.
[0027] Reference Figure 2 The core algorithm of the multimodal fusion-based intelligent monitoring method for cat health and emotions disclosed in this invention is "single-modal feature extraction + two-level multimodal fusion." By collaboratively modeling four types of data—cat's voice, facial expressions, fur condition, and body temperature—it achieves intelligent monitoring and accurate judgment of health status and emotional changes, and enables real-time early warning. The algorithm is based on Python and implemented using libraries such as PyTorch, OpenCV, and Scikit-learn.
[0028] Specifically, firstly, various professional devices were used to collect data on the cat's voice, facial and full-body images, and body temperature. For voice data acquisition, a high-sensitivity microphone array was used to record WAV audio files at a 16kHz sampling rate, effectively capturing subtle sound signals emitted by the cat, including purring and meowing. For facial and full-body image data, JPG or PNG images were captured using a high-definition camera (resolution no less than 1920×1080 pixels), equipped with autofocus and supplemental lighting. Furthermore, body temperature data was collected using a non-contact infrared thermometer, which measures the cat's surface temperature in real time and outputs a floating-point value, while simultaneously recording the ambient temperature for subsequent compensation calculations. All data acquisition devices were calibrated to ensure the accuracy and reliability of the acquired data, providing a high-quality data foundation for subsequent algorithm processing. After completing multimodal data acquisition, the data needed to be systematically organized, classified, and labeled to construct a data sample set for algorithm training and testing. The labeled data is divided into training, validation, and test sets according to a certain ratio, which provides support for algorithm training, parameter tuning, and performance evaluation, respectively.
[0029] Then, the single-modal feature extraction layer includes acoustic signal processing, facial image processing, hair image processing, and body temperature data processing.
[0030] In acoustic signal processing, wavelet denoising and spectral subtraction algorithms are used to denoise the original audio data. After denoising, the Librosa library is used to extract various quantization features to characterize health and emotional information in the acoustic signal. Mel-frequency cepstral coefficients (MFCCs) are used to capture the frequency characteristics of the sound.
[0031] The core of facial image processing lies in keypoint detection, a step achieved through a lightweight CNN model (MobileNet-v2). To train the keypoint detection model, a training dataset was first constructed: 3000 cat facial images covering various breeds and scenes were collected and manually labeled by professionals according to the annotation specifications for 28 core keypoints (eyes, ears, mouth, and whiskers). Based on this, data augmentation techniques such as random rotation, scaling, and brightness adjustment were used to expand the dataset to 15000 images to improve the model's generalization ability. A keypoint detection model was built based on a lightweight CNN (such as MobileNet-v2), and end-to-end regression training was performed using the augmented dataset. Mean squared error (MSE) was chosen as the loss function, and the Adam optimizer was used for parameter optimization, enabling the model to learn the mapping relationship from the image to the coordinates of the 28 keypoints. Testing showed that the model's mean keypoint error (NME) on the test set was less than 3.5%, and the inference speed per image reached 15ms, balancing accuracy and efficiency. Based on the 28 key points detected, quantitative features such as eye width, the angle between the ear and the head, mouth opening, and whisker extension are further extracted using a coordinate calculation algorithm. The changes in the values of these features enable the initial identification of feline emotions such as fear, stress, and calmness.
[0032] Hair image processing employs the Mask R-CNN image segmentation algorithm to extract hair regions. Feature maps are extracted from the image using a convolutional neural network. After image segmentation, the OpenCV library is called to perform image enhancement processing on the hair regions, including brightness and contrast correction.
[0033] In body temperature data processing, the core formula of this algorithm is: Core body temperature = Original body temperature - Ambient temperature deviation coefficient × (Ambient temperature - 25℃). First, ambient temperature data is collected and used as input parameters; then, based on historical data of ambient temperature and body temperature measurements, a linear regression model is established to solve for the ambient temperature deviation coefficient; finally, the coefficient is substituted into the compensation formula.
[0034] In the feature layer fusion, a feature normalization algorithm is first constructed to unify the dimensions of the acoustic quantization features, facial / hair visual quantization features, and body temperature numerical features extracted from the four modules mentioned above. Then, the normalized feature set is processed by PCA principal component analysis + t-SNE dimensionality reduction algorithm to extract the core fusion features across modalities, remove redundant information, and mine the correlation features between multimodal data, such as the correlation features of "elevated body temperature + high-frequency whimpering sound + ear pressed against head" and "sudden hair fluffing + narrowing of eye fissure + slight increase in body temperature".
[0035] In the decision-level fusion, a decision-making algorithm based on multi-feature weighting is constructed to dynamically allocate weights to the cross-modal fusion features output by the feature layer (the weight of single-modal features affected by environmental interference is reduced, and the weight of effective features is increased). Then, the weighted fusion features are analyzed by a classifier algorithm to achieve accurate determination of the cat's health status (such as fever, skin disease, respiratory abnormalities) and emotional status (such as pleasure, fear, stress, pain), and finally output clear monitoring results.
[0036] In the output, to provide users with intuitive and easy-to-understand monitoring results, the algorithm of this invention outputs the determination of the cat's health and emotional state in the form of a combination of numerical indicators and textual descriptions. Specifically, the numerical indicators generated by the algorithm after prediction by the classifier reflect the probability distribution of the cat's current health status (such as healthy, fever, skin disease, respiratory abnormality) and emotional state (such as happy, calm, stressed, pain, fear). These probability values are normalized to ensure that the sum of the probabilities of each state is 1, and are presented to the user in the form of percentages.
[0037] Reference Figure 3In the feature layer fusion, the Min-Max Scaler algorithm from the Scikit-learn library is called to map the four types of feature vectors—acoustic, facial, hair, and body temperature—to the [0,1] interval through linear transformation, thus achieving dimensionality unification. Subsequently, a two-stage dimensionality reduction strategy is employed to extract cross-modal core features: First, principal component analysis (PCA) is used to calculate the covariance matrix and its eigenvalues. The top k principal components with a cumulative variance contribution rate ≥95% are selected for preliminary dimensionality reduction to remove linear redundancy. Then, the PCA-reduced features are input into the t-SNE algorithm for nonlinear manifold learning. A Gaussian probability distribution is constructed in the high-dimensional space, and a t-distribution probability distribution is constructed in the low-dimensional space. With a perplexity of 30, a learning rate of 200, and 1000 iterations, gradient descent is used to minimize the KL divergence between the high- and low-dimensional probability distributions. After iterative optimization, a 16-dimensional cross-modal fusion feature set is finally output. This two-stage design can effectively mine deep correlation features such as "elevated body temperature + high-frequency whimpering sound + ear pressed against head", and solve the computational efficiency problem of single t-SNE on high-dimensional data, providing high-quality fused feature input for subsequent decision-level classification.
[0038] In the decision-level fusion, constructing a random forest-based classifier algorithm is a crucial step in the multimodal fusion algorithm for intelligent monitoring of cat health and emotion. First, the low-dimensional cross-modal fusion feature set, processed by the feature layer fusion, is used as input data and divided into training and validation sets. Then, the random forest model is initialized using the RandomForestClassifier class from the Scikit-learn library, and relevant hyperparameters are set.
[0039] Combination Figures 1 to 3 As can be seen, this invention aims to solve the core problems of "severe data silos, poor environmental robustness, and low accuracy of emotion recognition" in existing technologies. By constructing customized feature extraction algorithms for four types of data—voice, facial expressions, fur, and body temperature—and designing a two-level multimodal fusion algorithm architecture of feature layer + decision layer, it achieves accurate and real-time monitoring of cats' health and emotional states. At the same time, through data preprocessing, feature normalization, and dynamic weight allocation algorithms, it solves the problem of environmental interference with single-modal data. Ultimately, it provides pet owners and veterinarians with comprehensive and intuitive assessment of cats' physiological and psychological states, enabling early detection of health problems and accurate identification of emotional states, thus improving the practicality and accuracy of the algorithm in pet monitoring scenarios.
[0040] In some embodiments, step S102 above, which involves data cleaning and feature extraction of the multimodal raw data to obtain acoustic features, facial features, hair features, and body temperature features, specifically includes: The acoustic signal is denoised and preprocessed, and the Mel frequency cepstral coefficients, fundamental frequency, spectral entropy and snoring duration are extracted as acoustic features. Multiple facial key points are extracted from facial images using a key point detection model, and facial features such as eye width, ear angle, mouth opening, and beard spread are calculated based on these key points. The hair region was extracted by segmenting the full-body image, and the hair fluffiness, straightness, proportion of hair loss area and amount of dandruff were obtained by texture analysis as hair features. The body temperature value is compensated for by ambient temperature to obtain the corrected body temperature characteristics.
[0041] In this embodiment, for acoustic signals, a combined denoising algorithm of wavelet denoising and spectral subtraction is first used to preprocess the collected raw cat sound signals to filter out environmental noise such as television and human voices. Then, a feature extraction algorithm is used to calculate the cat's specific acoustic quantitative features, including Mel-frequency cepstral coefficients (MFCC), fundamental frequency, spectral entropy, and duration of purring. By analyzing the changes in the values of the above quantitative features, a preliminary judgment can be made on the cat's pain, pleasure, stress, respiratory abnormalities, and other states.
[0042] For facial images, a keypoint detection model is built based on a lightweight CNN deep learning algorithm to process cat facial images and accurately detect 28 core keypoints of the eyes, ears, mouth, and whiskers. Then, a coordinate calculation algorithm is used to extract quantitative features such as eye width, the angle between the ears and the head, mouth opening, and whisker extension. The numerical values of the above features are used to characterize the cat's facial expression state, and to achieve preliminary identification of emotions such as fear, stress, and calmness.
[0043] For full-body images, the cat's full-body image is first processed using an image segmentation algorithm to accurately extract the cat's fur area and remove background interference. Then, a texture analysis algorithm is used to extract quantitative features such as fur fluffiness, straightness, proportion of shed areas, and amount of dandruff. At the same time, an image enhancement algorithm is introduced to correct image deviations caused by lighting and fur color. By analyzing the changes in the values of the above quantitative features, the cat's stress, skin disease, malnutrition, and other conditions can be determined.
[0044] For body temperature values, an environmental temperature compensation algorithm is constructed to correct the raw body temperature values of the cats collected, eliminating errors caused by room temperature and sunlight exposure, and outputting accurate core body temperature values of the cats' body surfaces; at the same time, a body temperature threshold judgment algorithm is set to define healthy body temperature ranges and abnormal thresholds (such as body temperature ≥39.5℃ is judged as fever), providing quantitative numerical basis for judging health status.
[0045] Furthermore, the denoising preprocessing of the acoustic signal and the extraction of Mel frequency cepstral coefficients, fundamental frequency, spectral entropy, and snoring duration as acoustic features specifically include: A combined denoising algorithm that combines wavelet denoising and spectral subtraction is used to preprocess the acoustic signal, filter out environmental noise, and obtain the preprocessed acoustic signal. The preprocessed acoustic signal is subjected to frame-by-frame windowing, and the time-domain signal is converted into a time-frequency graph using short-time Fourier transform. Based on the time-frequency diagram, the Mel frequency cepstral coefficients are extracted using the Mel filter bank, the fundamental frequency of the sound is extracted using the fundamental frequency detection algorithm, the spectral entropy is calculated using spectral analysis, and the duration of snoring is statistically determined using endpoint detection technology. Mel frequency cepstral coefficients, fundamental frequency, spectral entropy, and duration of purring are used as acoustic features to characterize the acoustic state of cats.
[0046] In this embodiment, a combined denoising algorithm integrating wavelet denoising and spectral subtraction is used to preprocess the acquired raw cat acoustic signal. Wavelet denoising effectively separates and removes noise components based on the differences in wavelet transform characteristics between the signal and noise at different scales. It decomposes the signal into different wavelet scales, thresholds the wavelet coefficients at each scale, retains the main characteristic coefficients of the signal, suppresses noise coefficients, and finally performs wavelet reconstruction to obtain the denoised signal. Spectral subtraction is based on the difference in the spectrum between noise and speech signals. It subtracts the estimated spectrum of noise from the spectrum of the noisy signal to obtain a relatively clean speech signal spectrum, which is then recovered through inverse Fourier transform. Combining these two methods fully leverages their respective advantages, more comprehensively filtering out environmental noise and obtaining a preprocessed acoustic signal.
[0047] The preprocessed acoustic signal undergoes frame segmentation and windowing. Frame segmentation divides the continuous acoustic signal into short signal segments, as cat sounds are relatively stable over short periods, facilitating local signal analysis. Windowing reduces signal truncation caused by framing and avoids spectral leakage. Commonly used window functions include Hamming and Hanning windows; a suitable window function is selected for signal weighting. After frame segmentation and windowing, the time-domain signal is converted to a time-frequency graph using a short-time Fourier transform. The short-time Fourier transform applies a Fourier transform to each frame of the signal, obtaining the energy distribution of each frame at different frequencies, thus converting time-domain information into time-frequency information, providing a more intuitive display of the frequency variation of the sound signal over time.
[0048] Based on the converted time-frequency diagram, various acoustic features are extracted: The time-frequency graph is processed using a Mel filter bank. The Mel filter bank simulates the nonlinear perception of sound frequencies by the human ear, converting the linear frequency scale to a Mel frequency scale. It consists of a set of triangular filters uniformly distributed in the Mel frequency domain. Each filter weights and sums the energy of its corresponding frequency range in the time-frequency graph, resulting in a set of filter outputs. These outputs are then subjected to logarithmic operations and discrete cosine transforms to extract the Mel frequency cepstral coefficients. These coefficients effectively describe the spectral characteristics of sound, reflecting the unique pattern of a cat's meow.
[0049] A fundamental frequency detection algorithm is used to extract the fundamental frequency of a sound from a time-frequency graph. The fundamental frequency is the lowest frequency component in a sound signal, determining its pitch. The fundamental frequency detection algorithm determines the fundamental frequency value of a sound by analyzing the frequency regions where energy is concentrated in the time-frequency graph and combining certain algorithmic rules. For a cat's meow, the fundamental frequency reflects the basic characteristics of its vocalization and is closely related to the cat's mood and health.
[0050] Spectral entropy is calculated through spectral analysis. Spectral entropy is an indicator of the complexity of a sound signal's spectrum, reflecting the non-uniformity of energy distribution within the sound signal. To calculate spectral entropy, the energy at each frequency point in the time-frequency graph is first normalized to obtain a probability distribution. Then, the spectral entropy is calculated using the formula for information entropy. A higher spectral entropy value indicates a more complex sound signal spectrum, potentially suggesting different emotional or health states of the cat.
[0051] Endpoint detection technology is used to determine the duration of purring. This technology accurately identifies the start and end points of a sound signal. By setting appropriate thresholds and rules, the start and end times of purring are determined, thus calculating its duration. Purring duration is an important indicator of a cat's mood and health; different purring durations may represent different emotional or health conditions.
[0052] The extracted Mel-frequency cepstral coefficients, fundamental frequency, spectral entropy, and duration of purring are integrated to form acoustic features characterizing the cat's acoustic state. These acoustic features reflect the characteristics of the cat's vocalizations from different perspectives and can provide rich information support for subsequent identification of the cat's health and emotional state.
[0053] Furthermore, the facial image is processed by a key point detection model to extract multiple facial key points, and based on these key points, the width of the eye fissure, the angle of the ear corner, the mouth opening, and the beard extension are calculated as facial features, specifically including: Collect historical cat facial images of various breeds and scenarios, and manually label the positions of the eyes, ears, mouth and whiskers in each historical cat facial image to construct a first training dataset containing multiple core key points. The first training dataset is augmented by random rotation, scaling and brightness adjustment to increase the number of samples and form an augmented first training dataset. The lightweight convolutional neural network is trained end-to-end using the enhanced first training dataset, enabling it to learn the mapping relationship from facial images to the coordinates of multiple key points, thus forming a key point detection model. The real-time captured facial images are input into the key point detection model for processing, and the coordinates of multiple key points are output. Based on the coordinates of multiple key points, the eye width, the angle between the ear and the head, the mouth opening, and the whisker extension are calculated using a coordinate calculation algorithm to obtain facial features that characterize the cat's facial emotional state.
[0054] In this embodiment, historical cat facial images from various breeds and scenarios are collected. Different breeds of cats have different facial structures, and multiple scenarios cover the facial expressions of cats in different environments and states, making the collected data more representative and comprehensive. For each collected historical cat facial image, professionals manually calibrate the positions of the eyes, ears, mouth, and whiskers. During the calibration process, based on the obvious features and positional relationships of the cat's facial organs, several core key points are determined, such as the inner and outer corners of the eyes, the tips and bases of the ears, the corners of the mouth and the midpoints of the upper and lower lips, and the roots and tips of the whiskers. These calibrated images are then compiled and summarized to construct a first training dataset containing multiple core key points.
[0055] Since the actual number of collected data samples may be limited, data augmentation was performed on the first training dataset to expand the sample size and improve the generalization ability of the keypoint detection model. Random rotation was used to rotate each image randomly within a certain angle range, simulating cat facial images from different angles and increasing the model's adaptability to facial images at different angles. Scaling operations were performed to enlarge or reduce the images by different proportions, enabling the model to handle facial images of different sizes. Simultaneously, brightness adjustments were made to change the brightness values of the images, simulating cat facial images under different lighting conditions and enhancing the model's robustness to different lighting environments. After these data augmentation processes, the enhanced first training dataset was formed, with a significantly expanded sample size, better meeting the needs of model training.
[0056] A lightweight convolutional neural network (CNN) was chosen as the basic architecture for the keypoint detection model. Lightweight CNNs offer advantages such as fewer parameters, lower computational cost, and faster execution speed, making them suitable for deployment and application on resource-constrained devices. An enhanced first training dataset was used to perform end-to-end regression training on the lightweight CNN. During training, the facial image was used as input, and the coordinates of multiple keypoints were used as the output targets. By continuously adjusting the network parameters, the network learned the mapping relationship from the facial image to the coordinates of multiple keypoints. After multiple iterations of training, training was stopped when the network achieved a satisfactory performance on both the training and validation sets, thus forming the keypoint detection model. This model can accurately detect the coordinates of multiple keypoints in cat facial images.
[0057] The real-time captured facial images are input into a pre-trained keypoint detection model for processing. The keypoint detection model analyzes and calculates the input images, outputting the coordinate information of multiple keypoints. This coordinate information accurately locates the positions of various keypoints on the cat's face, providing the basic data for subsequent facial feature calculations.
[0058] Based on the coordinates of multiple keypoints output by the keypoint detection model, coordinate calculation algorithms are used to calculate the eye width, the angle between the ear and head, the mouth opening, and the whisker extension. For eye width, the horizontal distance between the inner and outer corners of the eye is calculated using coordinates, resulting in the eye width value. This value reflects the degree of eye opening and is related to the cat's emotional state. To calculate the angle between the ear and head, keypoints of the ear and reference points of the head are first determined. Vector operations are then used to calculate the angle between the ear and head, i.e., the ear angle. Changes in the ear angle can reflect the cat's alertness or emotional state. Mouth opening is calculated using the coordinates of the corners of the mouth. The vertical distance between the upper and lower corners of the mouth is calculated to obtain the mouth opening value. The size of the mouth opening reflects different emotional states such as calmness, surprise, or anger. Whisker extension is calculated based on the coordinates of keypoints at the root and tip of the whiskers, calculating the length and angle changes of the whiskers. This comprehensive assessment of whisker extension can also provide a reference for judging the cat's emotional state. These calculations yield facial features that characterize a cat's emotional state, providing crucial facial information for subsequent identification of a cat's health and emotional state.
[0059] Furthermore, the process of segmenting the full-body image to extract hair regions and obtaining hair fluffiness, straightness, proportion of bald areas, and amount of dandruff as hair features through texture analysis specifically includes: Collect full-body images of historical cats of various breeds and in various scenes, and perform pixel-level annotation on the hair regions in each full-body image of a historical cat to construct a second training dataset for hair segmentation; Based on the second training dataset, the Mask R-CNN algorithm was used to train the model to segment the cat's fur region from the full-body image, remove background and environmental interference, and obtain the trained fur segmentation model. The real-time acquired full-body image is input into the hair segmentation model for processing, and the binary mask of the hair region is output to extract the hair region image. Image enhancement processing is performed on the hair area image. By adjusting the brightness and contrast, the image deviation caused by changes in light or differences in hair color is corrected to obtain the enhanced hair image. Based on the enhanced hair image, texture analysis algorithms were used to extract the fluffiness region, straightness region, proportion of hair loss area to the whole body image, and amount of dandruff, obtaining hair features to characterize the health status of the cat's hair.
[0060] In this embodiment, historical full-body images of cats from various breeds and in various scenarios were collected. Different breeds of cats exhibit significant differences in their fur characteristics; for example, long-haired and short-haired cats differ in fur length and texture. Multiple scenarios cover the fur appearance of cats in different environments and states, such as indoor / outdoor conditions and varying lighting levels, making the collected data more representative and comprehensive. For each collected historical full-body image of a cat, professionals performed pixel-level annotation of the fur regions. During the annotation process, the fur on the cat's body was precisely delineated, excluding background and other environmental areas to ensure accuracy. These annotated images were then compiled and summarized to construct a second training dataset for hair segmentation, providing fundamental data support for the subsequent training of the hair segmentation model.
[0061] Based on the constructed second training dataset, the Mask R-CNN algorithm was used for training. Mask R-CNN is a powerful object detection and instance segmentation algorithm capable of accurately learning to segment specific target regions from full-body images. During training, the full-body image was used as input, and pixel-level labeled hair regions were used as the target output. By continuously adjusting the algorithm's parameters, it learned the features and patterns of segmenting cat hair regions from full-body images, eliminating background and environmental interference. After multiple iterations of training, training was stopped when the algorithm achieved a high level of segmentation accuracy on both the training and validation sets, resulting in a trained hair segmentation model. This model can accurately segment cat hair regions from full-body images.
[0062] Real-time captured full-body images are input into a pre-trained hair segmentation model for processing. The hair segmentation model analyzes and calculates the input images, outputting a binary mask of the hair region. A binary mask is an image with only two values, 0 and 1, where 1 represents a hair region and 0 represents a non-hair region. Based on this binary mask, the hair region image can be accurately extracted, providing clean hair image data for subsequent hair feature analysis.
[0063] Because the actual captured images of hair regions can be affected by factors such as changes in lighting and differences in hair color, resulting in inconsistent image quality and impacting the accuracy of subsequent texture analysis, image enhancement processing is performed on the extracted hair region images. This involves adjusting the image brightness and contrast to correct image deviations caused by changes in lighting or differences in hair color. For example, for images with low lighting, brightness is appropriately increased; for images with low contrast, contrast is enhanced to make the details of the hair region more clearly visible. After image enhancement processing, an enhanced hair image is obtained, providing higher-quality image data for subsequent texture analysis.
[0064] Based on the enhanced hair images, relevant hair features were extracted using texture analysis algorithms. For hair fluffiness, the density and three-dimensionality of the hair in the image were analyzed to determine fluffiness regions and calculate their proportion in the full-body image. For hair straightness, the direction and curvature of the hair were observed to determine straightness regions and calculate their proportion. For areas of hair loss, missing hair was detected in the image to determine these areas and calculate their proportion in the full-body image. For dandruff quantity, small white particles resembling dandruff were identified and their number was counted. Through these analyses and calculations, hair features characterizing the health of a cat's fur were obtained, providing important hair information for subsequent identification of the cat's health and emotional state.
[0065] Furthermore, the step of compensating the body temperature value for ambient temperature to obtain the corrected body temperature characteristics specifically includes: Multiple sets of historical cat body surface original body temperature values and corresponding ambient temperature data were collected. A linear regression model was established, and the ambient temperature deviation coefficient was solved by the least squares method. The ambient temperature deviation coefficient is used to characterize the degree of interference of ambient temperature on body temperature measurement values. Substitute the ambient temperature deviation coefficient into the pre-built ambient temperature compensation formula to correct the original body temperature value, eliminate the measurement error introduced by the ambient temperature, and output the corrected core body temperature value as the body temperature feature. The ambient temperature compensation formula is as follows: , Indicates core body temperature, Indicates the original body temperature. Indicates ambient temperature. Indicates the reference temperature. This represents the ambient temperature deviation coefficient.
[0066] In this embodiment, multiple sets of historical raw body temperature values of cats were collected using a body temperature measurement device, while the corresponding ambient temperature data for each set of body temperature measurements was recorded using an ambient temperature measurement instrument. This data covers different seasons, different indoor and outdoor environments, and different time periods to ensure data diversity and representativeness. After collecting a sufficient amount of data, a linear regression model was established using this data. A linear regression model is a statistical model used to study the linear relationship between two variables; in this embodiment, it is used to study the relationship between ambient temperature and the raw body temperature values of cats. The linear regression model was solved using the least squares method, which aims to find a set of parameters that minimizes the sum of squared errors between the predicted values calculated based on these parameters and the actual observed values. After calculation, the ambient temperature deviation coefficient was finally obtained, which accurately characterizes the degree of interference of ambient temperature on the body temperature measurement values. For example, when the ambient temperature is high, the cat's body temperature may rise due to heat dissipation, and this ambient temperature deviation coefficient can quantify the degree of this rise.
[0067] After obtaining the environmental temperature deviation coefficient, it is substituted into a pre-constructed environmental temperature compensation formula. This formula, derived from extensive experiments and data analysis, comprehensively considers the influence of factors such as the original body temperature, ambient temperature, and reference temperature on the cat's core body temperature. The original body temperature is the collected original body temperature value of the cat's skin; the ambient temperature is the recorded ambient temperature corresponding to each set of temperature measurements; and the reference temperature is a reference temperature value determined through numerous experiments. Substituting the environmental temperature deviation coefficient into the formula corrects the original body temperature value. This calculation process eliminates the measurement error introduced by ambient temperature, as it interferes with the cat's skin temperature, causing the measured value to not accurately reflect the cat's core body temperature. Through this formula correction, a value closer to the cat's actual core body temperature can be obtained.
[0068] After correction using the environmental temperature compensation formula, the corrected core body temperature value is output. This value eliminates the interference of environmental temperature and more accurately reflects the cat's health status. In the multimodal fusion-based intelligent monitoring method for cat health and emotion, this corrected core body temperature value is used as a body temperature feature, and is combined with other acoustic features, facial features, and fur features for subsequent data processing and analysis, providing important body temperature information for accurately identifying the cat's health and emotional state.
[0069] In some embodiments, step S103 above, which involves normalizing the acoustic features, facial features, hair features, and body temperature features, and then performing dimensionality reduction using principal component analysis and t-distributed random neighborhood embedding to extract a cross-modal core fusion feature set, specifically includes: Acoustic features, facial features, hair features, and body temperature features are combined into an initial high-dimensional multimodal feature vector. The Min-Max Scaler algorithm is then used to perform a linear transformation on each dimension of the initial high-dimensional multimodal feature vector, mapping all feature values to the [0,1] interval to obtain a normalized feature set. Principal component analysis is performed on the normalized feature set. By calculating the covariance matrix and its eigenvalues, the top k principal components with a cumulative variance contribution rate greater than or equal to a preset contribution rate threshold are selected for preliminary dimensionality reduction, resulting in the feature set after PCA dimensionality reduction. The feature set after PCA dimensionality reduction is input into the t-distribution random neighborhood embedding algorithm. By constructing a Gaussian probability distribution in the high-dimensional space and a t-distribution probability distribution in the low-dimensional space, the KL divergence between the probability distributions in the high-dimensional and low-dimensional spaces is minimized by the gradient descent method, and the cross-modal core fusion feature set after iterative optimization is output.
[0070] In this embodiment, the acoustic features, facial features, hair features, and body temperature features obtained through feature extraction are combined to form an initial high-dimensional multimodal feature vector. This vector contains various types of data features. Since the data range and scale of different features may vary significantly—for example, the numerical range of body temperature features may be tens of degrees Celsius, while some parameters in acoustic features may have smaller values—this difference can affect the accuracy and stability of subsequent analysis. Therefore, the Min-Max Scaler algorithm is used to linearly transform each dimension of the initial high-dimensional multimodal feature vector. The Min-Max Scaler algorithm is based on the principle of linear transformation. By scaling the value range of the original data according to a certain ratio, it maps the data to a specified interval, typically [0,1], but can also map to other intervals, such as [-1,1], depending on actual needs. Its core idea is to use the minimum and maximum values in the dataset as scaling benchmarks, recalculating the position of each data point relative to the minimum and maximum values to obtain normalized data. The Min-Max Scaler algorithm maps all feature values to the range [0,1] according to a certain rule, based on the maximum and minimum values of each feature. After this processing, a normalized feature set is obtained, where all features are within a relatively uniform numerical range, laying a good foundation for subsequent analysis and processing.
[0071] After obtaining the normalized feature set, principal component analysis (PCA) is performed. PCA is a commonly used data dimensionality reduction method. Its core idea is to find the direction of greatest variation in the data, i.e., the principal components. In practice, the covariance matrix of the normalized feature set is first calculated, reflecting the correlation between different features. Then, the eigenvalues of this covariance matrix are calculated; the magnitude of the eigenvalue represents the degree of data variation in the corresponding direction. Next, based on a pre-set contribution rate threshold, the top k principal components with a cumulative variance contribution rate greater than or equal to that threshold are selected. The cumulative variance contribution rate reflects the proportion of information in the original data that these principal components can explain. By selecting an appropriate number of principal components, preliminary dimensionality reduction can be achieved while retaining most of the important information in the data, resulting in the PCA-reduced feature set. For example, if the original feature set has 100 features, after PCA, it may only be necessary to select the first 20 principal components to retain more than 90% of the information, thus reducing the data dimensionality from 100 to 20.
[0072] The feature set after PCA dimensionality reduction is input into the t-distribution random neighborhood embedding algorithm. The goal of this algorithm is to establish a mapping between high-dimensional and low-dimensional spaces, allowing data to maintain a similar structure in the low-dimensional space as in the high-dimensional space. Specifically, it first constructs a Gaussian probability distribution in the high-dimensional space to describe the similarity between data points, and then constructs a t-distribution probability distribution in the low-dimensional space. Next, it continuously adjusts the positions of data points in the low-dimensional space using gradient descent to minimize the KL divergence between the probability distributions of the high and low-dimensional spaces. KL divergence is a metric that measures the difference between two probability distributions; minimizing it makes the data distribution in the low-dimensional space as close as possible to the data distribution in the high-dimensional space. After multiple iterations of optimization, the final output is an iteratively optimized cross-modal core fusion feature set. This feature set not only contains important information from multiple modalities but also has low dimensionality, making it more conducive to subsequent decision-making by a random forest classifier based on dynamic weight allocation, thereby accurately identifying the cat's health and emotional state.
[0073] Furthermore, the step of performing principal component analysis (PCA) on the normalized feature set involves calculating the covariance matrix and its eigenvalues, selecting the top k principal components whose cumulative variance contribution rate is greater than or equal to a preset contribution rate threshold for preliminary dimensionality reduction, and obtaining the PCA-reduced feature set. Specifically, this includes: The normalized feature set is used to construct a feature matrix. The feature matrix is then centered by subtracting the mean of each feature dimension to make the mean of each feature dimension zero, thus obtaining the centered feature matrix. Based on the centered feature matrix, the covariance matrix between features is calculated, and the covariance matrix is used to characterize the degree of linear correlation between features of different dimensions; The covariance matrix is decomposed into eigenvalues to obtain all eigenvalues and their corresponding eigenvectors. The eigenvalues are then sorted in descending order, and the order of the eigenvectors is adjusted accordingly. Calculate the variance contribution rate of each feature value, and calculate the cumulative variance contribution rate from front to back until the cumulative value reaches the preset contribution rate threshold. Determine the number of principal components k to be retained. The variance contribution rate is the proportion of a single feature value to the sum of all feature values. The eigenvectors corresponding to the top k largest eigenvalues are selected to form a projection matrix. The centered feature matrix is multiplied by the projection matrix to complete the linear transformation from the original high-dimensional space to the low-dimensional principal component space, and the dimensionality-reduced PCA feature set is output.
[0074] In this embodiment, the feature set obtained after normalization is constructed into a feature matrix. This feature matrix contains multi-dimensional information such as acoustic features, facial features, hair features, and body temperature features. For the accuracy and convenience of subsequent analysis, this feature matrix needs to be centered. Specifically, the mean of each dimension of the feature matrix is calculated, and then the mean of that dimension is subtracted from each value in each dimension. After this operation, the mean of each dimension becomes zero, thus obtaining the centered feature matrix.
[0075] The covariance matrix between features is calculated based on the centered feature matrix. The covariance matrix is a crucial concept, clearly representing the degree of linear correlation between features of different dimensions. For example, acoustic features and facial features may have a certain correlation, which can be quantified by the covariance matrix. A larger covariance value between two features indicates a stronger linear relationship; conversely, a smaller covariance value indicates a weaker linear relationship. Calculating the covariance matrix provides a comprehensive understanding of the relationships between different features, laying the foundation for subsequent principal component extraction.
[0076] The calculated covariance matrix is then subjected to eigenvalue decomposition. Eigenvalue decomposition is a mathematical method that solves for all eigenvalues of the covariance matrix and their corresponding eigenvectors. Eigenvalues reflect the degree of variation in the data along the direction of the corresponding eigenvector; the larger the eigenvalue, the more drastic the variation in that direction, and the more information it contains. After obtaining all eigenvalues and eigenvectors, the eigenvalues are sorted in descending order, and the order of the eigenvectors is adjusted accordingly to ensure that the eigenvalues and their corresponding eigenvectors maintain a consistent relationship. This step is helpful for subsequently selecting the most important principal components.
[0077] Calculate the variance contribution rate for each eigenvalue. The variance contribution rate is a key indicator, representing the proportion of a single eigenvalue to the total sum of all eigenvalues. By calculating the variance contribution rate, we can determine the proportion of information represented by each eigenvalue within the overall information. Then, calculate the cumulative variance contribution rate sequentially from front to back, that is, accumulate the variance contribution rates of the eigenvalues in descending order until the cumulative value reaches a preset contribution rate threshold. This preset contribution rate threshold is determined based on actual needs and data characteristics, and is generally chosen to retain most of the important information. When the cumulative variance contribution rate reaches this threshold, the number of principal components, k, that needs to be retained can be determined. For example, if we want to retain more than 90% of the information, the number of eigenvalues corresponding to a cumulative variance contribution rate of 90% is k.
[0078] The eigenvectors corresponding to the k largest eigenvalues are selected to form a projection matrix. This projection matrix transforms the original high-dimensional data into a low-dimensional principal component space. Multiplying the centered feature matrix by the projection matrix completes the linear transformation from the original high-dimensional space to the low-dimensional principal component space. After this transformation, the dimensionality of the data is reduced while retaining most of the important information, ultimately outputting a dimensionality-reduced PCA feature set. This PCA feature set will provide a more concise and effective data foundation for subsequent identification of cat health and emotional states.
[0079] Furthermore, the step of inputting the PCA-reduced feature set into the t-distributed random neighborhood embedding algorithm involves constructing a Gaussian probability distribution in the high-dimensional space and a t-distributed probability distribution in the low-dimensional space, minimizing the KL divergence between the probability distributions in the high and low dimensions using gradient descent, and outputting an iteratively optimized cross-modal core fusion feature set. Specifically, this includes: The feature set after PCA dimensionality reduction is used as the high-dimensional input space of the t-distributed random neighborhood embedding algorithm. The Euclidean distance between any two sample points in the high-dimensional input space is calculated, and the Euclidean distance is converted into a conditional probability to characterize the similarity between sample points. Based on conditional probability, a joint probability distribution of the high-dimensional input space is constructed through symmetry processing, so that sample points with high similarity in the high-dimensional input space have higher probability values and sample points with low similarity have lower probability values. In the low-dimensional target space, low-dimensional mapping points that correspond one-to-one with high-dimensional sample points are randomly initialized, and the joint probability distribution between any two mapping points in the low-dimensional target space is calculated based on the t-distribution. The t-distribution is a Cauchy distribution with 1 degree of freedom and is used to alleviate the crowding problem in high-dimensional data. The KL divergence between the joint probability distribution of the high-dimensional input space and the joint probability distribution of the low-dimensional target space is used as the loss function. Iterative optimization is performed by gradient descent. In each iteration, the gradient of the loss function with respect to the low-dimensional mapping point is calculated, and the coordinates of the low-dimensional mapping point are updated to gradually minimize the KL divergence. After reaching the preset maximum number of iterations or the loss function converges, the iterative optimization stops, and the final low-dimensional mapping point coordinates are output as the cross-modal core fusion feature set.
[0080] In this embodiment, the feature set obtained after PCA dimensionality reduction is used as the high-dimensional input space of the t-distributed random neighborhood embedding algorithm. Within this high-dimensional input space, the Euclidean distance between any two sample points is calculated. Euclidean distance is a common method for measuring the straight-line distance between two points, and the distance between each pair of sample points can be obtained through calculation. After obtaining these distance values, they are converted into conditional probabilities. These conditional probabilities characterize the similarity between sample points; the closer the sample points are, the higher their similarity, and the larger the converted conditional probability value; conversely, the farther the sample points are, the lower their similarity, and the smaller the conditional probability value.
[0081] Based on the conditional probabilities calculated earlier, a symmetry-based process is performed to construct the joint probability distribution of the high-dimensional input space. This symmetry-based process makes the measurement of similarity more reasonable and comprehensive. After this process, in the high-dimensional input space, sample points with high similarity will have higher probability values, while sample points with low similarity will have lower probability values. This allows for a more accurate description of the relationships between sample points in the high-dimensional space.
[0082] In the low-dimensional target space, low-dimensional mapping points are randomly initialized, each corresponding one-to-one with a high-dimensional sample point. These low-dimensional mapping points are the points corresponding to the sample points in the high-dimensional space after being mapped to the low-dimensional space using a specific dimensionality reduction algorithm. Then, the joint probability distribution between any two mapping points in the low-dimensional target space is calculated based on the t-distribution. The t-distribution used here is a Cauchy distribution with 1 degree of freedom. An important purpose of using this distribution is to alleviate the crowding problem in high-dimensional data. In high-dimensional data, data points are often relatively scattered, but when mapped to a low-dimensional space, overcrowding can easily occur. The t-distribution can effectively improve this problem, making the data distribution in the low-dimensional space more reasonable.
[0083] The t-distribution is a symmetric, bell-shaped probability distribution. Its shape is similar to the normal distribution, but its tails are thicker. The shape of the t-distribution is determined by the degree of freedom (df). The greater the degree of freedom, the closer the t-distribution is to a normal distribution; the smaller the degree of freedom, the thicker the tails. The t-distribution is often used to infer the population mean when the sample size is small and the population standard deviation is unknown.
[0084] The Cauchy distribution is also a symmetric probability distribution. When the t-distribution has 1 degree of freedom, its probability density function is exactly the same as that of the Cauchy distribution. In other words, the t-distribution with 1 degree of freedom is the Cauchy distribution.
[0085] The KL divergence between the joint probability distribution of the high-dimensional input space and the joint probability distribution of the low-dimensional target space is used as the loss function. KL divergence is a metric used to measure the difference between two probability distributions, indicating the degree of difference between the probability distributions in the high-dimensional and low-dimensional spaces. Based on this loss function, gradient descent is used for iterative optimization. In each iteration, the gradient of the loss function with respect to the low-dimensional mapping points is calculated. The gradient can be understood as the direction and extent of change of the loss function; by calculating the gradient, we can determine how to adjust the coordinates of the low-dimensional mapping points. Then, the coordinates of the low-dimensional mapping points are updated based on the calculated gradient. The purpose of this is to progressively minimize the KL divergence, making the probability distribution of the low-dimensional target space as close as possible to the probability distribution of the high-dimensional input space.
[0086] When the preset maximum number of iterations is reached or the loss function converges, it indicates that the iterative optimization process has achieved good results, and the iteration is stopped. The final low-dimensional mapping point coordinates are output, and these coordinates constitute the cross-modal core fusion feature set. This feature set can better integrate data features from different modalities, providing strong data support for accurately identifying the cat's health and emotional state.
[0087] In some embodiments, in step S104 above, the step of making decisions based on a cross-modal core fusion feature set, using a random forest classifier with dynamic weight allocation, identifying the cat's health and emotional states, and outputting monitoring results including health status, emotional category, and their confidence probabilities, specifically includes: The historical cross-modal core fusion feature set is divided into a training set and a validation set according to a preset ratio; A random forest classifier is constructed based on the RandomForestClassifier class in the Scikit-learn library, and hyperparameters such as the number of decision trees, maximum depth, and minimum number of samples required for splitting are initialized to obtain the initialized random forest model. A dynamic weight allocation mechanism is introduced. Based on interference factors such as the signal-to-noise ratio, light intensity, and temperature fluctuation of each single-modal data in the current environment, the confidence weight of each modal feature in the decision layer is calculated in real time. During the node splitting process of the random forest model, the feature dimension corresponding to the modal feature that is more affected by environmental interference is assigned a lower weight, and the effective feature is assigned a higher weight, thus forming a weighted random forest classifier. The weighted random forest classifier was trained using the training set. Through multiple iterations, the weighted random forest classifier learned the mapping relationship between the cross-modal core fusion feature set and the cat's health and emotional state. The weighted random forest classifier was then fine-tuned using the validation set to determine the optimal model parameters. The cross-modal core fusion feature set obtained in real time is input into the trained weighted random forest classifier to output the initial probability distribution of the cat's health status and emotion category at the current moment. The initial probability distribution is normalized to generate the confidence probability of each state. The health state and emotion category with the highest confidence probability are then used as the final health state and emotion category, and the monitoring results are generated and output.
[0088] In this embodiment, the historically accumulated cross-modal core fusion feature set is divided according to a preset ratio. This preset ratio is determined based on actual needs and the amount of data; generally, it ensures that the training set has sufficient data for the model to learn the relationship between features and states, while the validation set also has an appropriate amount of data for evaluating model performance. This division results in a training set for training the model and a validation set for evaluating and tuning the model.
[0089] This paper uses the `RandomForestClassifier` class from the Scikit-learn library to build a random forest classifier. During the construction process, several hyperparameters need to be initialized, such as the number of decision trees (which determines the number of decision trees in the random forest; too few trees may lead to insufficient model learning, while too many may increase computational cost and increase the risk of overfitting); the maximum depth limits the growth depth of each decision tree to prevent it from becoming overly complex; and the minimum number of samples required for splitting specifies the minimum number of samples a node must contain before splitting, which helps prevent the model from overfitting to noisy or outlier data. By setting these hyperparameters, the initialized random forest model is obtained.
[0090] A dynamic weight allocation mechanism is introduced. In actual monitoring environments, various interference factors exist, such as the signal-to-noise ratio (SNR) of each single modality data. A low SNR indicates that the data contains a lot of noise and has low reliability; light intensity affects the quality of image data; temperature fluctuations may affect body temperature data, etc. Based on these interference factors, the confidence weight of each modality feature in the decision layer is calculated in real time. During the node splitting process of the random forest model, feature dimensions corresponding to modalities that are more susceptible to environmental interference are assigned lower weights because these features may not be reliable in the current environment; while effective features, i.e., features that are less affected by interference and can accurately reflect the cat's state, are assigned higher weights. In this way, a weighted random forest classifier is formed, enabling the model to make more reasonable use of different features when making decisions.
[0091] The weighted random forest classifier is trained using the previously partitioned training set. Training is achieved through multiple iterations. In each iteration, the model learns the mapping relationship between the cross-modal core fusion feature set and the cat's health and emotional states. With each iteration, the model's understanding of this relationship becomes increasingly accurate. Simultaneously, the weighted random forest classifier is fine-tuned using the validation set. The model's performance is evaluated on the validation set, and the model parameters, such as the number of decision trees and maximum depth mentioned earlier, are adjusted based on the evaluation results to determine the optimal model parameters, ensuring the model achieves best performance on the validation set.
[0092] Once the model is trained, the real-time cross-modal core fusion feature set is input into a trained weighted random forest classifier. Based on the learned mapping relationships, the classifier outputs an initial probability distribution of the cat's current health status and emotional state. This initial probability distribution reflects the model's preliminary probability judgments of the various possible health states and emotional states the cat may be in.
[0093] The initial probability distribution is normalized. Normalization adjusts the initial probabilities of each state to a suitable range, making them comparable and generating confidence probabilities for each state. Then, the health state and mood category with the highest confidence probabilities are identified from these confidence probabilities and used as the final health state and mood category for the cat. The final judgment results are combined with the corresponding confidence probabilities to generate and output the monitoring results, allowing pet owners to clearly understand their cat's current health and mood status.
[0094] Reference Figure 4 An embodiment of the present invention provides a multimodal fusion intelligent monitoring system 4 for cat health and emotions, the system 4 specifically comprising: The data acquisition module 401 is used to acquire the cat's acoustic signals, facial images, full-body images and body temperature values to construct multimodal raw data; The feature extraction module 402 is used to perform data cleaning and feature extraction on multimodal raw data to obtain acoustic features, facial features, hair features and body temperature features; The feature fusion module 403 is used to normalize acoustic features, facial features, hair features and body temperature features, and combine principal component analysis and t-distribution random neighborhood embedding to perform dimensionality reduction and extract cross-modal core fusion feature set. The classification decision module 404 is used to make decisions based on the cross-modal core fusion feature set and a random forest classifier with dynamic weight allocation to identify the cat's health status and emotional state, and output monitoring results including health status, emotional category and their confidence probability.
[0095] It is understandable that, such as Figure 1 The content of the multimodal fusion cat health and emotion intelligent monitoring method embodiments shown are all applicable to the multimodal fusion cat health and emotion intelligent monitoring system embodiments. The specific functions implemented by the multimodal fusion cat health and emotion intelligent monitoring system embodiments are as follows: Figure 1 The multimodal fusion method for intelligent monitoring of feline health and emotions shown in the embodiment is the same, and the beneficial effects achieved are the same as those described above. Figure 1 The beneficial effects achieved by the multimodal fusion cat health and emotion intelligence monitoring method shown in the embodiment are the same.
[0096] It should be noted that the information interaction and execution process between the above systems are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0097] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0098] Reference Figure 5The present invention also provides a computer device 5, including: a memory 502 and a processor 501, and a computer program 503 stored in the memory 502. When the computer program 503 is executed on the processor 501, it implements the multimodal fusion intelligent monitoring method for cat health and emotion as described in any of the above methods.
[0099] The computer device 5 may be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device 5 may include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art will understand that... Figure 5 The computer device 5 is merely an example and does not constitute a limitation on the computer device 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components, such as input / output devices, network access devices, etc.
[0100] The processor 501 can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0101] In some embodiments, the memory 502 may be an internal storage unit of the computer device 5, such as a hard disk or memory of the computer device 5. In other embodiments, the memory 502 may be an external storage device of the computer device 5, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 5. Furthermore, the memory 502 may include both internal and external storage units of the computer device 5. The memory 502 is used to store the operating system, applications, boot loader, data, and other programs, such as the program code of the computer program. The memory 502 can also be used to temporarily store data that has been output or will be output.
[0102] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a multimodal fusion-based intelligent monitoring method for cat health and emotions as described in any of the above methods.
[0103] In this embodiment, if the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0104] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0105] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0106] In the embodiments disclosed in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0107] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. A multimodal fusion method for intelligent monitoring of feline health and emotions, characterized in that, The method specifically includes: Acoustic signals, facial images, full-body images, and body temperature values of cats were collected to construct multimodal raw data; Data cleaning and feature extraction were performed on the multimodal raw data to obtain acoustic features, facial features, hair features, and body temperature features; Acoustic features, facial features, hair features, and body temperature features are normalized and then dimensionality is reduced by combining principal component analysis and t-distribution random neighborhood embedding to extract cross-modal core fusion feature sets. Based on a cross-modal core fusion feature set, a random forest classifier with dynamic weight allocation is used to make decisions, identify the cat's health status and emotional state, and output monitoring results including health status, emotional category and their confidence probability.
2. The method according to claim 1, characterized in that, The process of cleaning and extracting features from the multimodal raw data to obtain acoustic features, facial features, hair features, and body temperature features specifically includes: The acoustic signal is denoised and preprocessed, and the Mel frequency cepstral coefficients, fundamental frequency, spectral entropy and snoring duration are extracted as acoustic features. Multiple facial key points are extracted from facial images using a key point detection model, and facial features such as eye width, ear angle, mouth opening, and beard spread are calculated based on these key points. The hair region was extracted by segmenting the full-body image, and the hair fluffiness, straightness, proportion of hair loss area and amount of dandruff were obtained by texture analysis as hair features. The body temperature value is compensated for by ambient temperature to obtain the corrected body temperature characteristics.
3. The method according to claim 2, characterized in that, The process of denoising and preprocessing the acoustic signal, and extracting Mel-frequency cepstral coefficients, fundamental frequency, spectral entropy, and snoring duration as acoustic features, specifically includes: A combined denoising algorithm that combines wavelet denoising and spectral subtraction is used to preprocess the acoustic signal, filter out environmental noise, and obtain the preprocessed acoustic signal. The preprocessed acoustic signal is subjected to frame-by-frame windowing, and the time-domain signal is converted into a time-frequency graph using short-time Fourier transform. Based on the time-frequency diagram, the Mel frequency cepstral coefficients are extracted using the Mel filter bank, the fundamental frequency of the sound is extracted using the fundamental frequency detection algorithm, the spectral entropy is calculated using spectral analysis, and the duration of snoring is statistically determined using endpoint detection technology. Mel frequency cepstral coefficients, fundamental frequency, spectral entropy, and duration of purring are used as acoustic features to characterize the acoustic state of cats.
4. The method according to claim 2, characterized in that, The process involves extracting multiple facial key points from the facial image using a key point detection model, and calculating facial features based on these key points, including eye width, ear angle, mouth opening, and beard extension. Specifically, this includes: Collect historical cat facial images of various breeds and scenarios, and manually label the positions of the eyes, ears, mouth and whiskers in each historical cat facial image to construct a first training dataset containing multiple core key points. The first training dataset is augmented by random rotation, scaling and brightness adjustment to increase the number of samples and form an augmented first training dataset. The lightweight convolutional neural network is trained end-to-end using the enhanced first training dataset, enabling it to learn the mapping relationship from facial images to the coordinates of multiple key points, thus forming a key point detection model. The real-time captured facial images are input into the key point detection model for processing, and the coordinates of multiple key points are output. Based on the coordinates of multiple key points, the eye width, the angle between the ear and the head, the mouth opening, and the whisker extension are calculated using a coordinate calculation algorithm to obtain facial features that characterize the cat's facial emotional state.
5. The method according to claim 2, characterized in that, The process of segmenting and extracting hair regions from a full-body image, and obtaining hair fluffiness, straightness, proportion of bald areas, and amount of dandruff through texture analysis as hair features, specifically includes: Collect full-body images of historical cats of various breeds and in various scenes, and perform pixel-level annotation on the hair regions in each full-body image of a historical cat to construct a second training dataset for hair segmentation; Based on the second training dataset, the Mask R-CNN algorithm was used to train the model to segment the cat's fur region from the full-body image, remove background and environmental interference, and obtain the trained fur segmentation model. The real-time acquired full-body image is input into the hair segmentation model for processing, and the binary mask of the hair region is output to extract the hair region image. Image enhancement processing is performed on the hair area image. By adjusting the brightness and contrast, the image deviation caused by changes in light or differences in hair color is corrected to obtain the enhanced hair image. Based on the enhanced hair image, texture analysis algorithms were used to extract the fluffiness region, straightness region, proportion of hair loss area to the whole body image, and amount of dandruff, obtaining hair features to characterize the health status of the cat's hair.
6. The method according to claim 2, characterized in that, The process of compensating for ambient temperature to obtain corrected body temperature characteristics specifically includes: Multiple sets of historical cat body surface original body temperature values and corresponding ambient temperature data were collected. A linear regression model was established, and the ambient temperature deviation coefficient was solved by the least squares method. The ambient temperature deviation coefficient is used to characterize the degree of interference of ambient temperature on body temperature measurement values. Substitute the ambient temperature deviation coefficient into the pre-built ambient temperature compensation formula to correct the original body temperature value, eliminate the measurement error introduced by the ambient temperature, and output the corrected core body temperature value as the body temperature feature. The ambient temperature compensation formula is as follows: , Indicates core body temperature, Indicates the original body temperature. Indicates ambient temperature. Indicates the reference temperature. This represents the ambient temperature deviation coefficient.
7. The method according to claim 1, characterized in that, The process involves normalizing acoustic features, facial features, hair features, and body temperature features, and then performing dimensionality reduction using principal component analysis and t-distributed random neighborhood embedding to extract a cross-modal core fusion feature set. Specifically, this includes: Acoustic features, facial features, hair features, and body temperature features are combined into an initial high-dimensional multimodal feature vector. The Min-Max Scaler algorithm is then used to perform a linear transformation on each dimension of the initial high-dimensional multimodal feature vector, mapping all feature values to the [0,1] interval to obtain a normalized feature set. Principal component analysis is performed on the normalized feature set. By calculating the covariance matrix and its eigenvalues, the top k principal components with a cumulative variance contribution rate greater than or equal to a preset contribution rate threshold are selected for preliminary dimensionality reduction, resulting in the feature set after PCA dimensionality reduction. The feature set after PCA dimensionality reduction is input into the t-distribution random neighborhood embedding algorithm. By constructing a Gaussian probability distribution in the high-dimensional space and a t-distribution probability distribution in the low-dimensional space, the KL divergence between the probability distributions in the high-dimensional and low-dimensional spaces is minimized by the gradient descent method, and the cross-modal core fusion feature set after iterative optimization is output.
8. The method according to claim 7, characterized in that, The step involves performing principal component analysis (PCA) on the normalized feature set. By calculating the covariance matrix and its eigenvalues, the first k principal components with a cumulative variance contribution rate greater than or equal to a preset contribution rate threshold are selected for preliminary dimensionality reduction, resulting in the PCA-reduced feature set. Specifically, this includes: The normalized feature set is used to construct a feature matrix. The feature matrix is then centered by subtracting the mean of each feature dimension to make the mean of each feature dimension zero, thus obtaining the centered feature matrix. Based on the centered feature matrix, the covariance matrix between features is calculated, and the covariance matrix is used to characterize the degree of linear correlation between features of different dimensions; The covariance matrix is decomposed into eigenvalues to obtain all eigenvalues and their corresponding eigenvectors. The eigenvalues are then sorted in descending order, and the order of the eigenvectors is adjusted accordingly. Calculate the variance contribution rate of each feature value, and calculate the cumulative variance contribution rate from front to back until the cumulative value reaches the preset contribution rate threshold. Determine the number of principal components k to be retained. The variance contribution rate is the proportion of a single feature value to the sum of all feature values. The eigenvectors corresponding to the top k largest eigenvalues are selected to form a projection matrix. The centered feature matrix is multiplied by the projection matrix to complete the linear transformation from the original high-dimensional space to the low-dimensional principal component space, and the dimensionality-reduced PCA feature set is output.
9. The method according to claim 7, characterized in that, The step involves inputting the PCA-reduced feature set into a t-distributed random neighborhood embedding algorithm. This algorithm constructs a Gaussian probability distribution in the high-dimensional space and a t-distributed probability distribution in the low-dimensional space, minimizing the KL divergence between the high- and low-dimensional probability distributions using gradient descent. The resulting iteratively optimized cross-modal core fusion feature set is then output. Specifically, this includes: The feature set after PCA dimensionality reduction is used as the high-dimensional input space of the t-distributed random neighborhood embedding algorithm. The Euclidean distance between any two sample points in the high-dimensional input space is calculated, and the Euclidean distance is converted into a conditional probability to characterize the similarity between sample points. Based on conditional probability, a joint probability distribution of the high-dimensional input space is constructed through symmetry processing, so that sample points with high similarity in the high-dimensional input space have higher probability values and sample points with low similarity have lower probability values. In the low-dimensional target space, low-dimensional mapping points that correspond one-to-one with high-dimensional sample points are randomly initialized, and the joint probability distribution between any two mapping points in the low-dimensional target space is calculated based on the t-distribution. The t-distribution is a Cauchy distribution with 1 degree of freedom and is used to alleviate the crowding problem in high-dimensional data. The KL divergence between the joint probability distribution of the high-dimensional input space and the joint probability distribution of the low-dimensional target space is used as the loss function. Iterative optimization is performed by gradient descent. In each iteration, the gradient of the loss function with respect to the low-dimensional mapping point is calculated, and the coordinates of the low-dimensional mapping point are updated to gradually minimize the KL divergence. After reaching the preset maximum number of iterations or the loss function converges, the iterative optimization stops, and the final low-dimensional mapping point coordinates are output as the cross-modal core fusion feature set.
10. A multimodal fusion intelligent monitoring system for cat health and emotions, characterized in that, The system specifically includes: The data acquisition module is used to collect acoustic signals, facial images, full-body images, and body temperature values of cats to construct multimodal raw data; The feature extraction module is used to clean and extract features from multimodal raw data to obtain acoustic features, facial features, hair features, and body temperature features. The feature fusion module is used to normalize acoustic features, facial features, hair features and body temperature features, and combine principal component analysis and t-distribution random neighborhood embedding to perform dimensionality reduction and extract cross-modal core fusion feature set. The classification decision module is used to make decisions based on cross-modal core fusion feature sets and a random forest classifier with dynamic weight allocation to identify the cat's health status and emotional state, and output monitoring results including health status, emotional category and their confidence probability.