Pet status recognition method and system based on motion capture and sound analysis
Patent Information
- Application Number
- CN202610480589.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-13
- Publication Date
- 2026-08-21
AI Technical Summary
然而该动物健康监测方法并没有考虑动物异常动作情况以及识别情感状态,难以对动物状态进行细化辨识,其动物状态辨识较为粗糙,有待改进优化
Smart Images

Figure CN122603783A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to a method and system for pet status recognition based on motion capture and sound analysis. Background Technology
[0002] In modern society, pets have become important members of many families, and their health and mental well-being are of great concern. However, because pets cannot directly express their feelings like humans, pet owners often find it difficult to accurately judge their pets' psychophysiological state, such as anxiety, comfort, or pain. Traditional methods mainly rely on observing the pet's external behavior and the pet owner's experience, which is highly subjective and struggles to capture subtle psychophysiological changes in pets.
[0003] With the development of sensor technology and artificial intelligence, motion capture and sound analysis technologies have offered new possibilities for identifying the psychological and physiological states of pets. Motion capture technology can track a pet's movement trajectory and posture changes in real time using sensors, while sound analysis technology can perform spectral analysis and emotion recognition on a pet's vocalizations. Combining these two technologies promises to achieve objective and accurate identification of a pet's psychological and physiological state, thereby providing pet owners with more scientific pet care advice.
[0004] Chinese invention patent application number CN201810750850.1 discloses an animal health monitoring method, device, and computer-readable storage medium. It utilizes sound and motion recognition from animal videos, combining these with animal sound and motion tags to predict the probability of disease and outputs the animal health monitoring results based on this probability. However, this animal health monitoring method does not consider abnormal animal movements or emotional states, making it difficult to achieve detailed identification of animal conditions. Its animal condition identification is relatively crude and requires improvement and optimization. Summary of the Invention
[0005] Based on this, in order to improve the recognition of animal states and perform personalized identification of animals, this invention provides a pet state recognition method and system based on motion capture and sound analysis, the specific technical solution of which is as follows:
[0006] A pet status identification method based on motion capture and sound analysis includes the following steps: Real-time acquisition of the target pet's movement and sound information; The motion information is processed to extract motion features, and a motion abnormality index is obtained based on the motion features to quantify the degree to which the pet's motion deviates from the normal baseline. The system processes incoming and outgoing sound information, extracts sound features, and obtains an emotional score based on these features to quantify the intensity of negative emotions in the vocalizations. The abnormal behavior index and the emotion score are fused together to obtain a comprehensive status score, and the target pet's status is identified based on the comprehensive status score.
[0007] The pet state identification method based on motion capture and sound analysis obtains a motion abnormality index to quantify the degree to which a pet's actions deviate from the normal baseline and an emotion score to quantify the intensity of negative emotions in its vocalizations. The motion abnormality index and the emotion score are then fused to obtain a comprehensive state score. Finally, the target pet's state is identified based on the comprehensive state score. This method can perform differentiated and personalized identification based on the different pets' behavioral habits and vocal characteristics, thus improving the precision of animal state identification.
[0008] Preferably, the specific method for obtaining the abnormal behavior index includes: Obtain the current acceleration of key monitoring points of the target pet; Obtain the absolute difference between the historical average acceleration and the current acceleration, and obtain the motion anomaly index based on the absolute difference value; Among these, motion characteristics include current acceleration.
[0009] Preferably, the specific methods for obtaining the sentiment score include: Initial sentiment scores are obtained based on Mel-spectral cepstral coefficients. Obtain the target pet's historical voice information and calculate the average emotion score benchmark based on the historical voice information; The initial sentiment score is offset and adjusted based on the average sentiment score benchmark to obtain the final sentiment score.
[0010] Preferably, the pet status identification method further includes the following steps: The motion features and voice features are fused together to obtain the fused features; Based on the fusion features, obtain the state probability distribution used to quantify the likelihood of the target pet being in different states.
[0011] Preferably, the abnormal movement index is expressed as: ; in, These represent the abnormal motion index, the current acceleration of the i-th key monitoring point in three-dimensional space, and the historical average acceleration, respectively. These represent the total value of key monitoring points and the weight coefficient of each monitoring point, respectively. This represents the Euclidean norm.
[0012] A pet state recognition system based on motion capture and sound analysis is used to implement the aforementioned pet state recognition method based on motion capture and sound analysis, comprising: The pet information acquisition module is used to acquire the target pet's action and sound information in real time. The abnormality index acquisition module is used to process motion information, extract motion features, and obtain a motion abnormality index based on the motion features to quantify the degree to which the pet's motion deviates from the normal baseline. The emotion score acquisition module is used to process incoming and outgoing sound information, extract sound features, and obtain an emotion score based on the sound features to quantify the intensity of negative emotions in the call. The pet status identification module is used to fuse the abnormal behavior index and the emotion score to obtain a comprehensive status score, and to identify the status of the target pet based on the comprehensive status score.
[0013] Preferably, the anomaly index acquisition module includes: The current acceleration acquisition unit is used to acquire the current acceleration of key monitoring points of the target pet; The motion anomaly index acquisition unit is used to obtain the absolute difference between the historical average acceleration and the current acceleration, and to obtain the motion anomaly index based on the absolute difference value. Among these, motion characteristics include current acceleration.
[0014] Preferably, the emotion score acquisition module includes: The initial sentiment score acquisition unit is used to acquire the initial sentiment score based on the Mel-spectrum cepstral coefficients. The emotional score benchmark acquisition unit is used to acquire the historical sound information of the target pet and obtain the average emotional score benchmark based on the historical sound information. The final sentiment score acquisition unit is used to offset and adjust the initial sentiment score based on the average sentiment score benchmark to obtain the final sentiment score.
[0015] Preferably, the pet state recognition system based on motion capture and sound analysis further includes: The fusion feature acquisition module is used to fuse action features and sound features to obtain fused features; The probability distribution acquisition module is used to obtain the state probability distribution based on the fusion features to quantify the likelihood of the target pet being in different states.
[0016] Preferably, the abnormal action index acquisition unit is based on the formula Obtain the abnormality index of the action; in, These represent the abnormal motion index, the current acceleration of the i-th key monitoring point in three-dimensional space, and the historical average acceleration, respectively. These represent the total value of key monitoring points and the weight coefficient of each monitoring point, respectively. This represents the Euclidean norm. Attached Figure Description
[0017] The invention will be further understood from the following description taken in conjunction with the accompanying drawings. The components in the drawings are not necessarily drawn to scale, but rather the emphasis is on illustrating the principles of the embodiments. In different views, the same reference numerals designate corresponding parts.
[0018] Figure 1 This is a schematic diagram of the overall process of a pet state recognition method based on motion capture and sound analysis in one embodiment of the present invention. Figure 2 This is a flowchart illustrating a specific method for obtaining an abnormal behavior index in one embodiment of the present invention. Figure 3 This is a flowchart illustrating a specific method for obtaining an emotion score in one embodiment of the present invention; Figure 4 This is a schematic diagram of the overall process of a pet state recognition method based on motion capture and sound analysis in another embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to its embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of the invention.
[0020] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly attached to the other element or there may be an intervening element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or there may be an intervening element. The terms "vertical," "horizontal," "left," "right," and similar expressions used herein are for illustrative purposes only and do not represent the only possible implementation.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0022] In this invention, "first" and "second" do not represent a specific quantity or order, but are merely used to distinguish names.
[0023] Before describing the specific embodiments of the present invention, a brief introduction to the prior art will be given first.
[0024] In modern society, pets have become important members of many families, and their health and mental well-being are of great concern. However, because pets cannot directly express their feelings like humans, pet owners often find it difficult to accurately judge their pets' psychophysiological state, such as anxiety, comfort, or pain. Traditional methods mainly rely on observing the pet's external behavior and the pet owner's experience, which is highly subjective and struggles to capture subtle psychophysiological changes in pets.
[0025] With the development of sensor technology and artificial intelligence, motion capture and sound analysis technologies have offered new possibilities for identifying the psychological and physiological states of pets. Motion capture technology can track a pet's movement trajectory and posture changes in real time using sensors, while sound analysis technology can perform spectral analysis and emotion recognition on a pet's vocalizations. Combining these two technologies promises to achieve objective and accurate identification of a pet's psychological and physiological state, thereby providing pet owners with more scientific pet care advice.
[0026] Existing technologies for identifying the psychological and physiological states of pets still have many problems, such as unreasonable sensor placement, inaccurate data processing algorithms, and low reliability of identification results. These problems include the following: 1. Lack of high-precision motion capture technology: In existing technologies, the accuracy and stability of motion capture sensors are insufficient, making it difficult to accurately capture subtle changes in a pet's movements.
[0027] 2. Inaccurate sound analysis algorithm: When processing pet barks, the sound analysis algorithm is easily affected by environmental noise, resulting in inaccurate identification results.
[0028] 3. Low data processing efficiency: A large amount of motion and sound data needs to be processed in real time, and the data processing efficiency of existing technologies is difficult to meet the needs.
[0029] 4. Insufficient personalized identification ability: Different pets have different behavioral habits and vocal characteristics, and current technology lacks the ability to identify individuals individually.
[0030] 5. Poor user interaction experience: The presentation of the identification results is not intuitive enough, making it difficult for pet owners to quickly understand their pets' psychological and physiological state.
[0031] For example, Chinese invention patent application number CN201810750850.1 discloses an animal health monitoring method, device and computer-readable storage medium. It uses sound recognition and action recognition of animal videos, combines animal sound tags and action tags to predict the probability of animal disease, and outputs the animal health monitoring results based on the disease probability. It effectively solves the problem in the prior art that the accuracy of animal health detection may be low due to the low accuracy of animal language recognition.
[0032] However, this animal health monitoring method does not take into account abnormal animal behavior or emotional state, making it difficult to identify animal conditions in detail. Its animal condition identification is rather crude and needs to be improved and optimized.
[0033] Therefore, developing a pet psychophysiological state identification technology based on motion capture and sound analysis to improve the accuracy of pet state identification has important practical significance and application prospects.
[0034] One objective of this invention is to improve the ability to identify animal states and to perform differentiated and personalized identification of animals. To this end, such as... Figure 1 As shown, an embodiment of the present invention provides a pet state recognition method based on motion capture and sound analysis, including the following steps: S1 acquires real-time information on the target pet's movements and sounds.
[0035] Here, IMU inertial measurement units and depth cameras can be deployed at key joints of the pet (such as the neck, back, and legs) with a sampling rate of over 100Hz to capture micro-movements. The motion trajectory is then calculated in real time using optical flow algorithms and attitude estimation models.
[0036] For audio information, a high-sensitivity microphone array with a sampling rate of 44.1kHz can be used, and directional beamforming technology can be applied to focus the pet's sound source.
[0037] S2 processes the motion information, extracts motion features, and obtains a motion abnormality index based on the motion features to quantify the degree to which the pet's motion deviates from the normal baseline.
[0038] As a preferred technical solution, such as Figure 2 As shown, the specific methods for obtaining the abnormal behavior index include: S21, Obtain the current acceleration of the target pet's key monitoring points.
[0039] S22, obtain the absolute difference between the historical average acceleration and the current acceleration, and obtain the motion anomaly index based on the absolute difference.
[0040] Among these, motion characteristics include current acceleration.
[0041] For example, the abnormal behavior index is represented as ;in, These represent the abnormal motion index, the current acceleration of the i-th key monitoring point in three-dimensional space, and the historical average acceleration, respectively. These represent the total value of key monitoring points and the weight coefficient of each monitoring point, respectively. This represents the Euclidean norm.
[0042] Key monitoring points include, but are not limited to, the neck, legs, and back.
[0043] Pet anxiety and stress manifest as fine tremors, which are continuous, high-frequency, small-amplitude vibrations across multiple frames. Here, the abnormal movement index employs a normalized design, which can effectively distinguish between a single unintentional limb twitch (normal) and a continuous low-frequency tremor (abnormal). By taking into account the temporal fluctuation characteristics of micro-movements, the false negative rate of core detection targets such as fine tremors can be reduced.
[0044] The neck and other key monitoring points are only applicable to anxious scenarios. Large neck movements during pet sleep and eating are normal behaviors. Here, the weight coefficients of the monitoring points can be dynamically adjusted in real time according to the sleep / eating / activity / anxiety state, decreasing the weight of the neck during sleep and automatically increasing it during anxiety.
[0045] Since pathological micro-movements in pets are concentrated along the Z-axis (vertical head nodding, up-and-down body shaking), while the X / Y axes are mostly routine translational swaying, a preferred technical solution is to differentiate the sensitivity coefficients across the three axes, focusing on the core tremor dimension of the pet to enhance Z-axis tremor detection and improve micro-movement sensitivity. For example, .in, The three-axis weighting coefficients are preferred. Increase the weight of vertical tremors and suppress irrelevant horizontal swaying. These represent the current acceleration of the i-th key monitoring point along the X, Y, and Z axes, respectively. These are the historical average accelerations along the X, Y, and Z axes, respectively.
[0046] When the abnormality index exceeds the preset abnormality threshold, it can be marked as an abnormal state. The abnormality threshold is based on a dynamic threshold adjustment function. Dynamic settings are implemented to adapt to different pets. Among these, These are represented in sequence as the base threshold (e.g., 0.5 m / s²), adjustment coefficient, and size factor. Size factor: Small pets < 1, Large pets > 1, which can be set according to actual needs.
[0047] In this way, by adjusting the threshold for abnormal movements using body shape factors, personalized identification capabilities can be improved.
[0048] S3 processes the incoming and outgoing sound information, extracts sound features, and obtains an emotional score based on the sound features to quantify the intensity of negative emotions in the call.
[0049] For the processing of sound information, short-time Fourier transform can be used to perform spectral analysis of the sound information, and the sound information can be dynamically adjusted based on the signal-to-noise ratio to suppress noise.
[0050] Specifically, the signal-to-noise ratio of the original sound information is first dynamically adjusted to suppress noise interference. Then, the MFCC features of the processed sound information, namely the Mel frequency cepstral coefficients, are extracted. Based on the Mel frequency cepstral coefficients, an emotional score is obtained to quantify the intensity of negative emotions in the call.
[0051] Here, the emotion score can be understood as the emotional inclination of a pet's vocalization. By obtaining the Mel-frequency cepstral coefficients, the emotion score can be derived based on the corresponding neural network model. This neural network model can be trained based on pre-acquired sound information and corresponding emotion score labels.
[0052] As a preferred technical solution, such as Figure 3 As shown, the specific methods for obtaining sentiment scores include: S31, the initial sentiment score is obtained based on the Mel-spectrum cepstral coefficients.
[0053] S32, obtain the target pet's historical voice information, and obtain the average emotion score benchmark based on the historical voice information.
[0054] S33, the initial sentiment score is offset and adjusted based on the average sentiment score benchmark to obtain the final sentiment score.
[0055] For example, .in, These are represented, in order, the learning rate, the initial sentiment score, and the average sentiment score baseline.
[0056] S4 integrates the abnormal behavior index and the emotion score to obtain a comprehensive status score, and identifies the target pet's status based on the comprehensive status score.
[0057] For example, the comprehensive state score is represented as a weighted sum of the action anomaly index and the emotion score, with a higher value indicating a higher probability of an abnormal / negative state. The weighting coefficient of the action anomaly index can be set as: action information confidence / (action information confidence + voice information confidence), and the corresponding weighting coefficient of the action anomaly index is set as: 1 - weighting coefficient of the action anomaly index.
[0058] In summary, the pet state identification method based on motion capture and sound analysis obtains a motion abnormality index to quantify the degree to which a pet's actions deviate from the normal baseline and an emotion score to quantify the intensity of negative emotions in its vocalizations. The motion abnormality index and emotion score are then fused to obtain a comprehensive state score. Finally, the target pet's state is identified based on the comprehensive state score. This method can perform differentiated and personalized identification based on the different behavioral habits and vocal characteristics of various pets, thus improving the precision of animal state identification.
[0059] In one embodiment, such as Figure 4As shown, the pet status identification method further includes the following steps: S5 performs fusion processing on motion features and voice features to obtain fused features.
[0060] Motion characteristics include, but are not limited to, joint acceleration variance and posture change frequency; sound characteristics include, but are not limited to, joint acceleration variance and posture change frequency.
[0061] S6. Obtain the state probability distribution based on the fusion features to quantify the possibility of the target pet being in different states.
[0062] Specifically, based on the CNN-LSTM model, the fused features are first used as input to the LSTM neural network in the CNN-LSTM model to obtain the hidden state vector at the current time step, and the hidden state sequence is obtained based on the hidden state vector at the current time step. The hidden state sequence is then used as input to the CNN neural network, and the final output is a state probability distribution representing the pet's psychophysiological state (such as the probability of anxiety, comfort, and pain).
[0063] For CNN-LSTM models, a training dataset can be constructed by collecting historical pet action and sound data along with corresponding psychophysiological state labels, and then the model can be trained based on the training dataset.
[0064] After obtaining the state probability distribution, in some cases, the maximum value of the state probability distribution can be used as the initial sentiment score, or the weighted value of the probabilities of each state in the state probability distribution can be used as the initial sentiment score.
[0065] It's important to note that the state probability distribution represents the probability of a specific psychological / physiological state of the target pet, such as anxiety = 0.7, comfort = 0.1, and pain = 0.2. The comprehensive state score can be understood as the probability of an abnormal / negative state of the target pet. Generally, the higher the value, the higher the probability of a negative state. The comprehensive state score serves as a lightweight intermediate output, reducing the computational load of subsequent models. For example, when the comprehensive state score is less than a preset state threshold, such as 0.3, the subsequent acquisition of fused features and deep learning inference of state probabilities can be skipped, and "normal state" can be directly output, improving system response speed. When the comprehensive state score is not less than the preset state threshold, the state probability distribution is then acquired for more refined and personalized identification of the target pet's state.
[0066] In other words, the comprehensive state score is used as a coarse screening layer, and the state probability distribution is used as a fine classification layer. The abnormal state of the target pet is first screened by the comprehensive state score, and then the abnormal state is finely classified by the calculation of the state probability distribution, thus balancing efficiency and accuracy.
[0067] As a preferred technical solution, the state probability distribution can also be expressed as .in, The output class label, class, and fusion feature of the Softmax classifier are represented in sequence, with k∈{1,2,3} corresponding to: 1=anxiety, 2=comfort, and 3=pain. This represents the classification weight of the k-th class. These represent the bias terms for the k-th and j-th classes, respectively, used to compensate for differences in baseline probabilities among the classes; K represents the total number of emotion categories, such as K=3 corresponding to anxiety / comfort / pain. These represent the matching degree between the k-th feature vector and the k-th class weight, and the matching degree between the j-th feature vector and the j-th class weight, respectively. The higher the value, the greater the probability of belonging to that class.
[0068] Here, the fused features of action and sound are mapped to a probability distribution by using the output category labels of the Softmax classifier, which can quantify the probability that the pet is in different states.
[0069] An embodiment of the present invention also provides a pet state recognition system based on motion capture and sound analysis, used to implement the pet state recognition method based on motion capture and sound analysis, which includes a pet information acquisition module, an abnormality index acquisition module, an emotion score acquisition module, and a pet state recognition module.
[0070] The pet information acquisition module is used to acquire the target pet's action and sound information in real time; the anomaly index acquisition module is used to process the action information, extract action features, and obtain an action anomaly index based on the action features to quantify the degree to which the pet's actions deviate from the normal baseline.
[0071] The emotion score acquisition module is used to process incoming and outgoing sound information, extract sound features, and obtain an emotion score based on the sound features to quantify the intensity of negative emotions in the vocalizations; the pet status identification module is used to fuse the abnormal behavior index and the emotion score to obtain a comprehensive status score, and identify the status of the target pet based on the comprehensive status score.
[0072] As a preferred technical solution, the anomaly index acquisition module includes a current acceleration acquisition unit and an action anomaly index acquisition unit.
[0073] The current acceleration acquisition unit is used to acquire the current acceleration of key monitoring points of the target pet; the action anomaly index acquisition unit is used to acquire the absolute difference between the historical average acceleration and the current acceleration, and to acquire the action anomaly index based on the absolute difference; among them, the action features include the current acceleration.
[0074] For example, the action anomaly index acquisition unit obtains the index according to the formula. Obtain the abnormality index of the action; among which, These represent the abnormal motion index, the current acceleration of the i-th key monitoring point in three-dimensional space, and the historical average acceleration, respectively. These represent the total value of key monitoring points and the weight coefficient of each monitoring point, respectively. This represents the Euclidean norm.
[0075] The emotion score acquisition module includes an initial emotion score acquisition unit, an emotion score benchmark acquisition unit, and a final emotion score acquisition unit.
[0076] The initial emotion score acquisition unit is used to obtain the initial emotion score based on the Mel spectrum cepstral coefficients; the emotion score benchmark acquisition unit is used to obtain the historical sound information of the target pet and obtain the average emotion score benchmark based on the historical sound information; the final emotion score acquisition unit is used to offset and adjust the initial emotion score based on the average emotion score benchmark to obtain the final emotion score.
[0077] For example, .in, These are represented, in order, the learning rate, the initial sentiment score, and the average sentiment score baseline.
[0078] In one embodiment, the pet state recognition system based on motion capture and sound analysis further includes a fusion feature acquisition module and a probability distribution acquisition module.
[0079] The fusion feature acquisition module is used to fuse action features and sound features to obtain fusion features; the probability distribution acquisition module is used to obtain the state probability distribution based on the fusion features to quantify the possibility that the target pet is in different states.
[0080] For example, the state probability distribution can also be expressed as .in, The output class label, class, and fusion feature of the Softmax classifier are represented in sequence, with k∈{1,2,3} corresponding to: 1=anxiety, 2=comfort, and 3=pain. This represents the classification weight of the k-th class. These represent the bias terms for the k-th and j-th classes, respectively, used to compensate for differences in baseline probabilities among the classes; K represents the total number of emotion categories, such as K=3 corresponding to anxiety / comfort / pain. These represent the matching degree between the k-th feature vector and the k-th class weight, and the matching degree between the j-th feature vector and the j-th class weight, respectively. The higher the value, the greater the probability of belonging to that class.
[0081] Here, the fused features of action and sound are mapped to a probability distribution by using the output category labels of the Softmax classifier, in order to quantify the probability that the pet is in different states.
[0082] The system also includes a user interaction module, which presents the identification results in an intuitive way, such as sending notifications and suggestions to pet owners via an app or smart device. In this way, a simple and intuitive user interface allows pet owners to easily understand their pets' psychological and physiological states.
[0083] In summary, the pet state recognition system based on motion capture and sound analysis described in this invention innovatively and objectively identifies the psychological and physiological states of pets using ingenious motion capture and sound analysis technologies. Motion capture technology utilizes sensor networks to capture subtle movements of pets in real time, accurately reflecting their activity status and behavioral habits. Sound analysis technology, through deep learning models, analyzes the frequency, rhythm, and other characteristics of pet vocalizations, providing insights into their emotional changes and health status.
[0084] In other words, the system described in this invention has the characteristic of highly reliable identification results, which can provide pet owners with scientific care basis, effectively guide daily care and health management, and has a strong personalized and differentiated identification capability. It can establish a unique file for each pet's unique behavioral habits and vocal characteristics to achieve accurate identification.
[0085] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0086] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.
Claims
1. A pet state recognition method based on motion capture and sound analysis, characterized in that, The pet status identification method includes the following steps: Real-time acquisition of the target pet's movement and sound information; The motion information is processed to extract motion features, and a motion abnormality index is obtained based on the motion features to quantify the degree to which the pet's motion deviates from the normal baseline. The system processes incoming and outgoing sound information, extracts sound features, and obtains an emotional score based on these features to quantify the intensity of negative emotions in the vocalizations. The abnormal behavior index and the emotion score are fused together to obtain a comprehensive status score, and the target pet's status is identified based on the comprehensive status score.
2. The pet state recognition method based on motion capture and sound analysis as described in claim 1, characterized in that, Specific methods for obtaining the abnormality index of actions include: Obtain the current acceleration of key monitoring points of the target pet; Obtain the absolute difference between the historical average acceleration and the current acceleration, and obtain the motion anomaly index based on the absolute difference value; Among these, motion characteristics include current acceleration.
3. The pet state recognition method based on motion capture and sound analysis as described in claim 2, characterized in that, Specific methods for obtaining an emotion score include: Initial sentiment scores are obtained based on Mel-spectral cepstral coefficients. Obtain the target pet's historical voice information and calculate the average emotion score benchmark based on the historical voice information; The initial sentiment score is offset and adjusted based on the average sentiment score benchmark to obtain the final sentiment score.
4. The pet state recognition method based on motion capture and sound analysis as described in claim 3, characterized in that, The pet status identification method also includes the following steps: The motion features and voice features are fused together to obtain the fused features; Based on the fusion features, obtain the state probability distribution used to quantify the likelihood of the target pet being in different states.
5. The pet state recognition method based on motion capture and sound analysis as described in claim 4, characterized in that, The abnormal behavior index is expressed as ; in, These represent the abnormal motion index, the current acceleration of the i-th key monitoring point in three-dimensional space, and the historical average acceleration, respectively. These represent the total value of key monitoring points and the weight coefficient of each monitoring point, respectively. This represents the Euclidean norm.
6. A pet state recognition system based on motion capture and sound analysis, used to implement the pet state recognition method based on motion capture and sound analysis as described in any one of claims 1-5, characterized in that, The pet state recognition system based on motion capture and sound analysis includes: The pet information acquisition module is used to acquire the target pet's action and sound information in real time. The abnormality index acquisition module is used to process motion information, extract motion features, and obtain a motion abnormality index based on the motion features to quantify the degree to which the pet's motion deviates from the normal baseline. The emotion score acquisition module is used to process incoming and outgoing sound information, extract sound features, and obtain an emotion score based on the sound features to quantify the intensity of negative emotions in the call. The pet status identification module is used to fuse the abnormal behavior index and the emotion score to obtain a comprehensive status score, and to identify the status of the target pet based on the comprehensive status score.
7. The pet state recognition system based on motion capture and sound analysis as described in claim 6, characterized in that, The abnormal index acquisition module includes: The current acceleration acquisition unit is used to acquire the current acceleration of key monitoring points of the target pet; The motion anomaly index acquisition unit is used to obtain the absolute difference between the historical average acceleration and the current acceleration, and to obtain the motion anomaly index based on the absolute difference value. Among these, motion characteristics include current acceleration.
8. The pet state recognition system based on motion capture and sound analysis as described in claim 7, characterized in that, The emotion score acquisition module includes: The initial sentiment score acquisition unit is used to acquire the initial sentiment score based on the Mel-spectrum cepstral coefficients. The emotional score benchmark acquisition unit is used to acquire the historical sound information of the target pet and obtain the average emotional score benchmark based on the historical sound information. The final sentiment score acquisition unit is used to offset and adjust the initial sentiment score based on the average sentiment score benchmark to obtain the final sentiment score.
9. The pet state recognition system based on motion capture and sound analysis as described in claim 8, characterized in that, The pet state recognition system based on motion capture and sound analysis also includes: The fusion feature acquisition module is used to fuse action features and sound features to obtain fused features; The probability distribution acquisition module is used to obtain the state probability distribution based on the fusion features to quantify the likelihood of the target pet being in different states.
10. The pet state recognition system based on motion capture and sound analysis as described in claim 9, characterized in that, The abnormality index acquisition unit obtains the index based on the formula. Obtain the abnormality index of the action; in, These represent the abnormal motion index, the current acceleration of the i-th key monitoring point in three-dimensional space, and the historical average acceleration, respectively. These represent the total value of key monitoring points and the weight coefficient of each monitoring point, respectively. This represents the Euclidean norm.
Citation Information
Patent Citations
An animal health monitoring method, device, and computer-readable storage medium
CN108922622B