Pet behavior monitoring system and method based on edge computing and cloud time series modeling

CN122818237APending Publication Date: 2026-09-25XINGRUI INNOVATION TECHNOLOGY (HANGZHOU) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611006451.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]目前市面上的宠物监控设备多以远程查看和视频录制为主,纯云端方案存在持续上传私密视频导致的隐私风险、带宽占用大以及网络波动引发的告警滞后等问题;同时,边缘端设备的算力有限,难以分析宠物的长期健康趋势,且单一视觉识别易受光照和遮挡的干扰

Benefits of technology

[0014]与现有技术相比,本发明的优点和积极效果在于:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818237A_ABST
    Figure CN122818237A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of behavior monitoring, in particular to a pet behavior monitoring system and method based on edge computing and cloud timing modeling, comprising a data acquisition synchronization module, a feature structuring module, a behavior feature fusion module, a timing feature modeling module and a behavior anomaly early warning module.In the present application, through the synchronous acquisition and time alignment of multi-modal data, the structured features are constructed by extracting the skeletal key points, acoustic features and motion energy, and the fusion of behavior features is completed, the long-term living habit baseline is generated, and the multi-dimensional space statistical deviation and dynamic adaptive alarm threshold are introduced for comparison and other technical means, realizing the efficient end-cloud collaborative processing of pet daily activities and the complementary fusion of multi-dimensional features, overcoming the interference of single visual environment, avoiding the continuous uploading of private video to the cloud, not only reducing the privacy risk and guaranteeing the timeliness of early warning and alarm, but also scientifically and effectively analyzing and evaluating the long-term health development trend of pets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of behavior monitoring technology, and in particular to a pet behavior monitoring system and method based on edge computing and cloud-based time-series modeling. Background Technology

[0002] Behavioral monitoring technology encompasses the continuous observation and recording of the posture changes and activity states of specific target objects. Its main purpose is to analyze the target object's daily routines and abnormal conditions by collecting data on changes in its movements at different times. It is widely used in security monitoring, smart healthcare, animal husbandry, and home care industries. Traditional pet behavioral monitoring refers to the long-term tracking and recording of companion animals' daily activity trajectories and physiological movements. It is mainly used to identify the animal's eating, sleeping, excretion, and exercise status. Typically, wearable sensors are used to perceive physical activity characteristics, or fixed-position cameras are used to record planar images. These images are then processed by a center according to predetermined rules to perform pattern matching and determine the specific category of the animal's movements.

[0003] Currently, most pet monitoring devices on the market focus on remote viewing and video recording. Pure cloud solutions have issues such as privacy risks caused by continuous uploading of private videos, high bandwidth consumption, and delayed alarms caused by network fluctuations. At the same time, edge devices have limited computing power, making it difficult to analyze the long-term health trends of pets, and single visual recognition is easily affected by lighting and occlusion. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a pet behavior monitoring system and method based on edge computing and cloud-based time-series modeling.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a pet behavior monitoring system based on edge computing and cloud-based time-series modeling, comprising: The data acquisition and synchronization module is used to acquire video keyframes, audio spectrum segments, and physical vibration signals during the continuous monitoring period of pet behavior monitoring, and to perform time-alignment processing on the timestamps to generate a multimodal raw data synchronization sequence. The feature structuring module is used to parse the multimodal raw data synchronization sequence, obtain pet video data stream, pet audio data stream and pet inertial data stream, extract normalized pet skeletal key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, pet environmental audio signal-to-noise ratio and pet motion energy statistics, and construct a multimodal structured feature set. The behavior feature fusion module is used to analyze the multimodal structured feature set, extract environmental noise evaluation coefficients, determine modal dynamic weights by combining visual recognition confidence scores, and use the modal dynamic weights to fuse the normalized pet skeleton key point coordinate sequence, pet acoustic feature vector, and pet motion energy statistics in the multimodal structured feature set to generate a multimodal behavior fusion feature vector. The temporal feature modeling module is used to obtain a sliding time window of the pet's historical behavior, extract the multimodal behavior fusion feature vector, construct a temporal feature sequence, analyze the temporal feature sequence to extract the temporal dependency distribution attributes, and construct a baseline of the pet's long-term living habits features. The abnormal behavior warning module is used to determine the multidimensional spatial statistical deviation between the multimodal behavior fusion feature vector and the pet's long-term living habit feature baseline, construct a dynamic adaptive alarm threshold, and generate a pet behavior monitoring warning instruction when the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold; and generate a pet behavior normal status indicator instruction when the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold.

[0006] As a further aspect of the present invention, the pet inertial data stream includes triaxial acceleration components and angular velocity components; The pet video data stream is analyzed using a multi-task lightweight AI backbone network to extract normalized pet skeletal key point coordinate sequences and pet visual recognition confidence scores; the pet audio data stream is analyzed to extract pet acoustic feature vectors and pet environmental audio signal-to-noise ratio; and the pet inertial data stream is analyzed to extract pet motion energy statistics.

[0007] As a further aspect of the present invention, when constructing the dynamic adaptive alarm threshold, a basic warning threshold is extracted from the medical pet behavior and health guidelines, and the warning parameters of the previous historical moment are extracted simultaneously. The basic warning threshold and the warning parameters of the previous historical moment are smoothly fitted using the exponential moving average algorithm, and the dynamic adaptive alarm threshold is output.

[0008] As a further aspect of the present invention, the data acquisition and synchronization module includes: The signal data acquisition submodule acquires visual keyframes collected by the visual sensor, acquires audio spectrum segments corresponding to the microphone array, acquires physical vibration signals corresponding to the three-axis inertial measurement unit, parses the visual keyframes to extract the first acquisition timestamp, parses the audio spectrum segments to extract the second acquisition timestamp, parses the physical vibration signals to extract the third acquisition timestamp, and combines the first acquisition timestamp, the second acquisition timestamp, and the third acquisition timestamp to establish an initial timetamp set. The time microsecond alignment submodule, for the initial set of timestamps, uses a hardware-level timestamp synchronization protocol to parse the first acquisition timestamp, the second acquisition timestamp, and the third acquisition timestamp and performs time comparison calculations, and performs time dimension alignment adjustment on the visual keyframes, audio spectrum segments, and physical vibration signals to generate an alignment signal correlation matrix. The synchronization sequence generation submodule, based on the alignment signal correlation matrix, extracts aligned visual keyframes to construct a pet video data stream, extracts aligned audio spectrum segments to construct a pet audio data stream, extracts aligned physical vibration signals to construct a pet inertial data stream containing triaxial acceleration components and angular velocity components, and combines them to generate a multimodal raw data synchronization sequence.

[0009] As a further aspect of the present invention, the feature structuring module includes: The video and audio parsing submodule, based on the multimodal raw data synchronization sequence, separates the pet video data stream and the pet audio data stream, uses a multi-task lightweight artificial intelligence backbone network to parse the pet video data stream, extracts the normalized pet skeleton key point coordinate sequence and pet visual recognition confidence score, parses the pet audio data stream, extracts the pet acoustic feature vector and pet environmental audio signal-to-noise ratio, and integrates and establishes audiovisual feature parsing parameters. The inertial motion statistics submodule separates the multimodal raw data synchronization sequence to obtain the pet inertial data stream based on the audiovisual feature parsing parameters, uses a multi-task lightweight artificial intelligence backbone network to parse the pet inertial data stream, extracts the pet motion energy change value corresponding to the pet inertial data stream and performs cumulative calculation to generate pet motion energy statistics. The structured feature construction submodule extracts the normalized pet skeletal key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, and pet environmental audio signal-to-noise ratio from the audiovisual feature parsing parameters, and summarizes them with the pet motion energy statistics to generate a multimodal structured feature set.

[0010] As a further aspect of the present invention, the behavioral feature fusion module includes: The environmental noise assessment submodule, based on the multimodal structured feature set, extracts the pet environment audio signal-to-noise ratio and pet visual recognition confidence score, compares the pet environment audio signal-to-noise ratio with the general acoustic environment interference evaluation standard library and performs matching mapping, and extracts environmental noise assessment coefficients based on the matching mapping action results; The dynamic weight allocation submodule uses a multimodal dynamic weight allocation model to analyze the pet visual recognition confidence score and the environmental noise evaluation coefficient, determine the distribution law of multimodal feature proportion, output the visual modal dynamic weight, acoustic modal dynamic weight and inertial modal dynamic weight, and establish the modal dynamic weight coefficient. The semantic feature fusion submodule, based on the modal dynamic weight coefficients, extracts visual semantic feature vectors through the normalized pet skeleton key point coordinate sequence, extracts acoustic semantic feature vectors through the pet acoustic feature vectors, and extracts inertial semantic feature vectors through the pet motion energy statistics. It then uses the corresponding modal dynamic weights to weight and combine the visual semantic feature vectors, acoustic semantic feature vectors, and inertial semantic feature vectors respectively to generate a multimodal behavior fusion feature vector.

[0011] As a further aspect of the present invention, the time-series feature modeling module includes: The sliding window capture submodule monitors and extracts the cloud-based time-series modeling status, constructs a sliding time window of pet historical behavior, covers the multimodal behavior fusion feature vector with the pet historical behavior sliding time window, performs data capture and filtering, and establishes a time-series feature sequence. The temporal dependency analysis submodule uses a long short-term memory network to analyze temporal feature sequences and calculate the numerical correlation between data over time spans. It then filters out target correlation values ​​that meet the set time span correlation judgment threshold, reads the step size parameter of the time node to which the target correlation value belongs, performs matrix coordinate mapping between the target correlation value and the step size parameter, outputs feature arrangement parameters, and obtains the temporal dependency distribution attributes. The lifestyle baseline establishment submodule performs a multivariate regression fitting operation on the temporal dependency distribution attributes and the temporal feature sequence to depict the trajectory curve of the change in multimodal data over time and establish a long-term lifestyle characteristic baseline for pets.

[0012] As a further aspect of the present invention, the abnormal behavior warning module includes: The spatial deviation measurement submodule analyzes the covariance relationship of the multidimensional variables of the baseline based on the long-term living habits of the pet, extracts the distribution characteristics of the historical covariance matrix, and uses the Mahalanobis distance measurement algorithm combined with the distribution characteristics of the historical covariance matrix to calculate the spatial distance difference between the multimodal behavior fusion feature vector and the long-term living habits of the pet, and summarizes and generates multidimensional spatial statistical deviation. The alarm threshold adaptive submodule obtains the medical pet behavior health guidelines and extracts the basic warning threshold based on the multidimensional spatial statistical deviation. It reads the warning parameters from the previous historical moment, uses the exponential moving average algorithm to allocate the corresponding weight ratio parameters of the basic warning threshold and the warning parameters from the previous historical moment, and performs a smooth fitting operation by combining the corresponding weight ratio parameters with the basic warning threshold and the warning parameters from the previous historical moment to generate a dynamic adaptive alarm threshold. The abnormal behavior warning determination submodule compares the multidimensional spatial statistical deviation with the dynamic adaptive alarm threshold. When the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold, a pet behavior monitoring warning instruction is established. When the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold, a pet behavior normal status indicator instruction is generated.

[0013] A pet behavior monitoring method based on edge computing and cloud-based time-series modeling includes the following steps: S1: Acquire video keyframes, audio spectrum segments, and physical vibration signals from continuous monitoring cycles during pet behavior monitoring, and perform time alignment processing to generate a multimodal raw data synchronization sequence; S2: Analyze the multimodal raw data synchronization sequence to obtain pet video data stream, pet audio data stream and pet inertial data stream, extract normalized pet skeletal key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, pet environmental audio signal-to-noise ratio and pet motion energy statistics, and construct a multimodal structured feature set; S3: Analyze the multimodal structured feature set, extract environmental noise evaluation coefficients, determine modal dynamic weights by combining visual recognition confidence scores, and use the modal dynamic weights to fuse the normalized pet skeleton key point coordinate sequence, pet acoustic feature vector, and pet motion energy statistics in the multimodal structured feature set to generate a multimodal behavior fusion feature vector. S4: Obtain the sliding time window of the pet's historical behavior, truncate the multimodal behavior fusion feature vector, construct a time-series feature sequence, analyze the time-series feature sequence to extract the time-series dependency distribution attributes, and construct a baseline of the pet's long-term living habits features; S5: Determine the multidimensional spatial statistical deviation between the multimodal behavior fusion feature vector and the pet's long-term living habit feature baseline, construct a dynamic adaptive alarm threshold, and generate a pet behavior monitoring early warning instruction when the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold; generate a pet behavior normal state indicator instruction when the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, by synchronously acquiring and time-aligning multimodal data (video keyframes, audio spectrum, physical vibration), key skeletal points, acoustic features, and motion energy are extracted to construct structured features. Visual confidence and environmental audio signal-to-noise ratio are used to determine dynamic weights to complete the fusion of behavioral features. Furthermore, a time-series feature sequence is constructed by combining a sliding time window to generate a long-term lifestyle baseline. Multidimensional spatial statistical deviation and dynamic adaptive alarm thresholds are introduced for comparison. These techniques enable efficient end-to-cloud collaborative processing of pets' daily activities and complementary fusion of multidimensional features. This overcomes interference from a single visual environment and avoids the continuous uploading of private videos to the cloud. It not only reduces privacy risks and ensures the timeliness of early warnings, but also scientifically and effectively analyzes and evaluates the long-term health development trend of pets. Attached Figure Description

[0015] Figure 1 This is a system flowchart of the present invention; Figure 2 This is a flowchart of the data acquisition and synchronization module of the present invention; Figure 3 This is a flowchart of the structured module features of the present invention; Figure 4 This is a flowchart of the behavioral feature fusion module of the present invention; Figure 5 This is a flowchart of the time-series feature modeling module of the present invention; Figure 6 This is a flowchart of the abnormal behavior warning module of the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0017] Please see Figure 1 A pet behavior monitoring system based on edge computing and cloud-based time-series modeling includes: The data acquisition and synchronization module is used to acquire video keyframes (acquired by the edge-integrated visual sensor), audio spectrum segments (acquired by the microphone array), and physical vibration signals (acquired by the three-axis inertial measurement unit) for continuous monitoring cycles during pet behavior monitoring, and to perform time-alignment processing on the timestamps to generate a multimodal raw data synchronization sequence. The feature structuring module is used to parse the multimodal raw data synchronization sequence, obtain pet video data stream, pet audio data stream and pet inertial data stream, extract normalized pet skeletal key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, pet environmental audio signal-to-noise ratio and pet motion energy statistics, and construct a multimodal structured feature set; The pet's inertial data stream includes triaxial acceleration components and angular velocity components; This study utilizes a multi-task lightweight AI backbone network to analyze pet video data streams, extracting normalized pet skeletal key point coordinate sequences and pet visual recognition confidence scores; analyzes pet audio data streams to extract pet acoustic feature vectors and pet environmental audio signal-to-noise ratio; and analyzes pet inertial data streams to extract pet motion energy statistics. The behavior feature fusion module is used to analyze the multimodal structured feature set. Based on the general acoustic environment interference evaluation standard library, it matches the pet environment audio signal-to-noise ratio to extract the environmental noise evaluation coefficient. Combined with the visual recognition confidence score, it determines the modal dynamic weight. The modal dynamic weight is used to fuse the normalized pet skeletal key point coordinate sequence, pet acoustic feature vector and pet motion energy statistics in the multimodal structured feature set to generate a multimodal behavior fusion feature vector. The multimodal dynamic weight allocation model is used to output the dynamic weights of the visual modality, acoustic modality, and inertial modality based on the pet visual recognition confidence score and the environmental noise evaluation coefficient. The temporal feature modeling module is used to obtain the sliding time window of pet's historical behavior, extract the multimodal behavior fusion feature vector based on the sliding time window, construct the temporal feature sequence, analyze the temporal feature sequence to extract the temporal dependency distribution attributes, and construct the long-term living habit feature baseline of pets; The abnormal behavior warning module is used to determine the multidimensional spatial statistical deviation between the multimodal behavior fusion feature vector and the pet's long-term living habit feature baseline, and to construct a dynamic adaptive alarm threshold. When the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold, a pet behavior monitoring warning instruction is generated; when the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold, a pet behavior normal status indicator instruction is generated. Basic warning thresholds are extracted from medical pet behavior and health guidelines, and historical warning parameters from the previous moment are extracted simultaneously. The exponential moving average algorithm is used to smoothly fit the basic warning thresholds and historical warning parameters from the previous moment, and output a dynamic adaptive alarm threshold.

[0018] Please see Figure 2 The data acquisition and synchronization module includes: The signal data acquisition submodule acquires visual keyframes collected by the visual sensor, acquires audio spectrum segments corresponding to the microphone array, acquires physical vibration signals corresponding to the three-axis inertial measurement unit, parses the visual keyframes to extract the first acquisition timestamp, parses the audio spectrum segments to extract the second acquisition timestamp, parses the physical vibration signals to extract the third acquisition timestamp, and combines the first acquisition timestamp, the second acquisition timestamp, and the third acquisition timestamp to establish an initial timetamp set. The signal data acquisition submodule utilizes a binocular wide-angle high-definition visual sensor mounted on top of the pet monitoring device to acquire key visual frame sequences at a fixed frame rate of 30 frames per second (where 30 represents the visual image sampling frequency). Leveraging the proximity advantage of edge computing frameworks, this submodule also controls a 4-channel omnidirectional microelectromechanical system (MEMS) microphone array deployed around the device. Hertz (i.e., sampling frequency) ,in This refers to the audio sampling rate (used to capture minute changes in sound), at which the microphone array acquires audio spectrum data corresponding to the audio spectrum segment. This submodule further uses a high-precision three-axis inertial measurement unit integrated inside the pet's collar to set the sampling frequency to [value missing]. Hertz (i.e.) ,in (Representing the sampling rate parameter of the inertial sensor), continuously acquiring data corresponding to the triaxial inertial measurement unit. Axis (where, (representing the horizontal spatial coordinate axis) and shaft (where (representing the horizontal and vertical spatial coordinate axes) and shaft (where Triaxial acceleration components (representing the directions perpendicular to the spatial coordinate axes) (in (representing linear acceleration measures along the corresponding coordinate axes) and angular velocity components (in The physical vibration signal (representing the three-dimensional rotation rate of the pet's limb movements) is used to aggregate multi-source heterogeneous initial dynamic features at the edge. After acquiring the multimodal raw signal, this submodule reads the header information of the video data streaming media transmission protocol packet and uses a system-level high-precision clock to analyze visual keyframes and extract the first acquisition timestamp accurate to the millisecond level. (in This represents the absolute network time value at the moment the visual stream was acquired. For audio data, this submodule extracts the second acquisition timestamp by reading the header register of the waveform audio file format data block and parsing the audio spectrum segment. (in This represents the absolute network time value of the audio signal sampling. For physical vibration signals, this submodule reads the data packets transmitted via the serial peripheral interface bus, parses the physical vibration signal, and extracts the third acquisition timestamp. (in This represents the absolute network time value indicating the trigger time of inertial data. Finally, this submodule establishes an array structure in memory containing visual, audio, and inertial timestamp fields, and aggregates the initial state set of the first, second, and third acquisition timestamps. , It represents a mathematical set of time records consisting of three independent sensing moments.

[0019] For example, when the first collection timestamp is 1632123456789, that is, the value is assigned. Milliseconds; the second collection timestamp is 1632123456795, i.e., the assigned value. Milliseconds; the third collection timestamp is 1632123456780, i.e., the assigned value. At millisecond intervals, this submodule stores each instance as a record in the set, providing a solid foundation for time-based data alignment for subsequent high-precision pet behavior monitoring.

[0020] The time microsecond alignment submodule, for the initial set of timestamps, uses a hardware-level timestamp synchronization protocol to parse the first, second and third acquisition timestamps and perform time comparison calculations. It then performs time dimension alignment and adjustment on the visual keyframes, audio spectrum segments and physical vibration signals to generate an alignment signal correlation matrix. The time microsecond alignment submodule, for each record in the initial timestamp set, aims to reduce multimodal asynchronous errors before transmitting the raw sensor signal to the cloud in an edge computing architecture. It calls a high-precision network time protocol layer interface and uses a hardware-level timestamp synchronization protocol to parse the internal clock deviation values ​​of the first, second, and third acquisition timestamps and perform time comparison calculations. This submodule sets the time comparison calculation logic, using the first acquisition timestamp as the reference time point, and calculates the absolute time difference between the second acquisition timestamp and the reference time point. ,in, This represents the absolute lag or lead of audio time relative to the visual reference time. This indicates an absolute value operation. For example, when the first collection timestamp is 1632123456789 and the second collection timestamp is 1632123456795, performing a subtraction operation yields the difference as: Milliseconds. When the third acquisition timestamp is 1632123456780, this submodule calculates the absolute time difference. ,in, This represents the degree of deviation between the inertial data trigger time and the visual reference time; the difference is calculated as follows: Milliseconds. Because... milliseconds and Even millisecond deviations exist. This submodule aligns and adjusts the actions by integrating visual keyframes, audio spectrum clips, and physical vibration signals along the time dimension, thereby supporting the synchronous recognition of micro-motion sequences by a highly sensitive pet behavior monitoring model. This submodule utilizes a spline interpolation algorithm. (in The input is a smooth piecewise polynomial function used to fit and reconstruct discrete signal data. (Using time index parameters) This submodule performs data resampling on audio and inertial data, inserting supplementary sampling points between two adjacent sampling points of the physical vibration signal, and shifting the effective time point of the data to a moment exactly consistent with the first acquisition timestamp. After the shift is complete, this submodule generates a dataset consisting of... Matrix structure consisting of rows and columns ,in, Representative dimension is Aligned data buffer array, The total number of synchronized sampling points in the time series is represented by the first row, which stores aligned visual keyframe data, the second row stores aligned audio data, and the third row stores aligned inertial data. This is used to generate an aligned signal correlation matrix to ensure that the multimodal state data captured by the edge device is in a consistent spatiotemporal observation dimension.

[0021] The synchronization sequence generation submodule extracts aligned visual keyframes to construct a pet video data stream based on the alignment signal correlation matrix, extracts aligned audio spectrum segments to construct a pet audio data stream, and extracts aligned physical vibration signals to construct a pet inertial data stream containing triaxial acceleration and angular velocity components. These components are then combined to generate a multimodal raw data synchronization sequence. The synchronization sequence generation submodule reads the first row of memory addresses from the alignment signal correlation matrix. Based on the alignment signal correlation matrix, it extracts aligned visual keyframes and encapsulates consecutive image frames according to the H.265 encoding standard to construct a pet video data stream. Next, this submodule reads the second row of the correlation matrix, extracts aligned audio spectrum segments, and encapsulates the frequency domain feature array into a fixed frame length, such as a specified frame length parameter. A byte-based audio stream, in which, This indicates the byte capacity limit of a single acoustic packet, constructing a pet audio data stream. Subsequently, this submodule reads the third line, extracts and aligns the physical vibration signals to construct a pet inertial data stream containing triaxial acceleration and angular velocity components. This submodule allocates a data buffer. , This is a physical memory space reserved by the edge system kernel specifically for temporarily caching high-speed concurrent stream segments. Pet video data stream segments, pet audio data stream segments, and pet inertial data stream segments are alternately pushed into the buffer in chronological order. This submodule adds a 32-bit synchronization identifier and a data payload length identifier field to the header of each synchronized segment. ,in This indicates the number of payload bytes actually carried within a specific slice. By cyclically packaging, this submodule generates a synchronized sequence of multimodal raw data, reducing the structural entropy increase during the transmission of multi-source data to the cloud-based time-series modeling engine. It is a crucial edge-side preprocessing link for achieving real-time pet behavior monitoring.

[0022] Please see Figure 3 The feature structuring module includes: The video and audio parsing submodule, based on the multimodal raw data synchronization sequence, separates the pet video data stream and the pet audio data stream. It uses a multi-task lightweight artificial intelligence backbone network to parse the pet video data stream, extract the normalized pet skeleton key point coordinate sequence and pet visual recognition confidence score, parse the pet audio data stream, extract the pet acoustic feature vector and pet environmental audio signal-to-noise ratio, and integrate and establish audiovisual feature parsing parameters. The video and audio parsing submodule identifies the synchronization code in the header of the multimodal raw data synchronization sequence and separates the pet video data stream from the pet audio data stream based on the multimodal raw data synchronization sequence. This submodule invokes a multi-task lightweight AI backbone network, fully utilizing the low-power neural computing power built into the edge computing hardware to achieve local ultra-fast inference. This network includes one input layer and four hidden layers containing depthwise separable convolutional kernels, with modified linear unit activation functions configured within the layers. ,in This represents the input feature map mapping value in the hidden layer of the network. This function acts to cut off negative activation responses and introduce nonlinearity. The output layer is divided into two paths: the first branch consists of fully connected layers, and the second branch uses a global average pooling layer to connect to the classifier. This submodule uses a multi-task lightweight AI backbone network to parse pet video data streams, extracting data through the first branch. A normalized sequence of key coordinates of pet skeleton ,in Representing the The horizontal and vertical two-dimensional absolute coordinate parameters of each pet joint in the normalized image plane. The output from the second branch is between... arrive Extract confidence scores for pet visual recognition from the parameters between them. ,in This represents the expected probability assessment value of the model for identifying the monitored target within the current frame as the target pet. This submodule converts it into a Mel-frequency cepstral coefficient feature map, inputs it into a convolutional network structure to parse the pet audio data stream, and extracts a length of... Pet acoustic feature vector ,in It is a set of 128 real-point floating-point numbers used to characterize the timbre attributes of sounds such as barking dogs and meowing cats. This submodule calculates the effective signal power. ,in This represents the energy density and background noise power carried by the analyzed target sound frequency band. ,in This represents the ratio of the average energy of the ambient noise floor within the extracted segment, substituted into the formula. ,in This parameter represents the signal-to-noise logarithmic ratio, a measure of sound purity, and is used to extract the signal-to-noise ratio of pet environment audio. If the effective power is... microwatts, i.e., parameter assignment The noise power is microwatts, i.e., parameter assignment The calculated ratio is Decibels, which are the result values ​​obtained by mapping the built-in logarithmic scaling constant. dB, where This module outputs the decibel measurement value. Ultimately, it integrates and establishes audiovisual feature parsing parameters, providing intuitive abstract semantic support for edge-side pet behavior monitoring.

[0023] Table 1. Audiovisual Feature Analysis Parameter Configuration Table As shown in Table 1, this table details the data attributes and quantization of the audiovisual feature analysis parameters.

[0024] The inertial motion statistics submodule analyzes parameters based on audiovisual features, separates the multimodal raw data synchronization sequence to obtain the pet inertial data stream, uses a multi-task lightweight artificial intelligence backbone network to analyze the pet inertial data stream, extracts the corresponding pet motion energy change values ​​from the pet inertial data stream and performs cumulative calculations to generate pet motion energy statistics. The inertial motion statistics submodule extracts the pet inertial data stream by separating the multimodal raw data synchronization sequence based on the time index information in the audiovisual feature parsing parameters. This submodule utilizes a multi-task lightweight artificial intelligence backbone network to parse the pet inertial data stream. The time series parsing branch of this network employs a long short-term memory network layer containing 64 neurons, combined with a tangent hyperbolic activation function. ,in This is the cumulative input stimulus to the neuron. The natural constant is used to smooth and compress oscillating numerical values ​​to a saturation range of -1 to 1. The output is connected to a regressor to calculate the square of the magnitude of the acceleration vector in three-dimensional space. ,in An indicator representing the instantaneous kinetic energy response of a pet at a single sampling moment. These correspond to the instantaneous acceleration sampling values ​​of the pet-wearing device in three mutually orthogonal directions in space. For example, when The shaft acceleration is 2, that is, substituting the parameters... , The shaft acceleration is 3, that is, substitute the parameters. , The shaft acceleration is 1, i.e., substitute the parameters. At that time, the submodule calculated the sum of squares to be 14, i.e., it executed... This submodule extracts the pet's motion energy change values ​​for a single sampling period from the pet inertial data stream and performs cumulative calculations using the formula... ,in The total motion integral energy within an evaluation period, The number of batches of sampling points included. For the first The instantaneous squared value of the step size. For 100 sampling points, i.e., the periodicity parameter. Assuming the energy at each point is 14, i.e., a constant setting. This submodule multiplies 14 by 100 to generate a pet exercise energy statistic with a total value of 1400, which is the result value. This allows for the quantification of the pet's active jumping behavior at the edge computing end, reducing the bandwidth overhead of transmitting the original high-frequency inertial waveform to the subsequent cloud-based time-series modeling center.

[0025] The structured feature construction submodule extracts the normalized pet skeleton key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, and pet environmental audio signal-to-noise ratio from the audiovisual feature parsing parameters, and summarizes them with the pet motion energy statistics to generate a multimodal structured feature set. The structured feature construction submodule analyzes the audiovisual features, sequentially extracting the normalized pet skeletal keypoint coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, and pet environmental audio signal-to-noise ratio. This submodule establishes a data fusion stack in memory, pushing the four audiovisual features onto the stack sequentially, and then summing them with the pet's motion energy statistics obtained through the data bus. This submodule calls a standardization algorithm. ,in This represents the standardized value after dimensionless conversion. This is the current raw feature scalar being read. and These are the minimum and maximum observed boundary values ​​of this type of feature in the historical sample set, respectively. Numerical domain mapping is performed on the non-interval variables in the stack, uniformly limiting the values ​​of each dimension to the interval between 0 and 1, i.e., limiting the value range to... Within this framework, a multimodal structured feature set is generated, thereby transforming discrete, multi-source heterogeneous data from the edge into a standardized data volume, and constructing a consistent pet behavior monitoring vector feature space.

[0026] Please see Figure 4 The behavioral feature fusion module includes: The environmental noise assessment submodule, based on a multimodal structured feature set, extracts the pet environment audio signal-to-noise ratio and pet visual recognition confidence score, compares the pet environment audio signal-to-noise ratio with a general acoustic environment interference evaluation standard library and performs matching mapping, and extracts environmental noise assessment coefficients based on the matching mapping results. The environmental noise assessment submodule locates specific byte segments based on a multimodal structured feature set and extracts the signal-to-noise ratio (SNR) of the pet environment audio and the confidence score for pet visual recognition. This submodule reads a pre-stored general acoustic environmental interference evaluation standard library, compares the pet environment audio SNR with the library, and performs a matching mapping. Assuming the extracted SNR is 50 dB, this submodule determines it falls within the 40 to 60 dB range, meaning the SNR value satisfies the comparison formula. The basic interference factor of 0.4 corresponding to this interval is extracted, that is, the mapping parameter is obtained by looking up a table. ,in This is specifically designed to quantify the degree of feature disruption penalty caused by the current auditory background noise. The extracted confidence score for pet visual recognition is 0.89, which is the value returned by the model. Then, based on the matching mapping action results, this submodule subtracts the visual confidence bias value of 0.1 from the basic interference factor of 0.4. This is the bias adjustment constant preset within the system. This constant is generated based on the visual high-definition compensation effect and is used to offset some of the adverse weighting effects of auditory background noise under good lighting conditions. The calculated difference is 0.3, which is obtained by substituting into the derived formula. Extract this 0.3 as the environmental noise assessment coefficient. ,in It is the core evaluation parameter that guides the subsequent edge sensing nodes to converge the participation rate of data from each modality. This process dynamically enhances the anti-interference and adaptive capabilities of the pet behavior monitoring system in noisy home edge environments.

[0027] The dynamic weight allocation submodule uses a multimodal dynamic weight allocation model to analyze the confidence score of pet visual recognition and the environmental noise assessment coefficient, determine the distribution law of multimodal feature proportion, output the dynamic weight of visual modality, dynamic weight of acoustic modality and dynamic weight of inertial modality, and establish the dynamic weight coefficient of modality. The dynamic weight allocation submodule utilizes a multimodal dynamic weight allocation model. This model includes an input layer with two neurons, a receiver confidence score of 0.89, and an environmental noise evaluation coefficient of 0.3. It is executed quickly using the lightweight computing power of the edge microcontroller. The hidden layer of 16 neurons employs modified linear unit activation, and the output layer of 3 neurons is connected to a Softmax function. ,in Represents the normalized i-th Modal allocation ratio, The initial response mapping values ​​of the neuron before scaling. Using the base of the natural logarithm, this function ensures that the sum of the three weights is exactly equal to 1. This model analyzes the confidence score for pet visual recognition and the environmental noise evaluation coefficient to determine the distribution pattern of multimodal features. The system outputs a value from the output layer. The visual modality dynamic weights, i.e., the allocation factors. ,in The allocation factor determines the authority of image modal features in the global state evaluation and has a dynamic weight of acoustic modality of 0.1. ,in The interference tolerance effect of controlling vocal and other sound modal data and the inertial modal dynamic weight with a value of 0.3, i.e., the allocation factor. ,in This submodule combines three sets of values ​​to establish a modal dynamic weighting coefficient array, reflecting the weighting of the movement intensity of the pet's neck-mounted device. This gives the monitoring system the ability to adaptively focus on high signal-to-noise ratio sensor data in environments where light or noise changes constantly, ensuring the purity of the pet behavior monitoring feature pool.

[0028] The semantic feature fusion submodule, based on modal dynamic weight coefficients, extracts visual semantic feature vectors through normalized pet skeleton key point coordinate sequences, acoustic semantic feature vectors through pet acoustic feature vectors, and inertial semantic feature vectors through pet motion energy statistics. It then uses corresponding modal dynamic weights to weight and combine the visual semantic feature vectors, acoustic semantic feature vectors, and inertial semantic feature vectors respectively to generate multimodal behavior fusion feature vectors. The semantic feature fusion submodule, based on the weight allocation strategy in modal dynamic weight coefficients, extracts visual semantic feature vectors from the normalized pet skeleton keypoint coordinate sequence through a mapping layer. ,in This represents a high-order image feature matrix that, after undergoing deep structural encoding, can express specific movements and postures. Acoustic semantic feature vectors are extracted from the pet's acoustic feature vectors using pooling layers. ,in This is the semantic information vector of the core laryngeal vocalization vibration extracted after dimensionality reduction. An inertial semantic feature vector is extracted through a broadcast mechanism that compares pet motion energy statistics with baseline values. ,in This represents the hidden features of the action state derived from the discrete kinetic energy integral result. This submodule uses corresponding modal dynamic weights to weight and combine visual semantic feature vectors, acoustic semantic feature vectors, and inertial semantic feature vectors respectively, completing the final efficient data convergence and aggregation within the edge computing node. Specifically, it multiplies the visual vector by 0.6, the acoustic vector by 0.1, and the inertial vector by 0.3, performs a bitwise addition operation on the weighted three-way matrix data, and uses the synthesis formula... ,in This represents the global fusion abstract representation vector produced after the complementary cancellation of errors by various heterogeneous modal information. It generates a multimodal behavior fusion feature vector, which serves as a high-quality, minimalist input carrier for subsequent transmission to high-computing centers and cloud-based time series modeling and analysis.

[0029] Table 2. Parameters for Multimodal Behavioral Feature Vector Calculation As shown in Table 2, this table records in detail the values ​​and weighting results of each item in the multimodal behavior feature vector fusion process.

[0030] Please see Figure 5 The time-series feature modeling module includes: The sliding window capture submodule monitors and extracts the cloud-based time-series modeling status, constructs a sliding time window of pet historical behavior, covers the multimodal behavior fusion feature vector with the pet historical behavior sliding time window, performs data capture and filtering, and establishes a time-series feature sequence. The sliding window capture submodule sends query commands to the server, gradually shifting the focus of pet behavior monitoring data processing and logical deduction from the edge perception side to the cloud decision-making center, monitoring the cloud time series modeling status and extracting cloud time series modeling status codes. ,in This is a hexadecimal protocol value command that indicates the current availability and occupancy of the resource queues from the remote server cluster. When the status code shows "idle," such as during parameter verification... At any given time, this submodule constructs a sliding time window of the pet's historical behavior in memory, with a step size of 500 milliseconds, which is configured as the step-sliding time difference constant. milliseconds, of which This is used to control the update overlap rate of the time window region between two adjacent extraction analyses, with a total length of 60,000 milliseconds, which defines the observation scale boundary constant. milliseconds, of which This represents the total time span of the complete long sequence that must be covered for a single intent inference. This submodule extends a sliding time window of the pet's historical behavior to the multimodal behavior fusion feature vector and performs data truncation and filtering. If 120 sets of valid data are extracted within the window, the length of the extracted vector sequence frame set is counted. If the set 90% data retention threshold is met, then the integrity control constant is defined. Requirements, i.e., compliance with the judgment criteria This submodule preserves these complete and compliant data sequences and establishes a time-series feature sequence consisting of a matrix of 120 feature columns, thereby ensuring that the data stream input to the cloud-based time-series modeling engine has a coherent and uninterrupted time context and time span integrity.

[0031] The temporal dependency analysis submodule uses a long short-term memory network to analyze temporal feature sequences and calculate the numerical correlation between data over time spans. It then filters out target correlation values ​​that meet the set time span correlation judgment threshold, reads the step size parameter of the time node to which the target correlation value belongs, performs matrix coordinate mapping between the target correlation value and the step size parameter, outputs feature arrangement parameters, and obtains the temporal dependency distribution attributes. The temporal dependency analysis submodule invokes a three-layer Long Short-Term Memory (LSTM) network, each containing 128 neurons. Utilizing the network's built-in tangent hyperbolic activation function and state gate control mechanism, and supported by the abundant computing power and distributed storage of the cloud cluster, it performs large-scale deep spatiotemporal computations on the daily patterns of pets. This submodule uses the LSM network to analyze temporal feature sequences and calculate the degree of correlation between data over time spans. Assuming 120 nodes are computed, representing a time-bound cross-section... The extracted fusion features and the first node represent the time-bound cross section. Pearson correlation coefficient between extracted fusion features ,in It is a parameter used to accurately measure the statistical co-occurrence tightness of two eigenvectors with extremely large time steps, where they are increasing or decreasing linearly. These represent the sample variables expanded along the feature vector dimensions of the corresponding nodes. The arithmetic mean of the historical feature distribution within the current sliding window is used as the basis for filtering. This submodule filters features that meet the set time span correlation threshold of 0.75, which is a comparison with the filter constant. This is used to forcibly truncate the correlation coefficient values ​​corresponding to weak correlations and random behavioral perturbations. Since 0.85 is greater than 0.75, it is used as the target correlation degree value. The time step parameter 119, corresponding to the target correlation degree value, is read to calculate the physical step distance. In frame units, the target correlation strength value of 0.85 is mapped to the time step parameter of 119 using matrix coordinate mapping. This is done in the distribution matrix, which is used to construct a topological adjacency matrix to store the connection strength of each local behavior. Write 0.85 to the position of column 119 in row 1, which is to perform an element-level addressing write operation. The system outputs feature arrangement parameters and obtains the temporal dependency distribution attributes of the distribution structure matrix, successfully achieving deep-level capture of the long-term hidden eating, sleeping, or anxiety behavior logic cycle of pets in the cloud-based temporal modeling center.

[0032] The submodule for establishing a baseline of living habits performs multivariate regression fitting operations on the distribution attributes of time-series dependencies and time-series feature sequences to depict the trajectory curve of the evolution of multimodal data over time and establish a baseline of long-term living habits characteristics for pets. The lifestyle habit baseline establishment submodule utilizes a least squares algorithm engine to transform the massive and complex cloud-based time-series modeling and analysis into clear and concrete normative health standards, providing highly valuable control anchors for long-term, cross-cycle pet behavior monitoring. Multiple regression fitting is performed on the time-series dependency distribution attributes and time-series feature sequences, iteratively reducing the model mapping error equation. ,in This represents the cumulative total squared residual loss between the baseline extrapolated predicted values ​​and the measured multidimensional fused values. The total number of valid sampling days included in the baseline calculation. These are the actual extracted normalized observation vector values. This represents the predicted trajectory under ideal conditions. This submodule uses the step size parameter as the independent variable. Using multimodal feature values ​​as dependent variables Solve the linear regression model equation ,in The slope rate of change is a core parameter representing the tendency of pet behavior to evolve. The intercept representing the intrinsic basic circadian rhythm intercept constant after eliminating the effects of time progression is 0.3, which is the value of the generated model parameters. The slope is 0.05, which represents the generated model parameters. This submodule uses these parameters to plot the trajectory curve of the multimodal data over time. It also collects the predicted values ​​and corresponding standard deviation ranges for each discrete point on the curve to establish a baseline for the long-term lifestyle characteristics of pets.

[0033] Please see Figure 6 The abnormal behavior warning module includes: The spatial deviation measurement submodule analyzes the covariance relationship of multidimensional variables of the baseline based on the long-term living habits of pets, extracts the distribution characteristics of the historical covariance matrix, and uses the Mahalanobis distance measurement algorithm combined with the distribution characteristics of the historical covariance matrix to calculate the spatial distance difference between the multimodal behavior fusion feature vector and the long-term living habits of pets, and summarizes and generates multidimensional spatial statistical deviation. The spatial deviation measurement submodule reads the pet data record library, analyzes the covariance relationship of the baseline multidimensional variables based on the long-term living habits of the pets, and extracts the distribution characteristics of the historical covariance matrix. ,in It is a set of symmetric positive definite matrices used to comprehensively characterize the historical joint fluctuation amplitude among multimodal feature variables of sound, vision, and motion. This submodule acquires the latest fusion features, fully coordinates the instantaneous burst feature vectors uploaded in real time by edge computing nodes with the full lifecycle historical baseline slowly accumulated by cloud time-series modeling, conducts cross-comparison verification in high-dimensional mathematical space, and uses the Mahalanobis distance metric algorithm combined with the distribution characteristics of historical covariance matrix to calculate the multimodal behavior fusion feature vector. ,in This consists of the observation index vector, which is currently parsed in real time from the monitoring stream and contains multimodal dynamic attributes, and the baseline mean vector of the long-term living habits characteristics of pets. ,in This represents the expected value of the spatial distance difference between the centroids of the pet under long-term, stable, and healthy living conditions. The Mahalanobis distance algorithm calculates the product of the difference vector and the inverse covariance matrix, and the computational model is based on... ,in This represents the spatial distance of true anomalies after successfully eliminating differences in data dimensions and strong correlations in internal attributes. For the linear algebraic transpose of the difference vector between the real-time and expected values, To characterize the inverse form of the historical covariance matrix of a discrete distribution, this submodule assumes a difference parameter containing 0.2 and 0.3, which represents the product of squared vector components along the directions of the independent principal components after expansion. This submodule then sums the results to generate a multidimensional spatial statistical deviation of 4.5. ,in It is a core scalar distance indicator that comprehensively measures the severity of a pet's current behavior that is different from usual, thereby keenly and accurately identifying and exposing abnormal pet behavior patterns such as hidden dangers or illnesses from a large amount of daily activities.

[0034] The adaptive alarm threshold submodule calculates the deviation in a multidimensional space, obtains the medical pet behavior and health guidelines, extracts the basic warning threshold, reads the warning parameters from the previous moment in history, uses the exponential moving average algorithm to assign the corresponding weight ratio parameters of the basic warning threshold and the warning parameters from the previous moment in history, and performs a smooth fitting operation by combining the corresponding weight ratio parameters with the basic warning threshold and the warning parameters from the previous moment in history to generate a dynamic adaptive alarm threshold. The alarm threshold adaptive submodule, based on the obtained multidimensional spatial statistical deviation, accesses an external interface to obtain medical pet behavior and health guidelines and extracts a basic warning threshold with an upper limit of 3.0. ,in This is a universally applicable health red line constant warning judgment line set by international authoritative veterinary academic institutions for different breeds of dogs and cats. This submodule retrieves the warning parameter 3.5 from local storage, which is the dynamic threshold temporary storage scalar value within the previous calculation update cycle. This submodule uses the exponential moving average algorithm to set the smoothing coefficient. ,in This is a damping parameter used to suppress sudden artifacts caused by a single intense movement, controlling the sensitivity step size during the transition between old and new benchmarks. Its value is strictly constrained between 0 and 1, and its weight in the basic warning threshold is 0.2, meaning the assigned value is equivalent to... The weighting parameter corresponding to the warning parameter at the previous historical moment is 0.8, which is the derived allocation value. This submodule combines the corresponding weight percentage parameters with the basic early warning threshold and the early warning parameters from the previous historical moment to perform a smooth fitting operation, that is, substituting them into the iterative update equation. ,in It is the latest periodic personality judgment threshold boundary intelligently generated after deeply integrating the long-term unique habits and general defense standards of this specific pet. Multiplying 3.5 by 0.8 gives 2.8, multiplying 3.0 by 0.2 gives 0.6, and adding 2.8 by 0.6 gives 3.4. Substituting these values ​​into the calculation yields... It generates dynamic adaptive alarm thresholds, and the self-balancing adjustment mechanism effectively eliminates the frequent false alarms caused by hard-coded fixed thresholds in traditional pet behavior monitoring solutions, thereby improving the individual adaptability of danger alarms.

[0035] The abnormal behavior warning and judgment submodule compares the multidimensional spatial statistical deviation with the dynamic adaptive alarm threshold. When the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold, a pet behavior monitoring warning instruction is established. When the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold, a pet behavior normal status indicator instruction is generated. The behavior anomaly warning and judgment submodule reads the deviation result of 4.5 from the above calculation and the threshold result of 3.4. This submodule calls an arithmetic comparator in the microprocessor to compare the multidimensional spatial statistical deviation. With dynamic adaptive alarm threshold This successfully closed the final evaluation and decision-making chain of this dual-track architecture for pet behavior monitoring based on edge computing and cloud-based time-series modeling. At that time, when the multidimensional spatial statistical deviation of 4.5 was greater than the dynamic adaptive alarm threshold of 3.4, the real-time logical state satisfied the alarm triggering condition equation. This submodule packages the current device code and establishes a pet behavior monitoring and early warning command to send an alarm. When the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold, for example, a deviation of 2.1 not greater than 3.4, the current state meets the health and safety survival conditions equation. This submodule generates a standard hexadecimal identifier corresponding to the pet's normal behavior status indication instruction and writes it to the system log, thereby seamlessly connecting to the owner's mobile terminal and providing the owner with intuitive feedback on scientific pet care around the clock, with high precision and low latency.

[0036] Table 3. Record of Abnormal Early Warning Judgment Status As shown in Table 3, this table records in detail the comparison process and judgment results between the multidimensional spatial statistical deviation and the dynamic adaptive alarm threshold.

[0037] A pet behavior monitoring method based on edge computing and cloud-based time-series modeling includes the following steps: S1: Acquire video keyframes, audio spectrum segments, and physical vibration signals from continuous monitoring cycles during pet behavior monitoring, and perform time alignment processing to generate a multimodal raw data synchronization sequence; S2: Analyze the multimodal raw data synchronization sequence to obtain pet video data stream, pet audio data stream and pet inertial data stream, extract normalized pet skeletal key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, pet environmental audio signal-to-noise ratio and pet motion energy statistics, and construct a multimodal structured feature set; S3: Analyze the multimodal structured feature set, extract the environmental noise evaluation coefficient, combine the visual recognition confidence score to determine the modal dynamic weight, and use the modal dynamic weight to fuse the normalized pet skeleton key point coordinate sequence, pet acoustic feature vector and pet motion energy statistics in the multimodal structured feature set to generate a multimodal behavior fusion feature vector. S4: Obtain the sliding time window of the pet's historical behavior, extract the multimodal behavior fusion feature vector, construct a time-series feature sequence, analyze the time-series feature sequence to extract the time-series dependency distribution attributes, and construct a baseline of the pet's long-term living habits features; S5: Determine the multidimensional spatial statistical deviation between the multimodal behavior fusion feature vector and the pet's long-term living habit feature baseline, construct a dynamic adaptive alarm threshold, and generate a pet behavior monitoring early warning instruction when the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold; generate a pet behavior normal status indicator instruction when the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold.

[0038] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A pet behavior monitoring system based on edge computing and cloud-based time-series modeling, characterized in that, include: The data acquisition and synchronization module is used to acquire video keyframes, audio spectrum segments, and physical vibration signals during the continuous monitoring period of pet behavior monitoring, and to perform time-alignment processing on the timestamps to generate a multimodal raw data synchronization sequence. The feature structuring module is used to parse the multimodal raw data synchronization sequence, obtain pet video data stream, pet audio data stream and pet inertial data stream, extract normalized pet skeletal key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, pet environmental audio signal-to-noise ratio and pet motion energy statistics, and construct a multimodal structured feature set. The behavior feature fusion module is used to analyze the multimodal structured feature set, extract environmental noise evaluation coefficients, determine modal dynamic weights by combining visual recognition confidence scores, and use the modal dynamic weights to fuse the normalized pet skeleton key point coordinate sequence, pet acoustic feature vector, and pet motion energy statistics in the multimodal structured feature set to generate a multimodal behavior fusion feature vector. The temporal feature modeling module is used to obtain a sliding time window of the pet's historical behavior, extract the multimodal behavior fusion feature vector, construct a temporal feature sequence, analyze the temporal feature sequence to extract the temporal dependency distribution attributes, and construct a baseline of the pet's long-term living habits features. The abnormal behavior warning module is used to determine the multidimensional spatial statistical deviation between the multimodal behavior fusion feature vector and the pet's long-term living habit feature baseline, construct a dynamic adaptive alarm threshold, and generate a pet behavior monitoring warning instruction when the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold; and generate a pet behavior normal status indicator instruction when the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold.

2. The pet behavior monitoring system based on edge computing and cloud-based time-series modeling according to claim 1, characterized in that: The pet inertial data stream includes three-axis acceleration components and angular velocity components; The pet video data stream is analyzed using a multi-task lightweight AI backbone network to extract normalized pet skeletal key point coordinate sequences and pet visual recognition confidence scores; the pet audio data stream is analyzed to extract pet acoustic feature vectors and pet environmental audio signal-to-noise ratio; and the pet inertial data stream is analyzed to extract pet motion energy statistics.

3. The pet behavior monitoring system based on edge computing and cloud-based time-series modeling according to claim 1, characterized in that: When constructing the dynamic adaptive alarm threshold, a basic warning threshold is extracted from the medical pet behavior and health guidelines, and the warning parameters from the previous historical moment are extracted simultaneously. The basic warning threshold and the warning parameters from the previous historical moment are smoothly fitted using the exponential moving average algorithm, and the dynamic adaptive alarm threshold is output.

4. The pet behavior monitoring system based on edge computing and cloud-based temporal modeling according to claim 1, characterized in that, The data acquisition and synchronization module includes: The signal data acquisition submodule acquires visual keyframes collected by the visual sensor, acquires audio spectrum segments corresponding to the microphone array, acquires physical vibration signals corresponding to the three-axis inertial measurement unit, parses the visual keyframes to extract the first acquisition timestamp, parses the audio spectrum segments to extract the second acquisition timestamp, parses the physical vibration signals to extract the third acquisition timestamp, and combines the first acquisition timestamp, the second acquisition timestamp, and the third acquisition timestamp to establish an initial timetamp set. The time microsecond alignment submodule, for the initial set of timestamps, uses a hardware-level timestamp synchronization protocol to parse the first acquisition timestamp, the second acquisition timestamp, and the third acquisition timestamp and performs time comparison calculations, and performs time dimension alignment adjustment on the visual keyframes, audio spectrum segments, and physical vibration signals to generate an alignment signal correlation matrix. The synchronization sequence generation submodule, based on the alignment signal correlation matrix, extracts aligned visual keyframes to construct a pet video data stream, extracts aligned audio spectrum segments to construct a pet audio data stream, extracts aligned physical vibration signals to construct a pet inertial data stream containing triaxial acceleration components and angular velocity components, and combines them to generate a multimodal raw data synchronization sequence.

5. The pet behavior monitoring system based on edge computing and cloud-based temporal modeling according to claim 1, characterized in that, The feature structuring module includes: The video and audio parsing submodule, based on the multimodal raw data synchronization sequence, separates the pet video data stream and the pet audio data stream, uses a multi-task lightweight artificial intelligence backbone network to parse the pet video data stream, extracts the normalized pet skeleton key point coordinate sequence and pet visual recognition confidence score, parses the pet audio data stream, extracts the pet acoustic feature vector and pet environmental audio signal-to-noise ratio, and integrates and establishes audiovisual feature parsing parameters. The inertial motion statistics submodule separates the multimodal raw data synchronization sequence to obtain the pet inertial data stream based on the audiovisual feature parsing parameters, uses a multi-task lightweight artificial intelligence backbone network to parse the pet inertial data stream, extracts the pet motion energy change value corresponding to the pet inertial data stream and performs cumulative calculation to generate pet motion energy statistics. The structured feature construction submodule extracts the normalized pet skeletal key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, and pet environmental audio signal-to-noise ratio from the audiovisual feature parsing parameters, and summarizes them with the pet motion energy statistics to generate a multimodal structured feature set.

6. The pet behavior monitoring system based on edge computing and cloud-based temporal modeling according to claim 1, characterized in that, The behavioral feature fusion module includes: The environmental noise assessment submodule, based on the multimodal structured feature set, extracts the pet environment audio signal-to-noise ratio and pet visual recognition confidence score, compares the pet environment audio signal-to-noise ratio with the general acoustic environment interference evaluation standard library and performs matching mapping, and extracts environmental noise assessment coefficients based on the matching mapping action results; The dynamic weight allocation submodule uses a multimodal dynamic weight allocation model to analyze the pet visual recognition confidence score and the environmental noise evaluation coefficient, determine the distribution law of multimodal feature proportion, output the visual modal dynamic weight, acoustic modal dynamic weight and inertial modal dynamic weight, and establish the modal dynamic weight coefficient. The semantic feature fusion submodule, based on the modal dynamic weight coefficients, extracts visual semantic feature vectors through the normalized pet skeleton key point coordinate sequence, extracts acoustic semantic feature vectors through the pet acoustic feature vectors, and extracts inertial semantic feature vectors through the pet motion energy statistics. It then uses the corresponding modal dynamic weights to weight and combine the visual semantic feature vectors, acoustic semantic feature vectors, and inertial semantic feature vectors respectively to generate a multimodal behavior fusion feature vector.

7. The pet behavior monitoring system based on edge computing and cloud-based temporal modeling according to claim 1, characterized in that, The time-series feature modeling module includes: The sliding window capture submodule monitors and extracts the cloud-based time-series modeling status, constructs a sliding time window of pet historical behavior, covers the multimodal behavior fusion feature vector with the pet historical behavior sliding time window, performs data capture and filtering, and establishes a time-series feature sequence. The temporal dependency analysis submodule uses a long short-term memory network to analyze temporal feature sequences and calculate the numerical correlation between data over time spans. It then filters out target correlation values ​​that meet the set time span correlation judgment threshold, reads the step size parameter of the time node to which the target correlation value belongs, performs matrix coordinate mapping between the target correlation value and the step size parameter, outputs feature arrangement parameters, and obtains the temporal dependency distribution attributes. The lifestyle baseline establishment submodule performs a multivariate regression fitting operation on the temporal dependency distribution attributes and the temporal feature sequence to depict the trajectory curve of the change in multimodal data over time and establish a long-term lifestyle characteristic baseline for pets.

8. The pet behavior monitoring system based on edge computing and cloud-based temporal modeling according to claim 1, characterized in that, The abnormal behavior warning module includes: The spatial deviation measurement submodule analyzes the covariance relationship of the multidimensional variables of the baseline based on the long-term living habits of the pet, extracts the distribution characteristics of the historical covariance matrix, and uses the Mahalanobis distance measurement algorithm combined with the distribution characteristics of the historical covariance matrix to calculate the spatial distance difference between the multimodal behavior fusion feature vector and the long-term living habits of the pet, and summarizes and generates multidimensional spatial statistical deviation. The alarm threshold adaptive submodule obtains the medical pet behavior health guidelines and extracts the basic warning threshold based on the multidimensional spatial statistical deviation. It reads the warning parameters from the previous historical moment, uses the exponential moving average algorithm to allocate the corresponding weight ratio parameters of the basic warning threshold and the warning parameters from the previous historical moment, and performs a smooth fitting operation by combining the corresponding weight ratio parameters with the basic warning threshold and the warning parameters from the previous historical moment to generate a dynamic adaptive alarm threshold. The abnormal behavior warning determination submodule compares the multidimensional spatial statistical deviation with the dynamic adaptive alarm threshold. When the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold, a pet behavior monitoring warning instruction is established. When the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold, a pet behavior normal status indicator instruction is generated.

9. A method for monitoring pet behavior based on edge computing and cloud-based temporal modeling, characterized in that, The method is used to implement the system according to any one of claims 1-8, and includes the following steps: S1: Acquire video keyframes, audio spectrum segments, and physical vibration signals from continuous monitoring cycles during pet behavior monitoring, and perform time alignment processing to generate a multimodal raw data synchronization sequence; S2: Analyze the multimodal raw data synchronization sequence to obtain pet video data stream, pet audio data stream and pet inertial data stream, extract normalized pet skeletal key point coordinate sequence, pet visual recognition confidence score, pet acoustic feature vector, pet environmental audio signal-to-noise ratio and pet motion energy statistics, and construct a multimodal structured feature set; S3: Analyze the multimodal structured feature set, extract environmental noise evaluation coefficients, determine modal dynamic weights by combining visual recognition confidence scores, and use the modal dynamic weights to fuse the normalized pet skeleton key point coordinate sequence, pet acoustic feature vector, and pet motion energy statistics in the multimodal structured feature set to generate a multimodal behavior fusion feature vector. S4: Obtain the sliding time window of the pet's historical behavior, truncate the multimodal behavior fusion feature vector, construct a time-series feature sequence, analyze the time-series feature sequence to extract the time-series dependency distribution attributes, and construct a baseline of the pet's long-term living habits features; S5: Determine the multidimensional spatial statistical deviation between the multimodal behavior fusion feature vector and the pet's long-term living habit feature baseline, construct a dynamic adaptive alarm threshold, and generate a pet behavior monitoring early warning instruction when the multidimensional spatial statistical deviation is greater than the dynamic adaptive alarm threshold; generate a pet behavior normal state indicator instruction when the multidimensional spatial statistical deviation is not greater than the dynamic adaptive alarm threshold.