Pet monitoring method and related equipment
By combining sound and image acquisition devices and using an environmental risk identification model for in-depth analysis, the problems of false alarms and slow response in traditional pet monitoring methods have been solved, enabling accurate identification and timely response to pet behavior and environmental risks.
Patent Information
- Application Number
- CN202510907529.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-28
AI Technical Summary
Traditional pet monitoring methods are prone to false alarms or missed alarms, and their response is not timely enough, making it difficult to effectively identify pet behavior and environmental risks.
By combining sound acquisition equipment to monitor target sound in real time, controlling image acquisition equipment to acquire video streams, and using environmental risk identification models for in-depth analysis, including sound feature extraction and image analysis, the identification threshold is dynamically adjusted to improve accuracy and response speed.
It enables precise location tracking of pet behavior and timely identification of environmental risks, reduces false alarm rates, and improves the reliability and response efficiency of the monitoring system in different scenarios.
Smart Images

Figure CN120853585A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer and communication technology, and more specifically, to a pet monitoring method and related equipment. Background Technology
[0002] Traditional monitoring methods typically rely on images or sounds to identify and capture pet movements. This can lead to false alarms or missed detections due to the reliance on a single technology. For example, sound acquisition may be misjudged due to environmental noise, and image acquisition may be affected by changes in lighting, thus making it difficult to respond to external changes in real time and miss important monitoring opportunities, especially in situations requiring rapid response, such as sudden pet behaviors (fighting, eating, etc.).
[0003] However, pets' behavior is complex and dynamic, and how to effectively identify environmental risks and assess whether pets are in danger through monitoring remains a technical challenge. Summary of the Invention
[0004] The embodiments of this application provide a pet monitoring method and related equipment, which can at least to some extent overcome the problems of frequent false alarms, slow positioning, and shallow analysis in traditional monitoring.
[0005] Other features and advantages of this application will become apparent from the following detailed description, or may be learned in part from practice of this application.
[0006] According to one aspect of the embodiments of this application, a pet monitoring method is provided, applied in a monitoring system including an image acquisition device, a sound acquisition device, and a server device, the pet monitoring method comprising:
[0007] The sound acquisition device monitors in real time whether the current ambient audio stream contains the target sound.
[0008] If the target sound is contained in the current ambient audio stream, then the image acquisition device is controlled to acquire the current target video stream based on the target sound;
[0009] The current target video stream is input into the environmental risk identification model to obtain the environmental risk identification result.
[0010] According to one aspect of the embodiments of this application, a pet monitoring device is provided, characterized in that it is applied in a monitoring system including an image acquisition device, a sound acquisition device, and a server device, the pet monitoring device comprising:
[0011] An environmental audio acquisition module is used to monitor in real time whether the current environmental audio stream contains the target sound through the sound acquisition device.
[0012] The target video acquisition module is used to control the image acquisition device to acquire the current target video stream based on the target sound if the current ambient audio stream contains the target sound;
[0013] The environmental risk identification module is used to input the current target video stream into the environmental risk identification model to obtain the environmental risk identification result.
[0014] According to one aspect of the embodiments of this application, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the pet monitoring method as described in the above embodiments.
[0015] According to one aspect of the embodiments of this application, an electronic device is provided, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the pet monitoring method as described in the above embodiments.
[0016] According to one aspect of the embodiments of this application, a computer program product is provided, including one or more computer programs that, when executed by one or more processors, implement the steps of the pet monitoring method as described in the above embodiments.
[0017] In some embodiments of this application, the technical solutions provide that, after initial sound screening to reduce false alarms, further image verification is performed to accurately locate the target pet. Finally, risk analysis and in-depth analysis are conducted, solving the core problems of traditional monitoring such as numerous false alarms, slow location, and shallow analysis. The dynamic threshold adjustment of the sound recognition module and the online learning of the image model enable the monitoring system to maintain high reliability in scenarios such as homes, farms, and outdoors.
[0018] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:
[0020] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.
[0021] Figure 2A flowchart illustrating a pet monitoring method provided in an embodiment of this application is shown.
[0022] Figure 3 It shows according to Figure 2 A flowchart illustrating a specific implementation of step S100 in the pet monitoring method shown in the corresponding embodiment.
[0023] Figure 4 It shows according to Figure 2 A flowchart illustrating a specific implementation of step S300 in the pet monitoring method shown in the corresponding embodiment.
[0024] Figure 5 A schematic diagram of the structure of a pet monitoring device provided in an embodiment of this application is shown.
[0025] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0026] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this application more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art.
[0027] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0028] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0029] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0030] Figure 1A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied is shown.
[0031] like Figure 1 As shown, the system architecture may include audio and video acquisition devices (such as...) Figure 1 The device shown includes one or more of the following: microphone 101, mobile phone 102, and camera 103 (which could also be a smart camera, etc.); network 104; and server 105. Network 104 serves as the medium for providing a communication link between the audio / video acquisition device and server 105. Network 104 can include various connection types, such as wired communication links, wireless communication links, etc.
[0032] It should be understood that Figure 1 The number of audio / video capture devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of audio / video capture devices, networks, and servers can be included. For example, server 105 could be a server cluster consisting of multiple servers.
[0033] Users can use audio and video capture devices to interact with server 105 via network 104 to receive or send messages, etc. Server 105 can be a server that provides various services. For example, users can use audio and video capture devices 103 (or 101 or 102) to upload real-time monitored audio streams of the current environment and video streams of the current target to server 105. Server 105 can monitor in real time whether the current environment audio stream contains the target sound; if the current environment audio stream contains the target sound, then based on the target sound, it controls the image capture device to acquire the current target video stream; and inputs the current target video stream into the environmental risk identification model to obtain the environmental risk identification result.
[0034] It should be noted that the pet monitoring method provided in this application embodiment is generally executed by server 105, and correspondingly, the pet monitoring device is generally installed in server 105. However, in other embodiments of this application, the audio and video acquisition device may also have similar functions to the server, thereby executing the pet monitoring scheme provided in this application embodiment.
[0035] The implementation details of the technical solutions in the embodiments of this application are described in detail below:
[0036] Figure 2 A flowchart of a pet monitoring method according to an embodiment of this application is shown. This pet monitoring method can be executed by a server, which may be... Figure 1 The server shown. (Refer to...) Figure 2 As shown, this pet monitoring method includes at least the following:
[0037] S100, the sound acquisition device monitors in real time whether the current ambient audio stream contains the target sound.
[0038] S200, if the current ambient audio stream contains the target sound, then control the image acquisition device to acquire the current target video stream based on the target sound.
[0039] S300, input the current target video stream into the environmental risk identification model to obtain the environmental risk identification result.
[0040] In the embodiments of this application, after initial sound screening to reduce false alarms, further image verification is performed to accurately locate the target pet. Finally, risk analysis and in-depth analysis are conducted to solve the core problems of traditional monitoring, such as high false alarm rates, slow location, and superficial analysis. The dynamic threshold adjustment of the sound recognition module and the online learning of the image model enable the monitoring system to maintain high reliability in scenarios such as homes, farms, and outdoors.
[0041] In the S100, a database of dog barking and cat meowing features is constructed. Noise is filtered by combining audio signal processing algorithms (such as MFCC). Features are compared in real time through deep learning models (such as convolutional neural networks) to accurately identify target sounds, reduce the false alarm rate, and avoid false triggering of traditional all-sound alarms.
[0042] Specifically, in some embodiments, the specific implementation of step S100 can be found in [reference needed]. Figure 3 . Figure 3 It is based on Figure 2 According to the detailed description of step S100 in the pet monitoring method shown in the corresponding embodiment, step S100 in the pet monitoring method may include the following steps:
[0043] S110, the current ambient audio stream is collected in real time through the sound acquisition device.
[0044] S120, extract the sound features of the current ambient audio stream, and determine whether the current ambient sound contains the target sound.
[0045] In this embodiment, the ambient audio stream is monitored in real time using a sound acquisition device to extract audio features and determine whether the target sound is present. This method can not only capture the sounds made by the pet but also identify the characteristics of the target sound (such as the pet's barking or howling), thus enabling more detailed monitoring of pet behavior. Simultaneously, by extracting the sound features of the current ambient audio stream, the system can analyze and determine whether the target sound is present. The audio feature extraction technique can use signal processing algorithms to reduce the influence of background noise and ensure accurate identification of the target sound (e.g., the pet's barking).
[0046] In S110, the sound acquisition device monitors the audio stream in the environment in real time and transmits it to the processing unit. The audio acquisition device typically uses a microphone or other sensors to continuously collect ambient sounds. The device may collect all sounds, including pet sounds, environmental noise, and human speech.
[0047] Specifically, a 4-microphone linear array (10cm spacing) can be used to support sound source localization and beamforming, enhancing the directional gain of the target sound. It also integrates an analog-to-digital converter (ADC) with a sampling rate of 44.1kHz and quantization accuracy of 16 bits, meeting CD-quality audio acquisition standards. The effective pickup distance of the aforementioned device can be 0-10 meters (adjusted according to ambient acoustic reflections) to cover scenarios such as family living rooms, pet store shelf areas, and small outdoor activity areas. Audio data can be transmitted to the edge computing module via an I2S bus and stored using a circular buffer (1024 frames in size) to ensure continuous acquisition without frame loss.
[0048] It should be noted that in some embodiments, an adaptive filtering algorithm (such as Minimum Variance Distortionless Response, MVDR) can be used to form a beam pointing towards the sound source to suppress noise from other directions (such as side TV noise), improving the signal-to-noise ratio by more than 10dB. At the same time, the microphone gain is dynamically adjusted to avoid signal clipping caused by loud sources at close range (such as a dog barking close to the microphone) while amplifying weak sounds at a distance (such as a cat meowing in the next room).
[0049] In S120, the system performs signal processing on the acquired audio stream to extract sound features. Commonly used audio feature extraction methods include spectral analysis and Mel-frequency cepstral coefficients (MFCC), which can reflect the time-domain and frequency-domain characteristics of the audio signal. These features help the system extract potential target sound features from complex audio streams.
[0050] Specifically, in some embodiments, the specific implementation of step S120 can be found in the following embodiments. This embodiment is based on... Figure 3 According to the detailed description of step S120 in the pet monitoring method shown in the corresponding embodiment, step S120 in the pet monitoring method may include the following steps:
[0051] Extract the Mel frequency cepstral coefficients of the current ambient audio stream.
[0052] Extract the short-time energy and zero-crossing rate of the current ambient audio stream.
[0053] Based on the Mel frequency cepstral coefficients, short-time energy, and zero-crossing rate, a sound feature vector is formed.
[0054] The sound feature vector is input into a pre-trained residual neural network, and the probability distribution of the corresponding sound type is determined by the Softmax function;
[0055] Based on the probability distribution of the sound type and a predetermined sound threshold, it is determined whether the current ambient sound contains the target sound.
[0056] The predetermined sound threshold is adjusted based on the duration at which the target sound is determined to be present.
[0057] In this embodiment, by extracting features such as Mel-frequency cepstral coefficients (MFCC), short-time energy, and zero-crossing rate from the ambient audio stream, different aspects of the audio signal can be effectively characterized. These features provide a more refined analysis of the audio signal, thereby improving the accuracy of target sound recognition. The Residual Neural Network (ResNet) combined with the Softmax function can further improve the accuracy of sound recognition, especially in complex environments. Through comprehensive analysis of Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate, effective features can be better extracted from complex audio signals, reducing the impact of noise. Classification using a pre-trained Residual Neural Network enables the system to accurately identify target sounds in different background noise environments. By adjusting a predetermined sound threshold based on the duration of the target sound, the system can flexibly adjust the recognition criteria according to the actual conditions of different environments, thereby improving the accuracy of target sound recognition.
[0058] First, the system converts the acquired ambient audio stream into a series of Mel-frequency cepstral coefficients (MFCCs). MFCCs are a commonly used audio feature that effectively describes the spectral characteristics of audio signals. By decomposing the audio signal in both the frequency and time domains, it can better capture key features in the audio, making it suitable for speech and sound recognition. Next, the system calculates the short-time energy and zero-crossing rate of the audio signal. Short-time energy describes the signal intensity, while the zero-crossing rate represents the waveform changes. Combining these two provides a detailed understanding of the dynamics of the sound waveform, helping to distinguish different types of sounds, especially with more precise recognition against the background of the target sound. Using the extracted Mel-frequency cepstral coefficients, short-time energy, and zero-crossing rate, the system combines these features into a sound feature vector. The feature vector is a set of values representing multiple dimensions of the audio signal, serving as input for subsequent sound recognition. The sound feature vector is then fed into a pre-trained residual neural network. ResNet is a deep learning model that effectively solves the vanishing gradient problem in deep networks through residual connections (skip connections), improving recognition accuracy. The network learns how to distinguish different sound types from the feature vector through training. The output of the neural network is converted into a probability distribution using a softmax function, representing the likelihood of each sound type. This step ensures that the system can probabilistically determine different sound types, thereby selecting the most probable target sound type. After detecting the presence of a target sound, the system considers its duration. If the duration of the target sound exceeds a certain threshold, the system adjusts the sound threshold to further optimize the sensitivity of sound recognition. This allows the system to handle different types of sound behaviors, such as the continuous barking of a pet, while avoiding brief, occasional noise interference.
[0059] Specifically, a pre-emphasis filter is applied to the audio signal, using the formula H(z) = 1 - αz⁻¹ (α = 0.97), to boost high-frequency components (the energy proportion of the high-frequency band in cat meows can reach 70%), compensating for high-frequency attenuation in sound transmission. The linear frequency is converted to Mel frequency using the formula fmel = 2595lg(1 + f / 700). Forty triangular filters are evenly spaced on the Mel scale, covering 20Hz-8kHz (canine hearing range 15Hz-50kHz, cat hearing range 70Hz-64kHz). The logarithm of the filter bank output is taken, and a Discrete Cosine Transform (DCT) is performed to extract the first 13 MFCC coefficients (retaining low-frequency cepstral coefficients and discarding high-frequency redundant information), and their first-order difference (ΔMFCC) is calculated, ultimately generating a 26-dimensional MFCC feature vector.
[0060] Short-time energy is used to reflect the energy intensity of each frame of sound, to distinguish between soft whimpers (low energy) and loud barks (high energy), and can be expressed using the following formula:
[0061]
[0062] Among them, E n Here, s represents the short-time energy value at time n, n represents the original signal, m represents the current time (or frame number), w(m) represents the time offset within the window, w(m) represents the Hamming window, m represents the frame data, and N = 256 points (corresponding to approximately 5.8 ms, aligned with the MFCC frame length).
[0063] The zero-crossing rate can be calculated using the following formula:
[0064]
[0065] Among them, ZCR n Let be the zero-crossing rate at time n, s represent the original signal, n be the current time (or frame number), m be the time offset within the window, and sign(·) be the sign function. If the ZCR is greater than 50 times / second, it is judged as a high zero-crossing rate (such as the continuous high-frequency vibration of a cat's meow); if the ZCR is less than 20 times / second, it is judged as a low zero-crossing rate (such as the pulsed sound of a dog's bark).
[0066] Then, the 26-dimensional MFCC vector, the 1-dimensional short-time energy, and the 1-dimensional zero-crossing rate are concatenated to generate a 28-dimensional feature vector. Z-score standardization is then applied to make the mean of each feature 0 and the standard deviation 1.
[0067] A pre-trained residual neural network consists of an input layer, convolutional blocks, a global average pooling layer, a fully connected layer, and a softmax layer.
[0068] The input layer is mapped to a 128-dimensional feature vector through a fully connected layer to adapt to the feature vector dimension and improve non-linear expressive power. Each convolutional block contains 1 to 3 layers, with two 3×3 convolutional layers (padding = 1, maintaining the same size) and skip connections across layers to extract acoustic features at different scales; shallower layers capture frequency components, while deeper layers capture syllable patterns. The output dimension of the global average pooling layer is 128 to reduce dimensionality and prevent overfitting of the fully connected layers. The fully connected layers have 3 neurons (corresponding to dog barks, cat meows, and others) to output classification logits.
[0069] The softmax layer is used to output the probability distribution. Convert logits to 0-1 probability values.
[0070] The initial preset sound threshold can be set to Thresh. init =0.9, and then update the threshold using exponential smoothing, specifically using the following formula:
[0071] Thresh new=λ·Thresh old +(1-λ)Thresh map
[0072] Among them, Thresh new Thresh old Thresh is the old preset sound threshold. map The mapping threshold has a coefficient λ = 0.8 to preserve historical threshold weights and avoid abrupt changes. The predetermined sound threshold has a minimum value of 0.5 and a maximum value of 0.95 to prevent oversensitivity or sluggishness in extreme scenarios.
[0073] If the probability of a dog barking or a cat meowing output by Softmax is greater than the current threshold, it is determined that the target sound is included, and the subsequent audio-visual linkage process is triggered; otherwise, monitoring continues.
[0074] For example, if a sound lasting 4 seconds is detected, the threshold is adjusted to 0.7. At this time, the model outputs a dog barking probability of 0.75 > 0.7, which is considered a valid trigger. If the same probability occurs in the initial state (threshold 0.9), it is considered invalid.
[0075] In the S200, the sound positioning module calculates the direction of the sound source using the TDOA algorithm, drives the gimbal to rotate at a predetermined rotation speed (e.g., 0.5πrad / s) and automatically focuses, and completes image capture within 3 seconds to achieve zero-delay linkage from sound triggering to image response, improve positioning efficiency, and avoid the problem of blind shooting with a fixed lens.
[0076] Specifically, in some embodiments, the specific implementation of step S200 can be found in the following embodiments. This embodiment is based on... Figure 2 According to the detailed description of step S200 in the pet monitoring method shown in the corresponding embodiment, step S200 in the pet monitoring method may include the following steps:
[0077] The location of the target pet is determined based on the target sound.
[0078] Based on the location of the target pet, the image acquisition device is controlled to acquire the current target video stream.
[0079] In this embodiment, based on the acoustic propagation law (sound speed 340m / s) and the geometric relationship of the sensor array, centimeter-level positioning accuracy is theoretically derived through time difference measurement. This is combined with engineering calibration (such as environmental reflection coefficient compensation) to reduce practical application errors. The gimbal motion is decomposed into a dual-objective optimization problem of angle tracking and trajectory smoothing. Feedforward control eliminates mechanical lag, ensuring no overshoot during high-speed rotation and improving image stability.
[0080] Specifically, for the preprocessed audio stream, the Generalized Cross-Correlation (GCC-PHAT) algorithm is used to calculate the time difference of arrival (TDOA) between each pair of microphones. Then, assuming the sound source is located in a plane (z = 0, ground coordinate system), a system of equations is established using triangulation, and the sound source coordinates are iteratively solved using the least squares method. Finally, a Kalman filter is introduced, and the optimal estimate is output by combining the localization result from the previous moment with the current measurement value.
[0081] The sound source coordinates are converted into gimbal rotation angles. Using the shortest path priority principle, the difference between clockwise and counterclockwise rotation angles is calculated, and a rotation direction less than 180° is selected. An S-shaped acceleration / deceleration algorithm is used to avoid mechanical shock during start-up and shutdown. The lens automatically adjusts its focus based on the horizontal distance to the sound source.
[0082] If the TDOA calculation error is greater than 2 meters, the backup plan will be triggered.
[0083] In the backup plan, the pan-tilt unit scans horizontally 360° at a speed of 5° / s, while the sound acquisition device continuously monitors and determines the approximate direction based on the peak sound energy. If the probability of a target sound is detected during the scan is greater than 0.9, the unit immediately stops rotating, zooms in, and initiates image recognition-assisted positioning. When the coordinates of multiple sound sources differ by less than 1.5 meters (e.g., multiple pets gathered together), the pan-tilt unit prioritizes aiming at the target with the highest sound probability, and marks multiple pets in the image with different colored boxes (e.g., yellow for dogs, blue for cats).
[0084] In the S300, the captured image stream is transmitted to a server device, where an environmental risk identification model is run. This model uses image analysis to identify potential risks to the pet's behavior or environment. Possible risks include whether the pet is injured, whether a fight has occurred, or whether it is in a dangerous environment. Based on the identification results, the system can generate a risk report or push notifications to the pet owner.
[0085] Specifically, in some embodiments, the specific implementation of step S300 can be found in [reference needed]. Figure 4 . Figure 4 It is based on Figure 2 According to the detailed description of step S300 in the pet monitoring method shown in the corresponding embodiment, the environmental risk identification model in the pet monitoring method includes a behavior analysis sub-model, an environmental analysis sub-model, and an environmental risk sub-model. Step S300 may include the following steps:
[0086] S310, input the current video stream into the behavior analysis sub-model to obtain target behavior features.
[0087] S320, input the current video stream into the environment analysis sub-model to obtain the target environment features.
[0088] S330, Input the target behavioral characteristics and the target environmental characteristics into the environmental risk sub-model to obtain the environmental risk identification result.
[0089] In this embodiment, by integrating behavioral analysis, environmental analysis, and environmental risk analysis, the system can comprehensively identify the pet's behavioral characteristics and environmental risks, ensuring that the system can react promptly and take appropriate countermeasures to prevent the pet from facing potential harm. Simultaneously, by comprehensively analyzing the pet's behavioral and environmental characteristics, the environmental risk sub-model can make more accurate risk predictions and decisions based on multi-faceted input information. The system can dynamically adjust the behavior of the monitoring equipment based on the environmental risk identification results input from the real-time video stream, enhancing its responsiveness to emergencies.
[0090] In S310, by analyzing the pet's behavior in the current video stream in real time, the system extracts features such as the pet's movements and postures to determine its behavioral patterns. For example, it determines whether the pet is resting, eating, playing, or exhibiting abnormal behavior (such as anxiety or aggression). These behavioral characteristics will form the basis for subsequent environmental risk identification.
[0091] Specifically, deep learning-based behavior recognition algorithms (such as convolutional neural networks CNN) or motion analysis-based behavior recognition models can be used to accurately classify and extract features from pets' movements.
[0092] Specifically, in some embodiments, the specific implementation of step S310 can be found in the following embodiments. This embodiment is based on... Figure 4 According to the detailed description of step S310 in the pet monitoring method shown in the corresponding embodiment, the behavior analysis sub-model in the pet monitoring method includes a target monitoring network and a target behavior analysis network. Step S310 may include the following steps:
[0093] The current target video stream is input into the target monitoring network to obtain a target-marked video stream, on which a target monitoring frame for the target pet is marked.
[0094] The target-labeled video stream is input into the target behavior analysis network to obtain target behavior features.
[0095] In this embodiment, by dividing the video stream into two steps—target monitoring and target behavior analysis—the system can more accurately identify pet behavior. This not only improves the system's behavior recognition capabilities in complex scenarios but also better addresses changes in dynamic environments. Based on real-time tagging of the target monitoring frame, this implementation ensures the system can track the pet's location and behavioral status during monitoring, thereby improving the system's reaction speed and decision-making accuracy. By processing the video stream in two steps—target monitoring and behavior analysis—the system can better separate information at different levels, enhancing the overall understanding of pet behavior.
[0096] Specifically, image information of the target pet is extracted from real-time monitoring video streams. These images are then fed into a target detection network. The network analyzes objects in the images using deep learning models (such as convolutional neural networks, CNNs), identifies the target pet, and generates bounding boxes in the video stream. These bounding boxes clearly show the pet's specific location, aiding in subsequent behavior analysis. The target detection network can use object detection algorithms (such as YOLO, Faster R-CNN, etc.) to detect and label the pet target.
[0097] The video stream, tagged with the target, is fed into a target behavior analysis network to analyze the pet's behavioral characteristics, such as movements and postures, in order to identify the pet's current behavioral state. In this way, the system can effectively extract specific behavioral features of the pet from the tagged video stream, such as whether it is eating, resting, playing, or exhibiting abnormal behavior. Target monitoring tags can use time-series-based behavior analysis models (such as LSTM or GRU) to analyze changes in pet behavior, or deep convolutional neural networks to process behavioral features in video frames.
[0098] In S320, the environmental analysis sub-model extracts environmental features by analyzing background information (such as temperature, lighting, obstacles, and other environmental factors) in the video stream. These features can reveal changes in the current environment, such as whether there is abnormal noise or whether the pet's environment is potentially dangerous (such as fragile items or disturbances from other animals).
[0099] Specifically, computer vision technology can be used to analyze scenes in real time, extract visual features of the environment, and combine external sensor data (such as temperature sensors, humidity sensors, etc.) to achieve comprehensive environmental analysis.
[0100] Specifically, in some embodiments, the specific implementation of step S320 can be found in the following embodiments. This embodiment is based on... Figure 4According to the detailed description of step S320 in the pet monitoring method shown in the corresponding embodiment, the environmental analysis sub-model in the pet monitoring method includes a scene recognition network and a key object recognition network. The target environmental features include scene recognition results and key object states. Step S320 may include the following steps:
[0101] The current video stream is input into the scene recognition network to obtain the scene recognition result.
[0102] The current video stream is input into the key object recognition network to obtain the key object status.
[0103] In this embodiment, by combining scene recognition and key object recognition, a more comprehensive understanding of the pet's environment can be achieved, including identifying important objects and the overall scene layout. This provides more contextual support for pet behavior analysis. The environmental analysis sub-model, by identifying the state of the scene and key objects, helps the system better understand the context of the pet's behavior. For example, a pet's behavior may differ in some environments from others; integrating environmental information helps improve the accuracy of behavior recognition. This embodiment not only detects the pet's own behavior but also performs a comprehensive analysis of its surrounding environment, enhancing the multi-dimensional processing capabilities of the monitoring system.
[0104] Specifically, scene recognition can employ a Two-Stream CNN. In the spatial stream, a ResNet50 network is used to extract spatial features (such as furniture layout and door / window positions) from the current frame image, outputting a spatial stream feature vector. In the temporal stream, a 3D ResNet18 backbone network is used to extract temporal features (such as object movement trajectories and door opening / closing actions) from five consecutive frames of optical flow maps, outputting a temporal stream feature vector. The spatial and temporal stream feature vectors are concatenated, dimensionality is reduced through a fully connected layer, and then input into a Softmax classifier for classification. The spatial stream branch outputs the scene probability distribution, while the temporal stream branch determines the scene's dynamic state, ultimately forming the scene recognition result.
[0105] The key object recognition network can be improved based on YOLOv8n by adding a dynamic feature head. After the video frame is input into the backbone network C3Ghost, the static branch of the detection head detects the object category and position, and the dynamic branch of the detection head predicts the changes in the object state to obtain the state of the key object.
[0106] In S330, once the target behavioral characteristics and target environmental characteristics are obtained, the system inputs these characteristics into the environmental risk sub-model for comprehensive analysis. Based on this input data, the model assesses the environmental risks currently faced by the pet, such as whether the pet is in a potentially dangerous behavioral state or whether there are external environmental factors threatening the pet.
[0107] Specifically, multimodal learning (such as deep neural networks that integrate behavioral and environmental information) can be used for risk assessment. By optimizing environmental risk models with a large amount of training data, they can accurately predict the environmental risks faced by pets and issue timely alerts or take appropriate measures. Environmental risks can also be analyzed by constructing knowledge graphs.
[0108] In some embodiments, a knowledge graph can be constructed first.
[0109] The entity types in a knowledge graph include pets, environmental entities, and behaviors; the relationship types include spatial relationships, behavioral relationships, and causal relationships.
[0110] Among them, pets include breed, age, health status, environmental entities include risk level, location, behavior includes risk factor, spatial relationship includes distance, direction, behavioral relationship includes biting, contact, and causal relationship, such as death caused by electric shock.
[0111] Based on knowledge graphs, a risk knowledge base can be formed to obtain the risk coefficients corresponding to each feature. Then, the risk level of each feature value can be calculated, specifically by weighted summation of the risks corresponding to each feature. The specific formula is as follows:
[0112]
[0113] Among them, Socre Risk For risk assessment score, f i w is the eigenvalue of the i-th feature. i Let be the weight of the i-th feature.
[0114] After obtaining the risk assessment score, the corresponding risk assessment level can be determined based on the risk assessment score, thereby obtaining the corresponding environmental risk identification result.
[0115] Specifically, the training methods for the aforementioned environmental risk identification model include:
[0116] Obtain the target video stream sample set, which contains multiple target video stream samples, each of which is labeled with a corresponding environmental risk identification tag.
[0117] The target video stream samples are input one by one into the environmental risk identification model to obtain the environmental risk identification results.
[0118] Based on the output environmental risk identification results and the environmental risk identification labels, the parameters of the environmental risk identification model are updated until the predetermined termination condition is met, thus ending the training and obtaining the trained environmental risk identification model.
[0119] In the embodiments of this application, during training, a target video stream sample set containing multiple target video stream samples can be obtained first, and each target video stream sample is labeled with a corresponding environmental risk identification label; then, the multiple target video stream samples are divided into a training set, a validation set, and a test set according to a predetermined ratio; then, the parameters of the encoder and decoder in the environmental risk identification model are adjusted and determined according to the target video stream samples included in the training set, validation set, and test set, so as to obtain the trained environmental risk identification model.
[0120] It should be noted that the target video stream samples mentioned above include target video stream samples with occlusions, in order to improve the model's ability to recognize various occlusions.
[0121] When training the model, the target video stream sample set can be divided into a training set, a validation set, and a test set. Then, the model is trained based on the training set, validated based on the validation set, and tested based on the test set to obtain a trained environmental risk identification model.
[0122] Before training on the training set, the video frame image samples in the training set can be preprocessed. Preprocessing includes image resizing, normalization, data augmentation, and class encoding.
[0123] The specific methods for data augmentation include performing geometric transformations (including flipping, rotating, deforming, scaling, etc.) on the data image, as well as color transformations (including noise reduction, blurring, color transformation, erasing, and filling). These operations can better expand the shape of the video frame image.
[0124] In some embodiments, GAN networks can also be used to generate images of target pets that are occluded, out of focus, or blurred, which are difficult to identify but not easy to obtain. This helps the algorithm model to accurately identify target pets, reduces false identification and missed identification, and improves the accuracy of the model in identifying target pets in special images.
[0125] After obtaining the enhanced training set, the environmental risk identification model can be trained based on the enhanced training set, and the parameters and weights in the network can be updated.
[0126] Specifically, video frame image samples from the training set are input into the environmental risk identification model to obtain the environmental risk identification result output by the model. The environmental risk identification result is compared with the environmental risk identification label, and the loss function is calculated. Then, the stochastic gradient descent method is used to minimize the loss function. The parameters and weights in the environmental risk identification model are updated by backpropagation until the loss function meets a predetermined condition, such as convergence or being less than a predetermined threshold.
[0127] In some embodiments, the environmental risk identification results include an environmental risk score, an environmental risk level, and environmental risk content. The environmental risk identification labels may include environmental risk score labels, environmental risk level labels, and environmental risk content labels. The loss function is calculated by comparing the environmental risk score with the environmental risk score label and / or comparing the environmental risk level with the environmental risk level label and / or comparing the environmental risk content with the environmental risk content label.
[0128] After training, the environmental risk identification network with updated parameters can be validated using a validation set. Specifically, the network is tuned based on the validation set data. When the loss function meets predetermined conditions, the model parameters for that stage are output. If the loss function does not meet the predetermined conditions, hyperparameters such as the learning rate are automatically adjusted, and the next round of network model training begins.
[0129] When the loss function computed on the validation set meets predetermined conditions, the parameters and weights can be retained, and then the retained parameters are tested based on the test set. Specifically, the input test set data enables the environmental risk identification network with retained parameters and weights to output environmental risk identification results and model weights. By comparing the model loss and corresponding weights across multiple rounds, the model weights with the minimum loss are output, thus determining the trained environmental risk identification model.
[0130] Once the trained environmental risk identification model is obtained, environmental risk identification can be completed based on the model.
[0131] Furthermore, data augmentation is not required for the input data in the validation and test sets.
[0132] In some embodiments of this application, after step S300 described above, the pet monitoring method further includes:
[0133] Based on the environmental risk identification results, push messages are generated according to pre-configured user preference data.
[0134] In this embodiment, push notifications are generated based on the output of the environmental risk identification model and the user's initial preference data. User preference data may include their level of concern regarding specific types of risks (e.g., whether they are particularly sensitive to pets near high-temperature equipment) and their preferred message types (e.g., text, audio, images). The system integrates these factors to generate personalized push notifications. For example, if a pet is detected near a power outlet, and the user has set their electrical safety awareness, the system will promptly send a warning. Specifically, push notifications can be generated and sent using a data processing and rule engine that matches the environmental risk identification results with the user's preference data. The generation of push notifications can be achieved through automated scripts or machine learning models to ensure they are as personalized as possible to the user's needs.
[0135] In some embodiments, users can set their desired message push format from multiple dimensions. For example, they can configure risk level filtering to select only to receive pushes with higher risk levels; they can also configure the level of detail in the push content, choosing between a concise text summary or a more detailed text and image format, or the most detailed text plus behavioral video clips and environmental annotation maps; they can also configure the receiving time and only push to scenarios that the user is interested in.
[0136] Furthermore, the system can analyze users' historical behavior to determine user preferences. For example, it can calculate click-through rates and push notifications based on user preferences for categories with high click-through rates. It can also calculate the duration of user engagement with each notification to push more detailed content for categories with longer engagement times.
[0137] When generating push notifications, key elements can be extracted from the environmental risk identification results. Then, the corresponding message template can be selected based on user preferences. After filling the key elements into the template, the details can be adjusted according to the user profile to highlight the information that the user is more concerned about, thus forming the push notification.
[0138] It should be noted that information of different risk levels can be pushed to customers in different ways. For example, for the highest risk level, Level 1, a combination of APP push notifications, SMS messages, and phone ringtones can be used to send reminders; for the relatively high risk level, Level 2, a combination of APP push notifications and internal messages can be used to send reminders; and for the lowest risk level, Level 3, a combination of silent APP push notifications and message center storage can be used to send reminders.
[0139] The following describes an embodiment of the apparatus described in this application, which can be used to execute the pet monitoring method described in the above embodiments of this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the pet monitoring method described above.
[0140] Figure 5A block diagram of a pet monitoring device according to an embodiment of this application is shown.
[0141] Reference Figure 5 As shown, the pet monitoring device 500 is used in a monitoring system that includes image acquisition equipment, sound acquisition equipment, and server equipment, including: an environmental audio acquisition module 510, a target video acquisition module 520, and an environmental risk identification module 530.
[0142] The environmental audio acquisition module 510 is used to monitor in real time whether the current environmental audio stream contains the target sound through the sound acquisition device; the target video acquisition module 520 is used to control the image acquisition device to acquire the current target video stream according to the target sound if the current environmental audio stream contains the target sound; and the environmental risk identification module 530 is used to input the current target video stream into the environmental risk identification model to obtain the environmental risk identification result.
[0143] In some feasible embodiments of this application, the environmental audio acquisition module 510 specifically includes: a current audio acquisition submodule, used to acquire the current environmental audio stream in real time through the sound acquisition device; and a target sound judgment submodule, used to extract the sound features of the current environmental audio stream and determine whether the current environmental sound contains a target sound.
[0144] In some feasible embodiments of this application, the target sound determination submodule specifically includes: a Mel frequency cepstral unit for extracting Mel frequency cepstral coefficients of the current ambient audio stream; a feature supplementation extraction unit for extracting short-time energy and zero-crossing rate of the current ambient audio stream; a sound vector formation unit for forming a sound feature vector based on the Mel frequency cepstral coefficients, short-time energy, and zero-crossing rate; a residual neural network unit for inputting the sound feature vector into a pre-trained residual neural network and determining the corresponding sound type distribution probability through a Softmax function; a target sound determination unit for determining whether the current ambient sound contains a target sound based on the sound type distribution probability and a predetermined sound threshold; and a sound threshold adjustment unit for adjusting the predetermined sound threshold based on the duration of the determined target sound.
[0145] In some feasible embodiments of this application, the target video acquisition module 520 specifically includes: a target pet positioning submodule, used to determine the location based on the target sound; and a target pet shooting submodule, used to control the image acquisition device to acquire the current target video stream based on the location of the target pet.
[0146] In some feasible embodiments of this application, the environmental risk identification model includes a behavior analysis sub-model, an environmental analysis sub-model, and an environmental risk sub-model; the risk identification module 530 specifically includes: a behavior analysis sub-module, used to input the current video stream into the behavior analysis sub-model to obtain target behavior features; an environmental analysis sub-module, used to input the current video stream into the environmental analysis sub-model to obtain target environmental features; and an environmental risk sub-module, used to input the target behavior features and the target environmental features into the environmental risk sub-model to obtain an environmental risk identification result.
[0147] In some feasible embodiments of this application, the behavior analysis sub-model includes a target monitoring network and a target behavior analysis network; the behavior analysis sub-module specifically includes: a target pet monitoring unit, used to input the current target video stream into the target monitoring network to obtain a target-marked video stream, wherein the target-marked video stream is marked with a target monitoring box for the target pet; and a target behavior analysis unit, used to input the target-marked video stream into the network to obtain target behavior features.
[0148] In some feasible embodiments of this application, the environment analysis sub-model includes a scene recognition network and a key object recognition network, and the target environment features include scene recognition results and key object states; the environment analysis sub-module includes: a current scene recognition unit, used to input the current video stream into the scene recognition network to obtain scene recognition results; and a key object recognition unit, used to input the current video stream into the key object recognition network to obtain key object states.
[0149] In some feasible embodiments of this application, after inputting the current target video stream into the environmental risk identification model to obtain the environmental risk identification result, the pet monitoring device further includes: a push message generation module, used to generate a push message according to the environmental risk identification result and pre-configured user preference data.
[0150] In the embodiments of this application, after initial sound screening to reduce false alarms, further image verification is performed to accurately locate the target pet. Finally, risk analysis and in-depth analysis are conducted to solve the core problems of traditional monitoring, such as high false alarm rates, slow location, and superficial analysis. The dynamic threshold adjustment of the sound recognition module and the online learning of the image model enable the monitoring system to maintain high reliability in scenarios such as homes, farms, and outdoors.
[0151] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the electronic device of the present application is shown.
[0152] It should be noted that, Figure 6The computer system of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0153] like Figure 6 As shown, the computer system includes a Central Processing Unit (CPU) 1801, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 1802 or programs loaded from storage portion 1808 into Random Access Memory (RAM) 1803, such as performing the methods described in the above embodiments. The RAM 1803 also stores various programs and data required for system operation. The CPU 1801, ROM 1802, and RAM 1803 are interconnected via a bus 1804. An Input / Output (I / O) interface 1805 is also connected to the bus 1804.
[0154] The following components are connected to I / O interface 1805: an input section 1806 including a keyboard, mouse, etc.; an output section 1807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1808 including a hard disk, etc.; and a communication section 1809 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 1809 performs communication processing via a network such as the Internet. A drive 1810 is also connected to I / O interface 1805 as needed. Removable media 1811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1810 as needed so that computer programs read from them can be installed into storage section 1808 as needed.
[0155] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1809, and / or installed from removable medium 1811. When the computer program is executed by central processing unit (CPU) 1801, it performs various functions defined in the system of this application.
[0156] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.
[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0158] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.
[0159] In another aspect, this application also provides a computer-readable medium, which may be included in the electronic device described in the above embodiments; or it may exist independently and not assembled into the electronic device. The computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods described in the above embodiments.
[0160] This specification also provides a computer program product that stores at least one instruction, said at least one instruction being loaded and executed by the processor as described above. Figures 1-4 The method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-4 The specific details of the illustrated embodiments will not be elaborated here.
[0161] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0162] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of this application.
[0163] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.
[0164] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A pet monitoring method, characterized in that, The pet monitoring method, applied in a monitoring system that includes image acquisition devices, sound acquisition devices, and server-side devices, comprises: The sound acquisition device monitors in real time whether the current ambient audio stream contains the target sound. If the target sound is contained in the current ambient audio stream, then the image acquisition device is controlled to acquire the current target video stream based on the target sound; The current target video stream is input into the environmental risk identification model to obtain the environmental risk identification result.
2. The pet monitoring method as described in claim 1, characterized in that, The step of monitoring in real time whether the target sound is contained in the current ambient audio stream through the sound acquisition device specifically includes: The sound acquisition device collects the current ambient audio stream in real time. Extract the sound features of the current ambient audio stream and determine whether the current ambient sound contains the target sound.
3. The pet monitoring method as described in claim 2, characterized in that, The step of extracting the sound features of the current ambient audio stream and determining whether the current ambient sound contains the target sound specifically includes: Extract the Mel-frequency cepstral coefficients of the current ambient audio stream; Extract the short-time energy and zero-crossing rate of the current ambient audio stream; Based on the Mel frequency cepstral coefficients, short-time energy, and zero-crossing rate, a sound feature vector is formed; The sound feature vector is input into a pre-trained residual neural network, and the probability distribution of the corresponding sound type is determined by the Softmax function; Based on the probability distribution of the sound type and the predetermined sound threshold, it is determined whether the current ambient sound contains the target sound; The predetermined sound threshold is adjusted based on the duration at which the target sound is determined to be present.
4. The pet monitoring method as described in claim 1, characterized in that, The step of controlling the image acquisition device to acquire the current target video stream based on the target sound specifically includes: Based on the target sound, determine the location of the target pet; Based on the location of the target pet, the image acquisition device is controlled to acquire the current target video stream.
5. The pet monitoring method as described in claim 1, characterized in that, The environmental risk identification model includes a behavioral analysis sub-model, an environmental analysis sub-model, and an environmental risk sub-model. The step of inputting the current target video stream into the environmental risk identification model to obtain the environmental risk identification result specifically includes: The current video stream is input into the behavior analysis sub-model to obtain the target behavior features; The current video stream is input into the environmental analysis sub-model to obtain the target environmental features; The target behavioral characteristics and the target environmental characteristics are input into the environmental risk sub-model to obtain the environmental risk identification result.
6. The pet monitoring method as described in claim 5, characterized in that, The behavior analysis sub-model includes a target monitoring network and a target behavior analysis network; The step of inputting the current video stream into the behavior analysis sub-model to obtain target behavior features specifically includes: The current target video stream is input into the target monitoring network to obtain a target-marked video stream, on which a target monitoring frame for the target pet is marked; The target-labeled video stream is input into the target behavior analysis network to obtain target behavior features.
7. A pet monitoring device, characterized in that, The pet monitoring device, applied in a monitoring system that includes image acquisition equipment, sound acquisition equipment, and server equipment, comprises: An environmental audio acquisition module is used to monitor in real time whether the current environmental audio stream contains the target sound through the sound acquisition device. The target video acquisition module is used to control the image acquisition device to acquire the current target video stream based on the target sound if the current ambient audio stream contains the target sound; The environmental risk identification module is used to input the current target video stream into the environmental risk identification model to obtain the environmental risk identification result.
8. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the pet monitoring method as described in any one of claims 1 to 6.
9. An electronic device, characterized in that, include: one or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the pet monitoring method as described in any one of claims 1 to 6.
10. A computer program product comprising one or more computer programs, characterized in that, When the one or more computer programs are executed by one or more processors, they implement the steps of the pet monitoring method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Pet behavior adjustment and human-pet interaction method and device based on multimedia information technology
CN115104548A
Animal monitoring identification method, device and equipment and storage medium
CN116168415A
Intelligent park intelligent monitoring system based on artificial intelligence
CN117423061A
Picture capture detection method and device based on family abnormal sound recognition
CN119182884A
Pet safety management method and system, computer equipment and storage medium
US11700837B1