Pet recall device and automatic recall method

Through the integrated multi-module pet recall device, the dynamic response strategy generation and optimization is performed using deep learning and reinforcement learning algorithms, the problems of low recall success rate and poor adaptability in the existing technology are solved, and efficient recall effect in complex scenarios is achieved.

CN120052283AActive Publication Date: 2025-05-30SHAOXING SOUND TECH CO LTD
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
CN202510126170.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

The existing pet automatic recall system has low recall success rate in complex and changing environments, making it difficult to adapt to changes in the environment and pet state, and fails to make full use of multi-source perception data and advanced intelligent algorithms.

Method used

The pet recall device with integrated multi-module is adopted, including a situational awareness module, a behavior recognition and prediction analysis module, a reinforcement learning strategy module and a self-learning and incremental optimization module. It realizes data transmission and control command interaction through wireless communication protocols, and uses deep learning and reinforcement learning algorithms to generate and optimize dynamic response strategies.

Benefits of technology

It realizes intelligent decision-making and recall instruction execution in complex scenarios, significantly improves recall success rate and response time, and significantly improves the recall effect in high noise, low light and weak signal areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120052283A_ABST
    Figure CN120052283A_ABST
Patent Text Reader

Abstract

The invention discloses a pet recall device and an automatic recall method. The method comprises the steps of collecting related data of a pet and related data of an environment where the pet is located in real time; preprocessing the collected data to obtain multi-modal data; processing the multi-modal data by using the trained deep learning model to obtain a feature vector; processing the feature vector by using the trained long-short-term memory network to obtain the classification probability of the current behavior state and the prediction of the future behavior trend; an optimal recall strategy is generated through hierarchical reinforcement learning, and the strategy is executed by an execution component; and collecting parameters of the optimal recall strategy, designing a loss function according to the parameters and the difference of actual feedback, and updating parameters of the deep learning model, the long-short-term memory network and the hierarchical reinforcement learning by adopting an optimization algorithm. According to the invention, real-time analysis and prospective decision making of pet behaviors and environment situations are realized, and intelligent decision making and recall instruction execution can be realized in various complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of intelligent pet devices and Internet of Things technology, and particularly relates to a pet recall device and an automatic recall method. Background Art

[0002] Existing pet automatic recall systems mostly rely on a single triggering mechanism, such as GPS / GPRS positioning or simple sound stimulation. In complex and changeable real-world scenarios (such as high noise, weak signals, insufficient light, and easily changeable pet behaviors), the recall success rate of these systems is low, and it is difficult to adapt to changes in the environment and pet states. For example, in the Chinese invention patent with the publication number CN109446373B and the invention name "Pet semi-scattered management system and method based on Internet of Things platform", it mainly focuses on GPS / GPRS positioning and simple feedback, lacking prediction of future pet behavior trends and optimization of contextual strategies. The Chinese invention patent application with the publication number CN115883610A and the invention name "An intelligent pet recall system and method" only performs passive recall based on simple sounds and stimuli, lacking forward-looking decision-making and personalized optimization for complex situations.

[0003] Currently, these traditional solutions do not fully utilize multi-source perception data and advanced intelligent algorithms, are powerless against upcoming escape or anxiety behaviors, and have not achieved true dynamic optimization in terms of energy management and data security. Therefore, there is an urgent need for a pet automatic recall system with predictive context awareness and reinforcement learning-driven dynamic adaptive strategies to improve the recall success rate and user experience. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: to provide a pet recall device and an automatic recall method, which can achieve comprehensive perception and intelligent decision-making of pet behaviors, emotions, and the environment by integrating multiple modules, significantly improve the intelligence level and practicality of the recall system, and solve the problems of insufficient intelligence, poor environmental adaptability, and low recall success rate in the prior art.

[0005] To solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A pet recall device includes:

[0007] A collar strap, a housing cover, a rope threading opening, an execution component, a pet-side intelligent device, and a master-side control device.

[0008] The housing cover is provided with rope threading openings at both ends, and one end of two collar straps is fixed to the rope threading openings, and the other ends of the two collar straps are connected by a buckle.

[0009] The execution component includes a speaker, a vibration device, and an LED. The LED is arranged on the outer surface of the housing cover, and the speaker and the vibration device are arranged inside the housing cover.

[0010] The intelligent device on the pet side includes a context awareness module, a behavior recognition and prediction analysis module, a reinforcement learning strategy module, and a self-learning and incremental optimization module, all of which are located inside the outer shell.

[0011] The context awareness module, the behavior recognition and prediction analysis module, the reinforcement learning strategy module, and the self-learning and incremental optimization module are integrated onto the same hardware platform.

[0012] The context awareness module includes a data acquisition unit and a data processing unit.

[0013] Furthermore, the intelligent device on the pet side and the control device on the owner side achieve data transmission and control instruction interaction through a wireless communication protocol.

[0014] Furthermore, the context awareness module and the behavior recognition and prediction analysis module perform data transmission through an internal bus, the behavior recognition and prediction analysis module and the reinforcement learning strategy module perform data transmission through an internal bus, and the self-learning and incremental optimization module and the behavior recognition and prediction analysis module, the reinforcement learning strategy module perform data transmission through an internal bus.

[0015] Furthermore, the data acquisition unit includes a GPS module, a Bluetooth beacon, an inertial measurement unit, a light sensor, a temperature sensor, a humidity sensor, a noise sensor, a heart rate sensor, and a body temperature sensor.

[0016] Furthermore, the control device on the owner side is a smartphone APP or a smartwatch.

[0017] Furthermore, the present invention also proposes an automatic recall method for a pet recall device, including:

[0018] S1. Use the data acquisition unit of the context awareness module to collect the pet's positioning information, Bluetooth beacon signal strength, acceleration and angular velocity of the pet's movement, the pet's heart rate and body temperature, the noise decibel, light intensity, temperature and humidity of the environment where the pet is located in real time.

[0019] S2. Use the data processing unit of the context awareness module to preprocess the data collected by the data acquisition unit to obtain initial dynamic motion features, and calculate the statistical features of the data within a time window; use a sliding window mechanism to segment the initial dynamic motion features into time periods of a fixed length, and normalize the features to obtain multi-modal data.

[0020] S3. In the behavior recognition and prediction analysis module, use the loss function to train the parameters of the deep learning model and the long short-term memory network through the backpropagation algorithm to obtain the trained deep learning model and the long short-term memory network; use the trained deep learning model to process the multimodal data to obtain feature vectors; use the trained long short-term memory network to process the feature vectors to obtain the classification probability of the current behavior state and the prediction of the future behavior trend.

[0021] S4. In the reinforcement learning strategy module, dynamically generate the optimal recall strategy through hierarchical reinforcement learning, hand over the strategy to the execution component for execution, and perform real-time adjustment and optimization of the strategy according to the real-time response of the pet.

[0022] S5. In the self-learning and incremental optimization module, collect the parameters of the optimal recall strategy, design a loss function according to the difference between the parameters and the actual feedback, use the backpropagation algorithm to calculate the gradients of the deep learning model, the long short-term memory network and the hierarchical reinforcement learning parameters, and use the optimization algorithm to update the model parameters.

[0023] S6. Repeat steps S3 - S5 to obtain the final optimal recall strategy.

[0024] S7. The master control device displays the data collected in the context awareness module and the instruction notification of the optimal recall strategy, and manually adjusts the optimal recall strategy according to the actual situation.

[0025] Further, in step S2, the preprocessing includes: using low-pass filtering to remove high-frequency noise, using high-pass filtering to obtain the motion signals whose standard deviation within 1 second exceeds the preset threshold, and then performing denoising and alignment.

[0026] The statistical features include the mean and variance.

[0027] Further, in step S3, use the convolutional layer in the trained deep learning model to extract the local spatio-temporal features of the multimodal data through local convolution operations, and use the pooling layer to reduce the dimension to obtain a feature map. This feature map is converted into a feature vector through the fully connected layer. This vector includes the statistical features and dynamic features of each window; use the trained long short-term memory network to process the feature vectors, capture the information transfer between the front and back time windows and the dynamic changes of the pet behavior pattern, and then perform fusion and dimensionality reduction through the fully connected layer, and output the classification probability of the current behavior state and the prediction of the future behavior trend through the Softmax classification layer.

[0028] Further, in step S4, the hierarchical reinforcement learning includes an upper policy network and a lower policy network. The multi-modal data is input into the upper policy network to obtain a parameter search range. Based on this parameter search range, the lower policy network combines the prediction results of the behavior recognition and prediction analysis module, and uses the reinforcement learning algorithm to perform weighted scoring on different parameter combinations to obtain the expected recall success rate, response time, and energy consumption. The parameter combination with the highest weighted score or reaching the set threshold within the parameter search range is selected as the optimal recall strategy.

[0029] Based on the optimal recall strategy, when the entropy value of the prediction distribution > 0.5, the recall strategy is immediately updated; otherwise, the recall strategy is updated at a set frequency.

[0030] Weighted score = w1 × recall success rate - w2 × response time - w3 × average energy consumption.

[0031] The calculation formula for the entropy value H(P) of the prediction distribution is:

[0032]

[0033] where p(i) represents the probability of the i-th behavior category.

[0034] Further, it also includes step S8: Adjust the parameters of the models in the behavior recognition and prediction analysis module and the reinforcement learning strategy module according to the characteristics of different pets.

[0035] Compared with the prior art, the present invention adopts the above technical solutions and has the following technical effects:

[0036] 1. Through the predictive context awareness and reinforcement learning-driven dynamic response strategy, the present invention realizes the real-time analysis and forward-looking decision-making of pet behavior and environmental context, and can achieve intelligent decision-making and execute recall instructions in various complex scenarios.

[0037] 2. The response time of the present invention is short, ensuring that the pet can be recalled in time.

[0038] 3. The recall success rate of the present invention is significantly improved in complex environments such as high-noise environments, low-light environments, and weak-signal areas.

[0039] 4. The present invention first applies deep learning and adaptive strategies to the field of pet recall, resulting in technological breakthroughs and innovations. Brief Description of the Drawings

[0040] Figure 1 is the overall structure diagram of the present invention.

[0041] Figure 2 is the internal module structure diagram of the pet terminal device of the present invention.

[0042] Figure 3 This is the overall implementation flowchart of the present invention. Specific implementation manner

[0043] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.

[0044] To achieve the above object, the present invention proposes a pet recall device, as Figure 1 shown, including:

[0045] A collar band, a housing cover, a rope threading opening, an execution component, a pet-side intelligent device and a master-side control device that implement data transmission and control instruction interaction through a wireless communication protocol (such as Bluetooth, Wi-Fi). Communication between devices requires two-way authentication to prevent illegal device access and data tampering, ensuring the security and reliability of the system operation.

[0046] Rope threading openings are provided at both ends of the housing cover, and the rope threading openings are used to fix one end of two collar bands, and the other ends of the two collar bands are connected by buckles.

[0047] The execution component includes a speaker, a vibration device and an LED. The LED is arranged on the outer surface of the housing cover, and the speaker and the vibration device are arranged inside the housing cover. The selection of multiple execution components can provide multi-sensory feedback and enhance the perceptibility and response rate of the recall instruction.

[0048] The pet-side intelligent device includes a context awareness module, a behavior recognition and prediction analysis module, a reinforcement learning strategy module and a self-learning and incremental optimization module, all of which are located inside the housing cover.

[0049] As Figure 2 shown, the context awareness module, the behavior recognition and prediction analysis module, the reinforcement learning strategy module and the self-learning and incremental optimization module are integrated onto the same hardware platform (such as the same main control board). The context awareness module and the behavior recognition and prediction analysis module perform data transmission through an internal bus (such as SPI, I 2 C), the behavior recognition and prediction analysis module and the reinforcement learning strategy module perform data transmission through an internal bus (such as SPI, I 2 C), and the self-learning and incremental optimization module and the behavior recognition and prediction analysis module, the reinforcement learning strategy module perform data transmission through an internal bus (such as SPI, I 2 C).

[0050] The context awareness module includes a data acquisition unit and a data processing unit; among them, the data acquisition unit includes a GPS module, a Bluetooth beacon, an inertial measurement unit (accelerometer and gyroscope), a light sensor, a temperature sensor, a humidity sensor, a noise sensor (microphone), a heart rate sensor, and a body temperature sensor.

[0051] The behavior recognition and prediction analysis module is used to process motion features using a deep learning model, such as a convolutional neural network, to obtain behavior recognition results and future behavior trends.

[0052] The reinforcement learning policy module is used to, based on the results of the behavior recognition and prediction analysis module, use hierarchical reinforcement learning to find the most suitable recall policy for the current situation, including recall content, volume, frequency, and feedback method, and hand over this policy to the execution component for execution.

[0053] The master control device is a smartphone APP or a smartwatch, so as to monitor the pet's status at any time, view the recall progress, and perform manual intervention.

[0054] The smartphone APP is used to display information such as the pet's location, behavior status, and physiological data; allows users to set parameters such as recall sound content, volume, and frequency; records the pet's behavior patterns and recall history for users to view and analyze; users can send instant recall instructions through the APP.

[0055] The smartwatch is used to provide real-time notifications of pet status changes and recall instructions; allows users to quickly adjust the recall policy while wearing the watch.

[0056] As Figure 3 shown, the method for automatically recalling a pet using a pet recall device includes:

[0057] S1. Use the data acquisition unit of the context awareness module to collect the pet's positioning information, Bluetooth beacon signal strength, acceleration and angular velocity of the pet's movement, the pet's heart rate and body temperature, the noise decibels, light intensity, temperature and humidity of the environment where the pet is located in real time.

[0058] S2. Use the data processing unit of the context awareness module to preprocess the data collected by the data acquisition unit to obtain initial dynamic motion features, specifically including: using low-pass filtering to remove high-frequency noise, using high-pass filtering to obtain motion signals whose standard deviation within 1 second exceeds a preset threshold, such as features of the pet running, jumping, etc., so as to better identify and distinguish different behavior patterns (such as stationary vs. accelerating and escaping), and then perform denoising and alignment.

[0059] Calculate the statistical features (including mean and variance) of the data within a time window. Use the sliding window mechanism to segment the initial dynamic motion features into time periods of fixed length (such as a 2-second window with a 1-second step size) to ensure the balance between real-time performance and data continuity, and perform normalization processing on this feature to obtain multi-modal data.

[0060] In the behavior recognition and prediction analysis module, use loss functions (such as cross-entropy or mean squared error) to train the parameters of the convolutional neural network and the long short-term memory network through the backpropagation algorithm to obtain the trained convolutional neural network and long short-term memory network. Use the trained convolutional neural network to process the multi-modal data to obtain feature vectors; use the trained long short-term memory network to process the feature vectors to obtain the classification probability of the current behavior state and the prediction of future behavior trends. The specific content is as follows:

[0061] Use the convolutional layer in the trained convolutional neural network to extract the local spatio-temporal features of the multi-modal data through local convolution operations, and use the pooling layer to reduce the dimension to obtain a feature map. This feature map is converted into a feature vector through a fully connected layer. This vector includes the statistical features and dynamic features of each window; use the trained long short-term memory network to process the feature vectors to capture long-term dependencies (information transfer between front and back time windows) and the dynamic changes of pet behavior patterns, and then perform fusion and dimensionality reduction through a fully connected layer. Output the classification probability of the current behavior state and the prediction of future behavior trends through the Softmax classification layer to accurately classify various behavior states of pets, such as static rest, free movement, wandering anxiety, escape tendency, and abnormal behaviors (illness, fatigue, cold), and identify potential escape or anxiety behaviors in advance.

[0062] In the reinforcement learning strategy module, dynamically generate the optimal recall strategy through hierarchical reinforcement learning, hand over this strategy to the execution component for execution, and perform real-time adjustment and optimization of the strategy according to the pet's real-time response (such as whether it returns to the owner's side). The specific content is as follows:

[0063] Hierarchical reinforcement learning includes an upper-layer policy network and a lower-layer policy network. Multimodal data (such as noise decibel, light intensity, GPS signal strength, uncertainty of behavior prediction, etc.) is input into the upper-layer policy network to obtain a parameter search range (such as volume 50% - 80%, frequency 1.5 kHz - 2 kHz, LED blinking frequency 1 - 5 times per second, etc.). Based on this parameter search range, the lower-layer policy network combines the prediction results of the behavior recognition and prediction analysis module, and uses a reinforcement learning algorithm (such as Q-learning or Policy Gradient) to perform weighted scoring on different parameter combinations, estimate the balance of recall success rate - energy consumption - response time, and obtain the expected recall success rate, response time, and energy consumption. The parameter combination with the highest weighted score or reaching 0.8 within the parameter search range is selected as the optimal recall strategy.

[0064] Based on the optimal recall strategy, when the entropy value of the prediction distribution > 0.5, the recall strategy is immediately updated; otherwise, the recall strategy is updated once every 5 seconds.

[0065] Weighted score = w1 × recall success rate - w2 × response time - w3 × average energy consumption.

[0066] The calculation formula for the entropy value H(P) of the prediction distribution is:

[0067]

[0068] Among them, p(i) represents the probability of the i-th behavior category. The higher the entropy value, the greater the uncertainty of the model.

[0069] In this embodiment, in the optimal recall strategy, the recall success rate ≥ 80%, the response time ≤ 2 seconds, the average energy consumption < 0.1% * total battery capacity, the volume = 70%, the frequency = 1.8 kHz, the LED blinks 3 times per second, and the vibration intensity is medium. w1 = 0.5, w2 = 0.3, w3 = 0.2.

[0070] S5. In the self-learning and incremental optimization module, collect the parameters of the optimal recall strategy, design a loss function (such as cross-entropy loss, mean squared error, etc.) according to the difference between the parameters and the actual feedback (including the user's manual recall operation, recall success rate, and pet response data), calculate the gradients of the deep learning model, long short-term memory network, and hierarchical reinforcement learning parameters using the backpropagation algorithm, use an optimization algorithm (such as Adam, SGD) to update the model parameters, complete online or offline fine-tuning, and after using the validation set or real-time monitoring the effect of the new strategy, confirm whether the update improves the system performance. If effective, apply the updated model.

[0071] When a large-scale data distribution change is detected, or when the system policy effect significantly decreases, a large-scale model update will be triggered (which can be carried out at night or when the battery is fully charged).

[0072] S6. Repeat steps S3 - S5 to obtain the final optimal recall strategy.

[0073] S7. The host - side control device displays the data collected by the context - awareness module and the instruction notification of the optimal recall strategy, and manually adjusts the optimal recall strategy according to the actual situation.

[0074] S8. According to the behavior habits and physiological characteristics of different pets, dynamically adjust and optimize the parameters of the models in the behavior recognition and prediction analysis module and the reinforcement learning strategy module (such as the convolution kernel weights of the convolutional neural network, the gating parameters of the long - short - term memory network, the value function or policy function of hierarchical reinforcement learning), and finally change the behavior recognition output and recall decision to ensure that the system shows higher accuracy and reliability during long - term use.

[0075] In the embodiment:

[0076] High - noise environment: In an environment with high noise, the reinforcement learning strategy module automatically increases the volume of the recalled sound or adjusts the audio frequency to overcome the interference of environmental noise and ensure the effective transmission of the recall instruction.

[0077] Low - light environment: At night or under low - light conditions, the reinforcement learning strategy module combines LED flashing prompts and sound recall to enhance the pet's perception of the recall instruction and avoid recall failures caused by insufficient visual recognition. Specifically: If the light intensity < 5 lx (extremely dark), the LED flashing frequency = 5 times / second, and the volume gain = + 30%; if 5 lx ≤ light intensity < 10 lx (dim), the LED flashing frequency = 3 times / second, and the volume gain = + 20%; if the light intensity > 10 lx (relatively bright), the normal strategy is maintained (LED frequency = 1 time / second or off, volume gain = 0% - 10%).

[0078] Weak - signal area: In an area where the GPS signal is weak or lost, the context - awareness module automatically enables Bluetooth beacon - assisted positioning and performs position compensation through an inertial measurement unit to improve the positioning accuracy and ensure reliable recall in indoor or areas with complex signals. Specifically: When the GPS signal strength < - 120 dBm or the number of consecutive packet losses is N times (such as 3 times), the inertial measurement unit + Bluetooth positioning compensation is enabled. In a short - distance range, the relative displacement can be calculated by integrating the acceleration and angular velocity output by the inertial measurement unit, and the distance between the collar and the Bluetooth base station (mobile phone) can be estimated by combining the Bluetooth wireless received signal strength for multi - point positioning fusion. If the Bluetooth wireless received signal strength > - 70 dBm, the positioning error can be controlled within 2 - 5 meters; when the detected acceleration bias exceeds 0.1 g or the gyroscope drift exceeds ± 0.5° / s, real - time correction is required.

[0079] In a high-noise environment, the recall success rate of the present invention is increased to 88% ± 3%; in a low-light environment, it is increased to 83% ± 2%; in a weak-signal area, it is increased to 80% ± 2%. The average response time is reduced to 1.5 seconds.

[0080] It can also perform expandability design of modules, specifically:

[0081] Support the future expansion of more sensors and feedback mechanisms, such as odor emitters, temperature adjustment devices, etc., to enhance the functions and applicability of the system.

[0082] Remote update and maintenance: Through OTA (Over-The-Air) technology, it supports remote software updates and maintenance to ensure that the system always runs the latest functions and optimizations.

[0083] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the technical principle of the present invention, several improvements and modifications can still be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. An automatic recall method for a pet recall device, characterized in that: include: S1. Using the data acquisition unit of the situational awareness module to collect the pet's location information, Bluetooth beacon signal strength, acceleration and angular velocity of the pet's movement, the pet's heart rate and body temperature, noise decibels, light intensity, temperature and humidity of the pet's environment in real time; S2, using the data processing unit of the situational awareness module to pre-process the data collected by the data collection unit to obtain initial dynamic motion features, and calculate the statistical features of the data within the time window; The sliding window mechanism is used to divide the initial dynamic motion features into time periods of fixed length, and the features are normalized to obtain multimodal data; S3. In the behavior recognition and prediction analysis module, the parameters of the deep learning model and the long short-term memory network are trained by the back propagation algorithm using the loss function to obtain the trained deep learning model and the long short-term memory network; the multimodal data are processed by the trained deep learning model to obtain the feature vector; the feature vector is processed by the trained long short-term memory network to obtain the classification probability of the current behavior state and the prediction of the future behavior trend; S4. In the reinforcement learning strategy module, the optimal recall strategy is generated through hierarchical reinforcement learning, and the strategy is handed over to the execution component for execution, and the strategy is adjusted and optimized in real time according to the real-time response of the pet; S5. In the self-learning and incremental optimization module, collect the parameters of the optimal recall strategy, design the loss function according to the difference between the parameters and the actual feedback, use the back propagation algorithm to calculate the gradient of the deep learning model, long short-term memory network and hierarchical reinforcement learning parameters, and use the optimization algorithm to update the model parameters; S6, repeat steps S3-S5 to obtain the final optimal recall strategy; S7. The host-side control device displays the data collected in the situational awareness module and the instruction notification of the optimal recall strategy, and manually adjusts the optimal recall strategy according to the actual situation.

2. The automatic recall method of the pet recall device according to claim 1, characterized in that: In step S2, the preprocessing includes: using low-pass filtering to remove high-frequency noise, using high-pass filtering to obtain motion signals whose standard deviation of the signal exceeds a preset threshold within 1 second, and then performing denoising and alignment; Statistical characteristics include mean and variance.

3. The automatic recall method of the pet recall device according to claim 2, characterized in that: In step S3, the convolution layer in the trained deep learning model is used to extract the local spatiotemporal features of the multimodal data through local convolution operations, and the pooling layer is used to reduce the dimension to obtain a feature map, which is converted into a feature vector through a fully connected layer. The vector includes statistical features and dynamic features of each window; The trained long short-term memory network is used to process the feature vector to capture the information transmission between the previous and next time windows and the dynamic changes of the pet's behavior pattern. It is then fused and reduced in dimension through the fully connected layer, and the classification probability of the current behavior state and the prediction of future behavior trends are output through the Softmax classification layer.

4. The automatic recall method of the pet recall device according to claim 1, characterized in that: In step S4, hierarchical reinforcement learning includes an upper strategy network and a lower strategy network. Multimodal data is input into the upper strategy network to obtain a parameter search range. The lower strategy network uses a reinforcement learning algorithm to weightedly score different parameter combinations based on the parameter search range and the prediction results of the behavior recognition and prediction analysis module to obtain the expected recall success rate, response time and energy consumption. The parameter combination with the highest weighted score or reaching the set threshold within the parameter search range is selected as the optimal recall strategy. Based on the optimal recall strategy, when the entropy value of the predicted distribution is greater than 0.5, the recall strategy is updated immediately; otherwise, the recall strategy is updated at the set frequency; Weighted score = w1 × recall success rate - w2 × response time - w3 × average energy consumption; The calculation formula for the predicted distribution entropy value H(P) is: Among them, p(i) represents the probability of the i-th behavior category.

5. The automatic recall method of the pet recall device according to claim 1, characterized in that: The method also includes step S8: adjusting the parameters of the models in the behavior recognition and prediction analysis module and the reinforcement learning strategy module according to the characteristics of different pets.

6. The pet recall device according to claim 1, characterized in that: include: Collar belt, housing cover, leash port, actuator, pet-side smart device and owner-side control device; Rope insertion openings are arranged at both ends of the outer shell, and one end of two collar straps is fixed at the rope insertion openings, and the other ends of the two collar straps are connected by buckles; The actuator comprises a speaker, a vibration device and an LED, wherein the LED is arranged on the outer surface of the outer shell, and the speaker and the vibration device are arranged inside the outer shell; The pet-side smart device includes a situational awareness module, a behavior recognition and prediction analysis module, a reinforcement learning strategy module, and a self-learning and incremental optimization module, all of which are located inside the outer shell; The situational awareness module, behavior recognition and prediction analysis module, reinforcement learning strategy module, and self-learning and incremental optimization module are integrated into the same hardware platform; The situation awareness module includes a data acquisition unit and a data processing unit.

7. The pet recalling device according to claim 6, characterized in that: The pet-side smart device and the owner-side control device realize data transmission and control command interaction through wireless communication protocols.

8. The pet recalling device according to claim 6, characterized in that: The situational awareness module and the behavior recognition and prediction analysis module transmit data through the internal bus, the behavior recognition and prediction analysis module and the reinforcement learning strategy module transmit data through the internal bus, and the self-learning and incremental optimization module and the behavior recognition and prediction analysis module and the reinforcement learning strategy module transmit data through the internal bus.

9. The pet recalling device according to claim 6, characterized in that: The data acquisition unit includes a GPS module, a Bluetooth beacon, an inertial measurement unit, a light sensor, a temperature sensor, a humidity sensor, a noise sensor, a heart rate sensor, and a body temperature sensor.

10. The pet recalling device according to claim 6, characterized in that: The host-side control device is a smartphone APP or a smart watch.

Citation Information

Patent Citations

  • A Pet Semi-Free-Range Management System and Method Based on an Internet of Things Platform

    CN109446373B

  • Intelligent pet recall system and method

    CN115883610A

  • Dynamic adaptive virtual reality dog training method and system, and storage medium

    CN118435880A

  • Smart bowl system, apparatus and method

    US10091972B1

  • Apparatus and Method for the Virtual Fencing of an Animal

    US20080035072A1