Pet food adding method and device, pet feeder and storage medium
By collecting environmental video data in the pet feeder and combining image segmentation and object detection models, we predict the amount of food intake of pets and generate accurate food addition plans, solving the problem of inaccurate food addition of pet feeding equipment and achieving more efficient food management.
Patent Information
- Application Number
- CN202510913170.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
AI Technical Summary
Existing pet feeding equipment lacks accuracy during the food addition process, and cannot flexibly deal with changes in pet food consumption, resulting in food accumulation or insufficient food.
By setting up a camera device in the pet feeder to collect environmental video data, filter out pet feeding video clips, combine feeding pot images and weighing data, use image segmentation and object detection models to predict the amount of pet food, and generate an accurate food addition plan.
Improves the accuracy of food addition, avoids food accumulation or inadequacy, and improves the flexibility and accuracy of pet feeding.
Smart Images

Figure CN120391353A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of pet intelligent technologies, and particularly relates to a pet food adding method, device, pet feeder and storage medium. Background Art
[0002] In recent years, with the continuous acceleration of the pace of life, people's life pressure has also been continuously increasing. The company of pets can, to a great extent, relieve people's mental stress and bring physical and mental pleasure to people, which has led to the rapid development of the pet supplies market in recent years.
[0003] Among them, the pet feeding equipment in pet supplies can automatically add food and water for pets by setting pet feeding plans, thereby reducing the pet feeding pressure of pet owners and enhancing the happiness of pet raising.
[0004] However, in the related art, the pet feeding equipment can only mechanically add food based on a specified pet feeding plan, resulting in insufficient accuracy of food addition. Summary of the Invention
[0005] The main purpose of the embodiments of the present application is to propose a pet food adding method, device, pet feeder and storage medium, aiming to improve the accuracy of pet feeding.
[0006] To achieve the above object, a first aspect of the embodiments of the present application proposes a pet food adding method, which is applied to a pet feeder, and the method includes: Obtain the environmental video data corresponding to the pet feeder, and screen out the pet eating video segments from the environmental video data; Extract the corresponding feeding bowl image sequence from the pet eating video segments; Predict the pet food intake according to the feeding bowl image sequence; Generate a pet food adding plan based on the pet food intake, and add pet food according to the pet food adding plan.
[0007] To achieve the above object, a second aspect of the embodiments of the present application proposes a pet food adding device, and the device includes: An obtaining unit, configured to obtain the environmental video data corresponding to the pet feeder, and screen out the pet eating video segments from the environmental video data; An extracting unit, configured to extract the corresponding feeding bowl image sequence from the pet eating video segments, and obtain the feeding bowl weighing data corresponding to each pet eating video segment; A predicting unit, configured to predict the pet food intake according to the feeding bowl image sequence and the feeding bowl weighing data; A generating unit is used to generate a pet food addition plan based on the pet's food intake, and add pet food according to the pet food addition plan.
[0008] Optionally, in some embodiments, the acquiring unit includes: an acquisition subunit, configured to acquire environmental video data captured by a camera device in the pet feeder; a detection subunit, configured to perform object detection on the environmental video data and filter out pet behavior video clips based on the object detection results; The screening subunit is used to perform behavior category recognition on the pet behavior video clips, and screen out the pet eating video clips according to the behavior category recognition results.
[0009] Optionally, in some embodiments, the detection subunit includes: A denoising module, configured to denoise the environmental video data to obtain a denoised video; A detection module, configured to perform motion detection on the denoised video based on a motion detection model, and extract motion video segments from the denoised video according to the motion detection result; The screening module is used to perform object detection on the objects in the motion video clips based on the object detection model, and screen out the pet behavior video clips from the motion video clips according to the object detection results.
[0010] Optionally, in some embodiments, the prediction unit includes: An extraction subunit, configured to extract key frames corresponding to a pet eating start node and a pet eating end node from the feeding bowl image sequence; an estimating subunit, configured to send the key frame to a pet management server to estimate the pet's food intake, and receive a first pet food intake prediction value returned by the pet management server; a calculation subunit, configured to obtain feeding bowl weighing data corresponding to each pet eating video clip, and calculate a second pet food intake prediction value based on the feeding bowl weighing data; The determining subunit is configured to determine the pet's food intake amount based on the first pet's food intake amount prediction value and the second pet's food intake amount prediction value.
[0011] Optionally, in some embodiments, the present application further provides a pet food intake prediction device, comprising: a segmentation subunit, configured to perform semantic segmentation on the key frame based on an image segmentation model, and determine the food area according to the semantic segmentation result; The prediction subunit is configured to predict a first pet food intake prediction value based on changes in the food area.
[0012] Optionally, in some embodiments, the prediction subunit includes: A determination module, configured to determine the corresponding pixel area of the food area and the stacking height corresponding to each pixel; A first calculation module, configured to calculate the food volume corresponding to each key frame according to the pixel area and the stacking height; A second calculation module, configured to calculate the change amount of the food volume, and determine a first pet food intake prediction value according to the change amount.
[0013] Optionally, in some embodiments, the determination subunit includes: An acquisition module, configured to acquire a first confidence level corresponding to the first pet food intake prediction value, and acquire a second confidence level corresponding to the second pet food intake prediction value; A determination module, configured to determine the pet food intake according to the first confidence level, the second confidence level, the first pet food intake prediction value, and the second pet food intake prediction value.
[0014] To achieve the above object, a third aspect of the embodiments of the present application provides a pet feeder, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the pet food addition method described in the first aspect is implemented.
[0015] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the pet food addition method described in the first aspect is implemented.
[0016] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer program product, which includes a computer program, and the computer program is read and executed by a processor of a computer device, so that the computer device executes the pet food addition method described in the first aspect.
[0017] The pet food addition method provided by the embodiments of the present application is applied to a pet feeder. The method includes: acquiring environmental video data corresponding to the pet feeder, and screening out pet eating video segments from the environmental video data; extracting a corresponding feeding bowl image sequence from the pet eating video segments; predicting the pet food intake according to the feeding bowl image sequence; generating a pet food addition plan based on the pet food intake, and adding pet food according to the pet food addition plan.
[0018] It can be seen from this that the pet food adding method provided by the embodiments of the present application can collect environmental videos of the pet feeder by setting a video data collection device near the pet feeder, and then comprehensively determine the accurate pet food intake based on the image difference of the food bowl before and after the pet eats and the weighing difference of the food bowl, and then add food based on the accurate pet food intake, which can improve the accuracy of food addition. Description of the Drawings
[0019] The drawings are used to provide a further understanding of the technical solutions of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solutions of the present application, and do not constitute a limitation to the technical solutions of the present application.
[0020] Figure 1 It is a schematic structural diagram of the system to which the pet food adding method provided by the present application is applied; Figure 2 It is a schematic flowchart of the pet food adding method provided by the present application; Figure 3 It is a schematic structural diagram of the pet food adding device provided by the present application; Figure 4 It is a schematic hardware structure diagram of the pet feeder provided by the embodiments of the present application. Detailed Embodiments
[0021] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0022] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations: Motion detection model: The motion detection model is one of the important technologies in the fields of computer vision and video analysis, and is mainly used to identify moving targets in videos or image sequences.
[0023] Image Segmentation Model: Image segmentation is the technology and process of dividing an image into several specific regions with unique properties and extracting the target of interest. It is a key step from image processing to image analysis. Existing image segmentation methods are mainly classified into the following categories: threshold-based segmentation methods, region-based segmentation methods, edge-based segmentation methods, and segmentation methods based on specific theories, etc. From a mathematical perspective, image segmentation is the process of dividing a digital image into non-overlapping regions. The process of image segmentation is also a labeling process, that is, pixels belonging to the same region are assigned the same number. An image segmentation model refers to using deep learning technology for image segmentation, such as Fully Convolutional Network (FCN), U-Net, Mask R-CNN, Transformer-based models.
[0024] Object Detection Model: Object detection aims to enable a computer to understand the content of an image, determine which objects exist in it, and mark their positions and categories. An object detection model refers to using artificial intelligence-based image processing technology to identify objects in a video or image.
[0025] In order to reduce the pressure of pet-raising for pet owners and bring a better pet-raising experience to them, various intelligent pet-raising devices have been continuously introduced in the pet market, such as intelligent pet feeders. An intelligent pet feeder can add food regularly and quantitatively according to the pet feeding plan set by the pet owner. In this way, it can not only avoid the pet-raising pressure brought by repeated food addition to the pet owner, but also avoid problems such as pet physical discomfort caused by forgetting to add food. However, the food intake of pets will change with various factors such as weather, pet physical condition, age, and pet mood. The method of adding pet food regularly and quantitatively according to a fixed plan lacks flexibility. For example, when a pet's gastrointestinal discomfort causes a decrease in food intake, if food is still added according to the set feeding plan, it may lead to the accumulation of food in the feeding bowl, and then the food gets damp, reducing the food quality and further affecting the pet's appetite and other problems. Or, when the pet grows older and its body size increases, but if the feeding plan is not adjusted in time, there may be a problem of insufficient feeding amount, which may further lead to bad habits such as the pet rummaging for food.
[0026] To solve the problems of food accumulation or food shortage caused by adding food according to the fixed pet feeding plan set by the pet owner, a scheme for intelligently regulating the feeding amount based on the pet's food intake has been proposed in related technologies. This scheme sets a weighing module for the feeding bowl of the pet feeder to measure the remaining amount in the pet feeder, and then calculates the pet's food intake based on the change in the weighed weight, and further adjusts the feeding amount. However, the weighing-based method will also have problems of inaccurate prediction caused by issues such as the pet owner actively adding food, the pet eating irregularly, the food being accidentally damp, and food overflow caused by environmental vibration.
[0027] Based on this, in order to improve the accuracy of estimating the food intake of pets, and further improve the accuracy of adding pet food. This application provides a pet food adding method, aiming to achieve accurate addition of pet food and avoid problems such as the decline in quality caused by the accumulation of pet food or the insufficient food affecting the growth of pets. Next, the pet food adding method provided by the embodiments of this application will be described.
[0028] Refer to Figure 1 , which is a schematic diagram of the system to which the pet food adding method provided by this application is applied. As shown in the figure, the system to which the pet food adding method provided by this application is applied includes a pet feeder 110, a pet management client 130, and a pet management server 120.
[0029] Among them, the pet feeder 110 can be a feeder for pets of any species and breed. For example, when the pet is a cat, the pet feeder 110 is an automatic cat food feeder; when the pet is a dog, the pet feeder 110 is an automatic dog food feeder, etc. The pet feeder 110 is equipped with a camera device, a feeding bowl, and a weighing device for weighing the feeding bowl. The camera device can collect the environmental video near the pet feeder 110, especially for real-time video collection at the feeding bowl in the pet feeder. The weighing device can collect the weight data of the feeding bowl and save the weighing data together with the corresponding weighing time.
[0030] The pet management client 130 is specifically a terminal loaded with a pet management application. This terminal can be a mobile terminal such as a smart phone, or various forms of terminals such as a tablet, a personal computer, a head-mounted device, or a vehicle-mounted terminal. In addition to managing the pet feeder 110, the pet management application can also manage other pet devices such as pet toilet devices and pet toy devices.
[0031] The pet management server 120 can be a server that provides management, data processing and interaction for multiple pet management application clients, and analyzes pet behavior data, etc. This server can be a physical server or a cloud server.
[0032] In the embodiments of the present application, a pet owner can set the feeding plan of the pet feeder 110 through the pet management client 130. For example, it can be set to add food according to a fixed time and quantity, or it can be set to an intelligent pet food addition plan. When the pet owner sets the feeding plan of the pet feeder 110 to an intelligent pet food addition plan, the pet feeder 110 continuously collects the environmental video data corresponding to the pet feeder, and then screens out the pet eating video segments from the environmental video data. Further, the pet feeder 110 can also extract the corresponding feeding bowl image sequence from the pet eating video segments, and further obtain the feeding bowl weighing data corresponding to each pet eating video segment. Then the pet feeder 110 can extract key frames from the feeding bowl image sequence, and send the key frames and the feeding bowl weighing data to the pet management server 120 for predicting the pet food intake. Then the pet management server 120 sends the predicted pet food intake to the pet feeder 110, and the pet feeder 110 generates a corresponding pet food addition plan and adds pet food according to the pet food addition plan.
[0033] It can be understood that the pet management server 120 can be connected to multiple pet feeders 110 and multiple pet management clients 130. The three pet management clients 130 shown in the figure are only for illustration and do not limit the number of pet feeders 110 and pet management clients 130.
[0034] Refer to Figure 2 , which is a schematic flowchart of the pet food addition method provided by the present application.
[0035] In some embodiments, the pet food addition method provided by the embodiments of the present application can be applied to Figure 1 the pet food addition system shown, and the pet food addition method includes but is not limited to steps S201 to S204.
[0036] Step S201, obtain the environmental video data corresponding to the pet feeder, and screen out the pet eating video segments from the environmental video data.
[0037] Among them, the pet food addition method provided by the embodiments of the present application is applied to the pet feeder provided by the present application. The pet feeder is equipped with a camera device, a feeding bowl, a grain storage tank, a weighing device for the feeding bowl, and a processor. The pet food addition method provided by the embodiments of the present application can be specifically applied to this processor.
[0038] In the embodiments of the present application, the pet feeder can first adopt the pet food adding method provided by the present application to determine an accurate pet food adding plan, and then add pet food according to the accurate food adding plan. Specifically, the environmental video data corresponding to the pet feeder can be collected first. The environmental video data here can be the video data collected by the imaging device since the last food addition. The imaging device can be used to fixedly capture the video at the position of the feeding bowl in the pet feeder, or the imaging device can also cooperate with the rotation module to collect video data in all directions of the pet feeder.
[0039] Among the environmental video data collected by the imaging device, there may be videos containing some body parts of the pet owner (such as when the pet owner manually adds or replaces pet food, or cleans the feeding bowl), videos containing pet eating videos, pet playing videos (the pet plays near the feeder without eating), or videos containing neither pets nor pet owners. Since the purpose of the method provided by the present application is to obtain the accurate pet food intake and then perform accurate pet food addition, the pet eating videos can be extracted from the environmental video data for analysis.
[0040] Specifically, obtaining the environmental video data corresponding to the pet feeder and screening out the pet eating video segments from the environmental video data includes: Obtaining the environmental video data collected by the imaging device in the pet feeder; Performing object detection on the environmental video data, and screening out the pet behavior video segments according to the object detection results; Performing behavior category recognition on the pet behavior video segments, and screening out the pet eating video segments according to the behavior category recognition results.
[0041] In the embodiments of the present application, after obtaining the environmental video data collected by the imaging device, object detection can be first performed on the environmental video data to detect the object information included in each video frame of the environmental video data. For example, objects such as the pet owner, pet, toy, and feeding bowl are included. When a pet is detected in the video, the pet category and pet identity can also be recognized. Since different types of pets are usually fed with different pet feeders, after performing object detection on the environmental video data to obtain the object detection results, the pet behavior video segments corresponding to the pet type corresponding to the pet feeder can be further screened out.
[0042] For example, when the pet feeder feeds cat food, if pets such as dogs and ducks are detected from the environmental video data, the video segments of these pets can be removed, and only the video segments containing cats are retained to obtain pet behavior video segments. For example, in some multi-pet families, such as families with not only pet cats but also pet dogs, if the pet dog accidentally eats the cat food of the pet cat, it will cause an abnormal excessive reduction in the cat food. If the pet food intake is determined only based on weighing and then cat food is added based on this intake, too much cat food will be added, which may lead to cat food accumulation and may also cause food spoilage over time. Based on the method provided in this application, by performing object detection on the video, when a pet is detected in the video, further pet identity recognition is performed, and then the video data of a specific pet is screened out to estimate the pet food intake, and the actual food intake of the pet can be accurately calculated, thereby improving the accuracy of feeding.
[0043] In some embodiments, object detection is performed on the environmental video data, and pet behavior video segments are screened out according to the object detection results, including: Denoise the environmental video data to obtain a denoised video; Perform motion detection on the denoised video based on a motion detection model, and extract motion video segments from the denoised video according to the motion detection results; Perform object detection on the objects in the motion video segments based on an object detection model, and screen out pet behavior video segments from the motion video segments according to the object detection results.
[0044] In the embodiments of this application, in order to improve the accuracy and efficiency of extracting pet eating video segments, after obtaining the environmental video data, the obtained environmental video data can be denoised first. Specifically, traditional image processing techniques, such as Gaussian filtering and median filtering, can be used to process the video frames to remove the influence of random pixel interference and uneven illumination. In addition, the intensity distribution of adjacent pixels can also be analyzed, and the pixel points significantly different from the surrounding environment are identified as noise points for removal, obtaining high-quality denoised video data.
[0045] Then, in order to reduce the data volume of object detection and improve the efficiency of object detection, the denoised video can first be subjected to motion detection using a simple motion detection AI model installed in the pet feeder. This motion detection model takes a video frame sequence as input and outputs a motion detection mask, which is used to indicate which regions in the video frame image have motion. Through this process, video frame images with significant scene changes can be identified, and redundant frames with no change or minor changes can be deleted to obtain a motion video clip, thereby effectively reducing the data volume and computational burden of subsequent processing. In some embodiments, motion detection can also be assisted by calculating inter-frame difference values (such as the Structural Similarity Index (SSIM) and the Mean Absolute Difference (MAD)) to obtain a more accurate motion detection result.
[0046] After the motion video clip is extracted, an object detection model can be further used to perform object detection on the motion video clip, and then pet behavior video clips can be screened out from the motion video clip according to the object detection results.
[0047] In the embodiments of the present application, since denoising the environmental video data can be achieved by using traditional image processing techniques without invoking a neural network model for processing, the resource consumption in the overall processing process is relatively small. To avoid excessive resource consumption when subsequently invoking the motion detection model and the object detection model for video detection, the environmental video data can be denoised first to avoid the consumption of detection resources by noise data and improve the detection efficiency. In addition, the process of performing motion detection on the denoised video based on a simple motion detection AI model requires relatively fewer detection resources compared to complex object detection. Therefore, the denoised video can first be subjected to motion detection using a simple AI model, and video clips with motion can be screened out according to the motion detection results and enter the next step of object detection, avoiding a large number of invalid videos without motion from entering the object detection model and wasting the computational resources of the object detection model.
[0048] Furthermore, the pet behavior video clip can contain video clips corresponding to various behaviors of the pet, and the present application only focuses on the food intake of the pet, that is, only the eating video of the pet is required. Thus, after the pet behavior video clip is screened out, the behavior category of the pet behavior video clip can be further identified. Specifically, a pet behavior category recognition model can be used to identify the behavior category of the pet behavior video clip, and then the pet eating video clip can be screened out from the pet behavior video clip according to the behavior category recognition result.
[0049] In the embodiments of the present application, the pet behavior category recognition model can be a neural network model trained separately for different types of pets. For example, for pets of different species such as cats, dogs, monkeys, and birds, the corresponding pet behavior category recognition models can be trained separately. Then, when it is necessary to recognize the behavior category of a pet behavior video segment, the species of the pet can be determined first, and then the pet behavior category recognition model corresponding to the species can be selected to recognize the pet behavior category, so as to improve the accuracy of pet behavior category recognition. Further, for different breeds of pets of the same species, the corresponding pet behavior category recognition models can be further trained separately. For example, for different breeds of cats such as British Shorthair, American Shorthair, Orange Cat, Persian Cat, and Ragdoll Cat, the pet behavior category recognition models corresponding to each breed can be trained separately. When it is necessary to recognize the behavior category of a pet behavior video segment, the species and breed of the pet can be determined first, and then the pet behavior category recognition model corresponding to the breed can be selected to recognize the pet behavior video segment to obtain the behavior category recognition result.
[0050] After obtaining the behavior category recognition for the pet behavior video segment, the video segments whose behavior category is not the eating behavior can be removed, so as to leave the pet eating video segments. The pet eating video segments can be one video segment or multiple video segments. The video segments not only contain the video image frame sequence but also contain the time data corresponding to each video image frame.
[0051] Step S202: Extract the corresponding feeding bowl image sequence from the pet eating video segments.
[0052] During the pet's eating process, the imaging device keeps collecting videos of the feeding bowl in the pet feeding device. Since the pet can eat from different angles, at some angles, the pet may block the feeding bowl during eating, causing the imaging device to be unable to capture the feeding bowl and only able to capture the pet. And the images that only capture the pet but not the feeding bowl are not helpful for analyzing the change of food during the pet's eating process. At this time, these image frames can be removed to obtain the feeding bowl image sequence. This can further reduce the amount of data to be processed subsequently and improve the accuracy of pet food intake prediction. Among them, the feeding bowl image not only captures the feeding bowl but also can include some areas near the feeding bowl. In this way, if the pet has abnormal eating behavior and spills the food outside the feeding bowl, the remaining food in the feeding bowl before eating, the remaining food in the feeding bowl after eating, and the spilled food amount can be detected through the feeding bowl image sequence, and then the accurate food intake can be calculated.
[0053] Step S203: Predict the pet food intake according to the feeding bowl image sequence.
[0054] After extracting the image sequence of the feeding bowl, the food residue in the image sequence of the feeding bowl can be detected, and the amount of food eaten by the pet can be determined by the change in the food residue in the feeding bowl before and after the pet eats.
[0055] Specifically, according to the image sequence of the feeding bowl, it includes: Extract the key frames corresponding to the start node and the end node of the pet's eating from the image sequence of the feeding bowl; Send the key frames to the pet management server for estimating the amount of food eaten by the pet, and receive the amount of food eaten by the pet returned by the pet management server.
[0056] In the embodiment of the present application, after the pet feeder extracts the image sequence of the feeding bowl, it can further extract the key frames corresponding to the start node and the end node of the pet's eating in the image sequence of the feeding bowl. Among them, the key frame is not necessarily the first frame and the last frame in the image sequence of the feeding bowl. The key frame representing the start node of the pet's eating can be the image frame of the feeding bowl that is the most complete and clear near the start node of the pet's eating (for example, the image frame of the feeding bowl not blocked by the pet before the pet starts eating). Similarly, the key frame representing the end node of the pet's eating can be the image frame of the feeding bowl that is the most complete and clear near the end node of the pet's eating (for example, the image frame after the pet leaves the feeding bowl after eating). In order to ensure accurate estimation of the food residue at the start and end of the pet's eating, multiple key frames can be extracted for the start node and the end node of the pet's eating respectively, and then sent to the pet management server for estimating the amount of food eaten by the pet.
[0057] In some embodiments, the process of the pet management server predicting the amount of food eaten by the pet based on the key frames includes: Perform semantic segmentation on the key frames based on the image segmentation model, and determine the food area according to the semantic segmentation result; Predict the amount of food eaten by the pet according to the change in the food area.
[0058] In the embodiment of the present application, after the pet management server receives the key frame data uploaded by the pet feeder, it can call the image segmentation model in the pet management server to perform semantic segmentation on the received key frames to determine the food area, and determine the remaining amount of food according to the food area. Then, the amount of food eaten by the pet can be predicted according to the change in the food area before and after eating.
[0059] Specifically, in some embodiments, predicting the first pet food intake prediction value according to the change in the food area includes: Determine the corresponding pixel area of the food area and the stacking height corresponding to each pixel; Calculate the food volume corresponding to each key frame according to the pixel area and the stacking height; Calculate the change in the volume of the food, and determine the predicted value of the first pet's food intake based on the change amount.
[0060] In the embodiment of the present application, the image segmentation model can specifically adopt a convolutional neural network with a U-Net architecture to accurately identify and label the pixels of the food area, and calculate the pixel area of the food area. In addition, the image segmentation model can further output the stacking height corresponding to each food pixel. Specifically, the feeding bowl image collected in the present application can be a depth image, so that the pixel-level depth information corresponding to each pixel in the food area can be obtained, and then the stacking height corresponding to each pixel can be calculated according to the pixel-level depth information and the shape of the pet feeding bowl. Further, after calculating the stacking height corresponding to all the pixels covered by the pixel area, the food volume corresponding to each key frame can be further estimated according to the pixel area and the stacking height. In addition, since most pet foods in the current industry are designed in a round flake shape, and their stacking looseness is relatively stable and can be measured. After calculating the food volume based on the pixel area and the stacking height corresponding to each pixel, the accurate food volume can be further calculated in combination with the stacking looseness. In this way, for the eating start node, the average value of the food volumes corresponding to multiple key frames can be calculated; for the eating end node, the average value of the food volumes corresponding to multiple key frames can also be calculated. Further calculating the difference between the average value of the food volume corresponding to the eating end node and the average value of the food volume corresponding to the eating start node can obtain the pet's food intake.
[0061] Among them, in the training process of the convolutional neural network model with a U-Net architecture, a large number of labeled image samples can be used, including various scenarios such as various lighting conditions, food types, and pet interference behaviors. To reduce the manual annotation cost, in the training data preparation stage of the embodiment of the present application, a preset image annotation model can be used for offline automatic annotation. After automatic annotation by the image annotation model, only a small amount of manual correction is required to complete the annotation work of a large-scale data set, which greatly improves the annotation efficiency. The data after annotation is used to train the convolutional neural network with a U-Net architecture, and at the same time, the model adaptability and accuracy are further improved through transfer learning and data augmentation techniques.
[0062] In some embodiments, the pet food intake determined based on the pet eating video analysis can be determined as the predicted value of the first pet's food intake. In this way, after predicting the predicted value of the first pet's food intake, the pet food adding method provided by the present application can further include: Obtain the weighing data of the feeding bowl corresponding to each pet eating video segment, and calculate the predicted value of the second pet's food intake based on the weighing data of the feeding bowl; Determine the pet food intake based on the first pet food intake prediction value and the second pet food intake prediction value.
[0063] That is, in the embodiments of the present application, for each pet food intake video segment, the corresponding pet food intake start time point and the pet food intake end time point can be obtained, and then based on these two time points respectively, the weighing data detected by the weighing device can be obtained, so as to obtain the weighing data of the feeding bowl corresponding to each pet food intake video segment. Then, based on the weighing data of the feeding bowl corresponding to the pet food intake segment, the difference between the weighing data of the feeding bowl before and after the pet food intake is calculated, so as to obtain the second pet food intake prediction value.
[0064] Then, determine the pet food intake comprehensively according to the first pet food intake prediction value and the second pet food intake prediction value.
[0065] Specifically, in some embodiments, determining the pet food intake according to the first pet food intake prediction value and the second pet food intake prediction value includes: Obtain the first confidence level corresponding to the first pet food intake prediction value, and obtain the second confidence level corresponding to the second pet food intake prediction value; Determine the pet food intake according to the first confidence level, the second confidence level, the first pet food intake prediction value, and the second pet food intake prediction value.
[0066] In the embodiments of the present application, a method for comprehensively determining the pet food intake by combining pet food intake video analysis and feeding bowl weighing method is provided. This method can determine the first confidence level corresponding to the first pet food intake prediction value obtained based on pet food intake video analysis and the second confidence level corresponding to the second pet food intake prediction value obtained based on the difference in feeding bowl weighing data before and after pet food intake according to the environmental information during pet food intake. For example, when the camera device can collect the image sequence of the feeding bowl corresponding to the pet food intake behavior and can effectively interact with the pet management server, it can be determined that the first confidence level is greater than the second confidence level. Because relying on the weighing data of the feeding bowl to estimate the pet food intake may lead to inaccurate estimation of the pet food intake due to the fact that the pet's non-standard behavior during food intake may cause food to spill outside the feeding bowl, or the pet may get the pet toy into the feeding bowl and other situations that cannot be recognized by weighing. At this time, determining the pet food intake based on the first pet food intake prediction value with a larger confidence level can obtain a more accurate pet food intake prediction result.
[0067] However, since the method based on pet feeding video analysis requires the camera device to be able to collect clear image data of the pet feeding bowl and can upload the key frame data to the server for estimating the remaining amount of food in the feeding bowl. Thus, if the camera of the camera device is contaminated and unable to collect clear images of the feeding bowl, or if there is a network failure in the connection between the pet feeder and the pet management server, at this time, it can be determined that the second confidence level of the second pet food intake prediction value determined based on the weighing difference of the feeding bowl is greater than the first confidence level. Therefore, the pet food intake can be determined based on the second pet food intake prediction value, and then a corresponding pet food addition plan can be generated and food addition can be carried out. In this way, it is possible to avoid the problem that the camera of the camera device is contaminated or the network is abnormal, resulting in the inability to determine the pet food intake solely relying on the method of pet feeding video analysis, and thus the inability to carry out pet food addition.
[0068] Step S204, generate a pet food addition plan based on the pet food intake, and add pet food according to the pet food addition plan.
[0069] After determining the accurate pet food intake, a pet food addition plan can be generated according to the accurate pet food intake, and pet food can be added according to the generated pet food addition plan.
[0070] Among them, the specific process of generating a pet food addition plan based on the pet food intake may include obtaining the pet feeding time set by the pet owner, and determining the pet feeding plan based on the sum of the predicted pet food intakes between the pet feeding time and the adjacent feeding times and the set pet feeding time. For example, if the set feeding times are 8 am and 3 pm, then before 3 pm, the predicted pet food intakes corresponding to each pet feeding during the period from 8 am to 3 pm can be obtained first, and then these pet food intakes are accumulated to obtain the total food intake. In this way, it can be determined that the pet food addition plan is to feed according to the total food intake at 3 pm. When 3 pm arrives, food addition can be triggered, and the specific added amount is the above total food intake.
[0071] Alternatively, in some embodiments, when the pet management server returns the predicted pet food intake, the remaining amount of food can also be returned. If the remaining amount of food is less than a preset value, a pet food addition plan for immediately adding food can be triggered. The added amount corresponding to this pet food addition plan can be the total food intake of the pet during the period from the last food addition to the current food addition.
[0072] Alternatively, in some embodiments, when the pet management server returns the predicted pet food intake and the remaining food amount, the predicted pet food intake can be compared with the remaining food amount. When the predicted pet food intake is greater than the remaining food amount, a pet food addition plan for immediately adding food can be triggered. The addition amount corresponding to the pet food addition plan can also be the total amount of food consumed by the pet during the period from the last food addition to the current food addition.
[0073] The pet food addition method provided by the embodiments of the present application screens out pet eating video segments from the environmental video data corresponding to the pet feeder, and then estimates the pet food intake based on the feeding bowl images in the pet eating video segments. Compared with the related art that only uses the weighing data of the feeding bowl to estimate the pet food intake, on the one hand, it can avoid the interference of the change in the weighing data of the feeding bowl caused by non-pet eating behaviors (such as manually adding pet food or foreign objects entering the feeding bowl) to the estimation of the pet food intake, and on the other hand, it can also avoid the problem of inaccurate prediction of the food intake based on weighing caused by food spillage due to the pet's bad eating behavior.
[0074] In addition, in some embodiments, the pet management server can summarize and statistically analyze the pet food intake obtained from previous analyses, and accordingly analyze whether there is an abnormality in the pet food intake. When it is determined based on the accurate analysis of the pet food intake that the pet has an abnormal eating situation, such as an abnormal decrease in food intake, a decrease in eating frequency, the pet not eating for a long time, or the pet eating quickly in a short period of time, etc., the pet management server can further send an abnormality prompt to the pet management client to remind the pet owner to perform a health check on the pet in a timely manner.
[0075] In summary, the pet food addition method proposed by the embodiments of the present application is applied to a pet feeder. The method includes: obtaining the environmental video data corresponding to the pet feeder, and screening out pet eating video segments from the environmental video data; extracting the corresponding feeding bowl image sequence from the pet eating video segments; predicting the pet food intake according to the feeding bowl image sequence; generating a pet food addition plan based on the pet food intake, and adding pet food according to the pet food addition plan.
[0076] It can be seen that the pet food addition method provided by the embodiments of the present application can set a video data acquisition device near the pet feeder to collect environmental video of the pet feeder, and then comprehensively determine the accurate pet food intake based on the image difference and weighing difference of the feeding bowl before and after the pet eats in the environmental video, and then add food based on the accurate pet food intake, which can improve the accuracy of food addition.
[0077] Please refer to Figure 3, in some embodiments, the embodiment of the present application further provides a pet food adding device 300, and the pet food adding device 300 includes: An obtaining unit 310, configured to obtain environmental video data corresponding to a pet feeder, and screen out a pet eating video segment from the environmental video data; An extracting unit 320, configured to extract a corresponding feeding bowl image sequence from the pet eating video segment; A predicting unit 330, configured to predict the pet food intake according to the feeding bowl image sequence; A generating unit 340, configured to generate a pet food adding plan based on the pet food intake, and perform pet food adding according to the pet food adding plan.
[0078] Optionally, in some embodiments, the obtaining unit includes: An obtaining subunit, configured to obtain environmental video data collected by a camera device in the pet feeder; A detecting subunit, configured to perform object detection on the environmental video data, and screen out a pet behavior video segment according to the object detection result; A screening subunit, configured to perform behavior category recognition on the pet behavior video segment, and screen out a pet eating video segment according to the behavior category recognition result.
[0079] Optionally, in some embodiments, the detecting subunit includes: A denoising module, configured to denoise the environmental video data to obtain a denoised video; A detecting module, configured to perform motion detection on the denoised video based on a motion detection model, and extract a motion video segment from the denoised video according to the motion detection result; A screening module, configured to perform object detection on the objects in the motion video segment based on an object detection model, and screen out a pet behavior video segment from the motion video segment according to the object detection result.
[0080] Optionally, in some embodiments, the predicting unit includes: An extracting subunit, configured to extract key frames corresponding to a pet eating start node and a pet eating end node from the feeding bowl image sequence; An estimating subunit, configured to send the key frames to a pet management server for estimating the pet food intake, and receive a first pet food intake prediction value returned by the pet management server; A calculating subunit, configured to obtain the feeding bowl weighing data corresponding to each pet eating video segment, and calculate a second pet food intake prediction value based on the feeding bowl weighing data; A determining subunit, configured to determine the pet food intake according to the first pet food intake prediction value and the second pet food intake prediction value.
[0081] Optionally, in some embodiments, the present application further provides a pet food intake prediction device, including: A segmentation subunit, configured to perform semantic segmentation on key frames based on an image segmentation model, and determine a food region according to the semantic segmentation result; A prediction subunit, configured to predict a first pet food intake prediction value according to the change of the food region.
[0082] Optionally, in some embodiments, the prediction subunit includes: A determination module, configured to determine the corresponding pixel area of the food region and the stacking height corresponding to each pixel; A first calculation module, configured to calculate the food volume corresponding to each key frame according to the pixel area and the stacking height; A second calculation module, configured to calculate the change amount of the food volume, and determine the first pet food intake prediction value according to the change amount.
[0083] Optionally, in some embodiments, the determination subunit includes: An acquisition module, configured to acquire a first confidence corresponding to the first pet food intake prediction value, and acquire a second confidence corresponding to the second pet food intake prediction value; A determination module, configured to determine the pet food intake according to the first confidence, the second confidence, the first pet food intake prediction value, and the second pet food intake prediction value.
[0084] Refer to Figure 4 , Figure 4 illustrates the hardware structure of a pet feeder according to another embodiment. The pet feeder includes: A processor 401, which can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application; A memory 402, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of the present specification through software or firmware, the relevant program codes are stored in the memory 402, and the processor 401 is called to execute some steps in the pet food addition method of the embodiments of the present application; The input / output interface 403 is used to implement information input and output; The communication interface 404 is used to implement the communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); The bus 405 transmits information between various components of the device (such as the processor 401, the memory 402, the input / output interface 403, and the communication interface 404); Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 are communicatively connected to each other inside the device through the bus 405.
[0085] The embodiment of the present application also provides a storage medium, which stores a computer program. When the computer program is executed by a processor, it implements some steps in the pet food adding method provided by the present application.
[0086] The embodiment of the present application also provides a computer program product, which includes a computer program. The processor of the computer device reads and executes the computer program, so that the computer device executes some steps in the above-mentioned pet food adding method.
[0087] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "comprise" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0088] It should be understood that in the present disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the associated objects before and after. "At least one (one)" or a similar expression below refers to any combination of these items, including any combination of a single item or plural items. For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0089] It should be understood that in the description of the embodiments of the present application, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.
[0090] In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0091] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0092] In addition, the functional units in each embodiment of the present disclosure can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0093] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.
[0094] It should also be understood that the various embodiments provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0095] The above is a specific description of the embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. A method for adding to pet food, characterized in that, The method is applied to a pet feeder, and the method includes: Obtain environmental video data corresponding to the pet feeder, and screen out pet eating video clips from the environmental video data; Extract the corresponding feeding bowl image sequence from the pet eating video clips; Predict the pet's food intake based on the feeding bowl image sequence; Generate a pet food addition plan based on the pet's food intake, and add pet food according to the pet food addition plan.
2. The method according to claim 1, characterized in that, The obtaining of the environmental video data corresponding to the pet feeder and screening out the pet eating video clips from the environmental video data includes: Obtain the environmental video data collected by the camera device in the pet feeder; Perform object detection on the environmental video data, and screen out pet behavior video clips according to the object detection results; Perform behavior category recognition on the pet behavior video clips, and screen out pet eating video clips according to the behavior category recognition results.
3. The method according to claim 2, wherein The performing of object detection on the environmental video data and screening out pet behavior video clips according to the object detection results includes: Denoise the environmental video data to obtain a denoised video; Perform motion detection on the denoised video based on a motion detection model, and extract motion video clips from the denoised video according to the motion detection results; Perform object detection on the objects in the motion video clips based on an object detection model, and screen out pet behavior video clips from the motion video clips according to the object detection results.
4. The method according to claim 1, characterized in that, The predicting according to the feeding bowl image sequence includes: Extract the key frames corresponding to the pet eating start node and the pet eating end node from the feeding bowl image sequence; Send the key frames to the pet management server for estimating the pet's food intake, and receive the first pet food intake prediction value returned by the pet management server; Obtain the feeding bowl weighing data corresponding to each pet eating video clip, and calculate the second pet food intake prediction value based on the feeding bowl weighing data; Determine the pet's food intake according to the first pet food intake prediction value and the second pet food intake prediction value.
5. The method according to claim 4, wherein The process of the pet management server predicting the pet's food intake based on the key frames includes: Perform semantic segmentation on the key frames based on an image segmentation model, and determine the food area according to the semantic segmentation results; Predict the first pet food intake prediction value according to the change of the food area.
6. The method according to claim 5, characterized in that The predicting the first pet food intake prediction value according to the change of the food area includes: Determine the corresponding pixel area of the food area and the stacking height corresponding to each pixel; Calculate the food volume corresponding to each key frame according to the pixel area and the stacking height; Calculate the change amount of the food volume, and determine the first pet food intake prediction value according to the change amount.
7. The method according to claim 4, characterized in that, The determining the pet's food intake according to the first pet food intake prediction value and the second pet food intake prediction value includes: Obtain the first confidence level corresponding to the first pet food intake prediction value, and obtain the second confidence level corresponding to the second pet food intake prediction value; Determine the pet food intake based on the first confidence level, the second confidence level, the first pet food intake prediction value, and the second pet food intake prediction value.
8. A pet food adding device, characterized in that, The device is applied to a pet feeder, and the device includes: An acquisition unit, configured to acquire the environmental video data corresponding to the pet feeder, and screen out the pet eating video segments from the environmental video data; An extraction unit, configured to extract the corresponding feeding bowl image sequence from the pet eating video segments; A prediction unit, configured to predict the pet food intake according to the feeding bowl image sequence; A generation unit, configured to generate a pet food addition plan based on the pet food intake, and add pet food according to the pet food addition plan.
9. A pet feeder, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the pet food addition method according to any one of claims 1 to 7.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the pet food addition method according to any one of claims 1 to 7.
11. A computer program product, the computer program product comprising a computer program, characterized in that, The computer program is read and executed by the processor of the computer device, so that the computer device executes the pet food addition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Pet feeding control method and device, equipment and storage medium
CN111869579A
Precise livestock and poultry feeding method, system and equipment based on image recognition
CN119580008A
Composition for treating infection of feline infectious peritonitis comprising chlortetracycline
KR1020230092754A