Pet condition evaluation system, pet camera, server, pet condition evaluation method, and program

By using a region detector and information generator in the pet status assessment system, and employing a learning model to identify pet postures and emotions, the system solves the problem of difficulty in assessing pet status in existing technologies, and achieves accurate identification and assessment of pet status.

CN115885313BActive Publication Date: 2025-11-25PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180050178.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-01
Filing Date
2021-08-20
Publication Date
2025-11-25
Estimated Expiration
2041-08-20

Smart Images

  • Figure CN115885313B_ABST
    Figure CN115885313B_ABST
Patent Text Reader

Abstract

The purpose of the present invention is to make the state of a pet more easily obtainable. A pet state evaluation system (1) includes a region detection unit (32), an information generation unit (33), and an evaluation unit (34). The region detection unit (32) detects, in image data, a specific region representing at least a part of an appearance of a pet as a subject. The information generation unit (33) generates pet information. The pet information includes posture information based on a learning model and the image data. The learning model is generated by learning a posture of a pet to recognize an image of a posture of the pet. The evaluation unit (34) evaluates a pet state based on the pet information, the pet state relating to at least one of an emotion or a behavior of the pet.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to a pet state evaluation system, a pet camera, a server, a pet state evaluation method, and a program. More specifically, the present disclosure relates to a pet state evaluation system for evaluating a state of a pet photographed as a subject in image data, a pet camera including the pet state evaluation system, a server, a pet state evaluation method, and a program. BACKGROUND

[0002] Patent Literature 1 discloses a detection device for recognizing and detecting animals and humans from an image. The detection device includes an animal detector for detecting an animal from an image and a human detector for detecting a human from an image. The detection device further includes a detection result output interface for outputting information that a target object is detected as a detection result when both an animal and a human are detected.

[0003] In the detection device, the animal detector scans an input image with reference to feature amount data reflecting animal features stored in an animal feature amount storage device. If the animal detector 21 finds any region matching or having a high degree of similarity with the animal feature amount data, the animal detector determines that the object photographed in the region is an animal.

[0004] Generally, a user (who can be a master of a pet, for example) can want to know a specific state of a pet (animal) photographed in image data or can want to be notified of the specific state if the pet photographed in the image data is in the specific state.

[0005] LIST OF CITATIONS

[0006] PATENT LITERATURE

[0007] Patent Literature 1: JP 2013-65110 A SUMMARY

[0008] In view of the above background, it is therefore an object of the present disclosure to provide a pet state evaluation system, a pet camera, a server, a pet state evaluation method, and a program, all of which are configured or designed to make a state of a pet more easily recognized.

[0009] A pet status evaluation system according to one aspect of the present disclosure includes a region detector, an information generator, and an evaluator. The region detector detects, in image data, a specific region representing at least a portion of an appearance of a pet as a subject. The information generator generates pet information. The pet information includes posture information about at least a posture of the pet. The posture information is based on a learning model and the image data. The learning model has been generated by learning a posture of the pet to recognize the posture of the pet on an image. The evaluator evaluates a pet status about at least one of an emotion of the pet appearing in the specific region or an action of the pet appearing in the specific region based on the pet information.

[0010] A pet camera according to another aspect of the present disclosure includes the pet status evaluation system described above and an image capturing device that captures image data.

[0011] A server according to another aspect of the present disclosure can communicate with a pet camera equipped with an information generator and an evaluator of the pet status evaluation system described above. The server is equipped with a region detector.

[0012] A server according to still another aspect of the present disclosure can communicate with a pet camera equipped with a region detector of the pet status evaluation system described above. The server is equipped with an information generator and an evaluator.

[0013] A pet status evaluation method according to still another aspect of the present disclosure includes a pet detecting step, an information generating step, and an evaluating step. The pet detecting step includes detecting, in image data, a specific region representing at least a portion of an appearance of a pet as a subject. The information generating step includes generating pet information. The pet information includes posture information about at least a posture of the pet. The posture information is based on a learning model and the image data. The learning model has been generated by learning a posture of the pet to recognize the posture of the pet on an image. The evaluating step includes evaluating a pet status about at least one of an emotion of the pet appearing in the specific region or an action of the pet appearing in the specific region based on the pet information.

[0014] A program according to another aspect of the present disclosure is designed to cause one or more processors to execute the pet status evaluation method described above. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1A A schematic configuration of a pet camera to which a pet status evaluation system according to one example embodiment is applied is shown.

[0016] Figure 1B A schematic configuration of a presentation device for communicating with a pet camera is shown.

[0017] Figure 2is a conceptual diagram showing the overall configuration of a pet management system including a pet condition assessment system.

[0018] Figures 3A-3C An example of image data to be subjected to assessment processing by the pet condition assessment system is shown.

[0019] Figures 4A-4C Another example of image data to be subjected to assessment processing by the pet condition assessment system is shown.

[0020] Figures 5A-5C Another example of image data to be subjected to assessment processing by the pet condition assessment system is shown.

[0021] Figure 6 Still another example of image data to be subjected to assessment processing by the pet condition assessment system is shown.

[0022] Figure 7A and 7B is a conceptual diagram showing a presentation device on the screen of which an assessment result made by the pet condition assessment system is presented.

[0023] Figure 8 is a flowchart showing exemplary operations of the pet condition assessment system.

[0024] Figure 9 is a flowchart showing exemplary operations of the pet condition assessment system; and

[0025] Figure 10 A schematic configuration diagram of a pet camera to which a variation example of the pet condition assessment system is applied is shown. DETAILED DESCRIPTION

[0026] (1) SUMMARY

[0027] The drawings referred to in the following embodiment description are schematic drawings. Therefore, the proportions of the dimensions (including thicknesses) of the constituent elements exemplified on the drawings do not always reflect the actual proportions of the dimensions thereof.

[0028] The pet condition assessment system 1 according to an exemplary embodiment includes a region detector 32, an information generator 33, and an assessor 34, as shown in Figure 1AThe pet status evaluation system 1 includes a computer system as its main constituent element, which includes one or more processors and one or more memories. In the following description, as an example, the respective constituent elements of the pet status evaluation system 1 (including the region detector 32, the information generator 33, and the evaluator 34) are assumed to be all gathered together in a single housing of the pet camera 100. However, this is only an example and should not be construed as limiting. In addition, the constituent elements of the pet status evaluation system 1 according to the present disclosure can also be distributed among multiple devices. For example, at least some of the constituent elements of the pet status evaluation system 1 can be provided outside the pet camera 100 (for example, in an external server such as the server 7). For example, the pet camera 100 can be equipped with the information generator 33 and the evaluator 34, while the server 7 having the communication capability with the pet camera 100 can be equipped with the region detector 32. Alternatively, the pet camera 100 can be equipped with the region detector 32, while the server 7 having the communication capability with the pet camera 100 can be equipped with the information generator 33 and the evaluator 34. As used herein, the "server" can be a single external device (which can also be a device installed in the house of the user 300) or be composed of multiple external devices, as appropriate.

[0029] The region detector 32 detects a specific region Al representing at least a part of the appearance of the pet 5 as the subject Hl in the image data Dl (refer to Figures 3B-6 ). In the present embodiment, the image data Dl is an image (data) captured (or generated) by the image capturing device 2 (refer to Figure 1A ) of the pet camera 100. The image data Dl can be a still picture or a picture constituting one frame of a moving picture, as appropriate. Alternatively, the image data Dl can also be an image captured by the image capturing device 2 and partially subjected to image processing. In the following description, the type of "pet" whose status is evaluated by the pet status evaluation system 1 is assumed to be a dog (as a kind of animal). However, the type of "pet" is not limited to any particular type and can be a cat or any other type of animal.

[0030] Also, in the following description, the dog (as the pet of interest) photographed in the image data Dl (designated as the subject) will be designated by the reference symbol "5", while the other dogs (as the pets of non-interest) will be mentioned without the reference symbol attached.

[0031] In the present embodiment, the specific region Al is a region surrounded with a rectangular frame in the image data Dl and represented by a "bounding box", as shown in Figures 3A-6As shown, the bounding box encloses the pet 5, which is photographed as the subject H1. The position of the pet 5 in the image data D1 can be defined by, for example, the X and Y coordinates of the upper left corner of the bounding box, as well as the horizontal width and height of the bounding box. However, a specific region A1 does not necessarily have to be represented by a bounding box, but can also be represented by, for example, segments that distinguish the subject H1 from the background on a pixel-by-pixel basis. In this disclosure, the XY coordinates used to determine the positions of the pet 5 and specific objects 6 other than the pet 5 in the image data D1 are assumed, for example, to be defined on a pixel-based basis.

[0032] Information generator 33 generates pet information. This pet information includes pose information about at least pet 5. The pose information is based on a learning model (hereinafter sometimes referred to as "first model M1") and image data D1. The learning model has been generated by learning the pet's pose in order to recognize the pet's pose in the image. The first model M1 is a model generated by machine learning and is stored in the model storage device P1 of the pet camera 100 (reference). Figure 1A )middle.

[0033] In this embodiment, the area detector 32 and the information generator 33 together constitute the pet detector X1 (see reference). Figure 1A The information generator 33 is used to detect dogs (as type 5 of pets) from image data D1. Optionally, at least some of the functions of the information generator 33 can be provided externally to the pet detector X1.

[0034] The evaluator 34 assesses the pet's state regarding the emotions and / or actions of the pet 5 present in a specific area A1, based on pet information. In this embodiment, the evaluator 34 bases its assessment on pet information and conditional information 9 regarding specific actions and / or specific emotions of the pet (see reference). Figure 1A The condition information 9 is stored in the condition storage device P2 of the pet camera 100 (see reference). Figure 1A )middle.

[0035] With this configuration, the evaluator 34 assesses the pet's state regarding its emotions and / or actions based on pet information, thereby making it easier to identify the pet's state.

[0036] The pet status evaluation method according to another embodiment of the exemplary embodiments includes a pet detection step, an information generation step, and an evaluation step. The pet detection step includes detecting, in the image data D1, a specific region A1 representing at least a portion of the appearance of the pet 5 as the subject H1. The information generation step includes generating pet information. The pet information includes posture information about at least a posture of the pet 5. The posture information is based on the learning model M1 and the image data D1. The learning model M1 has been produced by learning a posture of a pet so as to recognize the posture of the pet on an image. The evaluation step includes evaluating, based on the pet information, a pet status about an emotion and / or an action of the pet 5 appearing in the specific region A1.

[0037] According to the method, the evaluation step includes evaluating, based on the pet information, the pet status about the emotion and / or the action of the pet 5, thereby eventually making the status of the pet 5 more easily recognizable.

[0038] The pet status evaluation method is used on a computer system (the pet status evaluation system 1). That is, the pet status evaluation method can also be implemented as a program. The program according to the present embodiment is designed to cause one or more processors to execute the pet status evaluation method according to the present embodiment.

[0039] (2) Details

[0040] Next, a system to which the pet status evaluation system 1 according to the present embodiment is applied (hereinafter referred to as “pet management system 200”) will be described with reference to Figures 1A-9

[0041] (2.1) Overall Configuration

[0042] As shown in FIG. 2, the pet management system 200 includes one or more pet cameras 100, one or more presentation devices 4, and a server 7. In the following description, attention is paid to a user 300 who receives a pet 5 management (or viewing) service using the pet management system 200 (refer to FIG. 1). The user 300 can be, but is not necessarily, the owner of the pet 5. Figure 2 Figure 2 The user 300 installs the one or more pet cameras 100 at a predetermined location (or a plurality of predetermined locations) in a facility (for example, a residential facility in which the user 300 and the pet 5 reside). If a plurality of pet cameras 100 are to be installed, the user 300 can install one pet camera 100 in each room of the residential facility. The pet camera 100 does not necessarily have to be installed indoors, but can also be installed outdoors. In the following description, attention will be paid to one pet camera 100 for the sake of convenience of description.

[0043] The user 300 installs the one or more pet cameras 100 at a predetermined location (or a plurality of predetermined locations) in a facility (for example, a residential facility in which the user 300 and the pet 5 reside). If a plurality of pet cameras 100 are to be installed, the user 300 can install one pet camera 100 in each room of the residential facility. The pet camera 100 does not necessarily have to be installed indoors, but can also be installed outdoors. In the following description, attention will be paid to one pet camera 100 for the sake of convenience of description.

[0044] ​​The presentation device 4 is assumed to be, for example, a telecommunication device owned by the user 300. The telecommunication device is assumed to be, for example, a mobile telecommunication device such as a smartphone or a tablet. Alternatively, the presentation device 4 can also be a notebook or a desktop computer.

[0045] As shown in FIG. 1, the presentation device 4 includes a communication interface 41, a processing device 42, and a display device 43. Figure 1B

[0046] The communication interface 41 is a communication interface that allows the communication interface 41 to communicate with each of the pet camera 100 (refer to Figure 2 ) and the server 7 (refer to Figure 2 ). Alternatively, the communication interface 41 can be configured to communicate only with the pet camera 100 or the server 7.

[0047] The processing device 42 is implemented as a computer system including, for example, one or more processors (microprocessors) and one or more memories. That is, the respective functions of the processing device 42 are performed by causing the one or more processors to execute one or more programs (application programs) stored in the one or more memories. In the present embodiment, the programs are stored in the memory of the processing device 42. Alternatively, the programs can also be downloaded through a telecommunication line such as the Internet, or distributed after being stored in a non-transitory storage medium such as a memory card. The user 300 can cause his or her telecommunication device to function as the presentation device 4 by installing an application software program (hereinafter referred to as "pet application program") dedicated to presenting a graphical user interface (GUI) regarding the pet 5 to be watched and activating the pet application program.

[0048] The display device 43 can be implemented as a touch panel liquid crystal display or an organic electroluminescence (EL) display. As the presentation device 4 executes the pet application program, a screen image presenting information regarding the pet 5 is displayed (output) on the display device 43.

[0049] If a plurality of residents (family members) live together with the pet 5 in the living facility and receive the pet 5 management service as the user 300, the pet management system 200 includes a plurality of presentation devices 4 carried by the plurality of residents (the plurality of users 300) respectively. In the following description, attention will be paid to one presentation device 4 (smartphone) carried by one user 300 (resident).

[0050] The pet camera 100 is, for example, a device having the ability to capture images to watch a designated pet. In other words, the pet camera 100 includes Figure 1A ​The image capturing device 2 (camera) shown in FIG. 1 is installed in the pet camera 100. The user 300 installs the pet camera 100 so that the area where his or her own pet 5 mainly moves (for example, the place where the pet 5 is fed) falls within the angle of view of the image capturing device 2 within the residential facility (or outside). In this way, the user 300 can observe the state of the pet 5 through the image acquired by the image capturing device 2 even when, for example, he or she is away from the residential facility.

[0051] As described above, the type of pet to be observed is assumed to be a dog as an example. Although the plurality of frames of image data D1 representing a plurality of breeds of dogs are explained as an example in FIG. 1, these figures simply illustrate various postures that a dog can take to describe the operation of the pet state evaluation system 1 and should not be understood as limiting the breeds of dogs to which the pet state evaluation system 1 is applicable. For example, the pet state evaluation system 1 is configured to recognize the posture of any dog in the same manner to some extent regardless of its breed. However, this is just an example and should not be understood as limiting. Alternatively, the pet state evaluation system 1 can also recognize the posture of a specific dog on an individual basis depending on the breed of the dog. Figures 3A-6

[0052] As shown in FIG. 1, the pet camera 100 further includes a communication interface 11 in addition to the image capturing device 2. The communication interface 11 is a communication interface that allows the communication interface 11 to communicate with each of the presentation device 4 (refer to FIG. 2) and the server 7 (refer to FIG. 3). Alternatively, the communication interface 11 can have the ability to establish short-range wireless communication conforming to the Bluetooth (R) Low Energy (BLE) standard with the presentation device 4, for example. If the user 300 (refer to FIG. 2) who carries the presentation device 4 around stays at home, the communication interface 11 can transmit and receive data to / from the presentation device 4 by establishing short-range wireless communication directly with the presentation device 4. Figure 1A Figure 2 Figure 2 Figure 2

[0053] In addition, the communication interface 11 is also connected to a network NT1 (refer to FIG. 3), such as the Internet, through a router installed inside the residential facility. The pet camera 100 can communicate with the external server 7 through the network NT1 to acquire information from the server 7 and output information to the server 7. Figure 2

[0054] Alternatively, the pet camera 100 can be configured to recognize the posture of a specific dog on an individual basis depending on the breed of the dog. Figure 2 ​​​​​​The demonstration device 4 shown can also connect to network NT1 via, for example, a cellular phone network (carrier network) provided by a communications service provider or a public wireless local area network (LAN). Examples of cellular phone networks include third-generation (3G), long-term evolution (LTE), fourth-generation (4G), and fifth-generation (5G) networks. In an environment where the demonstration device 4 can connect to a cellular phone network, the demonstration device 4 can connect to network NT1 via the cellular phone network. For example, if the user 300 carrying the demonstration device 4 is not at home, when connected to network NT1 via the cellular phone network, the demonstration device 4 is allowed to communicate with each of the pet camera 100 and the server 7.

[0055] Alternatively, communication between the demonstration device 4 and the pet camera 100 can be established via network NT1 and server 7.

[0056] As described above, the pet condition assessment system 1 is provided for purposes such as Figure 1A The pet camera 100 shown. Specifically, as... Figure 1A As shown, the pet camera 100 further includes a processing unit 3, a model storage unit P1, and a condition storage unit P2, which together constitute the pet state assessment system 1. This pet state assessment system 1 will be described in further detail in the next section.

[0057] like Figure 2 As shown, server 7 is connected to network NT1. Server 7 can communicate with each of pet cameras 100 and demonstration devices 4 via network NT1. Server 7 manages information such as user information (e.g., his or her name, user ID, phone number, and email address), information about pet cameras 100 and demonstration devices 4 owned by user 300 (e.g., identification information), and information about pets 5 owned by user 300 (e.g., information about the breed of their dog). Furthermore, server 7 collects and accumulates various image data and processing results (especially processing errors) captured by the multiple pet cameras 100. Optionally, user 300 can access server 7 via demonstration device 4 to download pet applications.

[0058] In this embodiment, server 7 is assumed to be a single server device. However, this is merely an example and should not be construed as limiting. Optionally, server 7 may also consist of multiple server devices, for example, forming a cloud computing system. Optionally, at least some of the functions of pet status assessment system 1 may be provided within server 7.

[0059] (2.2) Pet condition assessment system

[0060] like Figure 1AAs shown, the pet camera 100 includes not only the image capturing device 2 and the communication interface 11, but also the processing device 3, the model storage device Pl, and the condition storage device P2 as described above with respect to the pet state evaluation system 1. The pet state evaluation system 1 performs an "evaluation process" to evaluate the pet state.

[0061] The model storage device Pl is configured to store data including a plurality of learning models. The model storage device Pl includes a rewritable memory such as an electrically erasable programmable read-only memory (EEPROM). Meanwhile, the condition storage device P2 is configured to store data including the condition information 9. The condition storage device P2 includes a rewritable memory such as an EEPROM. The model storage device Pl and the condition storage device P2 can be the same storage device (memory). Alternatively, the model storage device Pl and the condition storage device P2 can also be memories built in the processing device 3.

[0062] The processing device 3 is implemented as a computer system including one or more processors (microprocessors) and one or more memories, for example. That is, respective functions (to be described later) of the processing device 3 are performed by causing the one or more processors to execute one or more programs (applications) stored in the one or more memories. In the present embodiment, the programs are stored in the memories of the processing device 3. Alternatively, the programs can also be downloaded through a telecommunication line such as the Internet, or distributed after being stored in a non-transitory storage medium such as a memory card.

[0063] The processing device 3 has a function of a controller for performing overall control of the pet camera 100, that is, control of the image capturing device 2, the communication interface 11, the model storage device Pl, the condition storage device P2, and other components.

[0064] In the present embodiment, the processing device 3 includes an acquirer 31, a region detector 32, an information generator 33, an evaluator 34, an output interface 35, and an object detector 36, as shown in Figure 1A In the present embodiment, the region detector 32 and the information generator 33 collectively constitute a pet detector Xl for detecting a dog (as a type of the pet 5) from the above-described image data Dl.

[0065] The acquirer 31 is configured to acquire the image data Dl (e.g., a still picture) from the image capturing device 2. The acquirer 31 can acquire an image as a frame of a moving picture from the image capturing device 2 as the image data Dl. When the acquirer 31 acquires the image data Dl, the processing device 3 performs the evaluation process.

[0066] The region detector 32 of the pet detector X1 is configured to detect, in the image data D1, a specific region A1 representing at least a portion of the appearance of the pet 5 as the subject H1. In the present embodiment, the region detector 32 detects the specific region A1 based on a learning model (hereinafter sometimes referred to as "second model M2"). The second model M2 has been generated by learning appearance factors (feature amounts) of the appearance of a predetermined type of pet (e.g., a dog in the present example) so as to recognize the predetermined type of pet on an image. The second model M2 is stored in the model storage P1.

[0067] The second model M2 can include, for example, a model using a neural network or a model generated using deep learning of a multi-layer neural network. Examples of the neural network (including the multi-layer neural network) can include a convolutional neural network (CNN) and a Bayesian neural network (BNN). The second model M2 can be implemented by, for example, mounting a learned neural network into an integrated circuit such as an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). However, the second model M2 does not necessarily have to be a model generated by deep learning. Alternatively, the second model M2 can also be a model generated by a support vector machine or a decision tree, etc.

[0068] In short, the region detector 32 determines whether any dog (as a type of the pet 5) is present as the subject H1 in the acquired image data D1 using the second model M2. When it is determined that a dog (as a type of the pet 5) is present in the image data D1, the region detector 32 detects the specific region A1 defined by the bounding box that encloses the pet 5 (refer to Figures 3A-6 ). However, the specific region A1 does not necessarily have to be defined by the bounding box, but can be defined, for example, by segmentation.

[0069] The region detector 32 detects a head region A2 (refer to Figure 2 ) representing the head section 50 (refer to Figures 3A-6 ) of the subject H1 based on a learning model (hereinafter sometimes referred to as "third model M3"). The third model M3 has been generated by learning appearance factors (feature amounts) of the head section of a predetermined type of pet (e.g., a dog in the present example) so as to recognize the head section of the predetermined type of pet on an image. That is, the region detector 32 further has the function of a head section detector for detecting the head region A2 covering the face portion using the third model M3. Alternatively, the region detector 32 and the head section detector can be provided separately from each other. The third model M3 is stored in the model storage P1.

[0070] The third model M3 and the second model M2 can include, for example, a model using a neural network or a model generated using deep learning using a multi-layer neural network. Nonetheless, the third model M3 does not have to be a model generated by deep learning. Alternatively, the third model M3 can be the same model as the second model M2.

[0071] The region detector 32 determines whether or not a head section 50 of a dog (as a type of pet 5) is present in the image data Dl using the third model M3. When it is determined that the head section 50 of a dog (as a type of pet 5) is present in the image data Dl, the region detector 32 detects a head region A2 defined by a bounding box surrounding the head section 50 (refer to Fig. 6). However, the head region A2 does not have to be defined by a bounding box, but can be defined, for example, by segmentation. Figures 3A-6 ) However, the head region A2 does not have to be defined by a bounding box, but can be defined, for example, by segmentation.

[0072] If the image data Dl is an image representing a part (e.g., a face) of an appearance of a dog (as a type of pet 5) as a close-up, the detection of the specific region Al or the detection of the head region A2 can fail (i.e., the specific region Al or the head region A2 can be erroneously detected). Specifically, the image data Dl representing a face of a dog (as a type of pet 5) as a close-up is essentially an annotation about "a face of a dog". Thus, even if the region detector 32 successfully detects the head region A2 as a face of a dog (the head section 50), the head region A2 can fail to provide an annotation for the whole of the dog (i.e., its entire appearance). Thus, it is possible that the region detector 32 fails to detect the specific region Al of a dog. In the present embodiment, if the region detector 32 detects at least one of a dog or a face of a dog, the region detector 32 assumes that a dog (as a type of pet 5) is present in the acquired image data Dl. If the region detector 32 detects only the head region A2, the region detector 32 sets a region substantially equal in size to the head region A2 as the specific region Al. Note that if the region detector 32 fails to detect the head region A2, even if the region detector 32 has detected the specific region Al, the processing device 3 can end the evaluation process on the image data Dl.

[0073] The information generator 33 of the pet detector Xl generates pet information based not only on a learning model (the first model Ml) generated by learning a posture of a pet (e.g., a dog in the present embodiment) to recognize the posture of the pet on an image but also on image data Dl in which the specific region Al has been detected. The pet information includes posture information about a posture of the pet 5 present at least in the specific region Al.

[0074] Specifically, the information generator 33 includes a posture determiner 331, a direction determiner 332, and a distance determiner 333.

[0075] The posture determiner 331 is configured to determine (evaluate) the posture of a dog (as a type of pet 5) based on the first model Ml and the information about the specific area Al. The first model Ml has been generated by learning appearance factors (characteristic amounts) of the posture of a dog to recognize the posture of a dog on an image.

[0076] The first model Ml as well as the second model M2 and the third model M3 can include, for example, a model using a neural network or a model generated by deep learning using a multi-layer neural network. Nonetheless, the first model Ml does not necessarily have to be a model generated by deep learning. Alternatively, the first model Ml can be the same model as the second model M2 and the third model M3.

[0077] Next, the posture of a dog (as a type of pet 5) will be described, in which each illustrates an exemplary frame of the image data Dl that can be subjected to evaluation processing by the pet state evaluation system 1. Figures 3A-6

[0078] Figure 3A is an exemplary frame of the image data Dl indicating how the pet 5 takes a posture (first posture) of standing on four legs to observe its surroundings.

[0079] Figure 3B is another exemplary frame of the image data Dl indicating how the pet 5 takes a posture (second posture) of lying on the floor, facing the front and observing its surroundings.

[0080] Figure 3C is another exemplary frame of the image data Dl indicating how the pet 5 takes a second posture of slightly facing to the right and observing its surroundings.

[0081] Figure 4A is yet another exemplary frame of the image data Dl indicating how the pet 5 takes a posture (third posture) of running with its front legs stretched forward and its hind legs stretched backward. In Figure 4A , the pet 5 is holding its tail up.

[0082] Figure 4B is yet another exemplary frame of the image data Dl indicating how the pet 5 takes a posture (fourth posture) of walking with one of its front legs and one of its hind legs placed on the ground, and the other of its front leg and the other of its hind leg bent and off the ground. In Figure 4B , the tail of the pet 5 is kept down.

[0083] Figure 4C is yet another exemplary frame of the image data Dl indicating how the pet 5 takes a posture (fifth posture) of curling up on the floor and sleeping with its eyes closed.

[0084] ​Figure 5A is yet another exemplary frame of the image data D1, representing how the pet 5 assumes a pose of standing on its hind legs only to jump towards a person (e.g. the user 300) and to show him or her its love (sixth pose).

[0085] Figure 5B is yet another exemplary frame of the image data D1, representing how the pet 5 assumes a pose of sitting on the floor facing a person (e.g. the user 300) and showing him or her love (seventh pose).

[0086] Figure 5C is yet another exemplary frame of the image data D1, representing how the pet 5 assumes a pose of standing on three legs, with one of its front legs off the floor to play with a toy 63 (e.g. a ball in the example described in Figure 5C

[0087] Figure 6 is yet another exemplary frame of the image data D1, representing how the pet 5 assumes a pose of standing on four legs, with its head section 50 lowered to eat feed from a bowl 64 (ninth pose).

[0088] It is noted that the above first to ninth poses are merely exemplary poses that a dog (as a type of pet 5) can assume and should not be construed as limiting. Nonetheless, the first model M1 is generated by machine learning on poses of dogs, and more specifically, poses of dogs that are highly relevant to certain actions of the dog, in particular, actions related to certain types of emotions. Among the various poses of the dog, as to certain poses of the dog that require more accurate assessment, various mental and physical conditions of the dog are more finely distinguished by machine learning. As used herein, a “certain pose” refers to a pose that is relevant to an action that is deeply associated with a certain emotion of the dog. Examples of emotions of the dog that can be read from the actions of the dog include happiness, anger, loneliness, lightness, fear, and relaxation. Certain actions related to the certain pose of the dog can be associated with any of these emotions.

[0089] ​For example, if the pet 5 is posing the first posture, i.e., standing on four legs, the machine learning is made to evaluate the first posture of the pet 5 more finely by determining whether the pet 5 is showing teeth or a tongue and determining whether the ears of the pet 5 are up or down. For example, the first posture of the pet 5 showing teeth is associated with a "threatening" action. The first posture of the pet 5 with the ears up is associated with an "observing" action, more specifically, observing the environment around the pet 5, while the first posture of the pet 5 with the ears down is associated with a "not observing" action. The "threatening" action can be set as an action related to "angry", which is one of the emotions of a dog. The "observing" action can be set as an action related to "fear", which is another emotion of a dog. Further, the "not observing" action can be set as an action related to "lonely" or "relaxed", which are other emotions of a dog. In addition, if the pet 5 is posing the fifth posture, i.e., sleeping, the machine learning is made to evaluate the fifth posture of the pet 5 more finely by precisely determining how the pet 5 is sleeping, specifically, whether the pet 5 is sleeping with the back arched or the body straight, whether the pet 5 is sleeping with the eyes closed or open, and whether the pet 5 is sleeping with the tongue out or not.

[0090] When performing the annotation work by attaching a label to the image data (raw data) to prepare a learning data set for generating the first to third models M1-M3, i.e., at the time of determining the supervised data, a large amount of image data is used. The learning data set is selected from a large amount of image data, which is not limited in any way in terms of the breed of the dog, the color of the dog, the orientation of the dog, and the background in which the dog is photographed. The learning data set can include not only image data of a real dog but also image data of a stuffed dog and image data of a dog generated by CG. The machine learning is made using a combination of these types of image data.

[0091] The information on the posture of the pet 5 determined by the posture determiner 331 (including the determination result and the information on the specific region Al) is output to the distance determiner 333.

[0092] The facing direction determiner 332 is configured to determine (measure) the facing direction of the pet 5 in the image data D1 based on the image data D1 in which the specific region Al has been detected. That is, the pet information further includes the determination result made by the facing direction determiner 332. The facing direction determiner 332 receives the information on the detected specific region Al and the information on the detected head region A2 from the region detector 32. The facing direction determiner 332 can determine the facing direction of the pet 5 as the subject H1 based only on the information on the specific region Al detected by the region detector 32. In the present embodiment, the facing direction determiner 332 determines the facing direction of the pet 5 based on the information on the specific region Al and the information on the head region A2.

[0093] In particular, in the present embodiment, the facing direction determiner 332 determines the facing direction of the pet 5 based on at least the relative position of the head region A2 with respect to the specific region Al. Specifically, the facing direction determiner 332 obtains information about the position and size of the pet 5 in the image data Dl through the specific region Al detected by the region detector 32. In addition, the facing direction determiner 332 also obtains information about the position and size of the head section 50 of the pet 5 in the image data Dl through the head region A2 detected by the region detector 32.

[0094] For example, in the example shown in Fig. 27, the head region A2 is located at the right upper corner of the specific region Al, and thus the facing direction determiner 332 determines that the pet 5 is generally facing rightward. On the other hand, in the example shown in Fig. 28, the head region A2 is located at the top of the specific region Al, and thus the facing direction determiner 332 determines that the pet 5 is generally facing straight ahead. Figure 3A In the example shown in Fig. 27, the head region A2 is located at the right upper corner of the specific region Al, and thus the facing direction determiner 332 determines that the pet 5 is generally facing rightward. On the other hand, in the example shown in Fig. 28, the head region A2 is located at the top of the specific region Al, and thus the facing direction determiner 332 determines that the pet 5 is generally facing straight ahead. Figure 3B In the example shown in Fig. 27, the head region A2 is located at the right upper corner of the specific region Al, and thus the facing direction determiner 332 determines that the pet 5 is generally facing rightward. On the other hand, in the example shown in Fig. 28, the head region A2 is located at the top of the specific region Al, and thus the facing direction determiner 332 determines that the pet 5 is generally facing straight ahead.

[0095] The facing direction determiner 332 can consider not only the relative position of the head region A2 with respect to the specific region Al, but also, for example, the planar area ratio of the head region A2 with respect to the specific region Al, and the respective positions of the eyes, nose, mouth and other facial parts of the pet 5 in the head region A2, while determining the facing direction of the pet 5. This will further improve the reliability of the determination.

[0096] The information about the facing direction of the pet 5 determined by the facing direction determiner 332 (i.e., the result of the determination) is output to the evaluator 34.

[0097] The distance determiner 333 is configured to determine (measure) the relative distance between the pet 5 and an object region Bl (to be described later) (hereinafter sometimes referred to as "pet-object distance"). That is, the pet information further includes the result of the determination made by the distance determiner 333 (i.e., information about the pet-object distance). In other words, in the image data Dl, an object other than a dog (a specific object 6) as the type of the pet 5 can have been photographed as part of the subject.

[0098] In the example shown in Fig. 27, a leg 61 of a person has been photographed as the specific object 6. In the example shown in Fig. 28, the overall appearance 62 of a person sitting on the floor with the back turned has been photographed as the specific object 6. In the example shown in Fig. 29, a toy 63 of a dog has been photographed as the specific object 6. Figure 5A In the example shown in Fig. 27, a leg 61 of a person has been photographed as the specific object 6. In the example shown in Fig. 28, the overall appearance 62 of a person sitting on the floor with the back turned has been photographed as the specific object 6. In the example shown in Fig. 29, a toy 63 of a dog has been photographed as the specific object 6. Figure 5B In the example shown in Fig. 27, a leg 61 of a person has been photographed as the specific object 6. In the example shown in Fig. 28, the overall appearance 62 of a person sitting on the floor with the back turned has been photographed as the specific object 6. In the example shown in Fig. 29, a toy 63 of a dog has been photographed as the specific object 6. Figure 5C In the example shown in Fig. 27, a leg 61 of a person has been photographed as the specific object 6. In the example shown in Fig. 28, the overall appearance 62 of a person sitting on the floor with the back turned has been photographed as the specific object 6. In the example shown in Fig. 29, a toy 63 of a dog has been photographed as the specific object 6.Figure 6 In the example shown, the bowl 64 in which the dog's food is placed has been captured as the specific object 6.

[0099] Next, the object detector 36 will be described. The object detector 36 is configured to detect an object region B1 representing the specific object 6 other than the pet 5 in the image data D1. In the present embodiment, the object detector 36 detects the object region B1 based on a learning model (hereinafter sometimes referred to as "fourth model M4") generated by learning appearance factors (feature amounts) of specific objects of a predetermined type to recognize the specific objects of the predetermined type on an image.

[0100] The fourth model M4 and the first to third models M1-M3 can include, for example, a model using a neural network or a model generated by deep learning using a multi-layer neural network. Nonetheless, the fourth model M4 does not have to be a model generated by deep learning. Alternatively, the fourth model M4 can be the same model as the first model M1, the second model M2, or the third model M3.

[0101] In this case, the fourth model M4 is generated by machine learning on specific objects that are highly relevant to certain actions of the dog (in particular, its actions related to certain emotions). For example, if the specific object 6 is a part of a person (e.g., the leg 61) or the entirety of a person (e.g., the entire appearance 62), then the pet 5 is highly likely to make an action related to a certain emotion. On the other hand, if the specific object 6 is a toy 63 or a bowl 64, then the pet 5 is highly likely to make a "play" action or an "eat" action. In other words, as a learning data set for generating the fourth model M4, among a large amount of image data in which objects other than the dog have been captured, image data in which an object that the dog is often interested in has been captured as the specific object is selected. The learning data set includes not only image data of real objects but also image data of objects generated by CG. Machine learning is performed using these types of image data combinations. In this case, the specific object is defined as an object other than the dog. Therefore, the object that the dog is often interested in can also include other types of animals (e.g., a cat).

[0102] The object detector 36 determines whether any specific object 6 is present in the image data D1 using the fourth model M4. When it is determined that any specific object 6 is present in the image data D1, the object detector 36 detects the object region B1 defined by a bounding box that encloses the specific object 6 (see Fig. 6). Figures 5A-6 However, the object region B1 does not have to be defined by the bounding box but can also be defined by, for example, segmentation. Note that the object detector 36 regards an object that does not correspond to the specific object 6 as "background".

[0103] The object detector 36 outputs information on the detected object region B1 (including information on the type of the specific object 6) to the distance determiner 333. If there is no specific object 6 in the image data D1 and no object region B1 is detected there, the object detector 36 notifies the distance determiner 333 of this effect.

[0104] The distance determiner 333 determines the pet-object distance based on the information on the head region A2 detected by the region detector 32, the information on the object region B1 detected by the object detector 36, and the information on the posture of the pet 5 determined by the posture determiner 331.

[0105] Specifically, the distance determiner 333 determines which of the following three distance relationships the pet-object distance has based on the distance from the position of the object region B1 (which can be the position of its upper left corner or the position of its center of gravity) to the position of the pet 5, for example. In this case, the three distance relationships are assumed to represent a first distance state (corresponding to a very short distance), a second distance state (corresponding to a relatively short distance), and a third distance state (corresponding to a relatively long distance). The first, second, and third distance states can be classified based on the number of pixels, for example. In this example, three distance relationships are provided. However, this is just an example and should not be construed as limiting. Alternatively, the number of distance relationships can also be two or four or more. Also alternatively, the distance relationships can also be unbounded (in terms of pixels). In this example, the "position of the pet 5" is assumed to be defined by the position of the head region A2 (which can be the position of its upper left corner or the position of its center of gravity). However, this is just an example and should not be construed as limiting. Alternatively, the "position of the pet 5" can also be defined by the position of the specific region A1 (which can be the position of its upper left corner or the position of its center of gravity).

[0106] The distance determiner 333 preferably determines the pet-object distance while further taking into account the degree (or area) of overlap between the object region B1 and the head region A2 (or the specific region A1).

[0107] Meanwhile, even if the pet 5 is not actually interested in the specific object 6, the pet 5 and the specific object 6 can be arranged in the depth direction and can have been photographed in the image data Dl to overlap each other. If the distance determiner 333 determines the pet-object distance based only on the distance from the position of the specific object 6 to the position of the pet 5 in the image data Dl, the distance determiner 333 will determine the pet-object distance to be the first distance state even if the pet 5 has not actually made any action related to the specific object 6. Therefore, the distance determiner 333 determines which of the first, second, and third distance states the pet-object distance corresponds to while taking into account the information about the posture of the pet 5 determined by the posture determiner 331.

[0108] For example, even if the specific object 6 is the bowl 64 and the distance from the position of the bowl 64 to the position of the pet 5 is in the first distance state, the pet 5 has not taken a posture of lowering its head section 50, the distance determiner 333 can regard this frame of the image data Dl as representing the third distance state. Alternatively, the distance determiner 333 can regard this frame of the image data Dl as an outlier and end the evaluation process.

[0109] The distance determiner 333 outputs the determination result about the pet-object distance, the information about the head region A2, and the posture information to the evaluator 34.

[0110] If the object detector 36 does not detect the object region Bl, the distance determiner 333 skips the step of determining the pet-object distance and outputs the information about the head region A2 and the posture information to the evaluator 34.

[0111] In the present embodiment, the pet detector Xl performs the process of detecting the specific region Al using the region detector 32 and the process of generating the pet information using the information generator 33 in this order. However, this is only an example and should not be construed as limiting. Alternatively, the pet detector Xl can perform the detection process and the generation process substantially in parallel with each other.

[0112] The evaluator 34 is configured to evaluate the pet state about the emotion and / or action of the pet 5 appearing in the specific region Al based on the pet information. In this example, the evaluator 34 evaluates the pet state based on the pet information and the condition information 9.

[0113] As described above, the pet information includes the posture information about the posture of the pet 5 determined by the posture determiner 331, the information about the facing direction of the pet 5 determined by the facing direction determiner 332, and the information about the pet-object distance determined by the distance determiner 333.

[0114] As used herein, the condition information 9 is information on a specific action and / or a specific emotion of the pet 5, which has been specified in advance as a target to be extracted. For example, pieces of information on the correspondence shown in Tables 1-4 below (hereinafter sometimes referred to as "patterns") are exemplary pieces of condition information 9. A large number of such patterns are prepared and stored as a database in the condition storage device P2.

[0115] Table 1

[0116]

[0117] Table 2

[0118]

[0119] Table 3

[0120]

[0121] Table 4

[0122]

[0123] The evaluator 34 searches the condition information 9 for any condition combination (pattern) that matches the obtained pet information. Note that at this time, the evaluator 34 determines whether or not the pet 5 is facing a specific object 6 based on the information on the facing direction of the pet 5 and the information on the object region Bl (for example, whether or not the object region Bl is present in the line of sight of the pet 5), and also takes the determination result into account when searching the condition information 9.

[0124] For example, assume that the obtained pet information (also taking the determination result into account) includes three results, namely "first distance state", "four legs standing and lowering head section", and "facing bowl". The evaluator 34 searches the condition information 9 for any condition combination (pattern) that matches these results. In this example, there is one matching condition combination (pattern) as shown in Table 1, and it is associated with "eating / delicious" of "action / emotion". Thus, the evaluator 34 evaluates the state of the pet 5 appearing in the image data Dl as "eating / delicious".

[0125] On the other hand, assume that the obtained pet information includes three results, namely "first distance state", "standing on hind legs only", and "facing person". The evaluator 34 searches the condition information 9 for any condition combination (pattern) that matches these results. In this example, there is one matching condition combination (pattern) as shown in Table 2, and it is associated with "showing love / happy" of "action / emotion". Thus, the evaluator 34 evaluates the state of the pet 5 appearing in the image data Dl as "showing love / happy".

[0126] Further, assume that the obtained pet information includes three results, i.e., "the third distance state", "four legs standing and showing teeth", and "facing a person". The evaluator 34 searches for any condition combination (pattern) that matches these results in the condition information 9. In this example, there is one matching condition combination (pattern) as shown in Table 3, and it is associated with "action / emotion" of "threatening / angry". Therefore, the evaluator 34 evaluates the state of the pet 5 appearing in the image data Dl as "threatening / angry".

[0127] Further, assume that the obtained pet information includes three results, i.e., "the second distance state", "one front leg off the ground standing", and "facing a toy". The evaluator 34 searches for any condition combination (pattern) that matches these results in the condition information 9. In this example, there is one matching condition combination (pattern) as shown in Table 4, and it is associated with "action / emotion" of "playing / happy". Therefore, the evaluator 34 evaluates the state of the pet 5 appearing in the image data Dl as "playing / happy".

[0128] In the exemplary patterns shown in Tables 1-4, each condition combination (pattern) is associated with both an action and an emotion. However, this is only an example and should not be construed as limiting. Alternatively, each condition combination (pattern) can be associated with only one action or only one emotion. In addition, the condition combination does not have to be three conditions (i.e., distance from the object, posture of the pet, and facing direction of the pet) as long as it includes a condition regarding "posture of the pet". For example, the conditions can include a condition regarding "flat area" in which the head region A2 and the object region Bl overlap each other.

[0129] As can be seen, the condition information 9 according to the present embodiment includes facing direction information in which a plurality of directions in which the pet 5 is facing (i.e., directions facing the bowl 64, the person, and the toy 63) are associated with a plurality of pet states (i.e., eating / delicious, showing love / happy, and playing / happy). The evaluator 34 evaluates the pet state based on the determination result of the facing direction determiner 332 and the facing direction information. This increases the reliability of evaluating the state of the pet 5.

[0130] Further, the state information 9 according to the present embodiment further includes information in which a plurality of types of specific objects 6 (i.e., the bowl 64, the person, and the toy 63) are associated with a plurality of thresholds (for the first, second, and third distance states) regarding the distance between the pet 5 and the specific object 6. The evaluator 34 evaluates the pet state by comparing the determination result of the distance determiner 333 with the plurality of thresholds. This makes the evaluation of the state of the pet 5 more reliable. Note that if the pet-object distance determined by the distance determiner 333 is not any one of the first, second, and third distance states, but information expressed as a numerical value (e.g., a numerical value corresponding to the number of pixels), then the plurality of thresholds can also be information expressed as a numerical value.

[0131] In particular, if the specific object 6 appearing in the object region B1 detected by the object detector 36 is a bowl 64 and the distance determined by the distance determiner 333 is equal to or less than a predetermined threshold, the evaluator 34 evaluates the state of the pet 5 as "eating". This evaluation is based on the fact that if the specific object 6 is a bowl 64, the deeper the pet 5 sticks its nose into the bowl 64, the closer the pet 5 is to the specific object 6. Therefore, if the pet 5 photographed in the image data D1 is indeed eating something, the pet state is more likely to be evaluated as "eating".

[0132] According to the present embodiment, the evaluator 34 can evaluate the pet state even if there is no specific object 6 in the image data D1 and no object region B1 is detected from the image data D1. For example, the condition information 9 can include a pattern in which only the posture of the pet is related to at least one of a specific action or a specific emotion of the pet. Specifically, the posture of the pet "curling up on the ground with eyes closed" is associated with the "action / emotion" of "sleeping / resting". Therefore, the evaluator 34 evaluates the state of the pet 5 appearing in the image data D1 as "sleeping / resting" only by the posture of the pet.

[0133] The output interface 35 is configured to output the evaluation result made by the evaluator 34 (i.e., the evaluated pet state). In particular, in the present embodiment, the output interface 35 outputs the evaluation result made by the evaluator 34 in relation to the image data D1 in which the specific region A1 on which the evaluation result is based has been detected. The output interface 35 transmits the evaluation result (e.g., "sleeping / resting") and information (hereinafter referred to as "output information") correlating the evaluation result and the image data D1 to the presentation device 4 through the communication interface 11. If the user 300 carrying the presentation device 4 is not at home, the output information can be transmitted to the presentation device 4 through the server 7. The output information preferably further includes information on the time at which the image data D1 on which the evaluation result is based was captured by the image capturing device 2.

[0134] The output information is preferably stored in a memory built in, for example, the pet camera 100. Alternatively, the output information can also be transmitted to the server 7 or any other peripheral device and stored therein.

[0135] Upon receiving the output information from the pet camera 100, the presentation device 4 can replace the pet state included in the output information with, for example, a simple expression (message) and present the message on the screen as, for example, a push notification carrying the message. When the user 300 opens the received push notification, the presentation device 4 can launch the pet application and present the specific pet state including the image data D1 on the screen (refer to FIG. 6). The presentation device 4 can also present the specific pet state including the image data D1 on the screen when the user 300 launches the pet application. Figure 7A and 7BAlternatively, the output information can also be sent to the recipient as an email via an email server.

[0136] exist Figure 7A In the example shown, the demonstration device 4 displays image data D1 (reference) that forms the basis of the pet's condition assessment on the screen 430 of the display device 43. Figure 3C (The posture of lying on the floor). In this case, conditional information 9 includes a pattern where two conditions, "no specific object detected" and "the posture of lying on the floor," are associated with the emotion of "loneliness." Therefore, this is an example of a pet's state being assessed as "loneliness." Demonstration device 4 converts the pet 5's "loneliness" emotion into the more informal expression "miss you," and displays the string data including this expression as an overlay image in a balloon box on image data D1.

[0137] On the other hand, Figure 7B In the example shown, the demonstration device 4 displays image data D1 (reference) that forms the basis of the pet's condition assessment on the screen 430 of the display device 43. Figure 6 (Posture of standing on all fours with head lowered). In this case, conditional information 9 includes a pattern where three conditions, namely "first distance state," "standing on all fours with head lowered," and "facing the bowl," are associated with the "action / emotion" of "eating / delicious." Thus, this is an example of a pet's state being evaluated as "eating / delicious." Demonstration device 4 converts the pet 5's emotion "delicious" into the more informal expression "delicious" and displays the string data including the word "eat" and the informal expression as an overlay image in a balloon box on image data D1.

[0138] Note that the demonstration device 4 preferably further displays the date and time when the image data D1 was captured on the screen 430 of the display device 43.

[0139] Output interface 35 does not necessarily need to transmit output information including image data D1 (raw data) that forms the basis of the evaluation results; instead, it can transmit output information after the image data has been processed. Alternatively, output interface 35 can transmit output information after replacing image data D1 with an icon image corresponding to the evaluated pet state (e.g., an icon image representing a dog that looks lonely and is crying). Data processing and replacing image data D1 with icon images can be performed by demonstration device 4 or server 7, whichever is appropriate.

[0140] The evaluation result made by the evaluator 34 does not necessarily have to be output as a screen image, but can be output as voice information instead of a screen image.

[0141] The processing device 3 performs the evaluation process each time the acquirer 31 acquires the image data Dl. For example, if the image capturing device 2 captures still pictures at predetermined intervals (e.g., several minutes or tens of minutes), the processing device 3 can perform the evaluation process at the predetermined intervals in general. Alternatively, if the image capturing device 2 captures moving pictures at a predetermined frame rate, the processing device 3 can perform the evaluation process by acquiring a plurality of frame pictures as the image data Dl at constant time intervals (e.g., several minutes or tens of minutes), which are consecutive to each other in the moving pictures. The output interface 35 can transmit the output information to the presentation device 4 each time the evaluator 34 evaluates the pet state with respect to a single frame of the image data Dl. Alternatively, the output interface 35 can collectively transmit the output information after the output information is collected to a certain extent in a memory built in, for example, the pet camera 100.

[0142] In addition, if the evaluation results made by the evaluator 34 with respect to a plurality of frames of the image data Dl indicate that the pet 5 is continuously taking a posture facing the same direction a predetermined number of times (e.g., twice), the output interface 35 can restrict the output of the evaluation results made by the evaluator 34. Specifically, assume a case where the posture and the facing direction of the pet 5 with respect to one frame of the image data Dl have been determined as "four-leg standing and lowering the head section" and "facing the bowl" (i.e., making an "eating" action), and the output information indicating this effect has been output to the presentation device 4. In this case, if the posture and the facing direction of the pet 5 with respect to a plurality of subsequent frames of the image data Dl obtained after the one frame are determined to be the same as the posture and the facing direction with respect to the one frame of the image data Dl, the output interface 35 does not necessarily output this evaluation result. If a plurality of output information is collected in, for example, a built-in memory, and the output information successively indicates the same evaluation result a predetermined number of times, the output interface 35 can collectively transmit these output information as a single evaluation result. The setting of the "predetermined number of times" can be appropriately changed according to the operation instruction input to the user 300 of the pet camera 100 or the presentation device 4.

[0143] Restricting the output of the evaluation results in this way can reduce the chance of similar evaluation results being successively output, thereby contributing to the reduction of the processing load and the reduction of the communication congestion. Furthermore, this also reduces the chance of the user 300 being repeatedly notified of the same pet state (e.g., "eating") in a short time, thereby improving the user friendliness of the pet state evaluation system 1.

[0144] (2.3) Operation Explanation

[0145] Next, the operation of the pet state evaluation system 1 will be described with reference to Figure 8 and Figure 9Briefly describe how the pet management system 200 according to the present embodiment operates. Note that in the following operation description, the order in which the respective processing steps are executed is only an example and should not be construed as limiting. In the following description, an example in which the pet detector X1 executes the specific area Al detection processing step and the pet information generation processing step in this order will be described. However, this is only an example and should not be construed as limiting. In addition, these two processing steps can also be executed substantially simultaneously.

[0146] The pet camera 100 installed in the user 300's living facility mainly monitors the predetermined management area in which the pet 5 will act by capturing an image of the predetermined management area using the image capturing device 2. The pet camera 100 can capture a still picture of the management area at a predetermined time interval, or continuously shoot a moving picture of the management area for a predetermined period of time.

[0147] As shown in Figure 8 When acquiring the image data D1 (which can be a still picture or one frame of a moving picture) captured by the image capturing device 2 (in S1), the pet status evaluation system 1 of the pet camera 100 executes an evaluation process (in S2).

[0148] The pet status evaluation system 1 causes the area detector 32 to determine whether any dog (as a type of pet 5) has been photographed as the subject H1 in the image data D1 using the second model M2 (in S3). When it is found that a dog (as a type of pet 5) has been photographed (if the answer is "Yes" in S3), the pet status evaluation system 1 detects the specific area Al representing the pet 5 (in S4: pet detection step), and then the process proceeds to a step of determining whether the head section 50 has been photographed (in S5).

[0149] In the present embodiment, even if it has been decided that no dog (as a type of pet 5) has been photographed in the image data D1 (if the answer is "No" in S3), the process proceeds to a step of determining whether the head section 50 has been photographed (S5). This measure is taken in order to handle a detection failure with respect to a dog in the case where the image data D1 is a close-up of the dog's face, as described above.

[0150] The pet state assessment system 1 uses the region detector 32 to determine, using the third model M3, whether the head segment 50 of a dog (as a type of pet 5) has been captured in image data D1 (in S5). When it is determined that the head segment 50 has been captured (if the answer in S5 is "yes"), the pet state assessment system 1 detects the head region A2 representing the head segment 50 (in S6). In this embodiment, unless the head segment 50 has been captured (if the answer in S5 is "no"), the pet state assessment system 1 ends the assessment processing of the image data D1 and waits to acquire the next image data D1 (i.e., the process returns to S1). Nevertheless, as long as the specific region A1 has been detected, the pet state assessment system 1 can continue to perform the assessment processing even if the head region A2 has not been detected.

[0151] If the pet condition assessment system 1 detects a specific area A1 after detecting the head area A2 (if the answer in S7 is "yes"), the process proceeds to the step of determining whether any specific object 6 has been photographed (S9; see reference). Figure 9 On the other hand, if the pet status assessment system 1 does not detect the specific area A1 after detecting the head area A2 (if the answer in S7 is "no"), the pet status assessment system 1 sets the area that is substantially the same size as the head area A2 as the specific area A1 (in S8), and the process proceeds to the step of determining whether any specific object 6 has been photographed (S9).

[0152] like Figure 9 The pet state assessment system 1 shown enables the object detector 36 to determine, using the fourth model M4, whether any specific object 6 has been captured in the image data D1 (in S9). When it is determined that a specific object 6 has been captured (if the answer in S9 is "yes"), the pet state assessment system 1 detects the object region B1 representing that specific object 6 (in S10), and the process proceeds to the pose determination step (in S12). On the other hand, when it is determined that no specific object 6 has been captured (if the answer in S9 is "no"), the pet state assessment system 1 determines that no object region B1 has been detected (in S11), and the process proceeds to the pose determination step (in S12).

[0153] The pet state assessment system 1 enables the posture determiner 331 to determine the posture of the dog (as pet 5) in S12 based on the first model M1 and information about a specific region A1.

[0154] Next, the pet status assessment system 1 enables the facing direction determiner 332 to determine the direction the pet 5 is facing based on information about a specific area A1 and information about the head area A2 (in S13).

[0155] Further, the pet state evaluation system 1 also causes the distance determiner 333 to determine the distance between the pet and the object based on the information about the head region A2, the information about the object region Bl, and the posture information (in S14). Note that if the object region Bl is not detected, this processing step S14 is skipped.

[0156] The pet state evaluation system 1 generates pet information based on the determination results in the processing steps S12-S14 (in S15: information generation step).

[0157] Then, the pet state evaluation system 1 evaluates the pet state based on the pet information and the condition information 9 (in S16: evaluation step).

[0158] The pet state evaluation system 1 transmits output information in which the pet state thus evaluated and the image data Dl are associated with each other to the presentation device 4 and causes the presentation device 4 to present the output information (in S17).

[0159] [Advantages]

[0160] As can be seen from the foregoing description, in the pet state evaluation system 1, the evaluator 34 evaluates the pet state about the emotion and / or the action of the pet 5 based on the pet information, thereby eventually making it easier to recognize the state of the pet.

[0161] Further, the pet state evaluation system 1 according to the present embodiment includes the facing direction determiner 332 that determines the direction in which the pet 5 faces in the image data Dl. This can increase the reliability of the pet state evaluation by taking into account the direction in which the pet 5 faces. Further, the facing direction determiner 332 determines the direction in which the pet 5 faces based on the relative position of the head region A2 with respect to the specific region Al. This can increase the reliability of the determination about the direction in which the pet 5 faces.

[0162] Further, the output interface 35 outputs the evaluation result made by the evaluator 34 in association with the image data Dl in which the specific region Al that constitutes the basis of the evaluation result has been detected. This makes it easier to recognize the state of the pet.

[0163] The pet state evaluation system 1 allows the user 300 to more easily recognize the action and / or emotion of the pet 5 based on the state of the pet evaluated by the pet state evaluation system 1, thereby making it easier for the user 300 to more smoothly communicate with his or her pet 5. In addition, the pet state evaluation system 1 also makes it easier for the user 300 to recognize the action and / or emotion of the pet 5 in the living facility based on the notification issued by the demonstration device 4, thereby being able to lovingly manage (watch) the pet 5 even when the user 300 is not at home. This makes it possible to promptly notify the user 300 of such an emergency situation, particularly when the evaluated pet state indicates that the pet 5 is making a warning action (e.g., the pet 5 is not comfortable or is listless).

[0164] (3) Variations

[0165] Note that the above-described embodiments are merely exemplary ones of the various embodiments of the present disclosure and should not be construed as limiting. Rather, the exemplary embodiments can be modified in various ways at any time without departing from the scope of the present disclosure, in accordance with design choices or any other factors. Also, the functions of the pet state evaluation system 1 according to the above-described embodiments can also be implemented as, for example, a pet state evaluation method, a computer program, or a non-transitory storage medium on which a computer program is stored.

[0166] Next, variations of the exemplary embodiments will be described one by one. Note that some of the variations to be described below can be adopted in combination as appropriate. In the following description, the above-described exemplary embodiments will sometimes be referred to hereinafter as "basic examples".

[0167] The pet status evaluation system 1 according to the present disclosure includes a computer system. The computer system can include a processor and a memory as its main hardware components. The functions of the pet status evaluation system 1 according to the present disclosure can be executed by causing the processor to execute a program stored in the memory of the computer system. The program can be stored in the memory of the computer system in advance. In addition, the program can also be downloaded through a telecommunication line, or distributed after being recorded in some non-transitory storage medium such as a memory card, an optical disc, or a hard disk drive, any of which can be read by the computer system. The processor of the computer system can be composed of a single or multiple electronic circuits, including a semiconductor integrated circuit (IC) or a large-scale integrated circuit (LSI). As used herein, an "integrated circuit" such as an IC or an LSI is referred to by different names according to the degree of integration thereof. Examples of the integrated circuit include a system LSI, a very large scale integrated circuit (VLSI), and an ultra large scale integrated circuit (ULSI). Alternatively, a field programmable gate array (FPGA) that is programmed after the LSI is manufactured, or a reconfigurable logic device that allows the connections or circuit parts within the LSI to be reconfigured can also be employed as the processor. These electronic circuits can be integrated together on one chip, or distributed on multiple chips as appropriate. Those multiple chips can be gathered together in a single device, or distributed in multiple devices without limitation. As used herein, a "computer system" includes a microcontroller that includes one or more processors and one or more memories. Therefore, the microcontroller can also be implemented as a single or multiple electronic circuits, including a semiconductor integrated circuit or a large-scale integrated circuit.

[0168] In the above-described exemplary embodiment, the plurality of functions of the pet status evaluation system 1 are aggregated together in a single housing. However, this is not essential to the configuration of the pet status evaluation system 1. For example, the constituent elements of the pet status evaluation system 1 can be distributed in multiple housings. Specifically, at least some of the learning models in the first to fourth models M1-M4 of the pet status evaluation system 1 can be provided outside the pet camera 100 (for example, in an external server such as the server 7).

[0169] On the contrary, the plurality of functions of the pet status evaluation system 1 can be aggregated together in a single housing (i.e., the housing of the pet camera 100) as in the basic example described above. In addition, at least some of the functions of the pet status evaluation system 1 (for example, some of the functions of the pet status evaluation system 1) can also be implemented, for example, as a cloud computing system.

[0170] (3.1) First variant example

[0171] Next, the first variant example of the present disclosure will be described with reference to Figure 10 The first variant example of the present disclosure will be described. Figure 10The pet state evaluation system 1A according to this variant example is explained. In the following description, any constituent element of this first variant example that has substantially the same function as the counterpart according to the basic example of the pet state evaluation system 1 described above will be designated with the same reference numeral as the counterpart, and its description will be appropriately omitted here.

[0172] In the pet state evaluation system 1 according to the basic example described above, the pet detector X1 includes the region detector 32 and the information generator 33, and causes the region detector 32 to detect the pet 5, and then causes the posture determiner 331 of the information generator 33 to determine the posture of the pet 5 to generate the posture information. That is, the presence / absence of the pet 5 in / from the acquired image data D1 is first detected, and then the posture thereof is determined.

[0173] In the pet state evaluation system 1A according to this variant example, the region detector 32 has the function of the posture determiner 331 as shown in Figure 10

[0174] In this variant example, the region detector 32 detects the specific region Al of the pet 5 that takes a specific posture from the image data D1 based on the first model M1 generated by learning the posture of the pet to recognize the posture of the pet on the image. In this variant example, the region detector 32 determines whether the pet 5 that takes a specific posture has been photographed as the subject H1 in the image data D1 using, for example, the first to fourth models M1-M4, thereby detecting the specific region Al representing the pet 5 that takes a specific posture. As used herein, the "specific posture" refers to a posture related to the action related to the depth of emotion of a dog described above. Examples of the specific posture include sitting, lying, curling up on the floor, and standing on four legs.

[0175] The information on the specific region Al representing the pet 5 that takes a specific posture is input to the information generator 33, and is used by the facing direction determiner 332 to determine the facing direction, and is used by the distance determiner 333 to determine the distance.

[0176] In short, the pet detector X1 according to this variant example is configured to detect the pet 5 that takes a specific posture, rather than detecting the presence of the pet 5 and then determining the posture thereof.

[0177] The configuration of this variant example also makes it easier to recognize the state of the pet 5.

[0178] (3.2) Other variant examples

[0179] Next, other variant examples will be enumerated one by one.

[0180] ​In the basic example described above, the evaluator 34 evaluates the pet state based on the pet information and the condition information 9. However, this is only one example and should not be construed as limiting. In addition, instead of using the condition information 9, the evaluator 34 can evaluate the pet state using not only the pet information but also a learning model (classifier) generated by causing a machine to learn a specific action and / or a specific emotion of the pet. Upon receiving the pet information, the classifier classifies the pet information into the specific action and / or the specific emotion of the pet.

[0181] In the basic example described above, the number of dogs (as a type 5 of pet) photographed as the subject H1 in the single frame of the image data D1 is assumed to be one. However, naturally, two or more dogs (as the type 5 of pet) (for example, two dogs as a pair of a parent dog and / or its puppy) can be photographed as the subject H1 in the single frame of the image data D1. If the pet state evaluation system 1 detects a plurality of specific regions A1 in the single frame of the image data D1, the pet state evaluation system 1 generates pet information and evaluates the pet state for each of the plurality of specific regions A1.

[0182] In the basic example described above, the number of specific objects 6 photographed in addition to the pet 5 in the single frame of the image data D1 is assumed to be zero or one. However, naturally, the number of specific objects 6 photographed in the single frame of the image data D1 can also be two or more. If the pet state evaluation system 1 detects a plurality of object regions B1 in the single frame of the image data D1, the pet state evaluation system 1 determines the distance between the pet 5 and each of the plurality of object regions B1. In this case, the pet state evaluation system 1 can evaluate the pet state by selecting one object region B1 from the plurality of pet-object distances, which is located at the shortest distance from the pet 5.

[0183] In the basic example described above, the pet state evaluation system 1 has a function of determining the facing direction of the pet 5 (i.e., the facing direction determiner 332) and a function of determining the pet-object distance (i.e., the distance determiner 333). However, these functions are not essential functions and can be omitted as appropriate.

[0184] At least some of the first to fourth models M1-M4 in the basic example can be machine-learned by reinforcement learning. In this case, in consideration of the processing load of the reinforcement learning, those models are preferably provided outside the pet camera 100 (for example, in an external server such as the server 7).

[0185] (4) SUMMARY

[0186] As can be seen from the above description, the pet status evaluation system (1, 1A) according to the first aspect comprises a region detector (32), an information generator (33), and an evaluator (34). The region detector (32) detects, in the image data (D1), a specific region (A1) representing at least a part of an appearance of the pet (5) as the subject (H1). The information generator (33) generates pet information. The pet information includes posture information about at least a posture of the pet (5). The posture information is based on a learning model (first model M1) and the image data (D1). The learning model (first model M1) has been generated by learning a posture of a pet so as to recognize the posture of the pet on an image. The evaluator (34) evaluates a pet status based on the pet information, the pet status relating to at least one of an emotion of the pet (5) appearing in the specific region (A1) or an action of the pet (5) appearing in the specific region (A1). According to the first aspect, the evaluator (34) evaluates the pet status about the emotion and / or the action of the pet (5) based on the pet information, thereby eventually making the status of the pet (5) more easily recognizable.

[0187] In the pet status evaluation system (1, 1A) according to the second aspect, which can be implemented together with the first aspect, the evaluator (34) evaluates the pet status based on the pet information and condition information (9) about at least one of a specific action of the pet or a specific emotion of the pet. The second aspect enables the pet status evaluation system (1, 1A) to be implemented in a simpler configuration compared to, for example, a case where a learning model generated by machine learning is used to evaluate the pet status.

[0188] In the pet status evaluation system (1, 1A) according to the third aspect, which can be implemented together with the first or second aspect, the region detector (32) detects the specific region (A1) based on a learning model (second model M2). The learning model (second model M2) has been generated by learning appearance factors of pets of a predetermined type so as to recognize the pets of the predetermined type on an image. The third aspect increases the reliability of the detection of the specific region (A1), thereby eventually increasing the reliability of the evaluation of the status of the pet (5).

[0189] In the pet status evaluation system (1, 1A) according to the fourth aspect, which can be implemented in combination with any one of the first to third aspects, the region detector (32) detects a head region (A2) representing a head section (50) of the subject (H1) based on a learning model (third model M3). The learning model (third model M3) has been generated by learning appearance factors of head sections of pets of a predetermined type so as to recognize the head sections of the pets of the predetermined type on an image. The fourth aspect increases the reliability of the detection of the head region (A2), thereby eventually increasing the reliability of the evaluation of the status of the pet (5).

[0190] In a pet status evaluation system (1, 1A) according to the fifth aspect, which can be implemented together with the fourth aspect, the information generator (33) comprises a facing direction determiner (332) which determines a facing direction of the pet (5) in the image data (D1) based on the image data (D1) in which the specific region (A1) has been detected. The pet information further comprises the result of the determination made by the facing direction determiner (332). The fifth aspect can increase the reliability of the evaluation of the status of the pet (5) by taking into account the direction in which the pet (5) is facing.

[0191] In a pet status evaluation system (1, 1A) according to the sixth aspect, which can be implemented together with the fifth aspect, the facing direction determiner (332) determines the direction in which the pet (5) is facing based on at least the relative position of the head region (A2) with respect to the specific region (A1). The sixth aspect can increase the reliability of the determination of the direction in which the pet (5) is facing.

[0192] In a pet status evaluation system (1, 1A) according to the seventh aspect, which can be implemented together with the fifth or sixth aspect, the evaluator (34) evaluates the pet status based on the pet information and condition information (9) about at least one of a specific action of the pet or a specific emotion of the pet. The condition information (9) comprises facing direction information in which a plurality of directions in which the pet (5) is facing and a plurality of pet statuses are correlated with each other. The evaluator (34) evaluates the pet status based on the result of the determination made by the facing direction determiner (332) and the facing direction information. The seventh aspect can increase the reliability of the evaluation of the status of the pet (5).

[0193] A pet status evaluation system (1, 1A) according to the eighth aspect, which can be implemented together with any one of the fifth to seventh aspects, further comprises an output interface (35) which outputs the result of the evaluation made by the evaluator (34). The output interface (35) restricts the output of the result of the evaluation made by the evaluator (34) when the results of the evaluations made by the evaluator (34) for a plurality of frames of the image data (D1) indicate that the pet (5) has been facing in one and the same direction for a predetermined number of times in succession. The eighth aspect can reduce the chance of similar evaluation results being output in succession, thereby for example helping to reduce the processing load.

[0194] The pet state evaluation system (1, 1A) according to the ninth aspect, which can be implemented in combination with any one of the first to eighth aspects, further includes an object detector (36) that detects an object region (B1) representing a specific object (6) other than the pet (5) in the image data (D1). The information generator (33) includes a distance determiner (333) that determines a relative distance of the pet (5) with respect to the object region (B1). The pet information further includes a determination result made by the distance determiner (333). The evaluator (34) evaluates the pet state based on the pet information and condition information (9) about at least one of a specific action of the pet or a specific emotion of the pet. The condition information (9) includes information in which a plurality of types of specific objects and a plurality of threshold values about distances between the pet and the plurality of types of specific objects are associated with each other. The evaluator (34) evaluates the pet state by comparing the determination result made by the distance determiner (333) with the plurality of threshold values. The ninth aspect can increase reliability of evaluation of the pet (5) state by taking into account the relative distance of the specific region (A1) with respect to the object region (B1).

[0195] In the pet state evaluation system (1, 1A) according to the tenth aspect, which can be implemented in combination with the ninth aspect, the object detector (36) detects the object region (B1) based on a learning model (fourth model M4). The learning model (fourth model M4) has been generated by learning appearance factors of a predetermined type of specific object so as to recognize the predetermined type of specific object on an image. The tenth aspect increases reliability of detection of the object region (B1).

[0196] In the pet state evaluation system (1, 1A) according to the eleventh aspect, which can be implemented in combination with the ninth or tenth aspect, when the specific object (6) appearing in the object region (B1) detected by the object detector (36) is a bowl (64) and the relative distance determined by the distance determiner (333) is equal to or smaller than a predetermined threshold value, the evaluator (34) evaluates the pet state about an action of the pet (5) to eat. The eleventh aspect can increase a chance that the pet state is evaluated to be eating when the pet (5) appearing in the image data (D1) is actually eating.

[0197] The pet state evaluation system (1, 1A) according to the twelfth aspect, which can be implemented in combination with any one of the first to eleventh aspects, further includes an output interface (35). The output interface (35) outputs an evaluation result made by the evaluator (34) while associating the evaluation result with the image data (D1) in which a specific region (A1) that constitutes a basis for the evaluation result has been detected. The twelfth aspect can make it easier to recognize the state of the pet (5).

[0198] In a pet state evaluation system (1, 1A) according to a thirteenth aspect, which can be implemented together with any one of the first to twelfth aspects, the region detector (32) detects a specific region (A1) in which the pet (5) assumes a specific posture in the image data (D1) based on a learning model (first model M1). The learning model (first model M1) has been generated by learning the posture of the pet so as to recognize the posture of the pet on an image. The thirteenth aspect can make it easier to recognize the state of the pet (5).

[0199] A pet camera (100) according to a fourteenth aspect includes a pet state evaluation system (1, 1A) according to any one of the first to thirteenth aspects, and an image capturing device (2) that captures image data (D1). The fourteenth aspect can provide a pet camera (100) that can make it easier to recognize the state of the pet (5).

[0200] A server (7) according to a fifteenth aspect can communicate with a pet camera (100) equipped with an information generator (33) and an evaluator (34) of a pet state evaluation system (1, 1A) according to any one of the first to thirteenth aspects. The server (7) is equipped with a region detector (32). The fifteenth aspect can provide a server (7) that can make it easier to recognize the state of the pet (5).

[0201] A server (7) according to a sixteenth aspect can communicate with a pet camera (100) equipped with a region detector (32) of a pet state evaluation system (1, 1A) according to any one of the first to thirteenth aspects. The server (7) is equipped with an information generator (33) and an evaluator (34). The sixteenth aspect can provide a server (7) that can make it easier to recognize the state of the pet (5).

[0202] A pet state evaluation method according to a seventeenth aspect includes a pet detection step, an information generation step, and an evaluation step. The pet detection step includes detecting a specific region (A1) representing at least a portion of the appearance of a pet (5) as a subject (H1) in image data (D1). The information generation step includes generating pet information. The pet information includes posture information about at least a posture of the pet (5). The posture information is based on a learning model (first model M1) and the image data (D1). The learning model (first model M1) has been generated by learning the posture of the pet so as to recognize the posture of the pet on an image. The evaluation step includes evaluating a pet state based on the pet information about at least one of an emotion of the pet (5) appearing in the specific region (A1) or an action of the pet (5) appearing in the specific region (A1). The seventeenth aspect can provide a pet state evaluation method that can make it easier to recognize the state of the pet (5).

[0203] The program according to the eighteenth aspect is designed to cause one or more processors to execute the pet state evaluation method according to the seventeenth aspect. The eighteenth aspect can provide a function that can make the (5) state of the pet more easily identifiable.

[0204] Note that the constituent elements according to the second to thirteenth aspects are not essential constituent elements of the pet state evaluation system (1, 1A), and can be omitted as appropriate.

[0205] List of Reference Signs

[0206] 100 Pet camera

[0207] 1, 1A Pet state evaluation system

[0208] 2 Image capturing device

[0209] 32 Region detector

[0210] 33 Information generator

[0211] 332 Directional determiner

[0212] 333 Distance determiner

[0213] 34 Evaluator

[0214] 35 Output interface

[0215] 36 Object detector

[0216] 5 Pet

[0217] 50 Head section

[0218] 6 Specific object

[0219] 7 Server

[0220] 64 Bowl

[0221] 9 Condition information

[0222] A1 Specific region

[0223] A2 Head region

[0224] B1 Object region

[0225] D1 Image data

[0226] H1 Subject

[0227] M1-M4 First to fourth models (learning models)

Claims

1. A pet state evaluation system comprising: a region detector configured to detect, in image data, a specific region representing at least a portion of an appearance of a pet as a subject; an information generator configured to generate pet information including posture information on at least a posture of the pet, the posture information being based on a learning model and the image data, the learning model having been generated by learning a posture of the pet to recognize the posture of the pet on an image; an evaluator configured to evaluate a pet state related to at least one of an emotion of the pet appearing in the specific region or an action of the pet appearing in the specific region, based on the pet information; and an object detector configured to detect, in the image data, an object region representing a specific object other than the pet, the information generator including a distance determiner configured to determine a relative distance of the pet with respect to the object region, based on information on the specific region detected by the region detector, information on the object region detected by the object detector, and the posture information, the pet information further including a determination result made by the distance determiner, the evaluator being configured to evaluate the pet state based on the pet information and condition information on at least one of a specific action of the pet or a specific emotion of the pet, the condition information including information in which a plurality of types of specific objects and a plurality of thresholds on a distance between the pet and the specific object are associated with each other, and the evaluator being configured to evaluate the pet state by comparing the determination result made by the distance determiner with the plurality of thresholds.

2. The pet state evaluation system according to claim 1, wherein the region detector is configured to detect the specific region based on a learning model that has been generated by learning appearance factors of a predetermined type of pet so as to recognize the predetermined type of pet on the image.

3. The pet state evaluation system according to claim 1, wherein the region detector is configured to detect a head region representing a head section of the subject based on a learning model that has been generated by learning appearance factors of a head section of a predetermined type of pet so as to recognize the head section of the predetermined type of pet on the image.

4. The pet state evaluation system according to claim 3, wherein the information generator includes a facing direction determiner configured to determine a direction in which the pet is facing in the image data based on the image data in which the specific region has been detected, and the pet information further includes a determination result made by the facing direction determiner.

5. The pet state evaluation system according to claim 4, wherein the facing direction determiner is configured to determine the direction in which the pet is facing based on at least a relative position of the head region with respect to the specific region.

6. The pet state evaluation system according to claim 4, wherein ​ the condition information includes direction-of-face information in which a plurality of directions in which the pet faces and a plurality of pet states are associated with each other, and the evaluator is configured to evaluate the pet state based on a result of determination made by the direction-of-face determiner and the direction-of-face information.

7. The pet state evaluation system according to any one of claims 4 to 6, further comprising an output interface configured to output a result of evaluation made by the evaluator, wherein the output interface is configured to limit output of the result of evaluation made by the evaluator when the result of evaluation made by the evaluator on a plurality of frames of the image data indicates that the pet continuously faces the same direction for a predetermined number of times.

8. The pet state evaluation system according to claim 1, wherein the object detector is configured to detect the object region based on a learning model that has been generated by learning appearance factors of specific objects of a predetermined type so as to recognize the specific objects of the predetermined type on the image.

9. The pet state evaluation system according to claim 1 or 8, wherein the evaluator is configured to evaluate a pet state on an action of the pet to eat when the specific object appearing in the object region detected by the object detector is a bowl and the relative distance determined by the distance determiner is equal to or less than a predetermined threshold value.

10. The pet state evaluation system according to any one of claims 1 to 6, further comprising an output interface configured to output a result of evaluation made by the evaluator while associating the result of evaluation with the image data in which the specific region on which the evaluation result is based has been detected.

11. The pet state evaluation system according to any one of claims 1 to 6, wherein the region detector is configured to detect a specific region in which the pet takes a specific posture in the image data based on a learning model that has been generated by learning postures of the pet so as to recognize the posture of the pet on the image.

12. A pet camera comprising: the pet state evaluation system according to any one of claims 1 to 11; and an image capturing device configured to capture the image data.

13. A server configured to communicate with a pet camera equipped with the information generator and the evaluator of the pet state evaluation system according to any one of claims 1 to 11, the server being equipped with the region detector.

14. A server configured to communicate with a pet camera equipped with the region detector of the pet state evaluation system according to any one of claims 1 to 11, the server being equipped with the information generator and the evaluator.

15. A pet state evaluation method comprising a pet detection step of detecting a specific region representing at least a portion of an appearance of a pet as a subject in image data; the information generation step includes a distance determination step that determines a relative distance of the pet with respect to the object region based on information about the specific region detected by the pet detection step, information about the object region detected by the object detection step, and the posture information, the pet information further includes a determination result made by the distance determination step, the evaluation step evaluates the pet state based on the pet information and condition information about at least one of a specific action of the pet or a specific emotion of the pet, the condition information includes information in which a plurality of types of specific objects and a plurality of threshold values about distances between the pet and the specific objects are associated with each other, and the evaluation step evaluates the pet state by comparing the determination result made by the distance determination step with the plurality of threshold values.

16. A non-transitory storage medium storing a program designed to cause one or more processors to execute the pet state evaluation method according to claim 15. ​ ​ ​ ​

Citation Information

Patent Citations

  • Detection device, display control device and imaging control device provided with the detection device, object detection method, control program, and recording medium

    JP2013065110A

  • Animal behavior detection device based on distance measuring sensor

    CN109566451A

  • Intelligent pet training and accompanying method and device, equipment and storage medium

    CN111597942A

  • Pet monitoring method and pet monitoring system

    US20200205382A1