Image classification device, image classification method, and program
The image classification device uses a machine-learned model to classify images based on photographer heart rate and other biometric data to identify preferred images, addressing the inefficiency of existing systems and enabling easy retrieval of favorite pet images.
Patent Information
- Application Number
- JP2024507433
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-03-18
AI Technical Summary
Existing image classification systems, such as described in Patent Document 1, fail to effectively classify images that a pet owner prefers from a large collection of pet photos and videos, making it time-consuming to find favorite images.
An image classification device that utilizes a machine-learned model to classify images based on predetermined conditions, including heart rate of the photographer, to determine if a predetermined state of the target subject is satisfied, and outputs images that match the user's preferences.
Enables efficient classification of images that a user prefers, allowing pet owners to easily find and retrieve favorite images without manually sorting through numerous photos and videos.
Smart Images

Figure 0007729466000001 
Figure 0007729466000002 
Figure 0007729466000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a technique for classifying captured images. [Background technology]
[0002] There can be a huge amount of photos and videos of pets (hereinafter referred to as "pets") and other subjects, and among these photos and videos, there may be photos and videos that do not suit the pet owner's preferences, such as when the pet is facing away from them. It is time-consuming for pet owners to sort through such a huge amount of photos and videos to find their favorite ones. For example, Patent Document 1 describes a device that identifies and classifies the movements of a subject from multiple image data of the subject. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2005-267604 Summary of the Invention [Problem to be solved by the invention]
[0004] However, even with Patent Document 1, it is difficult to classify photos and videos that pet owners prefer.
[0005] One object of the present disclosure is to provide an image classification device that can classify images that a user prefers from a plurality of images. [Means for solving the problem]
[0006] In order to solve the above problem, in one aspect of the present disclosure, an image classification device includes: image acquisition means for acquiring an image including a target subject; a condition determination means for determining whether a predetermined condition for the target subject to be photographed is satisfied; and Images containing the target subject and a determination result as to whether a predetermined state occurrence condition is satisfied; anda predetermined state of the target subject; 、 Using a machine-learned model, The image acquired by the image acquisition means The image and the determination result determined by the condition determination means. image classification means for classifying images showing a predetermined state of the target subject from the image; an output means for outputting the image and the classification result; Equipped with The condition determination means determines whether or not the occurrence condition of the predetermined state is satisfied based on the heart rate of a photographer of the target subject. do.
[0007] In another aspect of the present invention, a computer implemented method for image classification comprises: Acquire an image containing the target subject Perform image acquisition processing , performing a condition determination process for determining whether a predetermined condition for the target subject to be photographed is satisfied; Images containing the target subject and a determination result as to whether a predetermined state occurrence condition is satisfied; and a predetermined state of the target subject; 、 Using a machine-learned model, The image acquired by the image acquisition process The image and the determination result determined by the condition determination process. Then, the images showing the target subject in a predetermined state are classified. Image classification processing is performed , Output the image and the classification results Output processing is performed. , The condition determination process determines whether or not the occurrence condition of the predetermined state is satisfied based on the heart rate of a photographer of the target subject. do.
[0008] In yet another aspect of the invention, a program includes: Acquire an image containing the target subject Perform image acquisition processing , performing a condition determination process for determining whether a predetermined condition for the target subject to be photographed is satisfied; Images containing the target subject and a determination result as to whether a predetermined state occurrence condition is satisfied; and a predetermined state of the target subject; 、 Using a machine-learned model, The image acquired by the image acquisition process The image and the determination result determined by the condition determination process. Then, the images showing the target subject in a predetermined state are classified. Image classification processing is performed , Output the image and the classification results Output processing is performed. , The condition determination process determines whether or not the occurrence condition of the predetermined state is satisfied based on the heart rate of a photographer of the target subject. The computer is caused to execute the process. [Effects of the Invention]
[0009] According to the present disclosure, it is possible to classify images that a user prefers from among a plurality of images. [Brief explanation of the drawings]
[0010] [Figure 1] 1 shows the overall configuration of an image classification system according to a first embodiment. [Figure 2] FIG. 2 is a block diagram showing the configuration of a server and a user terminal. [Figure 3] FIG. 2 is a block diagram showing the functional configuration of a server. [Figure 4] FIG. 2 is a block diagram showing the functional configuration of the learning device. [Figure 5] 1 is a flowchart of an image classification system. [Figure 6] FIG. 10 is a block diagram showing a functional configuration of a first modified example of the first embodiment. [Figure 7] FIG. 10 is a block diagram showing the functional configuration of an information processing apparatus according to a second embodiment. [Figure 8] 10 is a flowchart of a process performed by an information processing apparatus according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] First Embodiment [Overall configuration] 1 shows the overall configuration of an image classification system to which an image classification device according to the present disclosure is applied. The image classification system 1 includes a server 200 and a user terminal 300 used by a pet owner. The server 200 is an example of an image classification device. The server 200 and the owner's user terminal 300 are capable of wireless communication.
[0012] As a basic operation, the server 200 acquires an image showing a predetermined state of the pet based on a video transmitted from the owner's user terminal 300. Specifically, the owner sets the user terminal 300 to continuous recording mode and shoots a video when playing with the pet P, for example. The user terminal 300 then transmits the shot video (hereinafter also referred to as a "shot video") to the server 200. The server 200 extracts still images for each frame from the video shot by the user terminal 300 and classifies the images as showing a predetermined state of the pet through image analysis using AI (artificial intelligence). Here, an image showing a predetermined state of the pet (hereinafter also referred to as a "good shot") is an image of the pet that the pet owner finds pleasing, such as an image showing the pet's face, an image of the pet jumping, or an image of the pet playing. The server 200 then classifies still images (hereinafter also referred to as "extracted images") extracted from the video captured by the user terminal 300 as good shots or not, associates them with the owner, and stores them in a database. The owner then accesses the server 200 from the user terminal 300 or a terminal other than the user terminal 300, and checks only the good shots using a slideshow or the like. This allows the owner to capture images of their pet without missing a photo opportunity. Furthermore, by using smart glasses as the user terminal 300, the owner can capture good shots while interacting with their pet. Note that instead of smart glasses, other eyeglass-type wearable devices such as AR (Augmented Reality) glasses, MR (Mixed Reality) glasses, or VR (Virtual Reality) glasses may be used.
[0013] Note that images classified as GOOD shots are not limited to still images, but may also be videos. In this case, the server 200 extracts videos at predetermined time intervals from the videos captured by the user terminal 300. The server 200 then classifies the videos as to whether they contain GOOD shots, and saves the extracted videos (also referred to as "extracted images") with the classification result of whether they are GOOD shots or not.
[0014] [server] 2A is a block diagram showing the configuration of the server 200. The server 200 mainly includes a communication unit 211, a processor 212, a memory 213, a recording medium 214, and a database (DB) 215.
[0015] The communication unit 211 transmits and receives data to and from external devices. Specifically, the communication unit 211 transmits and receives information to and from the user terminal 300 of the owner.
[0016] The processor 212 is a computer such as a CPU (Central Processing Unit), and executes a prepared program to control the entire server 200. The processor 212 may be a GPU (Graphics Processing Unit), an FPGA (Field-Programmable Gate Array), a DSP (Demand-Side Platform), an ASIC (Application Specific Integrated Circuit), or the like.
[0017] The memory 213 is composed of a ROM (Read Only Memory), a RAM (Random Access Memory), etc. The memory 213 is also used as a working memory while the processor 212 is executing various processes. The memory 213 also temporarily stores a series of videos captured by the user terminal 300 under the control of the processor 212. These videos are stored in the memory 213 in association with, for example, the owner's identification information, timestamp information, etc.
[0018] The recording medium 214 is a non-volatile, non-transitory recording medium such as a disk-shaped recording medium or semiconductor memory, and is configured to be detachable from the server 200. The recording medium 214 stores various programs executed by the processor 212.
[0019] Database (DB) 215 stores extracted images with classification results of whether they are GOOD shots or not. DB 215 may include an external storage device such as a hard disk connected to or built into server 200, or may include a storage medium such as a removable flash memory. Note that instead of providing DB 215 in server 200, DB 215 may be provided in an external server or the like, and extracted images with classification results of whether they are GOOD shots or not may be stored in the server via communication.
[0020] The server 200 may also include an input unit such as a keyboard and a mouse for an administrator to give instructions and input, and a display unit such as a liquid crystal display.
[0021] [User device] 2(B) is a block diagram showing the internal configuration of a user terminal 300 used by a pet owner. The user terminal 300 is a terminal device such as smart glasses or a smartphone. The user terminal 300 includes a communication unit 311, a processor 312, a memory 313, a display unit 314, a camera 315, and a microphone 316.
[0022] The communication unit 311 transmits and receives data to and from an external device. Specifically, the communication unit 311 transmits and receives information to and from the server 200.
[0023] The processor 312 is a computer such as a CPU, and executes a program prepared in advance to control the entire user terminal 300. The processor 312 may be a GPU, FPGA, DSP, ASIC, or the like. The processor 312 executes the program prepared in advance to transmit video captured by the camera 315 to the server 200.
[0024] The memory 313 is composed of a ROM, a RAM, etc. The memory 313 stores various programs executed by the processor 312. The memory 313 is also used as a working memory while the processor 312 is executing various processes. The video captured by the camera 315 is stored in the memory 313 and then transmitted to the server 200. The display unit 314 is, for example, a liquid crystal display device, and displays the video captured by the camera 315, extracted images of good shots stored in the server 200, etc.
[0025] The camera 315 includes a camera that captures the user's field of view (also referred to as an "out-camera") and a camera that captures the user's eyeballs (also referred to as an "eye camera"). The out-camera is mounted on the outside of the user terminal 300. The out-camera captures the user's field of view, including subjects such as pets, and transmits the images to the server 200. This allows the server 200 to acquire images of subjects such as pets. The eye camera is mounted on the inside of the user terminal 300 to capture images of the user's eyeballs. The eye camera captures images of the user's eyeballs and transmits the images to the processor 312. The processor 312 detects the user's line of sight and other information based on the images of the user's eyeballs captured by the eye camera. This allows the user terminal 300 to acquire information such as the user's line of sight.
[0026] The microphone 316 collects the user's voice and surrounding sounds and transmits them to the server 200. For example, based on the user's voice and the pet's cries, the server 200 can estimate that the user has uttered a predetermined word or that the user has given an instruction or command to the pet.
[0027] [Function Configuration] 3 is a block diagram showing the functional configuration of the server 200. Functionally, the server 200 includes an image acquisition unit 411 and an image classification unit 412.
[0028] Video captured by the user terminal 300 is input to the server 200. The video captured by the user terminal 300 is input to the image acquisition unit 411. The image acquisition unit 411 extracts a still image or a video from the video captured by the user terminal 300 as an extracted image. The image acquisition unit 411 outputs the extracted image to the image classification unit 412.
[0029] The image classification unit 412 classifies the extracted image acquired from the image acquisition unit 411 as a good shot or not using a pre-prepared image recognition model or the like. This image recognition model is a machine learning model that has been trained in advance to classify images as good shots or not, and will hereinafter also be referred to as an "image classification model." If the extracted image is classified as a good shot by the image classification model, the image classification unit 412 attaches additional information indicating that it is a good shot to the extracted image. On the other hand, if the extracted image is classified as a bad shot by the image classification model, i.e., not a good shot, the image classification unit 412 attaches additional information indicating that it is a bad shot to the extracted image. A bad shot is an image other than a good shot, such as an image that does not show a pet's face. The image classification unit 412 outputs the extracted image with the attached additional information to the DB 215.
[0030] [Image classification model training] Next, we will explain the learning of the image classification model used by the image classification unit 412. The image classification model is generated by so-called supervised learning. Figure 4 is a block diagram showing the learning method of the image classification model, and includes learning data 511 and a learning device 512.
[0031] The learning data 511 is image data (hereinafter also referred to as "teaching data") that has been labeled in advance as to whether it is a GOOD shot or not. The labeling of image data is performed based on criteria such as whether a predetermined part of the pet is captured, whether the pet is performing a predetermined action, etc. The predetermined part of the pet refers to the pet's face, etc. For example, an image that captures the pet's face is labeled as a GOOD shot. On the other hand, an image that does not capture the pet, an image in which the pet is facing away, or an image in which only the pet's torso or legs are captured is labeled as a BAD shot. Furthermore, the predetermined action of the pet refers to an eye-catching action of the pet, etc. For example, an image in which the pet is jumping or holding a tool in its mouth is labeled as a GOOD shot.
[0032] Alternatively, pet owners can select multiple images of their pets as good or bad, and use the resulting labeled images as training data. This allows us to generate an image classification model that can classify images that better suit the pet owner's preferences.
[0033] Additionally, images of animals posted by pet owners or third parties on social networking services (SNS) can be collected and used as training data. In this case, images posted by pet owners or third parties on SNS are labeled as "good shots." This increases the amount of training data, making it possible to generate more accurate image classification models.
[0034] The learning device 512 learns patterns of good shots based on the learning data 511 and outputs an image classification model as a trained model. This generates an image classification model that has learned the relationship between images containing pets and the state of the pet that corresponds to a good shot.
[0035] [Classification using image classification model] The image classification unit 412 uses an image classification model to estimate whether an image is a good shot. Specifically, the image classification model estimates whether an input image is a good shot and calculates a score indicating the probability that the image is a good shot (referred to as a "good shot score") and a score indicating the probability that the image is a bad shot (referred to as a "bad shot score"). The image classification model calculates each score so that the sum of the good shot score and the bad shot score is "1." The image classification model then compares the good shot score and the bad shot score with a predetermined threshold TH and adopts the score greater than the threshold TH as the classification result. For example, for a certain image, the image classification model calculates a good shot score of "0.8" and a bad shot score of "0.2" and compares them with the predetermined threshold TH. If the threshold TH is "0.5," the image classification model estimates that the image is a good shot.
[0036] [Image classification processing] Next, the image classification process for performing the above-described image classification will be described. Fig. 5 is a flowchart of the image classification process performed in the server 200. This process is realized by the processor 212 shown in Fig. 2 executing a program prepared in advance and operating as each element shown in Fig. 3.
[0037] First, the image acquisition unit 411 acquires a shot video from the user terminal 300. Then, the image acquisition unit 411 acquires an image (a still image or a video) from the shot video (step S11). Next, the image classification unit 412 classifies the image acquired by the image acquisition unit 411 as whether it is a GOOD shot or not (step S12). Specifically, the image classification unit 412 calculates a score indicating the probability that the image is a GOOD shot and a score indicating the probability that the image is a BAD shot. The image classification unit 412 compares each calculated score with a threshold TH and classifies the image as either a GOOD shot or a BAD shot.
[0038] Next, the image classification unit 412 assigns classification results to the images acquired by the image acquisition unit 411 and stores them in the database (DB) 215 (step S13). For example, the image classification unit 412 assigns a flag such as "1" to an image classified as a GOOD shot and "0" to an image classified as a BAD shot, and stores them in the DB 215. Then, the image classification process ends.
[0039] As a result, images of good shots that match the user's preferences are extracted from the vast number of images taken by the user and stored in DB 215 of server 200. The user can access server 200 and view the images of good shots stored in DB 215. The user can also download the images of good shots from server 200 and save them in a terminal device such as user terminal 300.
[0040] [Variations] Next, a description will be given of modifications of the first embodiment. The following modifications can be applied to the first embodiment in appropriate combinations. (Variation 1) In the first embodiment described above, the server 200 classifies images based on extracted images extracted from a captured video. In addition to the above, the server 200 may determine whether a predetermined condition for occurrence is met and classify the images using the determination result. The predetermined condition for occurrence is a condition under which a good shot is estimated to have been captured, and is hereinafter also referred to as a "condition for occurrence of a good shot." The condition for occurrence of a good shot is determined, for example, based on biometric information and behavioral information of the photographer.
[0041] 6 shows the functional configuration of server 200a in Modification 1. As shown in the figure, in Modification 1, server 200a is provided with a condition determination unit 413. Condition determination unit 413 acquires biometric information of the photographer and the like together with a timestamp from user terminal 300. Then, using a trained model that has been trained in advance, condition determination unit 413 determines whether the biometric information of the photographer and the like meets predetermined conditions, and outputs the determination result to image classification unit 412.
[0042] The photographer's biometric information includes his / her gaze, voice, heart rate, etc. The photographer's biometric information is acquired by the user terminal 300. The user terminal 300 may acquire the biometric information from a camera, microphone, sensor, etc., built into the user terminal 300, or may acquire the biometric information from an external device by wirelessly communicating with the external device via Bluetooth (registered trademark) or Wi-Fi (registered trademark). Examples of predetermined conditions include the photographer's gaze at a pet, the photographer's voice volume exceeding a predetermined threshold, the photographer's utterance of a predetermined phrase such as "like," and the photographer's heart rate exceeding a predetermined threshold. If the photographer's biometric information satisfies the above conditions, it is estimated that a good shot was likely taken at that time and at times before and after that time. The condition determination unit 413 may determine that the photographer's biometric information satisfies the predetermined condition not only at the time when the photographer's biometric information satisfied the predetermined condition, but also at times before and after that time, and output the determination result to the image classification unit 412.
[0043] The condition determination unit 413 may also determine whether or not a condition for generating a GOOD shot is met based on behavioral information of the photographer and the pet. For example, if the photographer gives a signal and the pet acts in accordance with the signal, or if the photographer gives an instruction or command and the pet acts in accordance with the instruction or command, the condition determination unit 413 determines that the condition for generating a GOOD shot is met and outputs the determination result to the image classification unit 412. Note that the behavioral information of the photographer and the pet may be acquired from a microphone, sensor, etc. mounted on the user terminal 300, or may be acquired from videos captured by the user terminal 300.
[0044] The image classification unit 412 classifies the extracted image as to whether it is a GOOD shot or not, based on the extracted image input from the image acquisition unit 411 and the judgment result input from the condition judgment unit 413. In this case, the image classification model used by the image classification unit 412 is a trained model that has been trained in advance to estimate whether it is a GOOD shot or not, based on the extracted image and the judgment result.
[0045] As described above, by taking into account the photographer's biometric information and behavioral information and classifying the shot as good or not, it is possible to obtain with high accuracy an image of a pet captured at a moment that the photographer feels is good.
[0046] (Variation 2) Based on the GOOD shots classified by the first embodiment, training data for re-training the image classification model may be created. Specifically, the pet owner determines whether or not the GOOD shots classified by the server 200 are necessary. The server 200 determines that an image that the pet owner determines is necessary is a GOOD shot. On the other hand, the server 200 determines that an image that the pet owner determines is unnecessary is a BAD shot, and changes the label. The server 200 then uses the image data of the GOOD shots and the image data of the BAD shots as training data and re-trains the image classification model. This enables the server 200 to classify GOOD shots that better suit the pet owner's preferences.
[0047] (Variation 3) In the first embodiment described above, the user terminal 300 sets the camera to continuous recording mode and transmits the captured video to the server 200. Alternatively, the user terminal 300 may start recording when a subject is captured on the camera, stop recording when the subject is no longer captured on the camera, and transmit the captured video from the start to the end of recording to the server 200. Specifically, the user terminal 300 captures images captured on the camera at predetermined intervals and transmits them to the server 200. The server 200 determines whether a pet is captured on the camera of the user terminal 300 based on a pre-created image recognition model or the like. If a pet is captured on the camera of the user terminal 300, the server 200 sets the user terminal 300 to recording mode and starts recording. Thereafter, if the pet is no longer captured on the camera of the user terminal 300, the server 200 ends the recording mode of the user terminal 300. This reduces the amount of data of the captured video transmitted from the user terminal 300 to the server 200.
[0048] Note that the user terminal 300 may determine whether or not the pet is captured by the camera of the user terminal 300. In this case, the user terminal 300 may use a pre-created image recognition model or the like to determine whether or not the pet is captured by the camera of the user terminal 300. Then, the user terminal 300 may control the start and end of recording according to the determination result.
[0049] (Variation 4) In the first embodiment described above, the server 200 classifies GOOD shots based on videos shot with a pet as the subject, but the subject is not limited to a pet and may be another subject, such as a child, that often misses a photo opportunity.
[0050] (Variation 5) In the first embodiment described above, information acquired by the user terminal 300 is basically transmitted as is to the server 200, and the server 200 classifies the GOOD shots based on the received information. Alternatively, the user terminal 300 may perform a process for classifying the GOOD shots and transmit the processing results to the server 200. Alternatively, the server 200 may not be used, and the process for classifying the GOOD shots and the storage of the processing results may be performed by the user terminal 300. This reduces the communication load from the user terminal 300 to the server 200 and the processing load on the server 200. In these cases, the user terminal 300 is an example of an image classification device.
[0051] Second Embodiment 7 is a block diagram showing the functional configuration of an image classification device 50 according to the second embodiment. The image classification device 50 according to the second embodiment includes an image acquisition unit 51, an image classification unit 52, and an output unit 53.
[0052] 8 is a flowchart of processing by the image classification device 50. The image acquisition means 51 acquires images showing a target subject (step S51). The image classification means 52 classifies the images into images showing a predetermined state of the target subject using a machine-learned model of the relationship between images showing the target subject and a predetermined state of the target subject (step S52). The output means 53 outputs the images and the classification results (step S53).
[0053] According to the image classification device 50 of the second embodiment, it becomes possible to easily classify images that the user likes.
[0054] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.
[0055] (Appendix 1) image acquisition means for acquiring an image including a target subject; an image classification means for classifying, from the images, images showing a predetermined state of the target subject using a machine-learned model of the relationship between an image showing a target subject and a predetermined state of the target subject; an output means for outputting the image and the classification result; An image classification device comprising:
[0056] (Appendix 2) 2. The image classification device according to claim 1, wherein the image classification means classifies images that include a predetermined part of the target subject.
[0057] (Appendix 3) 3. The image classification device according to claim 1, wherein the image classification means classifies images in which the target subject is performing a predetermined action.
[0058] (Appendix 4) a condition determination means for determining whether or not the occurrence condition of the predetermined state is satisfied; 4. The image classification device according to claim 1, wherein the image classification means classifies the images based on the determination results of the images and the occurrence conditions.
[0059] (Appendix 5) 5. The image classification device according to claim 4, wherein the condition determination means determines whether the occurrence condition is satisfied based on a gaze direction of a photographer of the target subject.
[0060] (Appendix 6) 6. The image classification device according to claim 4, wherein the condition determination means determines whether the occurrence condition is met based on a heart rate of a photographer of the target subject.
[0061] (Appendix 7) 7. The image classification device according to claim 4, wherein the condition determination means determines whether the occurrence condition is met based on the voice of a photographer of the target subject.
[0062] (Appendix 8) The image classification device according to any one of appendices 4 to 7, wherein the condition determination means detects the voice of the photographer and determines that the target subject has acted in response to the voice of the photographer as the occurrence condition.
[0063] (Appendix 9) 9. The image classification device according to any one of claims 1 to 8, wherein the image acquisition means starts acquiring images containing the target subject when the target subject is captured by a camera of the terminal device, and stops acquiring images containing the target subject when the target subject is no longer captured by the camera of the terminal device.
[0064] (Appendix 10) An image classification device according to any one of claims 1 to 9, further comprising a learning means for re-training the model using images that have been judged necessary or not by a user from among the results output by the output means as training data.
[0065] (Appendix 11) Acquire an image containing the target subject, classifying, from the images, images showing a predetermined state of the target subject using a machine-learned model of the relationship between images showing the target subject and a predetermined state of the target subject; The image classification method outputs the image and the result of the classification.
[0066] (Appendix 12) Acquire an image containing the target subject, classifying, from the images, images showing a predetermined state of the target subject using a machine-learned model of the relationship between images showing the target subject and a predetermined state of the target subject; A recording medium on which a program for causing a computer to execute a process for outputting the image and the classification results is recorded.
[0067] Although the present disclosure has been described above with reference to the embodiments and examples, the present disclosure is not limited to the above-described embodiments and examples. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. [Explanation of symbols]
[0068] 200 servers 215 Database (DB) 300 User Terminals 411 Image acquisition unit 412 Image Classification Unit 413 Condition judgment section 511 training data 512 Learning Device
Claims
1. image acquisition means for acquiring an image including a target subject; a condition determination means for determining whether a predetermined condition for the target subject to be photographed is satisfied; and an image classification means for classifying images of a target subject that show a predetermined state based on the images acquired by the image acquisition means and the determination result made by the condition determination means, using a machine-learned model of the relationship between an image of the target subject, a determination result as to whether a condition for occurrence of a predetermined state is satisfied, and the predetermined state of the target subject; an output means for outputting the image and the classification result; Equipped with The condition determination means determines whether or not the occurrence condition of the predetermined state is satisfied based on the heart rate of a photographer of the target subject.
2. The image classification device according to claim 1 , wherein the image classification means classifies images that include a predetermined part of the target subject.
3. The image classification device according to claim 1 or 2, wherein the image classification means classifies images in which the target subject is performing a predetermined action.
4. The image classification device according to claim 1 , wherein the condition determination means determines whether the occurrence condition is satisfied based on a gaze direction of a photographer of the target subject.
5. The image classification device according to claim 1 , wherein the condition determination means determines whether the occurrence condition is met based on a voice of a photographer of the target subject.
6. The image classification device according to claim 1 , wherein the condition determination means detects a voice of a photographer and determines that the target subject has taken an action in response to the voice of the photographer as the occurrence condition.
7. 1. A computer-implemented method for image classification, comprising: An image acquisition process is performed to acquire an image containing the target subject. performing a condition determination process for determining whether a predetermined condition for the target subject to be photographed is satisfied; performing an image classification process to classify images of the target subject that show the predetermined state based on the images acquired by the image acquisition process and the determination result determined by the condition determination process, using a model in which the relationship between the image in which the target subject is captured, the determination result of whether or not a condition for occurrence of a predetermined state is satisfied, and the predetermined state of the target subject has been machine-learned; performing an output process for outputting the image and the classification result; The condition determination process determines whether or not the occurrence condition of the predetermined state is satisfied based on the heart rate of a photographer of the target subject.
8. An image acquisition process is performed to acquire an image containing the target subject. performing a condition determination process for determining whether a predetermined condition for the target subject to be photographed is satisfied; performing an image classification process to classify images of the target subject that show the predetermined state based on the images acquired by the image acquisition process and the determination result determined by the condition determination process, using a model in which the relationship between the image in which the target subject is captured, the determination result of whether or not a condition for occurrence of a predetermined state is satisfied, and the predetermined state of the target subject has been machine-learned; performing an output process for outputting the image and the classification result; The condition determination process is a program that causes a computer to execute a process of determining whether or not the occurrence condition of the predetermined state is satisfied based on the heart rate of a photographer of the target subject.
Citation Information
Patent Citations
Monitoring video recording system
JP1999205773A
Operation classification support device and operation classifying device
JP2005267604A
Information processing apparatus and information processing method, and computer program
JP2009059257A
Image processor, image processing method, and image processing program
JP2009259122A
Information processing device, information processing method, and program
JP2019047234A