Information processing apparatus, information processing method, and program
The information processing device enhances object recognition by using candidate area and estimated position detection to accurately identify a person's belongings, addressing the issue of incomplete body part detection in existing technologies.
Patent Information
- Application Number
- JP2025036119
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2038-04-26
AI Technical Summary
Existing object recognition technologies fail to accurately detect a person's belongings when certain body parts, such as arms, are not recognized in the captured image due to obstacles or being outside the camera's range.
An information processing device that detects candidate areas and estimated positions based on image features and person areas to identify object regions with high accuracy, using a combination of candidate area detection, estimated position detection, and identification units.
Enables precise detection of a person's belongings even when parts of the body are not visible, enhancing the accuracy of object recognition in images.
Smart Images

Figure 2025078813000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to object recognition. [Background technology]
[0002] Technologies for detecting objects from captured images generated by a camera have been developed. For example, Patent Document 1 discloses a device that detects multiple objects from a captured image and associates the detected objects. Specifically, when a captured image contains an object (e.g., a bag) and multiple people, the device of Patent Document 1 associates the object with the person who owns it.
[0003] For this purpose, the device of Patent Document 1 recognizes and connects the parts of a person in order using predefined connection relationships. For example, recognition and connection are performed in the order of face->neck->torso->arms. Furthermore, the device of Patent Document 1 recognizes objects that are predefined as objects that frequently exist around the recognized parts. For example, a bag is defined as an object that frequently exists around an arm. Therefore, as described above, when the person's arm is recognized, the bag is recognized. As a result, it is found that the connection is "face->neck->torso->arms->bag". Therefore, the device of Patent Document 1 associates the connected face and bag (i.e., associates the person and the bag).
[0004] Here, Patent Document 1 specifies information for estimating the approximate location of objects that frequently exist around a person's features relative to the features. Patent Document 1 also describes that this information may be used to limit the image area in which objects are recognized. For example, once the device of Patent Document 1 detects a person's arm in the above-mentioned manner, it uses information indicating the approximate location of the bag relative to the person's arm to limit the image area in which the bag is recognized. Then, the bag is recognized in the limited image area. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2010-086482 A [Non-patent literature]
[0006] [Non-Patent Document 1] Zhe Cao, 3 others, "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", CoRR, November 24, 2016 Summary of the Invention [Problem to be solved by the invention]
[0007] In the technology of Patent Document 1, on the premise that a person's features are recognized, objects that frequently exist around the features are recognized. Therefore, if a certain part of a person is not recognized, objects that frequently exist around that part are not recognized. For example, in the above example, if the person's arm is not recognized, the bag is not recognized.
[0008] In this regard, not all parts of a person are necessarily included in a captured image. For example, if an obstacle is standing in front of a person's arm, or if the person's arm is outside the camera's imaging range, the person's arm will not be recognized in the captured image, and the bag will not be recognized either.
[0009] The present invention has been made in consideration of the above-mentioned problems, and has an object to provide a technique for detecting a person's belongings from a captured image with high accuracy. [Means for solving the problem]
[0010] The information processing device of the present invention has: 1) a candidate area detection unit that detects one or more candidate areas, which are image areas estimated to represent a target object from a captured image based on image features of the target object to be detected; 2) an estimated position detection unit that detects a person area representing a person from the captured image and detects an estimated position of the target object in the captured image based on the person area; and 3) an identification unit that identifies an object area, which is an image area representing the target object, from among the one or more candidate areas based on the one or more candidate areas and the estimated position.
[0011] The control method of the present invention is a control method executed by a computer, and includes: 1) a candidate area detection step of detecting one or more candidate areas, which are image areas estimated to represent a target object, from a captured image based on image features of the target object to be detected, 2) an estimated position detection step of detecting a person area representing a person from the captured image and detecting an estimated position of the target object in the captured image based on the person area, and 3) an identification step of identifying an object area, which is an image area representing the target object, from the one or more candidate areas based on the one or more candidate areas and the estimated positions.
[0012] The program of the present invention causes a computer to execute each step of the control method of the present invention. Effect of the Invention
[0013] According to the present invention, a technique is provided for detecting a person's belongings from a captured image with high accuracy. [Brief description of the drawings]
[0014] The above objects, as well as other objects, features and advantages, will become more apparent from the following preferred embodiments and the accompanying drawings.
[0015] [Figure 1] 2 is a diagram conceptually illustrating a process performed by the information processing device of the present embodiment. FIG. [Diagram 2]FIG. 2 is a diagram illustrating an example of the functional configuration of the information processing apparatus according to the first embodiment. [Diagram 3] FIG. 1 is a diagram illustrating a computer for implementing an information processing device. [Figure 4] 4 is a flowchart illustrating a flow of processing executed by the information processing apparatus of the first embodiment. [Diagram 5] FIG. 13 is a diagram illustrating a candidate area including an estimated position. [Figure 6] 13 is a diagram illustrating an example of a first score calculated based on the number of estimated positions included in a candidate area. FIG. [Figure 7] 13 is a diagram illustrating an example of a first score calculated in consideration of the presence probability of a target object calculated for an estimated position. FIG. [Figure 8] FIG. 11 is a block diagram illustrating a functional configuration of an information processing apparatus according to a second embodiment. [Figure 9] 11 is a flowchart illustrating the flow of processing executed by an information processing apparatus according to a second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0016] Hereinafter, the embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are given the same reference numerals, and the description will be omitted as appropriate. In addition, unless otherwise specified, in each block diagram, each block represents a functional configuration, not a hardware configuration.
[0017] [Embodiment 1] <Summary> FIG. 1 is a diagram conceptually illustrating a process performed by an information processing device 2000 of this embodiment. The information processing device 2000 detects an object area 30, which is an image area representing a target object, from a captured image 20 generated by a camera 10. The target object is a person's belongings. Note that the "person's belongings" mentioned here is not limited to an object held by a person in his / her hand (such as a handbag or a walking stick), but generally includes an object held by a person in some form. For example, a person's belongings include an object carried on a person's shoulder (such as a shoulder bag), an object hung from a person's neck (such as an identification card), an object carried on a person's back (such as a backpack), an object worn on a person's head (such as a hat or helmet), an object worn on a person's face (such as glasses or sunglasses), and an object worn on a person's hand (such as a watch).
[0018] The information processing device 2000 detects one or more candidate regions 22 from the captured image 20 based on the image features of the target object. The candidate region 22 is an image region that is estimated to represent the target object. For example, if the target object is a hat, the information processing device 2000 detects an image region that is estimated to represent the hat based on the image features of the hat, and sets the detected image region as the candidate region 22. The candidate region 22 is, for example, an image region that is estimated to have a probability of representing the target object equal to or greater than a predetermined value.
[0019] Moreover, the information processing device 2000 detects a person area 26 from the captured image 20, and detects an estimated position 24 based on the detected person area 26. The person area 26 is an area estimated to represent a person. The estimated position 24 is a position in the captured image 20 where a target object is estimated to exist.
[0020] Here, the relative positional relationship between a person's belongings and the person can be predicted depending on the type of the object. For example, the position of a hat is highly likely to be on the person's head. For other examples, the position of sunglasses is highly likely to be on the person's face. For other examples, the position of a backpack is highly likely to be on the person's back.
[0021] Therefore, based on the relative positional relationship between the target object and the person that can be predicted in this way, the information processing device 2000 detects the estimated position 24. For example, if the target object is a hat, the information processing device 2000 detects the position where the hat is estimated to exist based on the relative positional relationship between the person represented by the person area 26 and the hat, and sets the detected position as the estimated position 24.
[0022] The information processing device 2000 then identifies an object region 30 based on the candidate region 22 and the estimated position 24. For example, the information processing device 2000 identifies, among the multiple detected candidate regions 22, a candidate region 22 that includes the estimated position 24 as the object region 30. However, as will be described later, the object region 30 identified based on the candidate region 22 and the estimated position 24 is not limited to the candidate region 22 that includes the estimated position 24.
[0023] <Actions and Effects> According to the information processing device 2000 of this embodiment, an object region 30 representing a target object is specified using a candidate region 22 detected based on the image features of the target object and an estimated position 24 detected based on a person region 26. In this way, not all of the candidate regions 22 detected based on the image features of the target object are specified as object regions 30 (image regions representing the target object), but the candidate regions 22 specified as object regions 30 are limited by the estimated position 24 detected based on the person region 26. For example, a candidate region 22 in a position where the probability that a target object exists is low is not specified as an object region 30. In this way, by specifying an image region representing a target object using two criteria, namely, the image features of the target object and the image region representing a person, the image region representing the target object can be specified with high accuracy compared to a case where the image region representing the target object is specified based on a single criterion, namely, the image features of the target object.
[0024] Here, the estimated position 24 of the target object is detected by using an image area representing a person. Therefore, even if some parts of the person (such as an arm) are not detected from the captured image 20, the estimated position 24 can be detected. Therefore, according to the information processing device 2000, even if some parts of the person are not included in the captured image 20, the object area 30 can be specified.
[0025] 1 is an example for facilitating understanding of the information processing device 2000, and does not limit the functions of the information processing device 2000. Hereinafter, the information processing device 2000 of the present embodiment will be described in further detail.
[0026] <Example of Functional Configuration of Information Processing Device 2000> 2 is a diagram illustrating an example of a functional configuration of an information processing device 2000 according to the first embodiment. The information processing device 2000 includes a candidate area detection unit 2020, an estimated position detection unit 2040, and an identification unit 2060. The candidate area detection unit 2020 detects one or more candidate areas 22 from the captured image 20 based on image features of a target object to be detected. The estimated position detection unit 2040 detects a person area 26 from the captured image 20. Furthermore, the estimated position detection unit 2040 detects an estimated position 24 based on the detected person area 26. The identification unit 2060 identifies an object area 30 based on the candidate area 22 and the estimated position 24.
[0027] <Hardware configuration of information processing device 2000> Each functional component of the information processing device 2000 may be realized by hardware that realizes each functional component (e.g., a hardwired electronic circuit, etc.), or may be realized by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it, etc.). The case where each functional component of the information processing device 2000 is realized by a combination of hardware and software will be further described below.
[0028] FIG. 3 is a diagram illustrating a computer 1000 for realizing the information processing device 2000. The computer 1000 is any computer. For example, the computer 1000 is a stationary computer such as a personal computer (PC) or a server machine. Alternatively, the computer 1000 is a portable computer such as a smartphone or a tablet terminal. Alternatively, the computer 1000 may be a camera 10 that generates a captured image 20. The computer 1000 may be a dedicated computer designed to realize the information processing device 2000, or may be a general-purpose computer.
[0029] The computer 1000 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. The bus 1020 is a data transmission path for the processor 1040, the memory 1060, the storage device 1080, the input / output interface 1100, and the network interface 1120 to transmit and receive data to and from each other. However, the method of connecting the processor 1040 and the like to each other is not limited to bus connection.
[0030] The processor 1040 is one of various processors such as a central processing unit (CPU), a graphics processing unit (GPU), or a field-programmable gate array (FPGA). The memory 1060 is a main storage device realized using a random access memory (RAM) or the like. The storage device 1080 is an auxiliary storage device realized using a hard disk, a solid state drive (SSD), a memory card, or a read only memory (ROM) or the like.
[0031] The input / output interface 1100 is an interface for connecting the computer 1000 to an input / output device. For example, an input device such as a keyboard and an output device such as a display device are connected to the input / output interface 1100. The network interface 1120 is an interface for connecting the computer 1000 to a communication network. This communication network is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The network interface 1120 may be connected to the communication network by wireless connection or by wired connection.
[0032] The storage device 1080 stores program modules that realize each functional component of the information processing device 2000. The processor 1040 reads each of these program modules into the memory 1060 and executes them to realize the function corresponding to each program module.
[0033] <About Camera 10> The camera 10 is any imaging device that captures images and generates image data as a result. For example, the camera 10 is a surveillance camera installed at a surveillance location.
[0034] As described above, the computer 1000 that realizes the information processing device 2000 may be the camera 10. In this case, the camera 10 identifies the object region 30 by analyzing the captured image 20 that it generates. As the camera 10 having such a function, for example, a camera called an intelligent camera, a network camera, or an IP (Internet Protocol) camera can be used.
[0035] <Examples of use of the information processing device 2000> The information processing device 2000 can be used in any situation where the process of "detecting a specific object from image data generated by a camera" is useful. For example, the information processing device 2000 is used to analyze surveillance video generated by a surveillance camera. In this case, the camera 10 is a surveillance camera that generates surveillance video. Also, the captured image 20 is a video frame that constitutes the surveillance video.
[0036] The information processing device 2000 identifies an image area representing a specific object (i.e., an object area 30 representing a target object) from video frames constituting a surveillance video. In this way, it is possible to grasp the presence of the target object in the surveillance location. It is also possible to detect a person holding the target object.
[0037] More specifically, the information processing device 2000 can use surveillance video to grasp the presence of dangerous objects or suspicious people (people carrying dangerous objects or people hiding their faces with sunglasses, helmets, etc.) In addition, when an abandoned object is found in a facility to be monitored, the information processing device 2000 can analyze past surveillance videos generated by surveillance cameras installed in various places in the facility to identify the route the abandoned object was taken and detect the person who carried the abandoned object.
[0038] <Processing flow> 4 is a flowchart illustrating a process flow executed by the information processing device 2000 of the first embodiment. The information processing device 2000 acquires a captured image 20 (S102). The candidate area detection unit 2020 detects one or more candidate areas 22 from the captured image 20 based on image features of a target object (S104). The estimated position detection unit 2040 detects a person area 26 from the captured image 20 (S106). The estimated position detection unit 2040 detects an estimated position 24 based on the detected person area 26 (S108). The identification unit 2060 identifies an object area 30 based on the candidate areas 22 and the estimated positions 24 (S110).
[0039] It is not necessary that all the processes are executed sequentially as shown in Fig. 4. For example, the process executed by the candidate area detection unit 2020 (S104) and the process executed by the estimated position detection unit 2040 (S106 and S108) may be executed in parallel.
[0040] The information processing device 2000 executes the series of processes shown in Fig. 4 at various times. For example, every time a captured image 20 is generated by the camera 10, the information processing device 2000 acquires the captured image 20 and executes the series of processes shown in Fig. 4. As another example, the information processing device 2000 acquires a plurality of captured images 20 generated by the camera 10 at a predetermined timing, and executes the series of processes shown in Fig. 4 for each captured image 20 (so-called batch processing). As another example, the information processing device 2000 accepts an input operation for designating a captured image 20, and executes the series of processes for the designated captured image 20.
[0041] <Acquisition of captured image 20: S102> The information processing device 2000 acquires the captured image 20 (S102). The captured image 20 may be the image data itself generated by the camera 10, or may be image data generated by the camera 10 that has been subjected to some processing (e.g., color correction, cropping, etc.).
[0042] The method by which the information processing device 2000 acquires the captured image 20 is arbitrary. For example, the information processing device 2000 acquires the captured image 20 by accessing a storage device in which the captured image 20 is stored. The storage device in which the captured image 20 is stored may be provided inside the camera 10, or may be provided outside the camera. Alternatively, for example, the information processing device 2000 may acquire the captured image 20 by receiving the captured image 20 transmitted from the camera 10. Note that, as described above, when the information processing device 2000 is realized as the camera 10, the information processing device 2000 acquires the captured image 20 generated by itself.
[0043] <Detection of candidate region 22: S104> The candidate area detection unit 2020 detects a candidate area 22 from the captured image 20 based on the image features of the target object (S104). Here, existing technology can be used for detecting an image area (i.e., the candidate area 22) that is estimated to represent the object from the image data based on the image features of the object to be detected. For example, a detector that has been trained in advance to detect an image area that is estimated to represent the target object from the image data can be used for detecting the candidate area 22. Any model such as a neural network (e.g., a convolutional neural network) or a support vector machine (SVM) can be used for the detector model.
[0044] Here, the candidate area detection unit 2020 detects an image area that is estimated to represent a target object with a probability equal to or greater than a threshold as a candidate area 22. Here, if this threshold is made large, false negatives (missed detections) are more likely to occur, whereas if this threshold is made small, false positives (misdetection) are more likely to occur.
[0045] In this regard, in the information processing device 2000, the object region 30 is specified not only by the candidate region detection unit 2020, but also by using the estimated position 24 detected by the estimated position detection unit 2040. Therefore, it can be said that it is preferable to set the threshold small and cause erroneous detection, rather than setting the threshold large and causing detection omission. This is because the object region 30 representing the target object can be specified with high accuracy by a method of setting the threshold small to detect more candidate regions 22 and narrowing down the candidate regions 22 using the estimated positions 24 detected by the estimated position detection unit 2040.
[0046] Therefore, it is preferable that the above threshold value used by the candidate area detection unit 2020 be a value equal to or lower than the threshold value set when identifying the object area 30 based only on the image features of the target object (i.e., when the estimated position detection unit 2040 is not used).
[0047] The candidate area detection unit 2020 generates data representing the detection result of the candidate area 22. This data is data that identifies the detected candidate area 22, and indicates, for example, a specific position (e.g., the coordinates of the upper left corner) and size (e.g., width and height) of the candidate area 22.
[0048] <Detection of person area 26: S106> The estimated position detection unit 2040 detects the person area 26 from the captured image 20 (S106). Here, existing technology can be used as a technique for detecting an image area representing a person from image data. For example, a detector that has been trained in advance to detect an image area representing a person from image data can be used. For example, any model such as a neural network can be used as a model for this detector.
[0049] Here, in order to detect the estimated position 24, it is preferable to detect parts of the human body (head, face, torso, hands, feet, etc.) from the person area 26. The parts of the human body can also be detected by detection using the above-mentioned detectors or by detection using a template image or local features.
[0050] Alternatively, for example, the estimated position detection unit 2040 may detect a collection of characteristic points of a person, such as the positions of the person's joints, as the person area 26. For example, the technology shown in Non-Patent Document 1 can be used to detect the positions of characteristic points of a person, such as joints.
[0051] <Detection of estimated position 24: S106> The estimated position detection unit 2040 detects the estimated position 24 based on the person area 26. As described above, the estimated position 24 is a position where a target object is estimated to exist in the captured image 20. The estimated position 24 may be represented by a single point on the captured image 20, or may be represented by an image area.
[0052] For example, a detector that has been trained in advance to detect the position where the target object is estimated to exist in image data in response to input of image data in which the position of an image area representing a person is specified can be used to detect the estimated position 24. Any model such as a neural network can be used as the detector model.
[0053] The detector is trained using training data consisting of a combination of, for example, "image data, a person area in the image data, and the position of a target object in the image data." By using such training data, the detector can learn the relative positional relationship between the target object and the person. Furthermore, it is preferable that the training data includes information indicating the position of each part of the person in the person area.
[0054] The estimated position detection unit 2040 detects a position where the probability that the target object exists is equal to or higher than a predetermined value as the estimated position 24. At this time, the estimated position detection unit 2040 may output the probability that the target object exists at the estimated position 24 together with the estimated position 24.
[0055] For example, the estimated position detection unit 2040 generates matrix data of the same size as the captured image 20 as data representing the detection result of the estimated position 24. This matrix data indicates, for example, 1 at the position of the estimated position 24 and 0 at other positions. Furthermore, when outputting the probability that a target object exists at the estimated position 24, this matrix data indicates the probability that a target object exists at each position. However, the data representing the detection result of the estimated position 24 may be in any format and is not limited to matrix data.
[0056] <<Limiting detection range>> The estimated position detection unit 2040 may limit the image area in which the estimated position 24 is detected by using the candidate area 22. That is, the estimated position 24 is detected not from the entire captured image 20 but from a part of the image area limited based on the candidate area 22. In this way, the time and computer resources required for detecting the estimated position 24 can be reduced.
[0057] For example, the estimated position detection unit 2040 sets only the inside of the candidate area 22 as the detection range for the estimated position 24. Alternatively, for example, the estimated position detection unit 2040 detects the estimated position 24 from a predetermined range including the candidate area 22. For example, this predetermined range is a range obtained by enlarging the candidate area 22 by a predetermined magnification factor greater than 1.
[0058] The estimated position detection unit 2040 may also limit the image area in which the person area 26 is detected by using the candidate area 22. For example, the estimated position detection unit 2040 detects the person area 26 from a predetermined range including the candidate area 22 (for example, a range obtained by enlarging the candidate area 22).
[0059] <Identifying object area 30> The identification unit 2060 identifies an object region 30 based on the candidate regions 22 and the estimated position 24. Conceptually, the identification unit 2060 uses the estimated position 24 to identify, from among the candidate regions 22, which are image regions suspected to include the target object, one that is particularly likely to include the target object, and identifies the identified candidate region 22 as an object region 30. However, as will be described later, the object region 30 does not need to completely match any one of the candidate regions 22, and may be an image region that is a part of the candidate regions 22.
[0060] The identification unit 2060 identifies the object region 30 by focusing on the overlap between the candidate region 22 and the estimated position 24. Various specific methods can be used for this. Examples of specific methods are given below.
[0061] <<Specific method 1>> The identification unit 2060 identifies the candidate region 22 including the estimated position 24 as the object region 30. FIG. 5 is a diagram illustrating an example of a candidate region 22 including the estimated position 24. In FIG. 5, a plurality of candidate regions 22 are detected from the captured image 20. Also, one estimated position 24 is detected. Here, the estimated position 24 is included in the candidate region 22-1. Therefore, the identification unit 2060 identifies the candidate region 22-1 as the object region 30.
[0062] <<Specific method 2>> Here, it is assumed that a plurality of estimated positions 24 are calculated. Then, the identification unit 2060 calculates a score (hereinafter, a first score) representing the degree to which each candidate area 22 includes the estimated position 24. The identification unit 2060 identifies the object area 30 based on the first score.
[0063] There are various methods for identifying the object region 30 based on the first score. For example, the identification unit 2060 identifies the candidate region 22 with the largest first score as the object region 30. Alternatively, for example, the identification unit 2060 identifies the candidate region 22 with a first score equal to or greater than a predetermined value as the object region 30. In the latter case, multiple object regions 30 may be identified.
[0064] There are various ways to determine the first score. For example, the identification unit 2060 calculates the number of estimated positions 24 included in a candidate area 22 as the first score for that candidate area 22. As another example, the identification unit 2060 calculates the number of estimated positions 24 included in a candidate area 22 normalized by the size of the candidate area 22 (for example, the value obtained by dividing the number of estimated positions 24 by the area of the candidate area 22) as the first score for that candidate area 22.
[0065] 6 is a diagram illustrating an example of a first score calculated based on the number of estimated positions 24 included in the candidate area 22. The candidate area 22 includes three estimated positions 24. Therefore, for example, the identification unit 2060 sets the first score of the candidate area 22 to 3. Here, it is assumed that the area of the candidate area 22 is S. In this case, the identification unit 2060 may normalize the first score of the candidate area 22 by the area of the candidate area 22 to 3 / S, and set this as the first score.
[0066] The method of calculating the first score is not limited to the above example. For example, it is assumed that the probability that a target object exists is calculated for each estimated position 24. In this case, the specification unit 2060 may calculate the sum of the existence probabilities calculated for each estimated position 24 included in the candidate area 22 as the first score for that candidate area 22.
[0067] 7 is a diagram illustrating a first score calculated in consideration of the presence probability of a target object calculated for an estimated position 24. The candidate area 22 includes three estimated positions 24, and the presence probabilities calculated for each are p1, p2, and p3. Therefore, the first score of the candidate area 22 is p1+p2+p3.
[0068] In this way, by calculating the first score taking into consideration the probability that a target object exists at the estimated position 24, it is possible to specify the object region 30 representing the target object with higher accuracy. For example, a candidate region 22 including one estimated position 24 with a target object existence probability of 0.6 is more likely to be an image region representing a target object than a candidate region 22 including three estimated positions 24 with a target object existence probability of 0.1. According to a calculation method in which the sum of the existence probabilities is used as the first score, the first score of the latter candidate region 22 is larger than the first score of the former candidate region 22. Therefore, the latter candidate region 22 is more likely to be specified as an object region 30.
[0069] <<Specific method 3>> Here, it is assumed that the candidate area detection unit 2020 calculates, for each candidate area 22, the probability that the candidate area 22 represents the target object. Also, it is assumed that the identification unit 2060 calculates the above-mentioned first score for each candidate area 22. The identification unit 2060 calculates a second score as the product of the probability that the candidate area 22 represents the target object and the first score. Then, the identification unit 2060 identifies the object area 30 based on the second score.
[0070] There are various methods for identifying the object region 30 based on the second score. For example, the identification unit 2060 identifies the candidate region 22 having the largest second score as the object region 30. As another example, the identification unit 2060 identifies the candidate region 22 having a second score equal to or greater than a predetermined value as the object region 30.
[0071] <<Specific method 4>> The identification unit 2060 calculates a third score based on the distance between the representative point of the candidate region 22 and the estimated position 24, and identifies the object region 30 based on the calculated third score. There are various methods for identifying the object region 30 based on the third score. For example, the identification unit 2060 identifies the candidate region 22 for which the smallest third score has been calculated as the object region 30. As another example, the identification unit 2060 identifies the candidate region 22 for which a third score equal to or less than a predetermined value has been calculated as the object region 30.
[0072] The representative point of the candidate region 22 may be any point included in the candidate region 22. For example, the representative point of the candidate region 22 is the center of the candidate region 22.
[0073] When there are multiple estimated positions 24, the identification unit 2060 may calculate the distance between each of the multiple estimated positions 24 and the representative point of the candidate area 22, or may calculate the distance between any one of the estimated positions 24 and the representative point of the candidate area 22. In the former case, for example, the identification unit 2060 calculates the third score based on a statistical value (such as a minimum value, a mode value, or an average value) of the multiple calculated distances. In the latter case, the identification unit 2060 calculates the distance between one estimated position 24 and the representative point of the candidate area 22, and calculates the third score based on the distance.
[0074] Here, when calculating the distance between only one estimated position 24 and the representative point of the candidate area 22, there are various methods for specifying the estimated position 24. For example, the specification unit 2060 calculates the center of an image area composed of a plurality of estimated positions 24, specifies the estimated position 24 closest to the center, and calculates the distance between the specified estimated position 24 and the representative point of the candidate area 22. In another example, when the existence probability of a target object is calculated for each estimated position 24, the specification unit 2060 calculates the distance between the estimated position 24 with the highest existence probability of the target object and the representative point of the candidate area 22.
[0075] Furthermore, when the estimated position 24 is expressed as an image area, the identification unit 2060 calculates the third score based on the distance between a representative point of the image area and a representative point of the candidate area 22. The representative point of the estimated position 24 expressed as an image area is, for example, the center position of the image area.
[0076] There are various methods for calculating the third score based on the distance between the estimated position 24 and the representative point of the candidate region 22. For example, the identification unit 2060 determines the distance between the representative point of the candidate region 22 and the estimated position 24 itself as the third score.
[0077] Alternatively, for example, the identification unit 2060 may determine, as the third score, a value obtained by multiplying the distance between the representative point of the candidate region 22 and the estimated position 24 by a correction coefficient based on the probability that a target object exists at the estimated position 24. The correction coefficient is set to be smaller as the probability that a target object exists at the estimated position 24 increases. For example, the correction coefficient is the reciprocal of the probability that a target object exists at the estimated position 24.
[0078] In this way, by considering the probability that a target object exists at the estimated position 24, the object region 30 representing the target object can be specified with higher accuracy. For example, a candidate region 22 with a distance of 2 from an estimated position 24 with a probability of 0.6 of the target object existing is considered to be more likely to be an image region representing a target object than a candidate region 22 with a distance of 1 from an estimated position 24 with a probability of 0.1 of the target object existing. According to the method using the correction coefficient described above, the latter candidate region 22 has a larger third score than the former candidate region 22. Therefore, the latter candidate region 22 is more likely to be specified as an object region 30.
[0079] <Result output> The information processing device 2000 outputs information (hereinafter, output information) that specifies the object region 30. There are various methods for outputting the output information. For example, the information processing device 2000 stores the output information in an arbitrary storage device. Alternatively, for example, the information processing device 2000 stores the output information in a display device.
[0080] For example, the output information indicates an identifier of the captured image 20, a specific position of the object area 30 (e.g., the coordinates of the upper left corner of the object area 30), and a size (e.g., width and height) of the object area 30. When an object area 30 is identified from the captured image 20, the output information indicates the position and size of each of the multiple object areas 30. Alternatively, for example, the output information may be the captured image 20 with information indicating the object area 30 (e.g., a frame) superimposed.
[0081] [Embodiment 2] 8 is a block diagram illustrating the functional configuration of an information processing device 2000 according to the second embodiment. Except for the points described below, the information processing device 2000 according to the second embodiment has the same functions as the information processing device 2000 according to the first embodiment.
[0082] The information processing device 2000 of the second embodiment handles a plurality of types of target objects. Specifically, the information processing device 2000 acquires type information indicating the type of object to be detected, and sets the object of the type indicated in the type information as the target object. To this end, the information processing device 2000 of the second embodiment has a type information acquisition unit 2080 that acquires the type information.
[0083] The type information may indicate one or more types of objects. When the type information indicates multiple types of objects, the information processing device 2000 specifies an object region 30 for each target object, with each type of object being a target object. For example, when the type information indicates three types, "hat, sunglasses, and white cane," the information processing device 2000 specifies, from the captured image 20, an object region 30 representing a hat, an object region 30 representing sunglasses, and an object region 30 representing a white cane.
[0084] There are various methods for the type information acquisition unit 2080 to acquire type information. For example, the type information acquisition unit 2080 acquires type information from a storage device in which the type information is stored. As another example, the type information acquisition unit 2080 acquires type information by receiving type information transmitted from another device. As another example, the type information acquisition unit 2080 acquires type information by accepting input of type information from a user.
[0085] The candidate area detection unit 2020 of the second embodiment detects the candidate area 22 for an object of a type indicated in the type information. Here, existing technology can be used for detecting a specific type of object from image data. For example, a detector trained to detect an object of that type from image data is prepared for each type of object. The candidate area detection unit 2020 detects the candidate area 22 for an object of that type by inputting the captured image 20 to a detector trained to detect the candidate area 22 for an object of the type indicated by the type information.
[0086] The estimated position detection unit 2040 of the second embodiment detects the estimated position 24 of an object of a type indicated in the type information, based on the person area 26. For example, the estimated position detection unit 2040 also prepares a detector for detecting the estimated position 24 for each type of object. That is, the detector is made to learn the positional relationship between the object and the person for each type of object. The estimated position detection unit 2040 inputs information specifying the captured image 20 and the person area 26 to a detector that has been made to learn to detect the estimated position 24 of an object of a type indicated by the type information, thereby detecting the estimated position 24 of that type of object.
[0087] The identification unit 2060 of the second embodiment identifies the object region 30 based on the candidate region 22 and estimated position 24 detected for the target object of the type indicated by the type information as described above. Output information is generated for each type of object.
[0088] <Action and effect> According to the information processing device 2000 of the embodiment, an object region 30 is specified for an object of a type indicated by type information. In this way, the information processing device 2000 can be set to detect a specified object from among a plurality of types of objects from the captured image 20. Therefore, it is possible to detect each of a plurality of types of objects from the captured image 20, and to change the type of object to be detected at each time. This improves the convenience of the information processing device 2000.
[0089] For example, when information on the belongings of a suspicious person is obtained, the captured image 20 can be set to detect the belongings of the suspicious person. Also, when an abandoned object is found, the information processing device 2000 can be set to detect the abandoned object.
[0090] <Example of hardware configuration> The hardware configuration of the computer that realizes the information processing device 2000 of the second embodiment is shown in Fig. 3, for example, as in the first embodiment. However, the storage device 1080 of the computer 1000 that realizes the information processing device 2000 of this embodiment further stores a program module that realizes the functions of the information processing device 2000 of this embodiment.
[0091] <Processing flow> 9 is a flowchart illustrating a process flow executed by the information processing device 2000 of the second embodiment. The type information acquisition unit 2080 acquires type information (S202). The information processing device 2000 acquires a captured image 20 (S204). The candidate area detection unit 2020 detects a candidate area 22 for an object of a type indicated in the type information (S206). The estimated position detection unit 2040 detects a person area 26 (S208). The estimated position detection unit 2040 detects an estimated position 24 for an object of a type indicated in the type information based on the person area 26 (S210). The identification unit 2060 identifies an object area 30 based on the detected candidate area 22 and estimated position 24.
[0092] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various configurations other than those described above can also be adopted.
Claims
1. a first detection unit that detects a plurality of image regions including a target object to be detected from a captured image using a detector that has been trained on images; a second detection unit that detects a body part of a person appearing in the captured image; an identification unit that identifies a part in which the target object exists among the plurality of body parts based on the image region and the body part; having The identification unit determines a probability that the target object is present in the body part based on the image region and the body part; the target object is a person's belongings; An information processing device, wherein the type of the target object differs for each of the image regions.
2. The first detection unit detects the image region using a detector that has been trained with an image representing the target object. The information processing device according to claim 1 .
3. The first detection unit detects the image region using a detector that has been trained with images representing the target objects for each type of the target objects. The information processing device according to claim 1 .
4. The information processing device according to claim 1 , wherein the person's belongings are wearable items.
5. 1. A computer-implemented information processing method, comprising: Using a detector that has been trained on images, a plurality of image regions including a target object to be detected are detected from the captured image; Detecting a body part of a person appearing in the captured image; determining a probability that the target object is present in the body part based on the image region and the body part; identifying a region of the body in which the target object is present among the plurality of body regions based on the image region and the body region; the target object is a person's belongings; the type of the target object varies for each of the image regions; Information processing methods.
6. 1. A computer-implemented information processing method, comprising: Detecting a body part of a person appearing in a captured image; detecting a plurality of image regions including a target object to be detected from the captured image using a detector that has been trained on images; determining a probability that the target object is present in the body part based on the image region and the body part; identifying a region of the body in which the target object is present among the plurality of body regions based on the image region and the body region; the target object is a person's belongings; the type of the target object varies for each of the image regions; Information processing methods.
7. The information processing method according to claim 5 or 6, wherein the person's belongings are wearable items.
8. On the computer, A process of detecting a plurality of image regions including a target object to be detected from a captured image using a detector that has been trained on images; A process of detecting a body part of a person appearing in the captured image; determining a probability that the target object is present in the body part based on the image region and the body part; identifying a part of the body in which the target object is present among the plurality of body parts based on the image region and the body part; Run the command, the target object is a person's belongings; the type of the target object varies for each of the image regions; program.
9. The program according to claim 8 , wherein the person's belongings are wearable items.
Citation Information
Patent Citations
Apparatus, and method for processing image, and program
JP2010252276A
Information processing device, information processing method, and program
JP2012190159A
Image analysis device and image evaluation apparatus
JP2013065156A
Semantic analysis of objects in the video
JP2013533563A
Information processor and information processing method
JP2015191334A