IMAGE PROCESSING SYSTEM, IMAGE PROCESSING METHOD, AND IMAGE PROCESSING PROGRAM
The image processing system addresses the limitations of existing techniques by incorporating posture estimation and object recognition to accurately determine image similarity, enhancing search precision for similar images.
Patent Information
- Application Number
- JP2023567489
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-17
AI Technical Summary
Existing image processing techniques struggle to accurately determine image similarity due to their focus solely on the posture of individuals, which can lead to incorrect identification of similar images.
An image processing system that includes posture estimation, object recognition, and similarity determination modules. This system acquires estimation results for the pose of individuals and recognition results for objects in images, then uses these results to determine similarity between images based on both posture and object features.
The system significantly improves the accuracy of image similarity determination by considering both the posture and object features, leading to more precise searches for similar images.
Smart Images

Figure 0007673837000001 
Figure 0007673837000002 
Figure 0007673837000003
Abstract
Description
[Technical field]
[0001] The present invention relates to an image processing system, an image processing method and a non-transitory computer readable medium. [Background technology]
[0002] In recent years, image processing techniques have been used to automatically classify and search for similar images from among a plurality of images. For example, Patent Document 1 is known as a related technique. Patent Document 1 discloses a technique for estimating a posture of a person from an image capturing the person, and searching for an image including a posture similar to the estimated posture.
[0003] In addition, Non-Patent Document 1 is known as a technology related to human behavior recognition. Also, Non-Patent Document 2 is known as a technology related to human skeleton estimation. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2019-091138 A [Non-patent literature]
[0005] [Non-Patent Document 1] Chen Gao, Yuliang Zou, Jia-Bin Huang, "iCAN: Instance-Centric Attention Network for Human-Object Interaction Detection", arXiv:1808.10437v1 [cs.CV],<URL:https: / / arxiv.org / abs / 1808.10437v1> , 30 Aug 2018 [Non-Patent Document 2] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299 Summary of the Invention [Problem to be solved by the invention]
[0006] In related techniques such as the above-mentioned Patent Document 1, feature quantities based on the characteristics of a person's posture are used to search for similar images. However, since the related techniques focus only on the posture of a person, there are cases where the similarity of images cannot be determined with high accuracy.
[0007] In view of such problems, an object of the present disclosure is to provide an image processing system, an image processing method, and a non-transitory computer-readable medium that are capable of improving the accuracy of image similarity determination. [Means for solving the problem]
[0008] The image processing system according to the present disclosure includes a posture estimation acquisition means for acquiring an estimation result of the posture of a person included in a first and second image, an object recognition acquisition means for acquiring a recognition result of recognizing an object other than the person included in the first and second images, and a similarity determination means for performing a similarity determination between the first image and the second image based on the estimation result of the posture of the person and the recognition result of the object.
[0009] The image processing method of the present disclosure obtains an estimation result of a posture of a person included in a first and second image, obtains a recognition result of an object other than the person included in the first and second images, and performs a similarity determination between the first image and the second image based on the estimation result of the posture of the person and the recognition result of the object.
[0010] A non-transitory computer-readable medium storing an image processing program according to the present disclosure is a non-transitory computer-readable medium storing an image processing program for causing a computer to execute a process of obtaining an estimation result of an estimation of a posture of a person included in a first and second image, obtaining a recognition result of an recognition of an object other than the person included in the first and second images, and determining a similarity between the first image and the second image based on the estimation result of the posture of the person and the recognition result of the object. Effect of the Invention
[0011] According to the present disclosure, it is possible to provide an image processing system, an image processing method, and a non-transitory computer-readable medium that are capable of improving the accuracy of image similarity determination. [Brief description of the drawings]
[0012] [Figure 1] FIG. 1 is a diagram for explaining a problem in the related technology. [Diagram 2] 1 is a configuration diagram showing an overview of an image processing system according to an embodiment; [Diagram 3] 1 is a configuration diagram showing an example of the configuration of an image processing device according to a first embodiment; [Figure 4] FIG. 4 is a configuration diagram showing another example of the configuration of the image processing device according to the first embodiment. [Diagram 5] 4 is a flowchart showing an example of the operation of the image processing device according to the first embodiment. [Figure 6] 4 is a diagram showing a skeleton structure used in an operation example of the image processing device according to the first embodiment; FIG. [Figure 7] 4 is a flowchart showing an example of the operation of the image processing device according to the first embodiment. [Figure 8] FIG. 4 is a diagram showing an example of a search performed by the image processing device according to the first embodiment. [Figure 9] FIG. 4 is a diagram showing an example of a search performed by the image processing device according to the first embodiment. [Figure 10A]FIG. 11 is a diagram for explaining a distance relationship feature amount according to the second embodiment. [Figure 10B] FIG. 11 is a diagram for explaining a distance relationship feature amount according to the second embodiment. [Figure 11A] FIG. 11 is a diagram for explaining a direction relationship feature amount according to the second embodiment. [Figure 11B] FIG. 11 is a diagram for explaining a direction relationship feature amount according to the second embodiment. [Figure 12A] FIG. 11 is a diagram for explaining a positional relationship feature amount according to the second embodiment. [Figure 12B] FIG. 11 is a diagram for explaining a positional relationship feature amount according to the second embodiment. [Figure 13] 13 is a flowchart showing an example of the operation of the image processing device according to the second embodiment. [Figure 14] 13 is a flowchart showing an example of the operation of the image processing device according to the second embodiment. [Figure 15] FIG. 11 is a configuration diagram showing an example of the configuration of an image processing device according to a third embodiment. [Figure 16] 13A to 13C are diagrams illustrating an example of detection by the image processing device according to the third embodiment. [Figure 17] 13 is a flowchart showing an example of the operation of the image processing device according to the third embodiment. [Figure 18] 13 is a flowchart showing an example of the operation of the image processing device according to the third embodiment. [Figure 19] FIG. 2 is a configuration diagram showing an overview of the hardware of a computer according to an embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0013] Hereinafter, an embodiment will be described with reference to the drawings. In the drawings, the same elements are denoted by the same reference numerals, and repeated explanations will be omitted as necessary.
[0014] (Considerations leading to the embodiment) As described above, the related technology estimates the posture of a person from an image and searches for an image that includes a posture similar to the estimated posture. However, if a search is performed based only on the posture of a person, it may not always be possible to find an image (scene) that a user desires.
[0015] For example, as shown in FIG. 1, when searching for a scene in which a wheelchair is moving by itself, an image of a person sitting in a wheelchair is set as a search query Q1, and similar images similar to the search query Q1 are searched for. In this case, in the related technology, images such as search target P1 and search target P2 are extracted as similar images from among images of search targets. In the related technology, since the search is performed based only on the posture of the person, not only an image of a person sitting in a wheelchair like the search target P1, but also an image of a person simply sitting in a chair like the search target P2 is extracted. In other words, if similarity judgment is performed based on the feature amount of posture only, an image of a person sitting in a chair will also be determined to be a similar image. For this reason, the related technology may not be able to accurately search for images similar to the image (scene) that the user wants to search.
[0016] Another possible method is to search for similar images using HOI (Human-Object-Interaction) detection described in Non-Patent Document 1. HOI detection makes it possible to detect pairs of related people and objects from images and detect the verbs (behaviors) of the people. By determining the similarity of images based on the verbs of people detected from the search query and the verbs of people detected from the search target, it is possible to search for similar images taking people and objects into consideration.
[0017] However, HOI detection requires prior preparation using machine learning. For this reason, it is necessary to learn a large number of images of people interacting with objects in advance. This makes it difficult to search for images of verbs that have not been learned in advance. For this reason, even in this case, it is not possible to accurately search for images that the user desires.
[0018] (Outline of the embodiment) Fig. 2 shows an overview of an image processing system 10 according to an embodiment. As shown in Fig. 2, the image processing system 10 includes a posture estimation acquisition unit 11, an object recognition acquisition unit 12, and a similarity determination unit 13. Note that the image processing system 10 may be configured by one device or may be configured by multiple devices.
[0019] The posture estimation acquisition unit 11 acquires an estimation result of estimating the posture of a person included in the first and second images. The posture estimation acquisition unit 11 may acquire an estimation result from a database or the like, or may perform a posture estimation process based on the first or second image to estimate the posture of the person included in the first or second image. For example, the posture estimation acquisition unit 11 estimates a person's skeletal structure as the posture of the person included in the first or second image based on the first or second image. The object recognition acquisition unit 12 recognizes an object other than a person included in the first and second images and acquires an estimation result. recognition The object recognition acquisition unit 12 may acquire the recognition result from a database or the like, or may perform object recognition processing based on the first or second image to recognize an object included in the first or second image. For example, the object recognition acquisition unit 12 recognizes an object class of an object included in the first or second image based on the first or second image.
[0020] The similarity determination unit 13 performs similarity determination between the first image and the second image based on the estimation result of the person's posture for the first and second images and the recognition result of the object for the first and second images. The similarity determination unit 13 may use the estimation result and the recognition result acquired from a database or the like, or may use the estimation result estimated by the posture estimation process based on the first or second image and the recognition result recognized by the object recognition process based on the first or second image. For example, the similarity determination unit 13 performs similarity determination based on the estimation result of the person's posture estimated based on the first image and the recognition result of the object recognized based on the first image, and the estimation result of the person's posture in the acquired second image and the recognition result of the object in the acquired second image. For example, the similarity determination unit 13 performs similarity determination between the first image and the second image based on the similarity of the posture feature based on the estimation result of the person's posture and the similarity of the object feature based on the recognition result of the object. The similarity determination is a determination of whether or not two images are similar. For example, if the similarity is higher than a predetermined value, the two images are determined to be similar, and if the similarity is lower than the predetermined value, the two images are determined to be dissimilar.
[0021] Furthermore, the first image may be a query image, the second image may be a plurality of search target images, and the similarity determination unit 13 may search for an image similar to the query image from among the plurality of search target images based on the result of the similarity determination.
[0022] In this manner, in the embodiment, in addition to the estimation result of the posture of a person, the recognition result of an object is also used to determine the similarity of images. This allows for more accurate similarity determination than when only the posture is used as in the related art. For example, in the example of FIG. 1, according to the embodiment, the search query Q1 and the search target P1 have high posture and object similarity, so the two images can be determined to be similar, whereas the search query Q1 and the search target P2 have high posture similarity but low object similarity, so the two images can be determined to be dissimilar.
[0023] (Embodiment 1) Hereinafter, the first embodiment will be described with reference to the drawings. Fig. 3 shows the configuration of an image processing device 100 according to the present embodiment.
[0024] The image processing device 100, together with a database (DB) 110, constitutes an image processing system 1. The image processing system 1 including the image processing device 100 is a system that searches for images (scenes) similar to a search query, based on a posture of a person estimated from an image and an object recognized from the image.
[0025] The image processing system 1 may also include an image providing device 200 that provides an image (search target) to the image processing device 100. For example, the image providing device 200 may be a camera that captures an image, or an image storage device in which images are stored in advance. The image providing device 200 generates (stores) a two-dimensional image including a person or object, and outputs the generated image to the image processing device 100. The image providing device 200 is directly connected to the image processing device 100 so as to be able to output an image (video) to the image processing device 100, or is connected via a network or the like. The image providing device 200 may also be provided inside the image processing device 100.
[0026] The database 110 is a database that stores information necessary for processing by the image processing device 100, data of processing results, and the like. The database 110 stores images (search targets) acquired by the image acquisition unit 101, estimation results by the posture estimation unit 102, recognition results by the object recognition unit 103, data for machine learning, features calculated by the feature calculation unit 104, search results by the search unit 105, and the like. The database 110 is directly connected to the image processing device 100 so as to be able to input and output data, or is connected via a network or the like. The database 110 may be provided inside the image processing device 100 as a non-volatile memory such as a flash memory, a hard disk device, or the like.
[0027] As shown in FIG. 3, the image processing device 100 includes an image acquisition unit 101, a posture estimation unit 102, an object recognition unit 103, a feature amount calculation unit 104, a search unit 105, an input unit 106, and a display unit 107. Note that the configuration of each unit (block) is an example, and the image processing device 100 may be configured with other units as long as the operation (method) described below is possible. In addition, the image processing device 100 is realized by a computer device such as a personal computer or a server that executes a program, but may be realized by one device or multiple devices on a network. For example, the posture estimation unit 102, the object recognition unit 103, etc. may be external devices.
[0028] The image acquisition unit 101 acquires an image from the image providing device 200. The image acquisition unit 101 acquires a two-dimensional image (a video including a plurality of images) including a person or an object generated (stored) by the image providing device 200. For example, the acquired image is an image to be searched, and the image acquisition unit 101 stores the acquired image in the database 110.
[0029] The posture estimation unit 102 estimates the posture of a person in an image based on the image. The posture estimation unit 102 may acquire an estimation result of estimating the posture of a person in an image in advance from an external device (such as the image providing device 200, the database 110, or the input unit 106). The posture estimation unit 102 estimates the posture of a person in the acquired image to be searched, and also estimates the posture of a person in an image of a search query during a search. It can also be said that the posture estimation unit 102 includes a first posture estimation unit that estimates the posture of a person in the search target, and a second posture estimation unit that estimates the posture of a person in the search query.
[0030] In this example, the posture estimation unit 102 detects the skeletal structure of a person from an image as the posture of the person. It is to be noted that the posture (posture label) of a person in an image may be estimated not only by detecting the skeletal structure but also by using other posture estimation engines using machine learning. The posture estimation unit 102 detects the two-dimensional skeletal structure of a person in an image based on a two-dimensional image. The posture estimation unit 102 detects the skeletal structure of all people recognized in the acquired image. The posture estimation unit 102 detects the skeletal structure of a person based on the features of the recognized person's joints and the like using a skeletal estimation technique using machine learning. The posture estimation unit 102 uses a skeletal estimation technique such as OpenPose in Non-Patent Document 2, for example. The posture estimation unit 102 outputs a certainty factor indicating the accuracy of the estimation together with the estimated posture (skeletal structure) of the person. The higher the certainty factor, the higher the possibility that the estimated posture of the person is correct (is a person). The posture estimation unit 102 stores the posture estimation result (skeletal structure and certainty factor) of the person to be detected in the database 110.
[0031] Skeleton estimation technologies such as OpenPose estimate a person's skeleton by learning various patterns of image data with annotated answers. The skeletal structure estimated by skeletal estimation technologies such as OpenPose consists of "keypoints," which are characteristic points such as joints, and "bones (bone links)," which indicate the links between keypoints. For this reason, in the following, the skeletal structure may be described using the terms "keypoints" and "bones," but unless otherwise specified, "keypoints" correspond to a person's "joints," and "bones" correspond to a person's "bones."
[0032] The object recognition unit 103 recognizes objects in an image based on the image. The object recognition unit 103 may obtain a recognition result of previously recognizing objects in an image from an external device (such as the image providing device 200, the database 110, or the input unit 106). The object to be recognized is an object other than a person, that is, an object other than a person including a person whose posture has been estimated (for example, an object whose class is other than a person). The object recognition unit 103 recognizes objects in the acquired image to be searched, and also recognizes objects in an image of a search query during a search. It can also be said that the object recognition unit 103 includes a first object recognition unit that recognizes an object to be searched, and a second object recognition unit that recognizes an object of the search query.
[0033] The object recognition unit 103 recognizes the class of an object in an image. The class of an object indicates the type or category of an object. The class of an object may be hierarchical (subdivided) according to search conditions, etc. The object recognition unit 103 recognizes the class of all objects in an acquired image. For example, the object recognition unit 103 may recognize the class of an object in an image by an object recognition engine using machine learning. An object can be recognized by machine learning the features (patterns) of an object's image and the class of the object. The object recognition unit 103 detects an object region in an image and recognizes the class of an object in the detected object region. In addition, the object recognition unit 103 outputs a confidence level indicating the accuracy of the recognition together with the class of the recognized object. The higher the confidence level, the higher the possibility that the class of the recognized object is correct. The object recognition unit 103 stores the object recognition result (object class and confidence level) of the search target in the database 110.
[0034] The object recognition unit 103 may recognize other information related to the characteristics of an object, not limited to the class of the object. As an example, the state of the object may be recognized from the characteristics of each part of the image of the object. For example, the state of the object may be whether a notebook PC (Personal Computer) is open / closed, whether the screen of a PC is displayed / off, whether the headlights and blinkers of a car are on / off, whether the door of a car is open / closed, etc. The state of the object of the target image may be stored, and it may be possible to search for an image similar to the state of the object in the search query.
[0035] The feature amount calculation unit 104 calculates a posture feature amount based on an estimation result of a posture of a person estimated (obtained) from an image, and calculates an object feature amount based on a recognition result of an object recognized (obtained) from an image. The feature amount calculation unit 104 also calculates a posture feature amount of a person and an object feature amount of an object in an image of a search target, and calculates a posture feature amount of a person and an object feature amount of an object in an image of a search query. It can be said that the feature amount calculation unit 104 includes a first feature amount calculation unit that calculates a posture feature amount and an object feature amount of a search target, and a second feature amount calculation unit that calculates a posture feature amount and an object feature amount of a search query. The feature amount calculation unit 104 stores the calculated posture feature amount and object feature amount of the search target (normalized value in case of normalization) in the database 110. Note that the feature amount calculation unit 104 may calculate both the posture feature amount and the object feature amount, or may calculate only the posture feature amount. For example, when the similarity of an object is determined using only information of the object recognition result (object class), the calculation of the object feature amount may be omitted. In this case, it can be said that the information on the object recognition result indicates the object feature amount.
[0036] The feature amount calculation unit 104 calculates the feature amount of the two-dimensional skeletal structure detected as the posture of the person. The feature amount (posture feature amount) of the skeletal structure indicates the feature of the skeleton (posture) of the person, and is an element for searching for an image based on the skeleton of the person. The feature amount of the skeletal structure may be the feature amount of the entire skeletal structure, may be the feature amount of a part of the skeletal structure, or may include a plurality of feature amounts like each part of the skeletal structure. As an example, the feature amount is a feature amount obtained by machine learning the skeletal structure, a size of the skeletal structure from the head to the foot on the image, etc. The size of the skeletal structure is the height in the vertical direction and the area of the skeletal region including the skeletal structure on the image, etc. The vertical direction (height direction or vertical direction) is the vertical direction (Y-axis direction) in the image, for example, a direction perpendicular to the ground (reference plane). The left-right direction (horizontal direction) is the left-right direction (X-axis direction) in the image, for example, a direction parallel to the ground.
[0037] The feature amount calculation unit 104 may also normalize the calculated posture feature amount. For example, the minimum value or maximum value of the skeletal region, the height of the person, or the like may be used as the normalization parameter. For example, the feature amount calculation unit 104 calculates the height (height pixel number) of the person when standing upright in the two-dimensional image, and normalizes the skeletal structure (skeletal information) of the person based on the calculated height pixel number of the person. The height pixel number is the height of the person in the two-dimensional image (the length of the whole body of the person in the two-dimensional image space). The feature amount calculation unit 104 obtains the height pixel number (number of pixels) from the length of each bone of the detected skeletal structure (length in the two-dimensional image space). The feature amount calculation unit 104 normalizes the height of each key point (feature point) included in the skeletal structure in the image by the height pixel number. The height of the key point can be obtained from the Y coordinate value (number of pixels) of the key point.
[0038] Alternatively, the height direction may be the direction of the vertical projection axis (vertical projection direction) obtained by projecting the direction of the vertical axis perpendicular to the ground (reference plane) in a three-dimensional coordinate space in the real world onto a two-dimensional coordinate space. In this case, the height of the key point can be obtained by calculating the vertical projection axis obtained by projecting the axis perpendicular to the ground in the real world onto a two-dimensional coordinate space based on the camera parameters, and from the value (number of pixels) along this vertical projection axis. The camera parameters are the imaging parameters of the image, and examples of the camera parameters include the camera's attitude, position, imaging angle, focal length, etc. An object whose length and position are known in advance is imaged by the camera, and the camera parameters can be obtained from the image.
[0039] When calculating the object feature amount, the feature amount calculation unit 104 calculates the feature amount of the object recognized from the image. The object feature amount indicates the feature of the object in the image, and is an element for searching for an image based on the object. For example, the object feature amount is the feature amount of the image of the recognized object. The object feature amount may be the feature amount of the entire object, or may be the feature amount of a part of the object, or may include multiple feature amounts such as each part of the object. As an example, the feature amount is a feature amount obtained by machine learning the object, or the size and shape of the recognized object on the image. The size of the object is the vertical height, horizontal width, area, etc. of the object area including the object on the image.
[0040] Furthermore, the feature amount calculation unit 104 may normalize the calculated object feature amount. For example, the minimum or maximum value of an object region corresponding to an object class, or the height or width of an object, may be used as a normalization parameter. For example, the feature amount calculation unit 104 calculates the area of an object region of an object in an image, and normalizes the area of the object region of the object based on the minimum or maximum value of the area corresponding to the object class.
[0041] The search unit (similarity determination unit) 105 searches for images that are highly similar to the search query image from among multiple images to be searched that are stored in the database 110. In this example, the search query (search conditions) are a person's posture and an object. The search unit 105 searches for images that match the search query based on the feature amount of the person's posture and the feature amount of the object (including the object class) in the image.
[0042] The search unit 105 performs a similarity determination of images based on the similarity between the posture feature of the search query and the posture feature of the search target, and the similarity between the object feature of the search query and the object feature of the search target, and extracts images similar to the search query. The search unit 105 searches for images having posture features that are highly similar to the posture feature of the search query and having object features that are highly similar to the object feature of the search query. The similarity between the features is the distance between the features. For example, the similarity determination may be performed based on the weight of the similarity of the posture feature and the similarity of the object feature. The similarity determination may also be performed based on the confidence of the person whose posture is estimated and the confidence of the estimated object.
[0043] When obtaining the similarity of postures, the search unit 105 may obtain the similarity of the feature amount of the entire skeletal structure, or may obtain the similarity of the feature amount of a part of the skeletal structure. For example, the search unit 105 may obtain the similarity of the feature amount of a first part (e.g., both hands) and a second part (e.g., both feet) of the skeletal structure. When obtaining the similarity of objects, the search unit 105 may obtain the similarity of the feature amount of the entire object, or may obtain the similarity of the feature amount of a part of the object. The search unit 105 may determine whether the object classes match as the similarity, or, when the object classes are hierarchical, may determine whether all or a part of the object classes match as the similarity.
[0044] The search unit 105 may perform a search based on the posture feature and object feature in each image, or may perform a search based on changes in posture feature and object feature (including object class) in multiple images (videos) that are consecutive in time series. That is, not only images but also acquired videos may be stored, and a video with similar human posture and object may be searched for from the video of the search query. The search unit 105 detects the similarity of the feature on a frame (image) basis. For example, key frames may be extracted from multiple frames, and the similarity may be determined using the extracted key frames. By searching for videos similar to the video of the search query, a search can be performed using changes in the posture of a person or the relationship between a person and an object as a search key. For example, a search can be performed using changes in an object as a search key, such as when a person holds a smartphone with a cup placed on it.
[0045] The input unit 106 is an input interface that acquires information input by a user operating the image processing device 100. The input unit 106 is, for example, a GUI (Graphical User Interface), and receives information according to a user's operation from an input device such as a keyboard, a mouse, or a touch panel. For example, the input unit 106 accepts a posture of a person and an object specified from among a plurality of images as a search query. In addition, the user may manually input the posture (skeleton) of a person and a class of an object that are to be the search query.
[0046] The display unit 107 is a display unit that displays the results of the operation (processing) of the image processing device 100, and is, for example, a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display. The display unit 107 displays the processing results of each unit, such as the search results of the search unit 105, on a GUI.
[0047] As shown in FIG. 4, the image processing device 100 may include a classification unit 108 that classifies images in addition to the search unit 105 or instead of the search unit 105. The classification unit 108 classifies (clusters) a plurality of images stored in the database 110 based on feature amounts. Like the search unit 105, the classification unit 108 performs similarity determination of images based on the similarity of posture feature amounts and object feature amounts between each image, and classifies similar images. The classification unit 108 classifies images having high similarity in posture feature amounts and high similarity in object feature amounts so that they are in the same cluster (group). Like the search unit 105, the classification unit 108 may classify images based on a specified query (classification condition).
[0048] FIG. 5 shows an example of the operation of the image processing device 100 according to this embodiment, illustrating the flow of processing for acquiring a search target image and storing it in a database.
[0049] As shown in Fig. 5, the image processing device 100 acquires an image from the image providing device 200 (S101). The image acquiring unit 101 acquires images that are search targets for performing a search based on a person's posture and an object from the image providing device 200, and stores the acquired images in the database 110. The image acquiring unit 101 may acquire multiple images captured by a camera during a predetermined period of time, or may acquire multiple images stored in a storage device. Subsequent processing is performed on the acquired multiple images.
[0050] Next, the image processing device 100 estimates the posture of the person based on the acquired image (S102a). For example, the acquired image to be searched includes multiple people, and the posture estimation unit 102 detects the skeletal structure as the posture of each person included in the image.
[0051] Fig. 6 shows the skeletal structure of the detected human body model 300. The posture estimation unit 102 detects the skeletal structure of the human body model (two-dimensional skeletal model) 300 as shown in Fig. 6 from a two-dimensional image using a skeletal estimation technique such as OpenPose. The human body model 300 is a two-dimensional model made up of key points such as a person's joints and bones connecting each key point.
[0052] For example, the posture estimation unit 102 extracts feature points that can be key points from the image, and detects each key point of the person by referring to information obtained by machine learning the image of the key points. In the example of Fig. 6, the head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71, left knee A72, right foot A81, and left foot A82 are detected as the key points of the person. Furthermore, the bones of the person connected by these key points are detected as bones, including bone B1 connecting the head A1 and neck A2, bones B21 and B22 connecting the neck A2 to the right shoulder A31 and left shoulder A32 respectively, bones B31 and B32 connecting the right shoulder A31 and left shoulder A32 to the right elbow A41 and left elbow A42 respectively, bones B41 and B42 connecting the right elbow A41 and left elbow A42 to the right hand A51 and left hand A52 respectively, bones B51 and B52 connecting the neck A2 to the right hip A61 and left hip A62 respectively, bones B61 and B62 connecting the right hip A61 and left hip A62 to the right knee A71 and left knee A72 respectively, and bones B71 and B72 connecting the right knee A71 and left knee A72 to the right foot A81 and left foot A82 respectively. The posture estimation unit 102 stores in the database 110 the human skeletal structure detected by the skeletal estimation technique and its confidence level.
[0053] Next, the image processing device 100 calculates posture feature amounts of the estimated posture of the person (S103a). For example, when the height or area of a skeletal region is used as the feature amount, the feature amount calculation unit 104 extracts a region including the skeletal structure, and calculates the height (number of pixels) and area (pixel area) of the region. The height and area of the skeletal region are calculated from the coordinates of the ends of the extracted skeletal region and the coordinates of the key points of the ends. The feature amount calculation unit 104 stores the calculated feature amount of the skeletal structure in the database 110.
[0054] In the example of Figure 6, a skeletal region including all bones is extracted from the skeletal structure of a person standing upright. In this case, the top end of the skeletal region is the head key point A1, the bottom end of the skeletal region is the right foot key point A81 or the left foot key point A82, the left end of the skeletal region is the right hand key point A51, and the right end of the skeletal region is the left hand key point A52. Therefore, the height of the skeletal region is calculated from the difference in the Y coordinate between key point A1 and key point A81 or A82. The width of the skeletal region is calculated from the difference in the X coordinate between key point A51 and key point A52, and the area is calculated from the height and width of the skeletal region.
[0055] In addition, when normalizing the posture feature amount, for example, the feature amount calculation unit 104 calculates normalization parameters such as the height pixel number based on the detected skeletal structure. The feature amount calculation unit 104 normalizes the feature amount such as the height or area of the skeletal region based on the height pixel number or the like.
[0056] In the example of Fig. 6, the height pixel count, which is the height of the skeletal structure of the person standing upright in the image, and the key point height, which is the height of each key point of the skeletal structure of the person in the image, are obtained. The height pixel count may be obtained by adding up the lengths of the bones from the head to the feet among the bones of the skeletal structure. If the posture estimation unit 102 (skeletal estimation technology) does not output the top of the head and the feet, a correction may be made by multiplying a constant as necessary.
[0057] Specifically, the feature amount calculation unit 104 obtains the length of the bones on the two-dimensional image from the head to the feet of the person, and obtains the height pixel number. Among the bones in FIG. 6, the length (number of pixels) of bone B1 (length L1), bone B51 (length L21), bone B61 (length L31), and bone B71 (length L41), or bone B1 (length L1), bone B52 (length L22), bone B62 (length L32), and bone B72 (length L42) is obtained. The length of each bone can be obtained from the coordinates of each key point in the two-dimensional image. The total of these, L1+L21+L31+L41, or L1+L22+L32+L42, is multiplied by a correction constant to calculate the height pixel number. When both values can be calculated, for example, the longer value is used as the height pixel number. That is, when each bone is imaged from the front, its length in the image is the longest, and when it is tilted in the depth direction with respect to the camera, it appears shorter. Therefore, it is more likely that the longer bone is imaged from the front, and is therefore closer to the true value. For this reason, it is preferable to select the longer value.
[0058] The height pixel count may be calculated by other calculation methods. For example, an average human body model showing the relationship (ratio) between the length of each bone and the height in the two-dimensional image space may be prepared in advance, and the height pixel count may be calculated from the length of each bone detected using the prepared human body model.
[0059] The feature amount calculation unit 104 calculates the height of each key point together with the height pixel number, identifies a reference point for normalization, and normalizes the height of each key point by the height pixel number. The feature amount calculation unit 104 stores the normalized posture feature amount in the database 110.
[0060] The key point height is the length (number of pixels) in the height direction from the lowest end of the skeletal structure (for example, a key point of one of the feet) to the key point. Here, as an example, the key point height is obtained from the Y coordinate of the key point in the image. The key point height may be obtained from the length in the direction along the vertical projection axis based on the camera parameters. The specified reference point is a point that serves as a reference for expressing the relative height of the key point. The reference point may be set in advance, or may be selected by the user. The reference point is preferably the center of the skeletal structure or higher than the center (upward in the vertical direction of the image), and for example, the coordinates of the neck key point are used as the reference point. The coordinates of the head or other key points may be used as the reference point, not limited to the neck. The reference point may be any coordinate (for example, the center coordinate of the skeletal structure, etc.) other than the key point. Each key point is normalized using the key point height, reference point, and height pixel number of each key point. Specifically, the feature amount calculation unit 104 normalizes the relative height of the key point with respect to the reference point by the height pixel number. Here, as an example focusing only on the height direction, only the Y coordinate is extracted, and normalization is performed with the reference point being the neck key point. The normalized value is the value obtained by subtracting the height of the reference point from the key point height, and dividing the subtracted value by the number of pixels of the height.
[0061] Furthermore, following S101, the image processing device 100 recognizes objects based on the acquired image (S104a). For example, the acquired image to be searched contains multiple objects in addition to people, and the object recognition unit 103 recognizes the class of each object contained in the image. The object recognition unit 103 detects an object region in the image using an object recognition engine, and recognizes the class of the object in the detected object region. The object recognition unit 103 stores the class of the object recognized by the object recognition engine and its confidence level in the database 110.
[0062] Next, when calculating an object feature amount, the image processing device 100 calculates the object feature amount of the recognized object (S105a). For example, when the size of an object region in which an object is recognized is used as the feature amount, the feature amount calculation unit 104 calculates the height (number of pixels), width (number of pixels), area (pixel area), etc. of the detected rectangular object region. The feature amount calculation unit 104 stores the calculated feature amount of the object in the database 110. When normalizing the object feature amount, the feature amount calculation unit 104 normalizes the calculated size of the object region by the minimum value or maximum value of the object region corresponding to the object class. For example, the area of the object region of the object in the image is calculated, and the value obtained by dividing the area of the object region by the minimum value or maximum value of the area corresponding to the class of the recognized object is used as the normalized value. The feature amount calculation unit 104 stores the normalized object feature amount in the database 110.
[0063] FIG. 7 shows an example of the operation of the image processing device 100 according to this embodiment, and shows the flow of a process for searching for images similar to a search query from among images to be searched and stored in a database by the process of FIG.
[0064] As shown in FIG. 7, when performing a search, the user inputs a search query to the image processing device 100 (S111). The search unit 105 accepts an input of a search query, which is a search condition, via the input unit 106 in response to a user's operation. For example, a plurality of images may be displayed on the display unit 107, and the user may select an image including the person's posture and object of the search query (search key). The image used for the search query may be an image stored in the database 110, or may be an image provided by the image providing device 200 or another image. For example, the skeleton of a person as a posture estimation result or the area and object class of an object as an object recognition result may be displayed in each image and made selectable.
[0065] The posture and object of the search query may be selected from one image, or the posture and object of the search query may be selected from different images. When the posture and object of the search query are selected from different images, the search unit 105 generates one search query image by merging the image of the posture and the image of the object selected, respectively. In addition, when one image includes multiple postures and objects, the user selects one posture and one object to be used as the search query. Note that the search query is not limited to one posture and one object, and may include any number of postures and any number of objects. For example, the confidence level of the posture estimation result and the confidence level of the object recognition result may be displayed on each image, and postures (skeleton) with high confidence levels and objects with high confidence levels may be recommended as search queries. Postures and objects with confidence levels equal to or higher than a predetermined value may be highlighted. In addition, the confidence level of a person's posture and the confidence level of an object may be input as a search query (search condition).
[0066] Furthermore, the user may input a person's posture (skeleton) and an object as a search query in other ways, not limited to an image. For example, as a search query, a posture may be input by moving each part of a skeletal structure according to a user's operation, or the user may input an object class. When a skeletal structure is input, the posture estimation process (S102b) may be omitted. When an object class is input, the object recognition process (S104b) may be omitted.
[0067] When a search query is input, the image processing device 100 estimates the posture of the person in the search query (S102b) and calculates posture features (S103b) in the same manner as when the search target is stored. The posture estimation unit 102 detects the skeletal structure of the person in the search query image (the person specified as the search query) and outputs the detected skeletal structure and its confidence level. The feature calculation unit 104 calculates the height, area, etc. of the skeletal region as the feature of the detected skeletal structure, and normalizes the feature such as the height and area of the skeletal region using a normalization parameter such as the height pixel count.
[0068] Furthermore, the image processing device 100 recognizes the object of the search query (S104b) and calculates the object feature amount (S105b) in the same manner as when the search target is stored. The object recognition unit 103 recognizes the class of the object in the image of the search query (the object specified as the search query), and outputs the class of the recognized object and its certainty. When calculating the object feature amount, the feature amount calculation unit 104 calculates the area of the object region as the feature amount of the recognized object, and normalizes the feature amount such as the area of the object region using normalization parameters such as the minimum and maximum values of the area.
[0069] Next, the image processing device 100 searches for images based on the search query (S112). The search unit 105 searches for images having high similarity in the feature amount of the person's posture and the feature amount of the object from among all images stored in the database 110 to be searched, using the posture of the person and the object specified by the user as the search query.
[0070] The search unit 105 calculates the similarity between each image of the search target stored in the database 110 and the search query. The search unit 105 calculates the similarity between the posture feature of the person of the search target stored in the database 110 and the calculated posture feature of the person of the search query. Also, the search unit 105 calculates the similarity between the object feature of the search target stored in the database 110 and the calculated object feature of the search query. The search unit 105 performs a similarity determination of the images based on the calculated similarity of the posture feature and the similarity of the object feature. For example, the search unit 105 extracts an image in which the calculated similarity of the posture feature and the similarity of the object feature are each greater than a threshold value as a similar image. The similarity determination may be performed by weighting either or both of the similarity of the posture feature and the similarity of the object feature. For example, the calculated similarity of the posture feature and the similarity of the object feature may be weighted (for example, 1.0, 0.8, etc.), and the sum of the weighted similarities may be compared with a threshold value to perform the similarity determination. Also, the threshold value for determining each similarity may be changed according to the weight.
[0071] Also, the certainty of the posture estimation may be reflected in the similarity of the posture feature, and the certainty of the object recognition may be reflected in the similarity of the object feature. For example, the similarity between the certainty of the posture of the person to be searched and the certainty of the posture of the person in the search query may be calculated, and the similarity between the certainty of the object to be searched and the certainty of the object in the search query may be calculated. Also, the similarity may be calculated by weighting the feature according to each certainty. For example, the posture feature of the person to be searched is multiplied by the certainty of the posture, and the posture feature of the person in the search query is multiplied by the certainty of the posture, and the similarity of the posture feature is calculated from the multiplication result. The object feature of the search target is multiplied by the certainty of the object, and the object feature of the search query is multiplied by the certainty of the object, and the similarity of the object feature is calculated from the multiplication result. Alternatively, the certainty may be compared with a threshold, and only features whose certainty exceeds the threshold may be used in the similarity calculation. For example, if the certainty of the recognition of the person and object in the search query and the person to be searched exceeds a threshold, but the certainty of the recognition of the object to be searched is below the threshold, the search may be performed based on the similarity of only the features related to the person's posture, without considering the similarity of the object.
[0072] Next, the image processing device 100 displays the image search results (S113). The search unit 105 acquires images (similar images) obtained as search results from the database 110, and displays them on the display unit 107. The similar images and the search query image may be displayed, and the posture (skeletal structure) and person area (skeletal area) of the person in each image, the class of the object, the object area, etc. may be displayed. When there are multiple similar images, the display of each image may be changed according to the similarity. The images may be displayed in descending order of similarity, or images with high similarity may be highlighted.
[0073] FIG. 8 shows a specific example of image search by the image processing device 100 according to the present embodiment. As shown in FIG. 8, for example, when searching for a scene (image) of a traffic accident, a person in a crouching posture and a car in an image capturing a traffic accident are selected and input as a search query Q2 to the image processing device 100. Then, the image processing device 100 estimates the skeleton of the person in the crouching posture from the image of the search query Q2, and recognizes a car of the object class from the image of the search query Q2. From the images to be searched in the database 110, images including a posture with a high similarity to the skeleton of the crouching posture and including an object of a class with a high similarity to a car are extracted. As a result, it is possible to extract images including a person in a crouching posture and a car, such as the search target P3 and the search target P4, and it is possible to search for a desired traffic accident scene.
[0074] As described above, in this embodiment, similar images are searched for using the posture feature of a person and the object feature of an object in an image as a search query. That is, the posture of a person in an image to be searched is estimated and the posture feature is calculated, and the object is recognized and the object feature is calculated. Furthermore, for the search query, the posture of a person is estimated and the posture feature is calculated, and the object is recognized and the object feature is calculated. Based on the similarity of each posture feature and object feature, images similar to the search query are extracted from the images to be searched for. This makes it possible to search for images with similar postures and similar objects, making it possible to search for images that are closer to the image (scene) to be searched for.
[0075] (Embodiment 2) Hereinafter, the second embodiment will be described with reference to the drawings. In this embodiment, an example of searching for similar images using the characteristics of the relationship between a person and an object in addition to the first embodiment will be described.
[0076] In the first embodiment, similar images are searched for by combining the characteristics of the person's posture and the characteristics of the object. This makes it possible to search for images with similar postures of people and objects, as described above. However, even in the first embodiment, there is a possibility that an image close to the image the user wants to search cannot be searched for in some cases.
[0077] For example, as shown in FIG. 9, when searching for a scene of a person operating a PC, a person in a sitting position in an image and a PC are selected as a search query Q3. Then, in the first embodiment, images containing a person in a sitting position and a PC are extracted in order to search for images in which the person's posture and object are similar. As a result, images containing not only a person operating a PC but also a person sitting away from the PC are extracted, as in the search target P5. That is, in the first embodiment, there are cases in which an image containing a similar posture and a similar object is detected by chance. Therefore, in the present embodiment, an image search that takes into account the relationship between a person and an object is made possible.
[0078] The configuration of the image processing device 100 is the same as in embodiment 1. In this embodiment, the image processing device 100 performs similarity determination based on the relationship between the person and the object in each image, and searches for similar images.
[0079] The feature calculation unit 104 calculates a relationship feature relating to the relationship between a person and an object, in addition to a posture feature of a person and an object feature of an object. The feature calculation unit 104 calculates a posture feature of a person, an object feature of an object, and a relationship feature between a person and an object in an image to be searched, and also calculates a posture feature of a person, an object feature of an object, and a relationship feature between a person and an object in an image of a search query.
[0080] The search unit 105 performs a similarity determination based on the similarity of the posture feature amount, the similarity of the object feature amount, and the similarity of the relationship feature amount. The similarity determination may be performed based on weights of the similarity of the posture feature amount, the similarity of the object feature amount, and the similarity of the relationship feature amount.
[0081] The relationship feature amount in this embodiment includes, for example, a distance relationship feature amount based on the distance between a person and an object, an orientation relationship feature amount based on the orientation of a person and an object, and a position relationship feature amount based on the positional relationship between a person and an object. The feature amount calculation unit 104 may calculate any one of the distance relationship feature amount, the orientation relationship feature amount, and the position relationship feature amount, or may calculate any combination of relationship feature amounts. An example of calculation of each relationship feature amount is shown below.
[0082] <Distance relationship feature value> The feature calculation unit 104 extracts the distance between a person and an object from the search query image or the search target image, and uses the extracted distance as a feature (distance relationship feature). Fig. 10A and Fig. 10B show examples of distance extraction used for the distance relationship feature. Fig. 10A shows an example of distance extraction for the search query Q3 in Fig. 9, and Fig. 10B shows an example of distance extraction for the search target P5 in Fig. 9.
[0083] The distance between a person and an object used in the distance relationship feature is, for example, the distance between the person region of the person whose posture is estimated and the object region of the recognized object. The person region is a rectangular region including the person whose posture is estimated, for example, a skeleton region including the skeleton of the person estimated by posture estimation as shown in the first embodiment. The person region may be a posture region including a person whose posture is detected when posture is detected by another method, or a person region including a person recognized when a person is image-recognized. The object region is a rectangular region including a recognized object, and is an object region including an object detected by object recognition.
[0084] The feature amount calculation unit 104 calculates the distance (number of pixels) of a line connecting the coordinates of an arbitrary point included in the person area and the coordinates of an arbitrary point included in the object area in the image. In the example of Fig. 10A and Fig. 10B, the distance between the center point of the person area and the center point of the object area is calculated. That is, the coordinate of the center point of the person area is calculated from the coordinates of each vertex of the rectangular person area, the coordinate of the center point of the object area is calculated from the coordinates of each vertex of the rectangular object area, and the distance between the center point of the person area and the center point of the object area is calculated.
[0085] The feature amount calculation unit 104 may also calculate the distance between the nearest points of the person area and the object area. For example, the nearest points of all points of the person area and all points of the object area may be calculated, and the distance between the nearest points may be calculated, or the distance between each vertex of the person area and the nearest vertex of each vertex of the object area may be calculated. The feature amount calculation unit 104 may also calculate the distance between the farthest points of the person area and the object area. For example, the farthest points of all points of the person area and all points of the object area may be calculated, and the distance between the farthest points may be calculated, or the distance between each vertex of the person area and the farthest vertex of each vertex of the object area may be calculated. Furthermore, the distance between any vertex of the person area and any vertex of the object area may be calculated.
[0086] Furthermore, the feature amount calculation unit 104 may normalize the calculated distance between the person and the object by a normalization parameter, and use the normalized distance as the feature amount. The normalization parameter may be, for example, the image size of the search query or the target image, the height of the person whose posture is estimated (the number of height pixels shown in the first embodiment), the average size (height, width, area, etc.) of the person region and the object region, or IoU (Intersection over Union) indicating the degree of overlap between the person region and the object region. The feature amount calculation unit 104 normalizes the distance by dividing the distance between the person and the object by the normalization parameter.
[0087] Such distance relationship feature amounts can be used to obtain feature amounts that indicate the characteristics of the relationship between an object and a person when the person is sitting near a PC as in the search query Q3 in Fig. 10A, or the relationship between an object and a person when the person is sitting away from a PC as in the search target P5 in Fig. 10B. Therefore, by making a similarity judgment based on the distance relationship feature amount, it can be determined that the search query Q3 and the search target P5 are dissimilar.
[0088] <Direction relationship feature> The feature calculation unit 104 obtains the orientation of a person from the search query image or the search target image and uses the obtained orientation as a feature (orientation relationship feature). Fig. 11A and Fig. 11B show examples of extracted orientations of people to be used as orientation relationship features. Fig. 11A shows an example of extracted orientations of people in the search query Q3 in Fig. 9, and Fig. 11B shows an example of extracted orientations of people in the search target P5 in Fig. 9.
[0089] The orientation of a person used in the orientation relationship feature may be extracted from, for example, the orientation of a person estimated by estimating the orientation of the person, as shown in Figs. 11A and 11B. That is, since the front, back, left and right of a person can be detected from the estimated skeletal structure, the forward direction of the person in the image is extracted as the orientation of the person. Even if the orientation is estimated by other methods other than the skeletal structure, the orientation of the person can be extracted in the same manner. Furthermore, the orientation of a person may be extracted from the orientation of the person's face, not limited to the orientation of the person. For example, the face of a person is recognized from an image, and the recognized orientation of the face is regarded as the orientation of the person. Furthermore, the orientation of a person may be extracted from the gaze of the person. For example, the gaze of a person is recognized from an image, and the recognized gaze direction is regarded as the orientation of the person.
[0090] Furthermore, the feature calculation unit 104 obtains, as a feature, for example, a similarity (relationship) between the orientation of the extracted person and the orientation of a line connecting the person and the object. As an example, the feature calculation unit 104 obtains a cosine similarity between the orientation of the person and the line connecting the person and the object. The line connecting the person and the object may be a line connecting the centers of rectangles or a line connecting any points of the rectangle, as in the case of the distance relationship feature described above.
[0091] In addition, in object recognition, if the orientation of an object can be detected from an image, the orientation of the detected object may be used as a feature. For example, if a PC is recognized as an object, the orientation of the PC screen may be used as the orientation of the object. If a car is recognized as an object, the forward direction of the car may be used as the orientation of the object. In this case, the similarity (relationship) between the orientation of the extracted object and the orientation of the person may be obtained as a feature.
[0092] Such orientation relationship features can be used to obtain features that indicate the characteristics of the relationship between an object and a person when, for example, a person is sitting facing toward a PC as in search query Q3 in Fig. 11A, or when a person is sitting facing away from a PC as in search target P5 in Fig. 11B. Therefore, by making a similarity judgment based on the orientation relationship features, it can be determined that search query Q3 and search target P5 are dissimilar.
[0093] <Location relationship feature value> The feature calculation unit 104 obtains the positional relationship between a person and an object from the search query image or the search target image, and uses the obtained positional relationship as a feature (positional relationship feature). Figures 12A and 12B show examples of extracted positional relationships used for the positional relationship feature. Figure 12A shows an example of extracted positional relationships in the search query Q3 in Figure 9, and Figure 12B shows an example of extracted positional relationships in the search target P5 in Figure 9.
[0094] The positional relationship used in the positional relationship feature can be extracted, for example, from multiple distances between the person whose posture is estimated and the recognized object. That is, the positional relationship between one point on one of the person whose posture is estimated and the estimated object and multiple points on the other is used. For example, it may be a one-to-many positional relationship between multiple points on the person whose posture is estimated and one point on the estimated object, or a one-to-many positional relationship between a point on the person whose posture is estimated and multiple points on the recognized object. Note that the positional relationship between multiple points in two regions may also be used.
[0095] The feature amount calculation unit 104 calculates the distances of multiple lines connecting the person area and the object area in the image. In the example of FIG. 12A and FIG. 12B, the distances between one point in the object area and multiple points in the person area are calculated. The one point in the object area may be the center point of the object area or any point in the object area, as in the case of the distance relationship feature amount described above. If the object can be recognized from the image of the object, a point of interest on the screen of a PC or the like may be set as one point on the object. The multiple points in the person area may be joint points (key points, parts) of the person included in the skeleton (posture) of the recognized person. In FIG. 12A and FIG. 12B, as an example, three points of the person's head (for example, key point A1), wrist (key point A51 or A52), and ankle (key point A81 or A82) are extracted. Not limited to the recognized skeleton (posture), points of each part such as the head, wrist, and ankle recognized by image recognition may be extracted.
[0096] The feature amount calculation unit 104 calculates the feature amount based on a plurality of distances between the obtained person region and object region, a normalized value of the plurality of distances (normalized in the same way as the distance relationship feature amount), or a plurality of vectors including distances and orientations. After the feature amount is calculated, in the similarity determination in the search unit 105, for example, the similarity of the feature amount of the query image and the feature amount of the search target image may be the similarity of the plurality of distances, the similarity of the normalized value of the plurality of distances, or the similarity of the plurality of vectors (Lk distance (Euclidean distance or Manhattan distance), cosine similarity). The search unit 105 may determine the distance or similarity between the feature amount of the query image and the feature amount of the search target image. For example, the similarity determination may be performed based on a plurality of distances, a normalized value of the plurality of distances, or the distance (Euclidean distance or Manhattan distance, etc.) or similarity (cosine similarity, etc.) between a plurality of vectors.
[0097] Also, the order of each point of a person or object according to the length of the calculated distances may be used as the feature. For example, when using the distance between one point of an object and multiple joint points, the order of the proximity of each joint point may be used as the feature. For example, in the search query Q3 of FIG. 12A, the distance of each joint point of the person from the PC is in the relationship of wrist<ankle<head, so the feature is (1 wrist, 2 ankle, 3 head). In the search target P5 of FIG. 12B, the distance of each joint point of the person from the PC is in the relationship of head<wrist<ankle, so the feature is (1 head, 2 wrist, 3 ankle).
[0098] Such positional relationship features can be used to obtain features that indicate the characteristics of the relationship between an object and a person when, for example, a person sits near a PC with his or her hands on the PC side as in search query Q3 in Fig. 12A, or when a person sits away from the PC with his or her hands on the opposite side to the PC as in search target P5 in Fig. 12B. Therefore, by determining similarity based on positional relationship features, it can be determined that search query Q3 and search target P5 are dissimilar.
[0099] Furthermore, the distance relationship feature, the orientation relationship feature, and the positional relationship feature may be feature of distance, orientation, and positional relationship in a three-dimensional space. By using the camera parameters that acquired the search query or the search target image, the posture of a person and the positional relationship (distance, orientation, and positional relationship) between a person and an object in a three-dimensional space may be estimated, and each feature may be calculated by the above-mentioned method. In this case, the image processing device 100 may include a camera parameter acquisition unit that acquires camera parameters from a camera or the like that captured the image. For example, as in the first embodiment, an object whose length and position are known in advance may be captured by a camera, and the camera parameters may be obtained from the image.
[0100] FIG. 13 shows an example of the operation of the image processing device 100 according to this embodiment, illustrating the flow of processing for acquiring a search target image and storing it in a database.
[0101] As shown in FIG. 13, similarly to the first embodiment, when the image processing device 100 acquires an image to be searched (S101), it estimates the posture of a person in the acquired image (S102a) and calculates posture features (S103a), and also recognizes an object in the acquired image (S104a) and calculates object features (S105a).
[0102] In this embodiment, following S103a and S105a, the image processing device 100 calculates a relationship feature amount relating to the relationship between the person and the object in the acquired image (S106a). The feature amount calculation unit 104 calculates the relationship feature amount based on the relationship between the person whose posture was estimated in S102a from the image to be searched and the object recognized in S104a. The feature amount calculation unit 104 calculates, as the relationship feature amount, for example, a distance relationship feature amount, a direction relationship feature amount, and a position relationship feature amount as described above. The feature amount calculation unit 104 stores the calculated relationship feature amount in the database 110.
[0103] FIG. 14 shows an example of the operation of the image processing device 100 according to this embodiment, and shows the flow of a process for searching for images similar to a search query from among images to be searched and stored in a database by the process of FIG.
[0104] As shown in FIG. 14, similarly to embodiment 1, when a search query is input (S111), the image processing device 100 estimates a posture of a person in the search query (S102b) and calculates posture features (S103b), and also recognizes an object in the search query (S104b) and calculates object features (S105b).
[0105] In this embodiment, following S103b and S105b, the image processing device 100 calculates a relationship feature amount relating to the relationship between the person and the object in the search query (S106b). The feature amount calculation unit 104 calculates the relationship feature amount based on the relationship between the person whose posture is estimated in S102b from the image of the search query and the object recognized in S104b. As in the case of storing the search target, the feature amount calculation unit 104 calculates, as the relationship feature amount, for example, a distance relationship feature amount, a direction relationship feature amount, and a position relationship feature amount.
[0106] Next, the image processing device 100 searches for an image based on the search query (S112). The search unit 105 calculates a similarity between the posture feature amount of the person in the search query and the posture feature amount of the person to be searched, and object The search unit 105 calculates the similarity between the feature and the object feature of the object to be searched, and further calculates the similarity between the relationship feature of the person and object of the search query and the relationship feature of the person and object to be searched. The search unit 105 performs similarity determination of the images based on the similarity of the posture feature, the similarity of the object feature, and the similarity of the relationship feature obtained. For example, the search unit 105 extracts an image in which the similarity of the posture feature, the similarity of the object feature, and the similarity of the relationship feature are each greater than a threshold value as a similar image. The similarity determination may be performed by weighting one of the similarity of the posture feature, the similarity of the object feature, and the similarity of the relationship feature, or a selected similarity. For example, the similarity determination may be performed by weighting each of the similarity of the posture feature, the similarity of the object feature, and the similarity of the relationship feature obtained (for example, 1.0, 0.8, 0.5, etc.), and comparing the sum of the weighted similarities with a threshold value. The threshold value for determining each similarity may be changed according to the weight.
[0107] Furthermore, the search unit 105 may calculate the similarity of the distance relationship feature amount, the similarity of the orientation relationship feature amount, and the similarity of the positional relationship feature amount (if any of the feature amounts is calculated, any of the similarities) as the similarity of the relationship feature amount. The search unit 105 performs similarity determination of the images including the similarity of the distance relationship feature amount, the similarity of the orientation relationship feature amount, and the similarity of the positional relationship feature amount obtained. For example, it is determined whether or not each of the similarity of the distance relationship feature amount, the similarity of the orientation relationship feature amount, and the similarity of the positional relationship feature amount is greater than a threshold value. The similarity determination may be performed by weighting any one of the similarity of the distance relationship feature amount, the similarity of the orientation relationship feature amount, and the similarity of the positional relationship feature amount, or a selected similarity. For example, the similarity determination may be performed by weighting each of the similarity of the distance relationship feature amount, the similarity of the orientation relationship feature amount, and the similarity of the positional relationship feature amount obtained, and comparing the sum of the weighted similarities with a threshold value. In addition, the threshold value for determining each similarity may be changed according to the weight.
[0108] As described above, in this embodiment, in addition to the configuration of the first embodiment, similar images are searched for using a relationship feature relating to the relationship between a person and an object. Furthermore, features relating to the distance between a person and an object, the orientation of a person and an object, and the positional relationship between a person and an object are used as the relationship feature. This makes it possible to search for images that are similar not only in terms of the person's posture and the object, but also in terms of the relationship between the person and the object, and therefore makes it possible to search for images that are closer to the image (scene) to be searched for.
[0109] (Embodiment 3) Hereinafter, a third embodiment will be described with reference to the drawings. In this embodiment, an example in which the first or second embodiment is further combined with HOI detection to search for similar images will be described.
[0110] Fig. 15 shows the configuration of an image processing device 100 according to this embodiment. As shown in Fig. 15, the image processing device 100 according to this embodiment includes an HOI detection unit 109 in addition to the configuration of the first or second embodiment.
[0111] The HOI detection unit 109 performs HOI detection as described in Non-Patent Document 1. The HOI detection unit 109 detects pairs of related people and objects and the person's verbs (for example, an action such as a person kicking a soccer ball) from an image by HOI detection. FIG. 16 shows a detection example of HOI detection. In the example of FIG. 16, a pair of related people and a mobile phone (object) is detected from an image, and a verb that the person is talking on the phone is detected. In addition, the HOI detection unit 109 generates a relevance score (certainty) of the person's verb detected by HOI detection. The higher the relevance score, the more likely it is that the detected person's verb (including the person and object pair) is correct.
[0112] Note that pairs of related people and objects and person verbs may be detected by detection techniques using machine learning other than HOI detection. For example, detection similar to HOI detection may be performed by machine learning images of pairs of related people and objects using labels of person verbs. Furthermore, the HOI detection unit 109 may obtain an HOI detection result obtained by performing HOI detection on an image in advance from an external device (such as the image providing device 200, the database 110, or the input unit 106). The search unit 105 may perform a similarity determination using the HOI detection result obtained from outside, or may perform a similarity determination using the HOI detection result detected by the HOI detection process of the HOI detection unit 109. For example, the search unit 105 may perform a similarity determination between the first image and the second image based on the HOI detection result obtained by performing HOI detection process on the first image and the HOI detection result of the second image obtained from outside.
[0113] 17 shows an example of the operation of the image processing device 100 according to this embodiment, and illustrates the flow of processing for acquiring a search target image and storing it in a database. Note that, although an example in which this embodiment is applied to the operation of embodiment 2 is shown here, this embodiment may also be applied to the operation of embodiment 1.
[0114] As shown in FIG. 17, similarly to the second embodiment, when the image processing device 100 acquires an image to be searched (S101), it estimates a posture of a person in the acquired image (S102a) and calculates posture features (S103a), recognizes an object in the acquired image (S104a) and calculates object features (S105a), and further calculates relationship features between the person and the object in the acquired image (S106a).
[0115] In this embodiment, following image acquisition (S101), the image processing device 100 performs HOI detection based on the acquired image (S201a). The HOI detection unit 109 performs HOI detection on the acquired image, detects related person-object pairs and person verbs (actions) in the image, and generates a relevance score (certainty) of the detected person verbs. The HOI detection unit 109 stores the detected person-object pairs, person verbs, and relevance score in the database 110.
[0116] Fig. 18 shows an example of the operation of the image processing device 100 according to this embodiment, and shows a process flow for searching for images similar to a search query from images to be searched and stored in a database by the process of Fig. 17. Note that, although an example in which this embodiment is applied to the operation of embodiment 2 is shown here, this embodiment may also be applied to the operation of embodiment 1.
[0117] As shown in FIG. 18, similarly to the second embodiment, when a search query is input (S111), the image processing device 100 estimates the posture of a person in the search query (S102b) and calculates posture features (S103b), recognizes an object in the search query (S104b) and calculates object features (S105b), and further calculates relationship features between the person and the object in the search query (S106b).
[0118] In this embodiment, the image processing device 100 performs HOI detection (S201b) following input of a search query (S111). As in the case of storing a search target, the HOI detection unit 109 performs HOI detection on the image of the search query, detects related person-object pairs and person verbs in the image, and generates a relevance score for the detected person verbs.
[0119] Next, the image processing device 100 searches for an image based on the search query (S112). a Similar images are searched for by combining a pose-object search (first similarity judgment) using the pose estimation results and object recognition results from S102b to S106b with a HOI search (second similarity judgment) using the HOI detection results from S201a and S201b.
[0120] The posture-object search is the search method described in the first or second embodiment. That is, for a search target image and a search query image, the similarity of the person posture feature amount and the similarity of the object feature amount (and further the similarity of the relationship feature amount) are calculated, a similarity determination is performed based on the calculated similarity amount, and similar images are searched for.
[0121] HOI search calculates the similarity between the HOI detection result for the search target image and the HOI detection result for the search query image, and performs a similarity judgment based on the calculated similarity to search for similar images. In other words, similarity judgment is performed based on the similarity between related person-object pairs and person verbs obtained by HOI detection, and similar images are searched for.
[0122] The search unit 105 may select either the pose-object search or the HOI search to perform the search. For example, the search unit 105 selects either the pose-object search or the HOI search based on the relevance score (certainty) of the HOI detection result to perform the search. When the relevance score of the HOI detection result of the search query is higher than a threshold, that is, when a verb with high confidence is estimated from the image of the search query, the search unit 105 searches for similar images by the pose-object search. When the relevance score of the HOI detection result of the search query is lower than a threshold, that is, when a verb with high confidence is not estimated from the image of the search query, the search unit 105 searches for similar images by the pose-object search.
[0123] The search unit 105 may also perform a search using both the posture-object search and the HOI search. For example, the search unit 105 performs a search by weighting the posture-object search and the HOI search based on the relevance score of the HOI search result. When the relevance score of the HOI detection result of the search query and the search target is higher than a threshold, that is, when a verb with high confidence is estimated from both the image of the search query and the image of the search target (such as calling: 0.8), the search unit 105 performs a search for similar images by weighting the HOI search. When the relevance score of the HOI detection result of the search query and the search target is lower than a threshold, that is, when a verb with high confidence is not estimated from both the image of the search query and the image of the search target (such as picking up: 0.03), the search unit 105 performs a search for similar images by weighting the posture-object search.
[0124] Either the pose-object search or the HOI search may be selected, or a weight may be assigned to the pose-object search and the HOI search, based on the confidence of the pose estimation and the object recognition (for example, the average of the confidence of the pose estimation and the confidence of the object recognition) instead of the relevance score of the HOI search result. Either the pose-object search or the HOI search may be selected, or a weight may be assigned to the pose-object search and the HOI search, based on the result of comparing the relevance score of the HOI search result and the confidence of the pose estimation and the object recognition.
[0125] Also, the user may manually adjust the weights of the pose-object search and the HOI search. For example, when weighting the pose-object search and the HOI search, weights may be applied to the similarity used in the pose-object search and the similarity used in the HOI search. Weights may be applied to the similarity of the person's pose features and the similarity of the object features (and further the similarity of the relationship features), and weights may be applied to the similarity of the person and object pairs related to the HOI detection and the person's verb, and images may be extracted in which any of the weighted similarities is greater than a threshold, or images in which the sum of the weighted similarities is greater than a threshold may be extracted.
[0126] As described above, in this embodiment, similar images are searched for using the detection result of HOI detection in addition to the configuration of embodiment 1 or 2. By searching for images using the pose-object search according to embodiment 1 or 2 and the HOI search using HOI detection, similar images can be effectively searched for.
[0127] HOI detection can only search for events that are in the training data. For example, if the verb "traffic accident" has not been learned, it cannot search for similar images. In addition, since HOI detection does not recognize posture, it may misjudge the similarity. For example, if a person and a soccer ball are close to each other, it may be judged as kicking the ball even if the person has not kicked it. On the other hand, HOI detection can exclude unrelated person-object pairs and narrow down the search to related person-object pairs. Therefore, by performing a search using either posture-object search or HOI search, or by weighting posture-object search and HOI search depending on the confidence level of HOI detection, it is possible to take advantage of the advantages of HOI detection while compensating for its disadvantages and search for similar images with high accuracy.
[0128] The present disclosure is not limited to the above-described embodiment, and can be modified as appropriate without departing from the spirit and scope of the present disclosure.
[0129] Each component in the above-described embodiment may be configured by hardware or software, or both, and may be configured by one piece of hardware or software, or may be configured by multiple pieces of hardware or software. Each device and each function (processing) may be realized by a computer 20 having a processor 21 such as a CPU (Central Processing Unit) and a memory 22 serving as a storage device, as shown in Fig. 19. For example, a program for performing the method in the embodiment (image processing method) may be stored in the memory 22, and each function may be realized by the processor 21 executing the program stored in the memory 22.
[0130] These programs include instructions (or software codes) that, when loaded into a computer, cause the computer to perform one or more functions described in the embodiments. The programs may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, computer-readable media or tangible storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technology, CD-ROM, digital versatile disc (DVD), Blu-ray® disk or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device. The programs may be transmitted on a transitory computer-readable medium or a communication medium. By way of example and not limitation, a transitory computer-readable medium or a communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
[0131] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-mentioned embodiments. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure.
[0132] A part or all of the above-described embodiments can be described as, but is not limited to, the following supplementary notes. (Appendix 1) a posture estimation acquisition means for acquiring an estimation result of estimating a posture of a person included in the first and second images; an object recognition and acquisition means for acquiring a recognition result of recognizing an object other than the person included in the first and second images; a similarity determination means for determining similarity between the first image and the second image based on the estimation result of the person's posture and the recognition result of the object; An image processing system comprising: (Appendix 2) the similarity determination means performs the similarity determination based on a similarity of a posture feature based on an estimation result of the posture of the person and a similarity of an object feature based on a recognition result of the object. 2. An image processing system as described in claim 1. (Appendix 3) the similarity determination means performs the similarity determination based on weights of the similarity of the posture feature amount and the similarity of the object feature amount. 3. An image processing system as described in claim 2. (Appendix 4) the similarity determination means performs the similarity determination based on a degree of certainty of the person whose pose has been estimated and a degree of certainty of the estimated object. 4. An image processing system according to claim 1. (Appendix 5) the first and second images each include a plurality of images that are successive in time series; the similarity determination means performs the similarity determination based on a change in the estimated posture of the person and a change in the recognized object. 5. An image processing system according to any one of claims 1 to 4. (Appendix 6) the similarity determination means performs the similarity determination based on a relationship between the person and the object based on a result of estimating a posture of the person and a result of recognizing the object. 6. An image processing system according to any one of claims 1 to 5. (Appendix 7) the similarity determination means performs the similarity determination based on a similarity of a posture feature amount of the posture of the person, a similarity of an object feature amount of the object, and a similarity of a relationship feature amount based on a relationship between the person and the object. 7. The image processing system according to claim 6. (Appendix 8) the similarity determination means performs the similarity determination based on weights of the similarity of the posture feature amount, the similarity of the object feature amount, and the similarity of the relationship feature amount. 8. The image processing system of claim 7. (Appendix 9) the relationship feature indicating the relationship between the person and the object includes any one of a distance relationship feature based on a distance between the person and the object, a direction relationship feature based on a direction between the person and the object, and a position relationship feature based on a positional relationship between the person and the object; 9. The image processing system according to claim 7 or 8. (Appendix 10) The distance between the person and the object used in the distance relationship feature is a distance between a person region including the person whose pose is estimated and an object region including the recognized object. 10. The image processing system of claim 9. (Appendix 11) the distance between the person and the object includes any one of a distance between a center point of the person region and a center point of the object region, a distance between a nearest point of the person region and the object region, a distance between a farthest point of the person region and the object region, and a distance between any vertex of the person region and any vertex of the object region. 11. The image processing system of claim 10. (Appendix 12) the distance relationship feature is a feature obtained by normalizing the distance between the person and the object using a normalization parameter; 12. The image processing system according to claim 10 or 11. (Appendix 13) The normalization parameters include any one of image sizes of the first and second images, a height of the person based on the estimated posture of the person, an average size of the person region and the object region, and an Intersection over Union (IoU) between the person region and the object region. 13. The image processing system of claim 12. (Appendix 14) The distance between the person and the object is a distance in a three-dimensional space calculated from camera parameters used to capture the first and second images. 14. An image processing system according to any one of claims 10 to 13. (Appendix 15) The orientation of the person used in the orientation relationship feature includes any one of a body orientation of the person based on the estimated posture of the person, a face orientation of the person recognized from an image of the person, and a gaze direction of the person recognized from an image of the person. 15. An image processing system according to any one of claims 9 to 14. (Appendix 16) The orientation relationship feature is a feature based on a similarity between an orientation of the person and an orientation of a line connecting the person and the object. 16. The image processing system of claim 15. (Appendix 17) The orientation of the person is an orientation in a three-dimensional space calculated from camera parameters used to capture the first and second images. 17. The image processing system according to claim 15 or 16. (Appendix 18) The positional relationship used in the positional relationship feature amount is a positional relationship between one point on one of the person whose posture is estimated and the estimated object and a plurality of points on the other of the person whose posture is estimated and the estimated object. 18. An image processing system according to any one of claims 9 to 17. (Appendix 19) the points on the person are joint points of the person based on the estimated person's pose. 19. The image processing system of claim 18. (Appendix 20) The positional relationship between the one point and the multiple points includes any one of distances of multiple lines connecting the person point and the object point, normalized values of the distances of the multiple lines, and vectors of the multiple lines. 20. The image processing system according to claim 18 or 19. (Appendix 21) the similarity determination means performs similarity determination based on any one of a similarity of distances between the plurality of lines, a similarity of normalized values of distances between the plurality of lines, and a similarity of vectors between the plurality of lines. 21. The image processing system of claim 20. (Appendix 22) the positional relationship feature amount indicates an order of a plurality of points of the person or the object according to distances of the plurality of lines; 21. The image processing system of claim 20. (Appendix 23) The positional relationship between the one point and the multiple points is a positional relationship in a three-dimensional space calculated from camera parameters used to capture the first and second images. 23. An image processing system according to any one of claims 18 to 22. (Appendix 24) a HOI detection acquisition means for acquiring a HOI detection result of HOI (Human Object Interaction) detection for the first and second images; the similarity determination means performs a first similarity determination based on the estimation result of the person's posture and the recognition result of the object, and a second similarity determination based on the HOI detection result; 24. An image processing system according to any one of claims 1 to 23. (Appendix 25) The HOI detection and acquisition means performs HOI detection on the first or second image based on the first or second image. 25. The image processing system of claim 24. (Appendix 26) the similarity determination means performs the second similarity determination based on a HOI detection result obtained by performing HOI detection based on the first image and a HOI detection result of the acquired second image. 26. The image processing system according to claim 24 or 25. (Appendix 27) the similarity determination means performs a similarity determination by weighting either the first similarity determination or the second similarity determination, or the first similarity determination and the second similarity determination, depending on a confidence level of the detection result of the HOI detection; 27. An image processing system according to any one of claims 24 to 26. (Appendix 28) the first image is a query image; the second image includes a plurality of search target images; the similarity determination means searches for an image similar to the query image from among the plurality of search target images based on a result of the similarity determination; 28. An image processing system according to any one of claims 1 to 27. (Appendix 29) a database for storing the estimation results of the person's posture and the recognition results of the object in the plurality of search target images; the similarity determination means refers to the database and searches the plurality of search target images for an image similar to the query image; 29. The image processing system of claim 28. (Appendix 30) the posture estimation acquisition means estimates a posture of a person included in the first or second image based on the first or second image; The object recognition and acquisition means recognizes an object included in the first or second image based on the first or second image. 30. An image processing system according to any one of claims 1 to 29. (Appendix 31) the posture estimation acquisition means estimates a bone structure of a person included in the first or second image as a posture of the person based on the first or second image; 31. The image processing system of claim 30. (Appendix 32) The object recognition and acquisition means recognizes an object class of an object included in the first or second image based on the first and second images. 32. The image processing system according to claim 30 or 31. (Appendix 33) the similarity determination means performs the similarity determination based on an estimation result of a posture of a person estimated based on the first image and a recognition result of an object recognized based on the first image, and based on an estimation result of a posture of a person in the acquired second image and a recognition result of an object in the acquired second image. 33. An image processing system according to any one of claims 30 to 32. (Appendix 34) Obtaining an estimation result of estimating a posture of a person included in the first and second images; obtaining a recognition result of recognizing an object other than the person included in the first and second images; determining a similarity between the first image and the second image based on the estimation result of the person's posture and the recognition result of the object; Image processing methods. (Appendix 35) Obtaining an estimation result of estimating a posture of a person included in the first and second images; obtaining a recognition result of recognizing an object other than the person included in the first and second images; determining a similarity between the first image and the second image based on the estimation result of the person's posture and the recognition result of the object; A non-transitory computer-readable medium having stored thereon an image processing program for causing a computer to execute a process. [Explanation of symbols]
[0133] 1, 10 Image Processing System 11 Posture estimation acquisition unit 12 Object recognition acquisition section 13 Similarity determination section 20. Computers 21 Processors 22 Memory 100 Image processing device 101 Image acquisition unit 102 Posture estimation section 103 Object recognition section 104 Feature Calculation Unit 105 Search Department 106 Input section 107 Display section 108 Classification Department 109 HOI detection unit 110 Database 200 Image providing device 300 human body models
Claims
1. a posture estimation acquisition means for acquiring an estimation result of estimating a posture of a person included in the first and second images; an object recognition and acquisition means for acquiring a recognition result of recognizing an object other than the person included in the first and second images; a similarity determination means for performing a similarity determination between the first image and the second image based on a similarity of a relationship feature that indicates a relationship between the person and the object based on an estimation result of the posture of the person and a recognition result of the object; Equipped with the relationship feature includes a direction relationship feature based on a direction between the person and the object, Image processing system.
2. the similarity determination means performs the similarity determination based on a similarity of a posture feature based on an estimation result of the posture of the person and a similarity of an object feature based on a recognition result of the object. The image processing system according to claim 1 .
3. the similarity determination means performs the similarity determination based on weights of the similarity of the posture feature amount and the similarity of the object feature amount. The image processing system according to claim 2 .
4. the similarity determination means performs the similarity determination based on a degree of certainty of the person whose pose has been estimated and a degree of certainty of the recognized object. The image processing system according to any one of claims 1 to 3.
5. the first and second images each include a plurality of images that are successive in time series; the similarity determination means performs the similarity determination based on a change in the estimated posture of the person and a change in the recognized object. The image processing system according to any one of claims 1 to 4.
6. The orientation of the person used in the orientation relationship feature includes any one of the body orientation of the person based on the estimated posture of the person, the face orientation of the person recognized from an image of the person, and the gaze direction of the person recognized from an image of the person. The image processing system according to any one of claims 1 to 5.
7. The orientation relationship feature is a feature based on a similarity between the orientation of the person and the orientation of a line connecting the person and the object. The image processing system according to claim 6.
8. The relationship feature includes a distance relationship feature based on a distance between the previous person and the object, or a positional relationship feature based on a positional relationship between the previous person and the object. The image processing system according to any one of claims 1 to 7.
9. Obtaining an estimation result of estimating a posture of a person included in the first and second images; obtaining a recognition result of recognizing an object other than the person included in the first and second images; performing a similarity determination between the first image and the second image based on a similarity between a relationship feature amount indicating a relationship between the person and the object based on the estimation result of the posture of the person and the recognition result of the object; the relationship feature includes a direction relationship feature based on a direction between the person and the object, Image processing methods.
10. Obtaining an estimation result of estimating a posture of a person included in the first and second images; obtaining a recognition result of recognizing an object other than the person included in the first and second images; performing a similarity determination between the first image and the second image based on a similarity between a relationship feature amount indicating a relationship between the person and the object based on the estimation result of the posture of the person and the recognition result of the object; the relationship feature includes a direction relationship feature based on a direction between the person and the object, An image processing program that causes a computer to carry out the processing.
Citation Information
Patent Citations
Attitude estimation device
JP2011113398A
Object identification device
JP2013232080A
Image retrieving apparatus, image retrieving method, and setting screen used therefor
JP2019091138A
Interaction Detection Model for Identifying Human-Object Interactions in Image Content
US20190286892A1
Image search device, image search method, electronic equipment, and control method
WO2019171803A1