Information processing device, information processing method, and program

The image selection device improves the reliability assessment of posture-based image searches by displaying target images based on the reliability of posture estimation, addressing the challenge of unreliable search results in existing technologies.

JP7775918B2Active Publication Date: 2025-11-26NEC CORP
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2024086244
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2025-11-26
Estimated Expiration
2040-05-14

AI Technical Summary

Technical Problem

Existing image search technologies struggle to accurately reflect the reliability of posture estimation results, making it difficult for users to assess the validity of search results when searching for images based on a person's posture.

Method used

An image selection device and method that acquires query and search posture information, selects similar posture information based on a criterion, and displays target images using the reliability of the posture estimation, allowing users to recognize the impact of reliability on search results.

Benefits of technology

Enhances the user's ability to understand how the reliability of posture estimation affects search results, improving the validity and accuracy of image searches based on posture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007775918000003
    Figure 0007775918000003
  • Figure 0007775918000004
    Figure 0007775918000004
  • Figure 0007775918000005
    Figure 0007775918000005
Patent Text Reader

Abstract

To make it easier for a user to recognize the influence of a degree of reliability of a pose estimation result of an image to be searched on appropriateness of a search result when an image including a person is searched with reference to a pose of the person.SOLUTION: A query acquisition unit acquires pose information (hereinafter, described as query pose information) that is information to be a query and indicates a pose of a person. A search information acquisition unit acquires a plurality of pieces of search pose information. The search pose information is pose information about an image (hereinafter, described as a target image) to be searched for, and is stored in a database for every multiple target images. A selection unit selects, from among the plurality of pieces of search pose information, two or more pieces of search pose information whose degree of similarity to the query pose information satisfies a reference. The selection processing is substantially a process to select a target image similar to an image (hereinafter, described as a query image) corresponding to the query pose information.SELECTED DRAWING: Figure 40
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image selection device, an image selection method, and a program. [Background technology]

[0002] In recent years, technologies for detecting and searching for a person's posture, behavior, and other states from images captured by a surveillance camera have been used in surveillance systems and the like. Related technologies include, for example, Patent Documents 1 and 2. Patent Document 1 discloses a technology for searching for similar postures of a person based on key joints of the person's head, limbs, and the like included in depth images. Patent Document 2 discloses a technology for searching for similar images using posture information such as tilt added to an image, although this information is not related to the posture of the person. In addition, Non-Patent Document 1 is known as a technology related to human skeleton estimation.

[0003] On the other hand, when searching for images similar to a query image, the search results often include multiple images. Patent Document 3 describes displaying images in descending order of similarity to the query image. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Special Publication No. 2014-522035 [Patent Document 2] Japanese Patent Application Laid-Open No. 2006-260405 [Patent Document 3] Japanese Patent Application Laid-Open No. 2012-003623 [Non-patent literature]

[0005] [Non-Patent Document 1] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299 Summary of the Invention [Problem to be solved by the invention]

[0006] The present inventors have considered that when searching for images including a person based on the posture of the person, the reliability of the posture estimation result of the image to be searched affects the validity of the search results. One of the problems that the present invention aims to solve is to make it easier for a user to recognize the impact that the reliability of the posture estimation result of the image to be searched for affects the validity of the search results when searching for images including a person based on the posture of the person. [Means for solving the problem]

[0007] According to the present invention, a query acquisition unit acquires query posture information indicating a posture of a person, which is information to be used as a query; a search information acquisition means for acquiring a plurality of pieces of search posture information that are generated for each of a plurality of target images to be searched and indicate the postures of people included in the target images; a selection means for selecting two or more pieces of search orientation information from the plurality of pieces of search orientation information, the two or more pieces of search orientation information having similarity to the query orientation information that satisfies a criterion; a display control means for displaying the target images corresponding to the two or more pieces of orientation information for search selected by the selection means on a display means, and setting the display positions of the target images using the reliability of the orientation information for search; An image selection device is provided, comprising:

[0008] According to the present invention, a computer Acquire query posture information that indicates the posture of a person. acquiring a plurality of pieces of search posture information that are generated for each of a plurality of target images to be searched and indicate postures of people included in the target images; selecting two or more pieces of search posture information from the plurality of pieces of search posture information whose similarity to the query posture information satisfies a criterion; An image selection method is provided in which the target images corresponding to the two or more selected search orientation information are displayed on a display means, and the display positions of the target images are set using the reliability of the search orientation information.

[0009] According to the present invention, a computer is provided with: a query acquisition function for acquiring query posture information indicating a posture of a person, which is information to be used as a query; a search information acquisition function that acquires a plurality of pieces of search posture information that are generated for each of a plurality of target images to be searched and indicate the postures of people included in the target images; a selection function for selecting two or more pieces of search posture information from the plurality of pieces of search posture information, the search posture information having a similarity to the query posture information that satisfies a criterion; a display control function that displays the target images corresponding to the two or more pieces of orientation information for search selected by the selection function on a display means, and sets the display positions of the target images using the reliability of the orientation information for search; A program to help them develop this skill will be provided. [Effects of the Invention]

[0010] According to the present invention, when searching for images containing a person based on the posture of the person, it becomes easier for a user to recognize the impact that the reliability of the posture estimation result of the image to be searched for has on the validity of the search results. [Brief explanation of the drawings]

[0011] The above-mentioned objects, as well as other objects, features and advantages, will become more apparent from the preferred embodiments described below and the accompanying drawings.

[0012] [Figure 1] 1 is a configuration diagram showing an overview of an image processing device according to an embodiment; [Figure 2] 1 is a configuration diagram showing a configuration of an image processing device according to a first embodiment. [Figure 3] 3 is a flowchart showing an image processing method according to the first embodiment. [Figure 4] 3 is a flowchart showing a classification method according to the first embodiment. [Figure 5] 3 is a flowchart showing a search method according to the first embodiment. [Figure 6] FIG. 3 is a diagram showing an example of detection of a skeletal structure according to the first embodiment. [Figure 7] FIG. 2 is a diagram showing a human body model according to the first embodiment. [Figure 8] FIG. 3 is a diagram showing an example of detection of a skeletal structure according to the first embodiment. [Figure 9] FIG. 3 is a diagram showing an example of detection of a skeletal structure according to the first embodiment. [Figure 10] FIG. 3 is a diagram showing an example of detection of a skeletal structure according to the first embodiment. [Figure 11] 4 is a graph showing a specific example of a classification method according to the first embodiment. [Figure 12] FIG. 4 is a diagram showing an example of displaying a classification result according to the first embodiment. [Figure 13] FIG. 2 is a diagram for explaining a search method according to the first embodiment. [Figure 14] FIG. 2 is a diagram for explaining a search method according to the first embodiment. [Figure 15] FIG. 2 is a diagram for explaining a search method according to the first embodiment. [Figure 16] FIG. 2 is a diagram for explaining a search method according to the first embodiment. [Figure 17] FIG. 10 is a diagram showing an example of display of search results according to the first embodiment. [Figure 18] FIG. 10 is a configuration diagram showing the configuration of an image processing device according to a second embodiment. [Figure 19] 10 is a flowchart showing an image processing method according to the second embodiment. [Figure 20] 10 is a flowchart showing a specific example 1 of a height pixel number calculation method according to the second embodiment. [Figure 21] 10 is a flowchart showing a specific example 2 of a height pixel number calculation method according to the second embodiment. [Figure 22] 10 is a flowchart showing a specific example 3 of a height pixel number calculation method according to the second embodiment. [Figure 23] 10 is a flowchart showing a normalization method according to the second embodiment. [Figure 24] FIG. 10 is a diagram showing a human body model according to a second embodiment. [Figure 25] FIG. 10 is a diagram showing an example of detection of a skeletal structure according to the second embodiment. [Figure 26] FIG. 10 is a diagram showing an example of detection of a skeletal structure according to the second embodiment. [Figure 27] FIG. 10 is a diagram showing an example of detection of a skeletal structure according to the second embodiment. [Figure 28] FIG. 10 is a diagram showing a human body model according to a second embodiment. [Figure 29] FIG. 10 is a diagram showing an example of detection of a skeletal structure according to the second embodiment. [Figure 30] 10 is a histogram for explaining a height pixel number calculation method according to the second embodiment. [Figure 31] FIG. 10 is a diagram showing an example of detection of a skeletal structure according to the second embodiment. [Figure 32] FIG. 10 is a diagram showing a three-dimensional human body model according to a second embodiment. [Figure 33] FIG. 10 is a diagram for explaining a height pixel number calculation method according to the second embodiment. [Figure 34] FIG. 10 is a diagram for explaining a height pixel number calculation method according to the second embodiment. [Figure 35] FIG. 10 is a diagram for explaining a height pixel number calculation method according to the second embodiment. [Figure 36]FIG. 10 is a diagram for explaining a normalization method according to the second embodiment. [Figure 37] FIG. 10 is a diagram for explaining a normalization method according to the second embodiment. [Figure 38] FIG. 10 is a diagram for explaining a normalization method according to the second embodiment. [Figure 39] FIG. 1 illustrates an example of a hardware configuration of an image processing apparatus. [Figure 40] FIG. 13 is a diagram showing an example of the functional configuration of a search unit 105 according to a search method 6. [Figure 41] 10 is a flowchart showing an example of processing performed by the search unit 105 according to a search method 6. [Figure 42] FIG. 10 is a diagram showing an example of a method for determining the display position of the target image, which is performed in step S340. [Figure 43] 42 is a flowchart showing a first modification of FIG. 41. [Figure 44] 42 is a flowchart showing a second modification of FIG. 41. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and their description will be omitted where appropriate.

[0014] (Considerations leading to the embodiment) In recent years, image recognition technology that utilizes machine learning such as deep learning has been applied to various systems. For example, it is being applied to surveillance systems that monitor images from surveillance cameras. By utilizing machine learning in surveillance systems, it is becoming possible to grasp to some extent the state of a person, such as their posture and behavior, from images.

[0015] However, such related technologies may not always be able to grasp the state of a person desired by a user on demand. For example, there are cases where the state of a person that a user wants to search and grasp can be specified in advance, and there are also cases where the state cannot be specifically specified, such as an unknown state. As a result, in some cases, the user cannot specify in detail the state of the person they want to search for. Furthermore, searches cannot be performed if part of the person's body is hidden. With related technologies, a person's state can only be searched for using specific search criteria, making it difficult to flexibly search for and classify the state of a desired person.

[0016] Therefore, the inventors have investigated a method of utilizing skeletal estimation technology such as that disclosed in Non-Patent Document 1 in order to recognize the state of a person desired by a user from an image on demand. Related skeletal estimation technologies, such as OpenPose disclosed in Non-Patent Document 1, estimate a person's skeletal structure by learning from image data with various patterns of correct answers. In the following embodiment, such skeletal estimation technology is utilized to enable flexible recognition of a person's state.

[0017] The skeletal structure estimated by a skeletal estimation technique such as OpenPose is composed of "key points," which are characteristic points such as joints, and "bones (bone links)," which indicate the links between the key points. Therefore, in the following embodiments, the skeletal structure will be described using the terms "key points" and "bones," but unless otherwise specified, "key points" correspond to the "joints" of a person, and "bones" correspond to the "bones" of a person. The positions of the "key points" are an example of joint information.

[0018] (Outline of the embodiment) FIG. 1 shows an overview of an image processing device 10 according to an embodiment. As shown in FIG. 1, the image processing device 10 includes a skeleton detection unit 11, a feature calculation unit 12, and a recognition unit 13. The skeleton detection unit 11 detects two-dimensional skeleton structures of multiple people based on two-dimensional images acquired from a camera or the like. The feature calculation unit 12 calculates feature amounts of the multiple two-dimensional skeleton structures detected by the skeleton detection unit 11. The recognition unit 13 performs recognition processing of the states of multiple people based on the similarity of the multiple feature amounts calculated by the feature calculation unit 12. The recognition processing includes classification processing and search processing (selection processing) of the states of people. Therefore, the image processing device 10 also functions as an image selection device.

[0019] In this way, in the embodiment, the two-dimensional skeletal structure of a person is detected from a two-dimensional image, and recognition processing such as classification and search of the person's state is performed based on the features calculated from this two-dimensional skeletal structure, thereby making it possible to flexibly recognize the state of a desired person.

[0020] (First Embodiment) Hereinafter, the first embodiment will be described with reference to the drawings. FIG. 2 shows the configuration of an image processing device 100 according to this embodiment. The image processing device 100, together with a camera 200 and a database (DB) 110, constitutes an image processing system 1. The image processing system 1 including the image processing device 100 is a system that classifies and searches for states such as posture and behavior of a person based on the skeletal structure of the person estimated from an image. The image processing device 100 also functions as an image selection device.

[0021] The camera 200 is an imaging unit such as a surveillance camera that generates two-dimensional images. The camera 200 is installed at a predetermined location and captures an image of a person or the like in the imaging area from the installation location. The camera 200 is directly connected to the image processing device 100 or connected via a network or the like so that the captured image (video) can be output to the image processing device 100. The camera 200 may also be provided inside the image processing device 100.

[0022] The database 110 is a database that stores information (data) necessary for processing by the image processing device 100, processing results, etc. The database 110 stores images acquired by the image acquisition unit 101, detection results by the skeletal structure detection unit 102, data for machine learning, features calculated by the feature calculation unit 103, classification results by the classification unit 104, search results by the search unit 105, etc. The database 110 is directly connected to the image processing device 100 or connected via a network or the like so as to be able to input and output data as needed. The database 110 may be provided inside the image processing device 100 as a non-volatile memory such as a flash memory, a hard disk drive, etc.

[0023] As shown in FIG. 2, the image processing device 100 includes an image acquisition unit 101, a skeletal structure detection unit 102, a feature calculation unit 103, a classification unit 104, a search unit 105, an input unit 106, and a display unit 107. Note that the configuration of each unit (block) is an example, and the image processing device 100 may be configured with other units as long as the method (operation) described below is possible. The image processing device 100 is realized by, for example, a computer device such as a personal computer or a server that executes a program, but may also be realized by a single device or multiple devices on a network. For example, the input unit 106, the display unit 107, etc. may be external devices. The image processing device 100 may also include both the classification unit 104 and the search unit 105, or only one of them. Either or both of the classification unit 104 and the search unit 105 is a recognition unit that performs a recognition process for recognizing a person's state.

[0024] The image acquisition unit 101 acquires a two-dimensional image including a person captured by the camera 200. For example, the image acquisition unit 101 acquires an image including a person (a video including a plurality of images) captured by the camera 200 during a predetermined monitoring period. Note that the image acquisition unit 101 is not limited to acquisition from the camera 200, and may acquire images including a person prepared in advance from the database 110 or the like.

[0025] The skeletal structure detection unit 102 detects the two-dimensional skeletal structure of a person in an image based on the acquired two-dimensional image. The skeletal structure detection unit 102 detects the skeletal structure of all people recognized in the acquired image. The skeletal structure detection unit 102 uses a skeletal estimation technology using machine learning to detect the skeletal structure of a person based on the features of the recognized person's joints, etc. The skeletal structure detection unit 102 uses a skeletal estimation technology such as OpenPose described in Non-Patent Document 1, for example.

[0026] The feature calculation unit 103 calculates feature amounts of the detected two-dimensional skeletal structure and stores the calculated feature amounts in the database 110, linked to the image to be processed. The feature amounts of the skeletal structure indicate the characteristics of a person's skeleton and serve as elements for classifying and searching a person's condition based on the person's skeleton. Typically, this feature amount includes multiple parameters (e.g., classification elements, described later). The feature amount may be the feature amount of the entire skeletal structure, the feature amount of a part of the skeletal structure, or may include multiple feature amounts for each part of the skeletal structure. The feature amount may be calculated using any method, such as machine learning or normalization, and the minimum or maximum value may be calculated as normalization. For example, the feature amount may be the feature amount obtained by machine learning the skeletal structure or the size of the skeletal structure on the image from the head to the feet. The size of the skeletal structure may be the vertical height or area of ​​the skeletal region on the image that includes the skeletal structure. The vertical direction (height direction or vertical direction) is the up-down direction (Y-axis direction) in the image, for example, the direction perpendicular to the ground (reference plane). The left-right direction (horizontal direction) is the left-right direction (X-axis direction) in the image, and is, for example, a direction parallel to the ground.

[0027] In order to perform the classification and search desired by the user, it is preferable to use features that are robust to the classification and search process. For example, if the user desires classification and search that are not dependent on the orientation or body shape of a person, features that are robust to the orientation and body shape of a person may be used. By learning the skeletons of people facing in various directions in the same posture or the skeletons of people with various body shapes in the same posture, or by extracting features only in the up-down direction of the skeleton, it is possible to obtain features that are not dependent on the orientation or body shape of a person.

[0028] The classification unit 104 classifies (clusters) the multiple skeletal structures stored in the database 110 based on the similarity of the feature amounts of the skeletal structures. It can also be said that the classification unit 104 classifies the states of multiple people based on the feature amounts of the skeletal structures as a recognition process for the state of a person. The similarity is the distance between the feature amounts of the skeletal structures. The classification unit 104 may classify based on the similarity of the feature amounts of the entire skeletal structure, the similarity of the feature amounts of a part of the skeletal structure, or the similarity of the feature amounts of a first part (e.g., both hands) and a second part (e.g., both feet) of the skeletal structure. The posture of a person may be classified based on the feature amounts of the skeletal structure of the person in each image, or the behavior of a person may be classified based on changes in the feature amounts of the skeletal structure of the person in multiple images in a time series. That is, the classification unit 104 can classify the state of a person, including the posture and behavior of the person, based on the feature amounts of the skeletal structure. For example, the classification unit 104 classifies multiple skeletal structures in multiple images captured during a predetermined monitoring period. The classification unit 104 calculates the similarity between the feature amounts of the objects to be classified, and classifies them so that skeletal structures with high similarity are grouped into the same cluster (a group of similar postures). As with search, the classification conditions may be specified by the user. The classification unit 104 stores the results of the skeletal structure classification in the database 110 and displays them on the display unit 107.

[0029] The search unit 105 searches for a skeletal structure that is highly similar to the feature of the search query (query state) from among multiple skeletal structures stored in the database 110. It can also be said that the search unit 105, as a person's state recognition process, searches for a person's state that matches the search criteria (query state) from among multiple human states based on the feature of the skeletal structure. As with classification, the similarity is the distance between the feature of the skeletal structure. The search unit 105 may search based on the similarity of the feature of the entire skeletal structure, the similarity of the feature of a part of the skeletal structure, or the similarity of the feature of a first part (e.g., both hands) and a second part (e.g., both feet) of the skeletal structure. It is also possible to search for a person's posture based on the feature of the person's skeletal structure in each image, or to search for a person's behavior based on changes in the feature of the person's skeletal structure in multiple images that are consecutive in time series. That is, the search unit 105 can search for a person's state, including the person's posture and behavior, based on the feature of the skeletal structure. For example, the search unit 105, like the classification target, searches for multiple skeletal structure feature of multiple images captured during a predetermined monitoring period. Furthermore, a skeletal structure (posture) specified by the user from the classification results displayed by the classification unit 104 is used as a search query (search key). Note that the search query may be selected from multiple unclassified skeletal structures, not limited to the classification results, or the user may input a skeletal structure to be used as a search query. The search unit 105 searches for feature quantities that are highly similar to the feature quantities of the skeletal structure of the search query from among the feature quantities to be searched. The search unit 105 stores the search results of the feature quantities in the database 110 and displays them on the display unit 107.

[0030] The input unit 106 is an input interface that acquires information input by a user operating the image processing device 100. For example, the user is a monitor who monitors a person in a suspicious state from images captured by a surveillance camera. The input unit 106 is, for example, a GUI (Graphical User Interface), and receives information in response to user operations from an input device such as a keyboard, mouse, or touch panel. For example, the input unit 106 accepts, as a search query, the skeletal structure of a specified person from the skeletal structures (postures) classified by the classification unit 104.

[0031] The display unit 107 is a display unit that displays the results of the operation (processing) of the image processing device 100, and is, for example, a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display. The display unit 107 displays the classification results of the classification unit 104 and the search results of the search unit 105 on a GUI according to the degree of similarity, etc.

[0032] 39 is a diagram showing an example of the hardware configuration of the image processing device 100. The image processing device 100 includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, an input / output interface 1050, and a network interface 1060.

[0033] The bus 1010 is a data transmission path for transmitting and receiving data among the processor 1020, memory 1030, storage device 1040, input / output interface 1050, and network interface 1060. However, the method of connecting the processor 1020 and the like to each other is not limited to bus connection.

[0034] The processor 1020 is implemented by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.

[0035] The memory 1030 is a main storage device realized by a RAM (Random Access Memory) or the like.

[0036] The storage device 1040 is an auxiliary storage device realized by an HDD (Hard Disk Drive), an SSD (Solid State Drive), a memory card, a ROM (Read Only Memory), or the like. The storage device 1040 stores program modules that realize each function of the image processing device 100 (for example, the image acquisition unit 101, the skeletal structure detection unit 102, the feature calculation unit 103, the classification unit 104, the search unit 105, and the input unit 106). The processor 1020 loads each of these program modules into the memory 1030 and executes them, thereby realizing each function corresponding to the program module. The storage device 1040 may also function as the database 110.

[0037] The input / output interface 1050 is an interface for connecting the image processing device 100 to various input / output devices. When the database 110 is located outside the image processing device 100, the image processing device 100 may be connected to the database 110 via the input / output interface 1050.

[0038] The network interface 1060 is an interface for connecting the image processing device 100 to a network. This network is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The network interface 1060 may be connected to the network wirelessly or by wire. The image processing device 100 may communicate with the camera 200 via the network interface 1060. If the database 110 is located outside the image processing device 100, the image processing device 100 may be connected to the database 110 via the network interface 1060.

[0039] Figures 3 to 5 show the operation of image processing device 100 according to this embodiment. Figure 3 shows the flow from image acquisition to search processing in image processing device 100, Figure 4 shows the flow of classification processing (S104) in Figure 3, and Figure 5 shows the flow of search processing (S105) in Figure 3.

[0040] 3, the image processing device 100 acquires images from the camera 200 (S101). The image acquisition unit 101 acquires images of people in order to classify and search based on their skeletal structures, and stores the acquired images in the database 110. The image acquisition unit 101 acquires, for example, multiple images captured during a predetermined monitoring period, and performs the following processing on all people included in the multiple images.

[0041] Next, the image processing device 100 detects the skeletal structure of the person based on the acquired image of the person (S102). Fig. 6 shows an example of skeletal structure detection. As shown in Fig. 6, an image acquired from a surveillance camera or the like includes multiple people, and the skeletal structure is detected for each person included in the image.

[0042] Fig. 7 shows the skeletal structure of the human body model 300 detected at this time, and Figs. 8 to 10 show examples of detected skeletal structures. The skeletal structure detection unit 102 detects the skeletal structure of the human body model (two-dimensional skeletal model) 300 as shown in Fig. 7 from a two-dimensional image using a skeletal estimation technique such as OpenPose. The human body model 300 is a two-dimensional model made up of key points such as the joints of a person and bones connecting each key point.

[0043] The skeletal structure detection unit 102, for example, extracts feature points that can be key points from an image and detects each key point of a person by referring to information obtained by machine learning of the image of the key points. In the example of Fig. 7, the following key points of the person are detected: head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71, left knee A72, right foot A81, and left foot A82. Furthermore, the bones of the person connected by these key points are detected as bones: bone B1 connecting head A1 and neck A2; bone B21 and bone B22 connecting neck A2 to right shoulder A31 and left shoulder A32, respectively; bone B31 and bone B32 connecting right shoulder A31 and left shoulder A32 to right elbow A41 and left elbow A42, respectively; bone B41 and bone B42 connecting right elbow A41 and left elbow A42 to right hand A51 and left hand A52, respectively; bone B51 and bone B52 connecting neck A2 to right hip A61 and left hip A62, respectively; bone B61 and bone B62 connecting right hip A61 and left hip A62 to right knee A71 and left knee A72, respectively; and bone B71 and bone B72 connecting right knee A71 and left knee A72 to right foot A81 and left foot A82, respectively. The skeletal structure detection unit 102 stores the detected skeletal structure of the person in the database 110.

[0044] Fig. 8 shows an example of detecting a person standing upright. In Fig. 8, the person standing upright is imaged from the front, and bones B1, B51 and B52, B61 and B62, and B71 and B72 are detected without overlapping, and bones B61 and B71 of the right foot are slightly more bent than bones B62 and B72 of the left foot.

[0045] Fig. 9 shows an example of detecting a person who is crouching. In Fig. 9, the image of the person crouching is captured from the right side, and bones B1, B51 and B52, B61 and B62, and B71 and B72 are detected as seen from the right side, with bones B61 and B71 of the right foot and bones B62 and B72 of the left foot being significantly bent and overlapping.

[0046] Fig. 10 shows an example of detecting a person who is lying down. In Fig. 10, the person lying down is imaged from the diagonal front left, and bones B1, B51 and B52, B61 and B62, and B71 and B72 as seen from the diagonal front left are detected, with bones B61 and B71 of the right foot and bones B62 and B72 of the left foot being bent and overlapping.

[0047] Next, as shown in FIG. 3, the image processing device 100 calculates the feature amounts of the detected skeletal structure (S103). For example, if the height or area of ​​the skeletal region is used as the feature amount, the feature amount calculation unit 103 extracts a region including the skeletal structure and calculates the height (number of pixels) and area (pixel area) of that region. The height and area of ​​the skeletal region are calculated from the coordinates of the ends of the extracted skeletal region and the coordinates of the key points at the ends. The feature amount calculation unit 103 stores the calculated feature amounts of the skeletal structure in the database 110.

[0048] The skeletal structure feature is also used as posture information indicating the posture of a person. The skeletal structure feature is stored in the database 110 together with the reliability of the feature. The reliability of a feature indicates the likelihood that it is the feature (i.e., the probability that the posture estimation result is correct). If the reliability of a certain feature is high, the probability that it is the feature increases. The reliability of a feature is calculated using, for example, the number of key points used when setting the skeletal structure feature and the reliability of the key points. For example, the reliability of a feature increases as the number of joints increases. Furthermore, the reliability of a feature increases as the reliability of the key points increases.

[0049] In the example of Figure 8, a skeletal region including all bones is extracted from the skeletal structure of a person standing upright. In this case, the top end of the skeletal region is key point A1 on the head, the bottom end of the skeletal region is key point A82 on the left foot, the left end of the skeletal region is key point A41 on the right elbow, and the right end of the skeletal region is key point A52 on the left hand. Therefore, the height of the skeletal region is calculated from the difference in the Y coordinates of key point A1 and key point A82. The width of the skeletal region is calculated from the difference in the X coordinates of key point A41 and key point A52, and the area is calculated from the height and width of the skeletal region.

[0050] In the example of Figure 9, a skeletal region including all bones is extracted from the skeletal structure of a crouching person. In this case, the top end of the skeletal region is key point A1 on the head, the bottom end of the skeletal region is key point A81 on the right foot, the left end of the skeletal region is key point A61 on the right hip, and the right end of the skeletal region is key point A51 on the right hand. Therefore, the height of the skeletal region is calculated from the difference in the Y coordinates of key point A1 and key point A81. The width of the skeletal region is calculated from the difference in the X coordinates of key point A61 and key point A51, and the area is calculated from the height and width of the skeletal region.

[0051] In the example of Figure 10, a skeletal region including all bones is extracted from the skeletal structure of a person lying down in the horizontal direction of the image. In this case, the top end of the skeletal region is left shoulder key point A32, the bottom end of the skeletal region is left hand key point A52, the left end of the skeletal region is right hand key point A51, and the right end of the skeletal region is left foot key point A82. Therefore, the height of the skeletal region is calculated from the difference in the Y coordinates of key points A32 and A52. The width of the skeletal region is calculated from the difference in the X coordinates of key points A51 and A82, and the area is calculated from the height and width of the skeletal region.

[0052] Next, as shown in FIG. 3, the image processing device 100 performs a classification process (S104). In the classification process, as shown in FIG. 4, the classification unit 104 calculates the similarity of the calculated feature amounts of the skeletal structures (S111) and classifies the skeletal structures based on the calculated similarity (S112). The classification unit 104 calculates the similarity of the feature amounts between all skeletal structures stored in the database 110 to be classified, and classifies (clusters) the skeletal structures (postures) with the highest similarity into the same cluster. Furthermore, the classification unit 104 calculates the similarity between the classified clusters and repeats the classification until a predetermined number of clusters are obtained. FIG. 11 shows an image of the classification result of the feature amounts of the skeletal structures. FIG. 11 shows an image of cluster analysis using two-dimensional classification elements, and two classification elements are, for example, the height and area of ​​the skeletal region. In FIG. 11, as a result of classification, the feature amounts of multiple skeletal structures are classified into three clusters C1 to C3. Clusters C1 to C3 correspond to postures such as standing, sitting, and lying down, and skeletal structures (people) are classified into similar postures.

[0053] In this embodiment, various classification methods can be used by classifying based on the feature amounts of a person's skeletal structure. The classification method may be set in advance, or may be set arbitrarily by the user. Furthermore, classification may be performed using the same method as the search method described below. In other words, classification may be performed using classification conditions similar to the search conditions. For example, the classification unit 104 performs classification using the following classification methods. Any of the classification methods may be used, or arbitrarily selected classification methods may be combined.

[0054] (Classification method 1) Classification by multiple levels Classification is performed by hierarchically combining classification based on the skeletal structure of the entire body, classification based on the skeletal structure of the upper and lower body, classification based on the skeletal structure of the arms and legs, etc. In other words, classification may be performed based on the feature amounts of the first and second parts of the skeletal structure, and further classification may be performed by weighting the feature amounts of the first and second parts.

[0055] (Classification method 2) Classification using multiple images in chronological order Classification is based on the feature values ​​of the skeletal structure in multiple consecutive images in a time series. For example, feature values ​​may be accumulated in the time series direction and classification may be based on the cumulative value. Furthermore, classification may be based on the change (amount of change) in the feature values ​​of the skeletal structure in multiple consecutive images.

[0056] (Classification method 3) Classification ignoring the left and right sides of the skeletal structure The skeletal structures of the right and left sides of a person are classified as the same skeletal structure.

[0057] Furthermore, the classification unit 104 displays the skeletal structure classification results (S113). The classification unit 104 acquires the necessary skeletal structure and person images from the database 110, and displays the skeletal structures and people for each similar posture (cluster) on the display unit 107 as the classification results. FIG. 12 shows a display example when postures are classified into three. For example, as shown in FIG. 12, posture areas WA1 to WA3 for each posture are displayed in the display window W1, and the skeletal structures and people (images) for the postures corresponding to the posture areas WA1 to WA3 are displayed in the posture areas WA1 to WA3. The posture area WA1 is, for example, a display area for a standing posture, and displays the skeletal structures and people similar to the standing posture, which are classified into cluster C1. The posture area WA2 is, for example, a display area for a sitting posture, and displays the skeletal structures and people similar to the sitting posture, which are classified into cluster C2. The posture area WA3 is, for example, a display area for a lying posture, and displays the skeletal structures and people similar to the lying posture, which are classified into cluster C3.

[0058] Next, as shown in FIG. 3, the image processing device 100 performs a search process (S105). In the search process, as shown in FIG. 5, the search unit 105 accepts input of search conditions (S121) and searches for a skeletal structure based on the search conditions (S122). The search unit 105 accepts input of a search query, which is a search condition, from the input unit 106 in response to a user operation. When inputting a search query from the classification results, for example, in the display example of FIG. 12, the user specifies (selects) a skeletal structure of a posture to be searched from posture areas WA1 to WA3 displayed in the display window W1. Then, the search unit 105 uses the skeletal structure specified by the user as a search query to search for a skeletal structure having a high similarity in feature amount from among all skeletal structures stored in the database 110 that are the search target. The search unit 105 calculates the similarity between the feature amount of the skeletal structure of the search query and the feature amount of the skeletal structure of the search target, and extracts skeletal structures for which the calculated similarity is higher than a predetermined threshold. The feature quantities of the skeletal structure of the search query may be pre-calculated feature quantities or feature quantities calculated at the time of the search. Note that the search query may be input by moving each part of the skeletal structure in response to a user's operation, or the posture demonstrated by the user in front of the camera may be used as the search query.

[0059] In this embodiment, similarly to the classification method, a variety of search methods can be used by searching based on the feature amounts of a person's skeletal structure. The search method may be set in advance or may be set arbitrarily by the user. For example, the search unit 105 performs a search using the following search methods. Any of the search methods may be used, or arbitrarily selected search methods may be combined. A search may be performed by combining multiple search methods (search conditions) using a logical expression (for example, AND (logical product), OR (logical sum), NOT (negation)). For example, a search may be performed using the search conditions "(posture with right hand raised) AND (posture with left leg raised)".

[0060] (Search method 1) Search using only height features By performing a search using only the height feature of a person, the influence of lateral changes in the person can be suppressed, improving robustness against changes in the person's orientation and body shape. For example, as in skeletal structures 501 to 503 in Fig. 13, even if the person's orientation or body shape is different, the height feature does not change significantly. Therefore, skeletal structures 501 to 503 can be determined to have the same posture during search (classification).

[0061] (Search Method 2) Partial Search: When a part of a person's body is hidden in an image, a search is performed using only information about the recognizable part. For example, as in the case of skeletal structures 511 and 512 in FIG. 14, even if the left foot keypoint cannot be detected because the left foot is hidden, a search can be performed using the feature values ​​of other detected keypoints. Therefore, for skeletal structures 511 and 512, it is possible to determine that the pose is the same during search (classification). In other words, classification and search can be performed using the feature values ​​of only some of the keypoints, rather than all of the keypoints. In the example of skeletal structures 521 and 522 in FIG. 15, although the orientations of the feet are different, it is possible to determine that the pose is the same by using the feature values ​​of the upper body keypoints (A1, A2, A31, A32, A41, A42, A51, A52) as the search query. Furthermore, a search can be performed by weighting the part (feature point) to be searched, or by changing the threshold for similarity determination. When a part of the body is hidden, the search can be performed by ignoring the hidden part, or by taking the hidden part into account. By including hidden parts in the search, it is possible to search for postures in which the same body part is hidden.

[0062] (Search method 3) Search ignoring the left and right sides of the skeletal structure Skeletal structures on opposite sides of a person's right and left sides are searched for as the same skeletal structure. For example, as in skeletal structures 531 and 532 in Figure 16, a posture with the right hand raised and a posture with the left hand raised can be searched for (classified) as the same posture. In the example of Figure 16, skeletal structures 531 and 532 differ in the positions of right hand key point A51, right elbow key point A41, left hand key point A52, and left elbow key point A42, but the positions of the other key points are the same. Of the key points A51 on the right hand and A41 on the right elbow of skeletal structure 531 and the key points A52 on the left hand and A42 on the left elbow of skeletal structure 532, if the key points of one of the skeletal structures are flipped left to right, they will be in the same position as the key points of the other skeletal structure.Furthermore, of the key points A52 on the left hand and A42 on the left elbow of skeletal structure 531 and the key points A51 on the right hand and A41 on the right elbow of skeletal structure 532, if the key points of one of the skeletal structures are flipped left to right, they will be in the same position as the key points of the other skeletal structure, and therefore they are determined to be the same posture.

[0063] (Search Method 4) Search by vertical and horizontal features After performing a search using only the feature values ​​in the vertical direction (Y-axis direction) of the person, the obtained results are further searched using the feature values ​​in the horizontal direction (X-axis direction) of the person.

[0064] (Search Method 5) Searching multiple images in chronological order Searches are performed based on the feature values ​​of the skeletal structure in multiple consecutive images in a time series. For example, feature values ​​may be accumulated in the time series direction and searches may be performed based on the cumulative value. Furthermore, searches may be performed based on the change (amount of change) in the feature values ​​of the skeletal structure in multiple consecutive images.

[0065] Furthermore, the search unit 105 displays the search results for the skeletal structures (S123). The search unit 105 acquires the necessary images of skeletal structures and people from the database 110, and displays the skeletal structures and people obtained as search results on the display unit 107. For example, if multiple search queries (search conditions) are specified, search results are displayed for each search query. FIG. 17 shows a display example when a search is performed using three search queries (postures). For example, as shown in FIG. 17, in the display window W2, the skeletal structures and people for the specified search queries Q10, Q20, and Q30 are displayed on the left edge, and the skeletal structures and people for the search results Q11, Q21, and Q31 for each search query are displayed side by side to the right of the search queries Q10, Q20, and Q30.

[0066] The order in which the search results are displayed next to the search query may be the order in which the corresponding skeletal structures were found, or may be in descending order of similarity. When a partial search is performed with weights assigned to the parts (feature points), the results may be displayed in order of similarity calculated using the weights. The results may also be displayed in order of similarity calculated only from the parts (feature points) selected by the user. Furthermore, the images (frames) before and after the search result image (frame) may be displayed by extracting a certain time period from the image (frame) in the chronological order.

[0067] (Search Method 6) In this search method, the searched images are set using the reliability of the posture information of the images.

[0068] 40 is a diagram showing an example of the functional configuration of the search unit 105 related to this search method. In the example shown in this diagram, the search unit 105 includes a query acquisition unit 610, a search information acquisition unit 620, a selection unit 630, and a display control unit 640.

[0069] The query acquisition unit 610 acquires posture information indicating the posture of a person (hereinafter referred to as query posture information), which is information that serves as a query.

[0070] The search information acquisition unit 620 acquires a plurality of pieces of search orientation information. The search orientation information is orientation information of images to be searched (hereinafter referred to as target images), and is stored in the database 110 for each of the plurality of target images.

[0071] The selection unit 630 selects two or more pieces of search posture information whose similarity to the query posture information satisfies a criterion from among the plurality of pieces of search posture information. This selection process is essentially a process of selecting a target image similar to an image corresponding to the query posture information (hereinafter referred to as a query image).

[0072] The display control unit 640 causes the display unit 107 to display target images corresponding to the two or more pieces of search orientation information selected by the selection unit 630, and sets the display positions of these target images using the reliability of the search orientation information. This reliability is stored in the database 110 as described above.

[0073] 41 is a flowchart showing an example of processing performed by the search unit 105 according to this search method. First, the query acquisition unit 610 acquires query posture information (step S300). As an example, the query acquisition unit 610 acquires the query posture information from the database 110. In this case, for example, a user provides an input to the query acquisition unit 610 to identify a query image. Then, the query acquisition unit 610 reads out query posture information corresponding to this query image from the database 110.

[0074] The query acquisition unit 610 may also acquire an image serving as a query from an external device. In this case, the skeletal structure detection unit 102 and the feature calculation unit 103 process the image to calculate the feature of the skeletal structure. The query acquisition unit 610 then acquires the feature of the skeletal structure as query posture information.

[0075] Next, the search information acquisition unit 620 acquires a plurality of pieces of posture information stored in the database 110 as search information (step S310). Next, the selection unit 630 selects two or more pieces of search posture information whose similarity to the query posture information satisfies a criterion from the plurality of pieces of search posture information (step S320). The process performed here is performed using, for example, at least one of the above-described search methods 1 to 5, but may also be performed using other methods.

[0076] Next, the display control unit 640 reads out a target image corresponding to the search posture information selected in step S320 from the database 110. Because two or more search posture information are selected in step S320, the display control unit 640 also reads out two or more target images from the database 110. At this time, the display control unit 640 also reads out the reliability of each target image from the database 110 (step S330).

[0077] Next, the display control unit 640 sets the display position of each target image using the reliability of that target image (step S340). As described above, the reliability indicates the probability that the pose estimation result is correct. Therefore, if the reliability is high, the target image is more likely to be similar to the query image. Therefore, setting the display position of the target image using the reliability of that target image allows the user to recognize the impact that the reliability of each target image has on the validity of the selection result in step S320.

[0078] Then, display control unit 640 displays the target images read out in step S330 on display unit 107. At this time, display control unit 640 displays each target image at the display position determined in step S340 (step S350).

[0079] The selection unit 630 also stores the selection result of step S320 in the database 110 (step S360). Here, the selection unit 630 may store the target images read out in step S330 in the database 110 as one cluster, or may associate a target image already stored in the database 110 with a flag indicating that the target image is similar to the query image.

[0080] 42 is a diagram showing an example of a method for determining the display position of a target image, which is performed in step S340. In the example shown in this figure, the display control unit 640 arranges the target images in descending order of reliability and displays them on the display unit 107. However, the arrangement order of the target images is not limited to this example.

[0081] Furthermore, the display control unit 640 displays the target image together with the reliability of the retrieval posture information of the target image on the display unit 107. In this way, the user can directly check the reliability of the selected target image, and can therefore understand the reliability required to obtain a reasonable selection result.

[0082] Fig. 43 is a flowchart showing Modification 1 of Fig. 41. The example shown in this figure is the same as Fig. 41 except that, when determining the display position of the target image, the display control unit 640 uses the similarity between the target image and the query image, i.e., the similarity between the search posture information and the query posture information, in addition to the reliability of the target image (step S342).

[0083] As an example, the display control unit 640 determines the sorting order of the target images using the reliability and similarity described above. For example, the display control unit 640 calculates an evaluation value using a function with the reliability and similarity described above as parameters, and displays the target images in descending (or descending) order of this evaluation value. An example of this function is, for example, a function that adds a predetermined weighting coefficient to each of the reliability and similarity and then calculates the sum of these.

[0084] Fig. 44 is a flowchart showing Modification 2 of Fig. 41. The example shown in this figure is the same as Fig. 41 except that when the display control unit 640 determines the display position of the target image on the condition that a predetermined condition is satisfied (step S332: Yes), it also uses the similarity between the target image and the query image in addition to or instead of the reliability of the target image (step S344).

[0085] The process when the similarity between the target image and the query image is used in addition to the reliability of the target image is the same as step S342 in Fig. 43. Also, the process when the similarity between the target image and the query image is used instead of the reliability of the target image is the same as step S340 in Fig. 41, except that the similarity is used.

[0086] As described above, this embodiment makes it possible to detect a person's skeletal structure from a two-dimensional image and perform classification and search based on the feature quantities of the detected skeletal structure. This allows classification into similar poses with high similarity, and also makes it possible to search for similar poses with high similarity to a search query (search key). By classifying and displaying similar poses from an image, the pose of a person in the image can be understood without the user having to specify a pose, etc. The user can specify a search query pose from the classification results, so that a desired pose can be searched for even if the user does not have a detailed understanding of the pose they want to search for in advance. For example, classification and search can be performed using the entire or part of a person's skeletal structure as a condition, enabling flexible classification and search.

[0087] Furthermore, according to search method 6, when at least two images (target images) are selected, the display positions of the target images are set using the reliability of the target images, which allows the user to recognize the influence that the reliability of each target image has on the validity of the selection result.

[0088] (Embodiment 2) Hereinafter, embodiment 2 will be described with reference to the drawings. In this embodiment, a specific example of feature calculation in embodiment 1 will be described. In this embodiment, feature amounts are calculated by normalizing using the person's height. Other aspects are the same as embodiment 1.

[0089] Fig. 18 shows the configuration of an image processing device 100 according to this embodiment. As shown in Fig. 18, the image processing device 100 further includes a height calculation unit 108 in addition to the configuration of embodiment 1. Note that the feature calculation unit 103 and the height calculation unit 108 may be integrated into one processing unit.

[0090] The height calculation unit (height estimation unit) 108 calculates (estimates) the height (referred to as height pixel count) of a person in a two-dimensional image when standing upright, based on the two-dimensional skeletal structure detected by the skeletal structure detection unit 102. The height pixel count can also be said to be the height of a person in a two-dimensional image (the length of the person's entire body in two-dimensional image space). The height calculation unit 108 obtains the height pixel count (number of pixels) from the length of each bone of the detected skeletal structure (length in two-dimensional image space).

[0091] In the following examples, specific examples 1 to 3 are used as methods for calculating the height pixel count. Note that any of the methods from specific examples 1 to 3 may be used, or a combination of multiple arbitrarily selected methods may be used. In specific example 1, the height pixel count is calculated by adding up the lengths of the bones from the head to the feet among the bones of the skeletal structure. If the skeletal structure detection unit 102 (skeleton estimation technology) does not output the top of the head and the feet, correction can be made by multiplying a constant as necessary. In specific example 2, the height pixel count is calculated using a human body model that indicates the relationship between the length of each bone and the length of the entire body (height in two-dimensional image space). In specific example 3, the height pixel count is calculated by fitting a three-dimensional human body model to the two-dimensional skeletal structure.

[0092] The feature amount calculation unit 103 in this embodiment is a normalization unit that normalizes the skeletal structure (skeletal information) of a person based on the calculated number of pixels of the person's height. The feature amount calculation unit 103 stores the feature amount (normalized value) of the normalized skeletal structure in the database 110. The feature amount calculation unit 103 normalizes the height of each key point (feature point) included in the skeletal structure on the image by the number of pixels of the height. In this embodiment, for example, the height direction is the up-down direction (Y-axis direction) in a two-dimensional coordinate (XY coordinate) space of the image. In this case, the height of the key point can be obtained from the Y coordinate value (number of pixels) of the key point. Alternatively, the height direction may be the direction of the vertical projection axis (vertical projection direction) obtained by projecting the direction of a vertical axis perpendicular to the ground (reference plane) in a three-dimensional coordinate space of the real world onto the two-dimensional coordinate space. In this case, the height of the keypoint can be determined from the value (number of pixels) along the vertical projection axis, which is obtained by projecting an axis perpendicular to the ground in the real world onto a two-dimensional coordinate space based on the camera parameters. The camera parameters are image capture parameters, such as the attitude, position, imaging angle, and focal length of the camera 200. An object whose length and position are known in advance is captured using the camera 200, and the camera parameters can be determined from the captured image. Distortion may occur at both ends of the captured image, causing the vertical direction in the real world and the up-down direction in the image to not match. In response to this, the parameters of the camera that captured the image can be used to determine the degree to which the vertical direction in the real world is tilted in the image. Therefore, by normalizing the value of the keypoint along the vertical projection axis projected onto the image based on the camera parameters by height, the keypoint can be converted into a feature quantity that takes into account the deviation between the real world and the image. The left-right direction (horizontal direction) refers to the left-right direction (X-axis direction) in the two-dimensional coordinate (XY coordinate) space of the image, or the direction parallel to the ground in the three-dimensional coordinate space of the real world projected onto the two-dimensional coordinate space.

[0093] Figures 19 to 23 show the operation of the image processing device 100 according to this embodiment. Figure 19 shows the flow from image acquisition to search processing in the image processing device 100, Figures 20 to 22 show the flow of specific examples 1 to 3 of the height pixel number calculation processing (S201) in Figure 19, and Figure 23 shows the flow of the normalization processing (S202) in Figure 19.

[0094] 19, in this embodiment, height pixel number calculation processing (S201) and normalization processing (S202) are performed as the feature amount calculation processing (S103) in Embodiment 1. The rest is the same as in Embodiment 1.

[0095] Following image acquisition (S101) and skeletal structure detection (S102), the image processing device 100 performs height pixel number calculation processing based on the detected skeletal structure (S201). In this example, as shown in Fig. 24, the height of the skeletal structure of a person standing upright in the image is defined as height pixel number (h), and the height of each key point of the skeletal structure in the state of the person in the image is defined as key point height (yi). Specific examples 1 to 3 of height pixel number calculation processing will be described below.

[0096] <Specific Example 1> In specific example 1, the height pixel number is calculated using the lengths of the bones from the head to the feet. In specific example 1, as shown in Fig. 20, the height calculation unit 108 acquires the length of each bone (S211) and sums up the acquired lengths of each bone (S212).

[0097] The height calculation unit 108 obtains the lengths of the bones in the two-dimensional image from the head to the feet of the person and calculates the height pixel count. That is, from the image in which the skeletal structure is detected, the length (number of pixels) of bone B1 (length L1), bone B51 (length L21), bone B61 (length L31), and bone B71 (length L41) or bone B1 (length L1), bone B52 (length L22), bone B62 (length L32), and bone B72 (length L42) among the bones in FIG. 24 is obtained. The length of each bone can be obtained from the coordinates of each key point in the two-dimensional image. The sum of these lengths, L1 + L21 + L31 + L41 or L1 + L22 + L32 + L42, is multiplied by a correction constant to calculate the height pixel count (h). If both values ​​can be calculated, for example, the longer value is used as the height pixel count. That is, each bone will be the longest in the image when photographed from the front, and will appear shorter when tilted in the depth direction relative to the camera. Therefore, longer bones are more likely to have been photographed from the front, and are considered closer to the true value. For this reason, it is preferable to select the longer value.

[0098] In the example of Figure 25, bone B1, bones B51 and B52, bones B61 and B62, and bones B71 and B72 are detected without overlapping. The totals of these bones, L1+L21+L31+L41 and L1+L22+L32+L42, are calculated, and the height pixel count is determined by multiplying the detected bone, for example, L1+L22+L32+L42 on the left leg side, which has the longer length, by a correction constant.

[0099] 26, bone B1, bones B51 and B52, bones B61 and B62, and bones B71 and B72 are detected, and bones B61 and B71 of the right foot overlap with bones B62 and B72 of the left foot. The totals of these bones, L1+L21+L31+L41 and L1+L22+L32+L42, are calculated, and the height pixel count is determined by multiplying the detected bone L1+L21+L31+L41 on the right foot, which has the longer length, by a correction constant.

[0100] 27, bone B1, bones B51 and B52, bones B61 and B62, and bones B71 and B72 are detected, and bones B61 and B71 of the right foot overlap with bones B62 and B72 of the left foot. The totals of these bones, L1+L21+L31+L41 and L1+L22+L32+L42, are calculated, and the height pixel count is calculated by multiplying the detected bone L1+L22+L32+L42 on the left foot, which has the longer length, by a correction constant.

[0101] In Example 1, height can be calculated by adding up the lengths of the bones from head to toe, so the number of pixels of height can be calculated in a simple manner. Furthermore, because skeleton estimation technology using machine learning only requires that the skeleton from head to toe be detected, the number of pixels of height can be estimated with high accuracy even when the entire person is not necessarily captured in the image, such as when the person is crouching.

[0102] <Specific Example 2> In specific example 2, the height pixel number is calculated using a two-dimensional skeleton model that shows the relationship between the length of the bones included in the two-dimensional skeleton structure and the length of the entire body of a person in two-dimensional image space.

[0103] FIG. 28 shows a human body model (two-dimensional skeletal model) 301 used in Example 2, which shows the relationship between the length of each bone in a two-dimensional image space and the length of the entire body in the two-dimensional image space. As shown in FIG. 28, the relationship between the length of each bone in an average person and the length of the entire body (the ratio of each bone's length to the length of the entire body) is associated with each bone in the human body model 301. For example, the length of the head bone B1 is the length of the entire body × 0.2 (20%), the length of the right hand bone B41 is the length of the entire body × 0.15 (15%), and the length of the right foot bone B71 is the length of the entire body × 0.25 (25%). By storing information about this human body model 301 in the database 110, the average entire body length can be calculated from the length of each bone. In addition to the human body model of an average person, human body models may be prepared for each person's attributes, such as age, gender, and nationality. This allows the entire body length (height) to be calculated appropriately according to the person's attributes.

[0104] In specific example 2, as shown in FIG. 21, height calculation unit 108 acquires the length of each bone (S221). Height calculation unit 108 acquires the lengths of all bones (lengths in two-dimensional image space) in the detected skeletal structure. FIG. 29 shows an example in which a person in a crouching position is imaged from diagonally behind to the right and the skeletal structure is detected. In this example, the person's face and left side are not captured, so the head bone and the bones of the left arm and left hand cannot be detected. Therefore, the length of each of the detected bones B21, B22, B31, B41, B51, B52, B61, B62, B71, and B72 is acquired.

[0105] Next, the height calculation unit 108 calculates the height pixel number from the length of each bone based on the human body model, as shown in Fig. 21 (S222). The height calculation unit 108 references a human body model 301, as shown in Fig. 28, which shows the relationship between each bone and the length of the entire body, and calculates the height pixel number from the length of each bone. For example, since the length of the right hand bone B41 is the length of the entire body × 0.15, the height pixel number based on bone B41 is calculated by dividing the length of bone B41 by 0.15. Furthermore, since the length of the right foot bone B71 is the length of the entire body × 0.25, the height pixel number based on bone B71 is calculated by dividing the length of bone B71 by 0.25.

[0106] The human body model referenced at this time is, for example, a human body model of an average person, but a human body model may be selected depending on the person's attributes such as age, gender, and nationality. For example, if a person's face is captured in a captured image, the person's attributes are identified based on the face, and a human body model corresponding to the identified attributes is referenced. By referencing information obtained by machine learning on faces for each attribute, the person's attributes can be recognized from the facial features in the image. Furthermore, if the person's attributes cannot be identified from the image, a human body model of an average person may be used.

[0107] Furthermore, the height pixel count calculated from the bone lengths may be corrected using camera parameters. For example, if a camera is placed in a high position and a person is photographed looking down on the camera, the horizontal length of bones such as shoulder width in the two-dimensional skeletal structure is not affected by the camera's inclination angle, but the vertical length of bones such as the neck-waist bones becomes smaller as the camera's inclination angle increases. As a result, the height pixel count calculated from the horizontal length of bones such as shoulder width tends to be larger than the actual height. Therefore, by utilizing camera parameters, the angle at which the camera is looking down on the person can be determined, and this inclination angle information can be used to correct the two-dimensional skeletal structure to look as if the person were photographed from the front. This allows for a more accurate calculation of the height pixel count.

[0108] Next, the height calculation unit 108 calculates the optimal height pixel count, as shown in FIG. 21 (S223). The height calculation unit 108 calculates the optimal height pixel count from the height pixel count calculated for each bone. For example, a histogram of the height pixel count calculated for each bone, as shown in FIG. 30, is generated, and the largest height pixel count is selected. That is, a height pixel count longer than the others is selected from the multiple height pixel counts calculated based on multiple bones. For example, the top 30% are considered valid values, and in FIG. 30, the height pixel counts for bones B71, B61, and B51 are selected. The optimal value may be the average of the selected height pixel counts, or the largest height pixel count. Since height is calculated from the length of the bones in the 2D image, if the bones are not captured from the front, i.e., if the bones are captured at an angle in the depth direction when viewed from the camera, the bone lengths will be shorter than when captured from the front. As a result, a larger height pixel count is more likely to have been captured from the front than a smaller height pixel count, and is therefore a more plausible value. Therefore, a larger height pixel count is set as the optimal value.

[0109] In Example 2, a human body model showing the relationship between bones in a two-dimensional image space and the length of the entire body is used to calculate the height pixel count based on the detected bones of the skeletal structure, so even if the entire skeleton from head to toe cannot be obtained, the height pixel count can be calculated from some of the bones. In particular, by adopting the larger value among the values ​​calculated from multiple bones, the height pixel count can be estimated with high accuracy.

[0110] <Specific Example 3> In specific example 3, a two-dimensional skeletal structure is fitted to a three-dimensional human body model (three-dimensional skeletal model), and the skeletal vectors of the whole body are obtained using the height pixel count of the fitted three-dimensional human body model.

[0111] 22, in specific example 3, height calculation unit 108 first calculates camera parameters based on images captured by camera 200 (S231). Height calculation unit 108 extracts an object whose length is known in advance from multiple images captured by camera 200, and calculates camera parameters from the size (number of pixels) of the extracted object. Note that camera parameters may be calculated in advance and acquired as needed.

[0112] Next, the height calculation unit 108 adjusts the position and height of the 3D human body model (S232). The height calculation unit 108 prepares a 3D human body model for calculating the number of height pixels for the detected 2D skeletal structure, and places it in the same 2D image based on the camera parameters. Specifically, the "relative positional relationship between the camera and the person in the real world" is identified from the camera parameters and the 2D skeletal structure. For example, if the camera position is assumed to be coordinates (0,0,0), the coordinates (x,y,z) of the position where the person is standing (or sitting) are identified. Then, by imagining an image that would be captured if the 3D human body model were placed at the same position (x,y,z) as the identified person, the 2D skeletal structure and the 3D human body model are superimposed.

[0113] FIG. 31 shows an example in which a crouching person is imaged from the diagonal front left and a two-dimensional skeletal structure 401 is detected. The two-dimensional skeletal structure 401 has two-dimensional coordinate information. It is preferable to detect all bones, but it is also possible that some bones are not detected. A three-dimensional human body model 402, as shown in FIG. 32, is prepared for this two-dimensional skeletal structure 401. The three-dimensional human body model (three-dimensional skeletal model) 402 has three-dimensional coordinate information and is a skeletal model with the same shape as the two-dimensional skeletal structure 401. Then, as shown in FIG. 33, the prepared three-dimensional human body model 402 is positioned and superimposed on the detected two-dimensional skeletal structure 401. Furthermore, while the two are superimposed, the height of the three-dimensional human body model 402 is adjusted to match the two-dimensional skeletal structure 401.

[0114] The 3D human body model 402 prepared at this time may be a model in a state close to the posture of the 2D skeletal structure 401, as shown in FIG. 33, or may be a model in an upright position. For example, a 3D human body model 402 in an estimated posture may be generated using a technology that uses machine learning to estimate a posture in 3D space from a 2D image. By learning information about the joints in the 2D image and the joints in 3D space, it is possible to estimate a 3D posture from a 2D image.

[0115] Next, the height calculation unit 108 fits the three-dimensional human body model to the two-dimensional skeletal structure (S233), as shown in FIG. 22. With the three-dimensional human body model 402 superimposed on the two-dimensional skeletal structure 401, as shown in FIG. 34, the height calculation unit 108 deforms the three-dimensional human body model 402 so that the postures of the three-dimensional human body model 402 and the two-dimensional skeletal structure 401 match. That is, the height, body orientation, and joint angles of the three-dimensional human body model 402 are adjusted and optimized to eliminate any difference with the two-dimensional skeletal structure 401. For example, the joints of the three-dimensional human body model 402 are rotated within the range of human movement, and the entire three-dimensional human body model 402 is also rotated and its overall size adjusted. The fitting of the three-dimensional human body model to the two-dimensional skeletal structure is performed in two-dimensional space (two-dimensional coordinates). In other words, a 3D human body model is mapped onto a 2D space, and the 3D human body model is optimized to a 2D skeletal structure, taking into consideration how the deformed 3D human body model changes in the 2D space (image).

[0116] Next, height calculation unit 108 calculates the height pixel count of the fitted three-dimensional human body model, as shown in FIG. 22 (S234). When the difference between three-dimensional human body model 402 and two-dimensional skeletal structure 401 disappears and the postures match, as shown in FIG. 35, height calculation unit 108 calculates the height pixel count of three-dimensional human body model 402 in that state. With optimized three-dimensional human body model 402 standing upright, the length of the entire body in two-dimensional space is calculated based on the camera parameters. For example, the height pixel count is calculated from the bone lengths (number of pixels) from the head to the feet when three-dimensional human body model 402 is standing upright. As in specific example 1, the bone lengths from the head to the feet of three-dimensional human body model 402 may be summed up.

[0117] In specific example 3, a 3D human body model is fitted to a 2D skeletal structure based on camera parameters, and the height pixel count is calculated based on the 3D human body model. This makes it possible to accurately estimate the height pixel count even when not all bones are viewed from the front, i.e., even when all bones are viewed at an angle, resulting in a large error.

[0118] <Normalization Process> As shown in FIG. 19, the image processing device 100 performs a normalization process (S202) following the height pixel count calculation process. In the normalization process, as shown in FIG. 23, the feature calculation unit 103 calculates keypoint heights (number of pixels) for all keypoints included in the detected skeletal structure (S241). The feature calculation unit 103 calculates the keypoint heights (number of pixels) for all keypoints included in the detected skeletal structure. The keypoint height is the height (number of pixels) from the lowest end of the skeletal structure (e.g., a keypoint on one of the feet) to that keypoint. Here, as an example, the keypoint height is calculated from the Y coordinate of the keypoint in the image. Note that, as described above, the keypoint height may also be calculated from the length along the vertical projection axis based on the camera parameters. For example, in the example of FIG. 24, the height (yi) of the neck keypoint A2 is the Y coordinate of the keypoint A2 minus the Y coordinate of the right foot keypoint A81 or the left foot keypoint A82.

[0119] Next, the feature calculation unit 103 identifies a reference point for normalization (S242). The reference point is a point that serves as a reference for expressing the relative height of the key point. The reference point may be set in advance or may be selectable by the user. The reference point is preferably the center of the skeletal structure or higher than the center (upper in the vertical direction of the image), and for example, the coordinates of the key point of the neck are used as the reference point. Note that the reference point is not limited to the neck, and the coordinates of the head or other key points may also be used. The reference point is not limited to a key point, and any coordinates (for example, the center coordinates of the skeletal structure) may also be used.

[0120] Next, the feature calculation unit 103 normalizes the keypoint height (yi) by the number of height pixels (S243). The feature calculation unit 103 normalizes each keypoint using the keypoint height, reference point, and height pixel count of each keypoint. Specifically, the feature calculation unit 103 normalizes the relative height of the keypoint with respect to the reference point by the number of height pixels. Here, as an example focusing only on the height direction, only the Y coordinate is extracted, and normalization is performed using the reference point as the neck keypoint. Specifically, the Y coordinate of the reference point (neck keypoint) is set as (yc), and the feature (normalized value) is calculated using the following equation (1). Note that when a vertical projection axis based on the camera parameters is used, (yi) and (yc) are converted into values ​​in the direction along the vertical projection axis.

number

[0121] For example, if there are 18 keypoints, the coordinates of the 18 keypoints (x0, y0), (x1, y1), ... (x17, y17) are converted into 18-dimensional features using the above formula (1) as follows:

number

[0122] FIG. 36 shows an example of the feature amounts of each key point calculated by the feature amount calculation unit 103. In this example, since the neck key point A2 is used as the reference point, the feature amount of key point A2 is 0.0, and the feature amounts of right shoulder key point A31 and left shoulder key point A32, which are at the same height as the neck, are also 0.0. The feature amount of head key point A1, which is higher than the neck, is -0.2. The feature amounts of right hand key point A51 and left hand key point A52, which are lower than the neck, are 0.4, and the feature amounts of right foot key point A81 and left foot key point A82 are 0.9. If the person raises their left hand from this state, the left hand will be higher than the reference point as shown in FIG. 37, and the feature amount of left hand key point A52 will be -0.4. However, since normalization is performed using only the Y-axis coordinate, the feature amount does not change even if the width of the skeletal structure changes, as shown in FIG. 38. That is, the feature amount (normalized value) of this embodiment indicates the feature in the height direction (Y direction) of the skeletal structure (key point), and is not affected by changes in the lateral direction (X direction) of the skeletal structure.

[0123] As described above, in this embodiment, a person's skeletal structure is detected from a two-dimensional image, and each key point of the skeletal structure is normalized using the height pixel count (height when standing upright in two-dimensional image space) calculated from the detected skeletal structure. Using this normalized feature can improve robustness when performing classification, search, etc. In other words, the feature of this embodiment is not affected by lateral changes in the person as described above, and is therefore highly robust to changes in the person's orientation and body shape.

[0124] Furthermore, in this embodiment, since this can be achieved by detecting a person's skeletal structure using a skeletal estimation technology such as OpenPose, there is no need to prepare training data for learning a person's posture, etc. Also, by normalizing the key points of the skeletal structure and storing them in a database, it becomes possible to classify and search for a person's posture, etc., so that classification and search can be performed even for unknown postures. Furthermore, normalizing the key points of the skeletal structure makes it possible to obtain clear and easy-to-understand features, which, unlike black-box algorithms such as machine learning, makes it highly likely that users will be satisfied with the processing results.

[0125] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations can also be adopted.

[0126] In addition, in the flowcharts used in the above description, multiple steps (processes) are described in order, but the order of execution of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the drawings can be changed to the extent that the content is not affected. Furthermore, the above-mentioned embodiments can be combined to the extent that the content is not contradictory.

[0127] Some or all of the above-described embodiments can be described as, but are not limited to, the following supplementary notes. 1. A query acquisition means for acquiring query posture information indicating a posture of a person, which is information to be used as a query; a search information acquisition means for acquiring a plurality of pieces of search posture information that are generated for each of a plurality of target images to be searched and indicate the postures of people included in the target images; a selection means for selecting two or more pieces of search orientation information from the plurality of pieces of search orientation information, the two or more pieces of search orientation information having similarity to the query orientation information that satisfies a criterion; a display control means for displaying the target images corresponding to the two or more pieces of orientation information for search selected by the selection means on a display means, and setting the display positions of the target images using the reliability of the orientation information for search; An image selection device comprising: 2. In the image selection device described in 1 above, The display control means displays the reliability of the orientation information for search on the display means together with the target image corresponding to the orientation information for search. 3. In the image selection device according to 1 or 2 above, The display control means further determines a display position of the target image using the similarity. 4. In the image selection device according to 1 or 2 above, The display control means is an image selection device that uses the similarity in place of or in addition to the search orientation information when determining the display position of the target image when a predetermined condition is met. 5. In the image selection device according to any one of the above items 1 to 4, When the query acquisition means acquires information specifying a query image that is an image including a person, the query acquisition means acquires posture information corresponding to the query image as the query posture information. 6. In the image selection device according to any one of the above items 1 to 5, the search posture information is set using positions of a plurality of joints, The reliability is calculated using the number of joints used when setting the search posture information and the reliability of the joints. 7. The computer a query acquisition process for acquiring query posture information indicating a posture of a person, the query information being information to be used as a query; a search information acquisition process for acquiring a plurality of pieces of search posture information that are information generated for each of a plurality of target images to be searched, and that indicate postures of people included in the target images; a selection process of selecting two or more pieces of search orientation information from the plurality of pieces of search orientation information, the two or more pieces of search orientation information having similarity to the query orientation information that satisfies a criterion; a display control process for displaying the target images corresponding to the two or more selected pieces of orientation information for search on a display means, and setting display positions of the target images using the reliability of the orientation information for search; Image selection method. 8. In the image selection method described in 7 above, In the display control process, the computer displays the reliability of the orientation information for search on a display means together with the target image corresponding to the orientation information for search. 9. In the image selection method described in 7 or 8 above, In the display control process, the computer further determines a display position of the target image using the similarity. 10. In the image selection method described in 7 or 8 above, In the display control process, when a predetermined condition is met, the computer uses the similarity in place of or in addition to the search posture information when determining the display position of the target image. 11. In the image selection method according to any one of the above items 7 to 10, In the query acquisition process, when the computer acquires information that identifies a query image that is an image including a person, the computer acquires posture information corresponding to the query image as the query posture information. 12. In the image selection method described in any one of 7 to 11 above, the search posture information is set using positions of a plurality of joints, The image selection method, wherein the reliability is calculated using the number of joints used when setting the search posture information and the reliability of the joints. 13. On the computer, a query acquisition function for acquiring query posture information indicating a posture of a person, which is information to be used as a query; a search information acquisition function that acquires a plurality of pieces of search posture information that are generated for each of a plurality of target images to be searched and indicate the postures of people included in the target images; a selection function for selecting two or more pieces of search posture information from the plurality of pieces of search posture information, the search posture information having a similarity to the query posture information that satisfies a criterion; a display control function that displays the target images corresponding to the two or more pieces of orientation information for search selected by the selection function on a display means, and sets the display positions of the target images using the reliability of the orientation information for search; A program that allows you to have 14. In the program according to 13 above, The display control function is a program that causes a display unit to display the reliability of the search orientation information together with the target image corresponding to the search orientation information. 15. In the program according to 13 or 14 above, The display control function further determines a display position of the target image using the similarity. 16. In the program according to 13 or 14 above, The display control function is a program that uses the similarity in place of or in addition to the search posture information when determining the display position of the target image when a predetermined condition is met. 17. In the program according to any one of the above items 13 to 16, The query acquisition function is a program that, when acquiring information that identifies a query image that is an image including a person, acquires posture information corresponding to the query image as the query posture information. 18. In the program according to any one of the above items 13 to 17, the search posture information is set using positions of a plurality of joints, The reliability is calculated using the number of joints used when setting the posture information for search and the reliability of the joints. [Explanation of symbols]

[0128] 1. Image processing system 10 Image processing device (image selection device) 11 Skeleton detection unit 12 Feature calculation unit 13 Recognition part 100 Image processing device (image selection device) 101 Image acquisition unit 102 Skeletal structure detection unit 103 Feature calculation unit 104 Classification Department 105 Search Department 106 Input section 107 Display section 108 Height Calculation Unit 110 databases 200 cameras 300, 301 Human body model 401 2D skeletal structure 610 Query Acquisition Unit 620 Search information acquisition unit 630 Selection Section 640 Display control unit

Claims

1. a query acquisition means for acquiring first posture information indicating a predetermined posture of a person; an extraction means for extracting two or more pieces of second posture information that satisfy a predetermined condition for the first posture information from a plurality of pieces of second posture information that are stored in advance with reliability associated with the posture; a display control means for displaying the extracted two or more pieces of second orientation information based on the reliability associated with each of the two or more pieces of second orientation information; An information processing device comprising:

2. the second posture information is information generated based on skeletal information of a person included in the image, the reliability indicates a reliability of the second posture information generated based on the skeleton information; The information processing device according to claim 1 .

3. the display control means determines a position of each of the two or more pieces of second attitude information extracted by the extraction means based on the reliability.

3. The information processing device according to claim 1 or 2.

4. the display control means causes a display means to display an image corresponding to each of the two or more pieces of second attitude information extracted by the extraction means.

4. The information processing device according to claim 2 or 3.

5. the predetermined condition includes that the similarity to the first posture information satisfies a predetermined criterion. The information processing device according to any one of claims 1 to 4.

6. the display control means causes the display means to display the reliability of the second attitude information together with an image corresponding to the second attitude information.

6. The information processing apparatus according to claim 4 or claim 5 which cites claim 4.

7. the display control means further determines a display position of an image corresponding to the second attitude information using the similarity.

7. The information processing apparatus according to claim 5, which cites claim 4, or claim 6, which cites claims 4 and 5.

8. When the query acquisition means acquires information identifying a query image that is an image including a person, the query acquisition means acquires, as the first orientation information, orientation information corresponding to the query image. The information processing device according to any one of claims 1 to 7.

9. The computer First posture information indicating a predetermined posture of the person is acquired. extracting two or more pieces of second posture information that satisfy a predetermined condition for the first posture information from a plurality of pieces of second posture information that are stored in advance with their reliabilities associated with the postures; displaying the two or more pieces of extracted second orientation information based on the reliability associated with each of the two or more pieces of second orientation information; Information processing methods.

10. On the computer, a query acquisition function for acquiring first posture information indicating a predetermined posture of a person; an extraction function of extracting two or more pieces of second posture information that satisfy a predetermined condition for the first posture information from a plurality of pieces of second posture information that are stored in advance with their reliabilities associated with the postures; a display control function of displaying the two or more pieces of extracted second orientation information based on the reliability associated with each of the two or more pieces of second orientation information; A program that allows you to have

Citation Information

Patent Citations

  • System and method for image retrieval

    JP2001265811A

  • Image information updating system, image inputting device, image processing device, image updating device, image information updating method, image information updating program, and recording medium recording the program

    JP2006260405A

  • Face image search device and face image search method

    JP2012003623A

  • Posture estimation device and posture estimation method

    JP2013125402A

  • Apparatus and method for retrieving object poses

    JP2014522035A