Image processing apparatus, image processing method, and program
The image processing apparatus enhances video search accuracy by using skeleton structure detection and feature amount analysis to flexibly recognize and classify human postures and behaviors, addressing limitations in existing systems.
Patent Information
- Application Number
- JP2023523758
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-05-25
AI Technical Summary
Existing image processing systems struggle to accurately search for videos containing desired scenes due to limitations in recognizing and classifying human postures and behaviors, especially when partial body parts are hidden or orientations vary.
An image processing apparatus and method that utilizes skeleton structure detection, feature amount calculation, and change calculation to enhance the search accuracy by detecting key points, calculating feature amounts, and analyzing their direction of change over time, enabling flexible classification and search of human postures and behaviors.
Improves the search accuracy for videos containing desired scenes by allowing flexible recognition and classification of human postures and behaviors, even when partial body parts are hidden or orientations vary, thereby enhancing the user's ability to specify and find desired postures.
Smart Images

Figure 0007708182000003 
Figure 0007708182000004 
Figure 0007708182000005
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.
Background Art
[0002] In recent years, in a surveillance system or the like, a technique for detecting and searching for states such as a person's posture and behavior from an image of a surveillance camera has been used. As related techniques, for example, Patent Documents 1 and 2 are known. Patent Document 1 discloses a technique for searching for a similar person's posture based on key joints such as a person's head and limbs included in a depth video. Patent Document 2 discloses a technique for searching for a similar image using posture information such as an inclination added to an image, although it is not related to a person's posture. In addition, Non-Patent Document 1 is known as a technique related to human skeleton estimation.
[0003] On the other hand, in recent years, using a video as a query and searching for a video similar to this query has also been studied. For example, Patent Document 3 describes that when a reference video serving as a query is input, a similar video is searched for using the number of faces of the persons appearing, and the position, size, and orientation of the face of each person appearing.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] It is difficult to improve the search accuracy of the process of searching for a video containing a desired scene. One of the objects of the present invention is to improve the search accuracy of the process of searching for a video containing a desired scene.
Means for Solving the Problems
[0007] According to the present invention, query acquisition means for acquiring a plurality of first frame images in time series, skeleton structure detection means for detecting key points of an object included in each of the plurality of first frame images, feature amount calculation means for calculating a feature amount of the detected key points for each of the first frame images, change calculation means for calculating a direction of change of the feature amounts along a time axis of the plurality of first frame images in time series, search means for searching for a video using the calculated direction of change of the feature amounts as a key, An image processing apparatus having the above is provided.
[0008] Also, according to the present invention, a computer performs a query acquisition step of acquiring a plurality of first frame images in time series, a skeleton structure detection step of detecting key points of an object included in each of the plurality of first frame images, For each of the first frame images, a feature amount calculation step of calculating the feature amount of the detected key points; A change calculation step of calculating the direction of change of the feature amount along the time axis of a plurality of the first frame images in time series; A search step of searching for a moving image using the calculated direction of change of the feature amount as a key; An image processing method for executing the above is provided.
[0009] Further, according to the present invention, A computer is caused to function as Query acquisition means for acquiring a plurality of first frame images in time series, Skeleton structure detection means for detecting key points of an object included in each of the plurality of first frame images, Feature amount calculation means for calculating the feature amount of the detected key points for each of the first frame images, Change calculation means for calculating the direction of change of the feature amount along the time axis of a plurality of the first frame images in time series, and Search means for searching for a moving image using the calculated direction of change of the feature amount as a key, A program is provided.
Effect of the Invention
[0010] According to the present invention, the search accuracy of the process of searching for a moving image including a desired scene can be improved.
Brief Description of the Drawings
[0011] The above-described object, and other objects, features and advantages will become more apparent from the public embodiments described below and the accompanying drawings. [Figure 1] It is a block diagram showing an outline of an image processing apparatus according to an embodiment. [Diagram 2] It is a block diagram showing a configuration of an image processing apparatus according to Embodiment 1. [Diagram 3] It is a flowchart showing an image processing method according to Embodiment 1. [Figure 4] It is a flowchart showing the classification method according to Embodiment 1. [Diagram 5] It is a flowchart showing the search method according to Embodiment 1. [Figure 6] It is a diagram showing a detection example of a skeletal structure according to Embodiment 1. [Figure 7] It is a diagram showing a human body model according to Embodiment 1. [Figure 8] It is a diagram showing a detection example of a skeletal structure according to Embodiment 1. [Figure 9] It is a diagram showing a detection example of a skeletal structure according to Embodiment 1. [Figure 10] It is a diagram showing a detection example of a skeletal structure according to Embodiment 1. [Figure 11] It is a graph showing a specific example of the classification method according to Embodiment 1. [Figure 12] It is a diagram showing a display example of the classification result according to Embodiment 1. [Figure 13] It is a diagram for explaining the search method according to Embodiment 1. [Figure 14] It is a diagram for explaining the search method according to Embodiment 1. [Figure 15] It is a diagram for explaining the search method according to Embodiment 1. [Figure 16] It is a diagram for explaining the search method according to Embodiment 1. [Figure 17] It is a diagram showing a display example of the search result according to Embodiment 1. [Figure 18] It is a block diagram showing the configuration of the image processing apparatus according to Embodiment 2. [Figure 19] It is a flowchart showing the image processing method according to Embodiment 2. [Figure 20] It is a flowchart showing a specific example 1 of the method for calculating the number of pixels of height according to Embodiment 2. [Figure 21] It is a flowchart showing a specific example 2 of the method for calculating the number of pixels of height according to Embodiment 2. [Figure 22]It is a flowchart showing a specific example 2 of the height pixel number calculation method according to Embodiment 2. [Figure 23] It is a flowchart showing the normalization method according to Embodiment 2. [Figure 24] It is a diagram showing the human body model according to Embodiment 2. [Diagram 25] It is a diagram showing a detection example of the skeletal structure according to Embodiment 2. [Figure 26] It is a diagram showing a detection example of the skeletal structure according to Embodiment 2. [Figure 27] It is a diagram showing a detection example of the skeletal structure according to Embodiment 2. [Figure 28] It is a diagram showing the human body model according to Embodiment 2. [Figure 29] It is a diagram showing a detection example of the skeletal structure according to Embodiment 2. [Diagram 30] It is a histogram for explaining the height pixel number calculation method according to Embodiment 2. [Diagram 31] It is a diagram showing a detection example of the skeletal structure according to Embodiment 2. [Diagram 32] It is a diagram showing the 3D human body model according to Embodiment 2. [Diagram 33] It is a diagram for explaining the height pixel number calculation method according to Embodiment 2. [Diagram 34] It is a diagram for explaining the height pixel number calculation method according to Embodiment 2. [Diagram 35] It is a diagram for explaining the height pixel number calculation method according to Embodiment 2. [Diagram 36] It is a diagram for explaining the normalization method according to Embodiment 2. [Figure 37] It is a diagram for explaining the normalization method according to Embodiment 2. [Figure 38] It is a diagram for explaining the normalization method according to Embodiment 2. [Figure 39] It is a diagram showing an example of the hardware configuration of the image processing apparatus. [Diagram 40] It is a block diagram showing the configuration of the image processing apparatus according to Embodiment 3. [Diagram 41] This is a diagram for explaining the query frame selection process according to Embodiment 3. [Diagram 42] This is a diagram for explaining the query frame selection process according to Embodiment 3. [Diagram 43] This is a diagram for explaining the calculation process of the direction of change of feature amounts according to Embodiment 3. [Diagram 44] This is a flowchart showing an example of the processing flow of the image processing apparatus according to Embodiment 3. [Diagram 45] This is a configuration diagram showing the configuration of the image processing apparatus according to Embodiment 3. [Figure 46] This is a flowchart showing an example of the processing flow of the image processing apparatus according to Embodiment 3.
Embodiments for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same reference numerals are given to the same components, and the description will be omitted as appropriate.
[0013] (Considerations Leading to the Embodiment) In recent years, image recognition technologies utilizing machine learning such as deep learning have been applied to various systems. For example, the application to surveillance systems that perform surveillance using images from surveillance cameras is being promoted. By utilizing machine learning in surveillance systems, it is becoming possible to grasp to some extent the states such as the postures and actions of people from images.
[0014] However, in such related technologies, there are cases where it is not always possible to grasp the state of a person desired by the user on demand. For example, there are cases where the state of the person the user wants to search for and grasp can be specified in advance, and there are also cases where it cannot be specifically specified as in an unknown state. Then, in some cases, the user cannot specify in detail the state of the person they want to search for. Also, in cases where a part of a person's body is hidden, searching and the like cannot be performed. In related technologies, since the state of a person can only be searched from specific search conditions, it is difficult to flexibly search for and classify the state of a desired person.
[0015] Therefore, the inventors considered a method of using a skeleton estimation technique such as Non-Patent Document 1 in order to recognize the state of a person desired by the user from an image on demand. In related skeleton estimation techniques such as OpenPose disclosed in Non-Patent Document 1, the skeleton of a person is estimated by learning image data with various patterns that have been labeled. In the following embodiments, by utilizing such a skeleton estimation technique, it becomes possible to flexibly recognize the state of a person.
[0016] Note that the skeleton structure estimated by a skeleton estimation technique such as OpenPose is composed of "keypoints", which are characteristic points such as joints, and "bones (bone links)" that indicate the links between the keypoints. Therefore, in the following embodiments, the terms "keypoint" and "bone" will be used to describe the skeleton structure, but unless otherwise specifically limited, "keypoint" corresponds to the "joint" of a person, and "bone" corresponds to the "bone" of a person.
[0017] (Outline of Embodiment) FIG. 1 shows an overview of the image processing apparatus 10 according to the embodiment. As shown in FIG. 1, the image processing apparatus 10 includes a skeleton detection unit 11, a feature amount calculation unit 12, and a recognition unit 13. The skeleton detection unit 11 detects the two-dimensional skeleton structures of a plurality of persons based on a two-dimensional image acquired from a camera or the like. The feature amount calculation unit 12 calculates the feature amounts of the plurality of two-dimensional skeleton structures detected by the skeleton detection unit 11. The recognition unit 13 performs recognition processing of the states of a plurality of persons based on the similarity of the plurality of feature amounts calculated by the feature amount calculation unit 12. The recognition processing is classification processing, search processing, or the like of the states of persons.
[0018] As described above, in the embodiment, by detecting the two-dimensional skeleton structure of a person from a two-dimensional image and performing recognition processing such as classification and search of the state of the person based on the feature amount calculated from this two-dimensional skeleton structure, the state of a desired person can be flexibly recognized.
[0019] (Embodiment 1) Hereinafter, Embodiment 1 will be described with reference to the drawings. FIG. 2 shows the configuration of the image processing apparatus 100 according to the present embodiment. The image processing apparatus 100 constitutes an image processing system 1 together with a camera 200 and a database (DB) 201. The image processing system 1 including the image processing apparatus 100 is a system that classifies and searches for the states of a person, such as the posture and behavior of the person, based on the skeleton structure of the person estimated from the image.
[0020] The camera 200 is an imaging unit such as a surveillance camera that generates a two-dimensional image. The camera 200 is installed at a predetermined location and images a person or the like in the imaging area from the installation location. The camera 200 is directly connected to be able to output the captured image (video) to the image processing apparatus 100, or is connected via a network or the like. Note that the camera 200 may be provided inside the image processing apparatus 100.
[0021] The database 201 is a database that stores information (data) necessary for the processing of the image processing apparatus 100, processing results, and the like. The database 201 stores the image acquired by the image acquisition unit 101, the detection result of the skeleton structure detection unit 102, the data for machine learning, the feature amount calculated by the feature amount calculation unit 103, the classification result of the classification unit 104, the search result of the search unit 105, and the like. The database 201 is directly connected to the image processing apparatus 100 so that data can be input and output as needed, or is connected via a network or the like. Note that the database 201 may be provided inside the image processing apparatus 100 as a non-volatile memory such as a flash memory or a hard disk device.
[0022] As shown in FIG. 2, the image processing apparatus 100 includes an image acquisition unit 101, a skeleton structure detection unit 102, a feature amount calculation unit 103, a classification unit 104, a search unit 105, an input unit 106, and a display unit 107. Note that the configuration of each unit (block) is an example, and as long as the method (operation) described later is possible, it may be configured by other units. Further, the image processing apparatus 100 is realized by a computer device such as a personal computer or a server that executes a program, for example, but it may be realized by one device or by a plurality of devices on a network. For example, the input unit 106, the display unit 107, and the like may be external devices. Further, both the classification unit 104 and the search unit 105 may be provided, or only one of them may be provided. Both or one of the classification unit 104 and the search unit 105 is a recognition unit that performs recognition processing of the state of a person.
[0023] The image acquisition unit 101 acquires a two-dimensional image including a person imaged by the camera 200. The image acquisition unit 101 acquires, for example, an image (video including a plurality of images) including a person imaged by the camera 200 during a predetermined monitoring period. Note that the acquisition is not limited to that from the camera 200, and an image including a person prepared in advance may be acquired from the database 201 or the like.
[0024] Based on the acquired two-dimensional image, the skeletal structure detection unit 102 detects the two-dimensional skeletal structure of the person in the image. The skeletal structure detection unit 102 detects the skeletal structure for all the persons recognized in the acquired image. The skeletal structure detection unit 102 uses a skeletal estimation technique based on machine learning to detect the skeletal structure of the recognized person based on features such as joints of the person. The skeletal structure detection unit 102 uses, for example, a skeletal estimation technique such as OpenPose in Non-Patent Document 1.
[0025] The feature amount calculation unit 103 calculates the feature amount of the detected two-dimensional skeletal structure, and associates the calculated feature amount with the image to be processed and stores it in the database 201. The feature amount of the skeletal structure indicates the features of the person's skeleton and is an element for classifying and searching the state of the person based on the person's skeleton. Usually, this feature amount includes a plurality of parameters (for example, classification elements described later). And the feature amount may be the overall feature amount of the skeletal structure, or the feature amount of a part of the skeletal structure, or may include a plurality of feature amounts such as each part of the skeletal structure. The method for calculating the feature amount may be any method such as machine learning or normalization, and the minimum value or the maximum value may be obtained as normalization. As an example, the feature amount is a feature amount obtained by machine learning the skeletal structure, or the size on the image from the head to the feet of the skeletal structure, etc. The size of the skeletal structure is the vertical height or area of the skeletal region including the skeletal structure on the image. The vertical direction (height direction or longitudinal direction) is the vertical direction (Y-axis direction) in the image, for example, the direction perpendicular to the ground (reference plane). Also, the left-right direction (horizontal direction) is the left-right direction (X-axis direction) in the image, for example, the direction parallel to the ground.
[0026] In addition, in order to perform the classification and search desired by the user, it is preferable to use a feature amount that is robust to the classification and search processing. For example, when the user desires a classification and search that does not depend on the orientation or body type of the person, a feature amount that is robust to the orientation and body type of the person may be used. By learning the skeletons of persons facing in various directions in the same posture and the skeletons of persons with various body types in the same posture, or by extracting only the features in the vertical direction of the skeleton, a feature amount that does not depend on the orientation and body type of the person can be obtained.
[0027] The classification unit 104 classifies (clusters) a plurality of skeleton structures stored in the database 201 based on the similarity of the feature amounts of the skeleton structures. It can also be said that the classification unit 104 classifies the states of a plurality of persons based on the feature amounts of the skeleton structures as a recognition process of the states of persons. The similarity is the distance between the feature amounts of the skeleton structures. The classification unit 104 may classify based on the similarity of the overall feature amounts of the skeleton structures, or may classify based on the similarity of a part of the feature amounts of the skeleton structures, and may classify based on the similarity of the feature amounts of the first part (for example, both hands) and the second part (for example, both feet) of the skeleton structure. Note that the posture of a person may be classified based on the feature amounts of the skeleton structure of the person in each image, or the action of a person may be classified based on the change in the feature amounts of the skeleton structure of the person in a plurality of images continuous in time series. That is, the classification unit 104 can classify the state of a person including the posture and action of the person based on the feature amounts of the skeleton structure. For example, the classification unit 104 targets a plurality of skeleton structures in a plurality of images captured during a predetermined monitoring period for classification. The classification unit 104 obtains the similarity between the feature amounts of the classification targets and classifies them so that the skeleton structures with high similarity belong to the same cluster (a group with a similar posture). Note that, similar to the search, the user may be able to specify the classification conditions. The classification unit 104 stores the classification result of the skeleton structure in the database 201 and displays it on the display unit 107.
[0028] The search unit 105 searches for a skeletal structure with a high similarity to the feature amount of the search query (query state) from among a plurality of skeletal structures stored in the database 201. It can also be said that the search unit 105 is performing recognition processing of the state of a person, and searching for the state of a person that corresponds to the search condition (query state) from among the states of a plurality of persons based on the feature amount of the skeletal structure. Similar to the classification, the similarity is the distance between the feature amounts of the skeletal structures. The search unit 105 may perform the search based on the similarity of the overall feature amount of the skeletal structure, or may perform the search based on the similarity of a partial feature amount of the skeletal structure, and may also perform the search based on the similarity of the feature amounts of the first part (for example, both hands) and the second part (for example, both feet) of the skeletal structure. Note that the posture of a person may be searched based on the feature amount of the skeletal structure of the person in each image, or the action of a person may be searched based on the change in the feature amount of the skeletal structure of the person in a plurality of images continuous in time series. That is, the search unit 105 can search for the state of a person including the posture and action of the person based on the feature amount of the skeletal structure. For example, the search unit 105 sets the feature amounts of a plurality of skeletal structures in a plurality of images captured during a predetermined monitoring period as the search target, similar to the classification target. Also, the skeletal structure (posture) specified by the user from among the classification results displayed by the classification unit 104 is used as the search query (search key). Note that not limited to the classification results, a search query may be selected from among a plurality of unclassified skeletal structures, or the user may input the skeletal structure to be the search query. The search unit 105 searches for a feature amount with a high similarity to the feature amount of the skeletal structure of the search query from among the feature amounts of the search target. The search unit 105 stores the search result of the feature amount in the database 201 and displays it on the display unit 107.
[0029] The input unit 106 is an input interface that acquires information input from a user who operates the image processing apparatus 100. For example, the user is a monitor who monitors a person in a suspicious state from the image of a surveillance camera. The input unit 106 is, for example, a GUI (Graphical User Interface), and information corresponding to the user's operation is input from an input device such as a keyboard, a mouse, or a touch panel. For example, the input unit 106 accepts the skeletal structure of a specified person as a search query from among the skeletal structures (postures) classified by the classification unit 104.
[0030] The display unit 107 is a display unit that displays the results of operations (processes) of the image processing apparatus 100, etc. For example, it is a display device such as a liquid crystal display or an organic EL (Electro Luminescence) display. The display unit 107 displays the classification result of the classification unit 104 and the search result of the search unit 105 on the GUI according to the similarity, etc.
[0031] FIG. 39 is a diagram showing a hardware configuration example of the image processing apparatus 100. The image processing apparatus 100 includes a bus 1010, a processor 1020, a memory 1030, a storage device 1040, an input / output interface 1050, and a network interface 1060.
[0032] The bus 1010 is a data transmission path for the processor 1020, the memory 1030, the storage device 1040, the input / output interface 1050, and the network interface 1060 to transmit and receive data to and from each other. However, the method of connecting the processor 1020, etc. to each other is not limited to bus connection.
[0033] The processor 1020 is a processor realized by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or the like.
[0034] The memory 1030 is a main storage device realized by a RAM (Random Access Memory) or the like.
[0035] The storage device 1040 is an auxiliary storage device realized by, for example, a HDD (Hard Disk Drive), an SSD (Solid State Drive), a memory card, or a ROM (Read Only Memory). The storage device 1040 stores program modules that implement the respective functions of the image processing apparatus 100 (for example, the image acquisition unit 101, the skeletal structure detection unit 102, the feature amount calculation unit 103, the classification unit 104, the search unit 105, and the input unit 106). When the processor 1020 reads and executes these program modules onto the memory 1030, the respective functions corresponding to the program modules are realized. Also, the storage device 1040 may also function as the database 201.
[0036] The input / output interface 1050 is an interface for connecting the image processing apparatus 100 and various input / output devices. When the database 201 is located outside the image processing apparatus 100, the image processing apparatus 100 may be connected to the database 201 via the input / output interface 1050.
[0037] The network interface 1060 is an interface for connecting the image processing apparatus 100 to a network. This network is, for example, a LAN (Local Area Network) or a WAN (Wide Area Network). The method by which the network interface 1060 connects to the network may be a wireless connection or a wired connection. The image processing apparatus 100 may communicate with the camera 200 via the network interface 1060. When the database 201 is located outside the image processing apparatus 100, the image processing apparatus 100 may be connected to the database 201 via the network interface 1060.
[0038] FIGs. 3 to 5 show the operations of the image processing apparatus 100 according to the present embodiment. FIG. 3 shows the flow from image acquisition to search processing in the image processing apparatus 100, FIG. 4 shows the flow of the classification processing (S104) in FIG. 3, and FIG. 5 shows the flow of the search processing (S105) in FIG. 3.
[0039] As shown in FIG. 3, the image processing apparatus 100 acquires an image from the camera 200 (S101). The image acquisition unit 101 acquires an image of a person taken for classification and search from the skeletal structure, and stores the acquired image in the database 201. The image acquisition unit 101 acquires, for example, a plurality of images captured during a predetermined monitoring period, and performs subsequent processing on all the persons included in the plurality of images.
[0040] Subsequently, the image processing apparatus 100 detects the skeletal structure of the person based on the acquired image of the person (S102). FIG. 6 shows an example of the detection of the skeletal structure. As shown in FIG. 6, the image acquired from a surveillance camera or the like includes a plurality of persons, and the skeletal structure is detected for each person included in the image.
[0041] FIG. 7 shows the skeletal structure of the human body model 300 detected at this time, and FIGS. 8 to 10 show examples of the detection of the skeletal structure. The skeletal structure detection unit 102 uses a skeletal estimation technique such as OpenPose to detect the skeletal structure of the human body model (2D skeletal model) 300 as shown in FIG. 7 from a 2D image. The human body model 300 is a 2D model composed of key points such as joints of a person and bones connecting the key points.
[0042] The skeletal structure detection unit 102 extracts, for example, feature points that can be key points from an image, and detects each key point of a person with reference to the information obtained by machine learning of the key point images. In the example of FIG. 7, as the key points of the person, the head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71, left knee A72, right foot A81, and left foot A82 are detected. Further, as the bones of the person connecting these key points, the bone B1 connecting the head A1 and the neck A2, the bones B21 and B22 connecting the neck A2 and the right shoulder A31 and the left shoulder A32 respectively, the bones B31 and B32 connecting the right shoulder A31 and the left shoulder A32 and the right elbow A41 and the left elbow A42 respectively, the bones B41 and B42 connecting the right elbow A41 and the left elbow A42 and the right hand A51 and the left hand A52 respectively, the bones B51 and B52 connecting the neck A2 and the right hip A61 and the left hip A62 respectively, the bones B61 and B62 connecting the right hip A61 and the left hip A62 and the right knee A71 and the left knee A72 respectively, and the bones B71 and B72 connecting the right knee A71 and the left knee A72 and the right foot A81 and the left foot A82 respectively are detected. The skeletal structure detection unit 102 stores the detected skeletal structure of the person in the database 201.
[0043] FIG. 8 is an example of detecting a person in an upright state. In FIG. 8, an upright person is imaged from the front, and the bones B1, B51 and B52, B61 and B62, B71 and B72 seen from the front are detected without overlapping each other, and the bones B61 and B71 of the right foot are slightly bent more than the bones B62 and B72 of the left foot.
[0044] FIG. 9 is an example of detecting a person in a crouched state. In FIG. 9, a crouched person is imaged from the right side, and the bones B1, B51 and B52, B61 and B62, B71 and B72 seen from the right side are detected respectively, and the bones B61 and B71 of the right foot and the bones B62 and B72 of the left foot are greatly bent and overlapped.
[0045] FIG. 10 is an example of detecting a person in a lying state. In FIG. 10, a person in a lying state is imaged from the front left diagonal. Bones B1, B51 and B52, B61 and B62, B71 and B72 are detected respectively from the front left diagonal, and the bones B61 and B71 of the right foot and the bones B62 and B72 of the left foot are bent and overlapped.
[0046] Subsequently, as shown in FIG. 3, the image processing apparatus 100 calculates a feature amount of the detected bone structure (S103). For example, when the height and area of the bone region are used as feature amounts, the feature amount calculation unit 103 extracts the region including the bone structure, and obtains the height (number of pixels) and area (pixel area) of the region. The height and area of the bone region are obtained from the coordinates of the ends of the extracted bone region and the coordinates of the key points at the ends. The feature amount calculation unit 103 stores the obtained feature amount of the bone structure in the database 201.
[0047] In the example of FIG. 8, a bone region including all bones is extracted from the bone structure of a standing person. In this case, the upper end of the bone region is the key point A1 of the head, the lower end of the bone region is the key point A82 of the left foot, the left end of the bone region is the key point A41 of the right elbow, and the right end of the bone region is the key point A52 of the left hand. Therefore, the height of the bone region is obtained from the difference in the Y coordinates between the key point A1 and the key point A82. Also, the width of the bone region is obtained from the difference in the X coordinates between the key point A41 and the key point A52, and the area is obtained from the height and width of the bone region.
[0048] In the example of FIG. 9, a bone region including all bones is extracted from the bone structure of a crouched person. In this case, the upper end of the bone region is the key point A1 of the head, the lower end of the bone region is the key point A81 of the right foot, the left end of the bone region is the key point A61 of the right waist, and the right end of the bone region is the key point A51 of the right hand. Therefore, the height of the bone region is obtained from the difference in the Y coordinates between the key point A1 and the key point A81. Also, the width of the bone region is obtained from the difference in the X coordinates between the key point A61 and the key point A51, and the area is obtained from the height and width of the bone region.
[0049] In the example of FIG. 10, a skeleton region including all bones is extracted from the skeleton structure of a person lying horizontally in the left - right direction of the image. In this case, the upper end of the skeleton region is the key point A32 of the left shoulder, the lower end of the skeleton region is the key point A52 of the left hand, the left end of the skeleton region is the key point A51 of the right hand, and the right end of the skeleton region is the key point A82 of the left foot. Therefore, the height of the skeleton region is obtained from the difference in the Y - coordinates between the key point A32 and the key point A52. Also, the width of the skeleton region is obtained from the difference in the X - coordinates between the key point A51 and the key point A82, and the area is obtained from the height and width of the skeleton region.
[0050] Subsequently, as shown in FIG. 3, the image processing apparatus 100 performs classification processing (S104). In the classification processing, as shown in FIG. 4, the classification unit 104 calculates the similarity of the feature amounts of the calculated skeleton structure (S111), and classifies the skeleton structure based on the calculated feature amounts (S112). The classification unit 104 obtains the similarity of the feature amounts between all the skeleton structures stored in the database 201 to be classified, and classifies (clusters) the skeleton structure (posture) with the highest similarity into the same cluster. Further, the similarity between the classified clusters is obtained and classified, and the classification is repeated until a predetermined number of clusters are obtained. FIG. 11 shows an image of the classification result of the feature amounts of the skeleton structure. FIG. 11 is an image of cluster analysis using two - dimensional classification elements, and the two classification elements are, for example, the height of the skeleton region and the area of the skeleton region, etc. In FIG. 11, as a result of classification, the feature amounts of a plurality of skeleton structures are classified into three clusters C1 to C3. The clusters C1 to C3 correspond to each posture, such as a standing posture, a sitting posture, and a lying posture, and the skeleton structures (persons) are classified according to similar postures.
[0051] In this embodiment, various classification methods can be used by classifying based on the feature amounts of the skeletal structures of a person. Note that the classification method may be set in advance, or may be set arbitrarily by the user. Further, classification may be performed by the same method as the search method described later. That is, classification may be performed according to the same classification conditions as the search conditions. For example, the classification unit 104 classifies according to the following classification methods. Any one of the classification methods may be used, or arbitrarily selected classification methods may be combined.
[0052] (Classification method 1) Classification by multiple hierarchies Classify by hierarchically combining classifications based on the skeletal structure of the whole body, classifications based on the skeletal structures of the upper body and the lower body, classifications based on the skeletal structures of the arms and legs, etc. That is, classification may be performed based on the feature amounts of the first part and the second part of the skeletal structure, and further, classification may be performed by weighting the feature amounts of the first part and the second part.
[0053] (Classification method 2) Classification by a plurality of images along a time series Classify based on the feature amounts of the skeletal structures in a plurality of images continuous in a time series. For example, the feature amounts may be stacked in the time series direction and classified based on the cumulative value. Further, classification may be performed based on the change (change amount) of the feature amounts of the skeletal structures in a plurality of consecutive images.
[0054] (Classification method 3) Classification ignoring the left and right of the skeletal structure Classify the skeletal structures with the right and left sides of a person reversed as the same skeletal structure.
[0055] Furthermore, the classification unit 104 displays the classification result of the skeletal structure (S113). The classification unit 104 acquires the necessary skeletal structures and images of people from the database 201, and displays the skeletal structures and people in the display unit 107 for each similar posture (cluster) as the classification result. FIG. 12 shows a display example when the postures are classified into three. For example, as shown in FIG. 12, in the display window W1, the posture regions WA1 to WA3 for each posture are displayed, and the skeletal structures and people (images) of the postures corresponding to the posture regions WA1 to WA3 are respectively displayed. The posture region WA1 is, for example, a display region for a standing posture, and displays the skeletal structure and people similar to the standing posture classified into the cluster C1. The posture region WA2 is, for example, a display region for a sitting posture, and displays the skeletal structure and people similar to the sitting posture classified into the cluster C2. The posture region WA3 is, for example, a display region for a lying posture, and displays the skeletal structure and people similar to the lying posture classified into the cluster C2.
[0056] Subsequently, as shown in FIG. 3, the image processing apparatus 100 performs a search process (S105). In the search process, as shown in FIG. 5, the search unit 105 receives an input of search conditions (S121) and searches for a skeletal structure based on the search conditions (S122). The search unit 105 receives an input of a search query, which is a search condition, from the input unit 106 in response to a user operation. When inputting a search query from the classification result, for example, in the display example of FIG. 12, the user designates (selects) the skeletal structure of the posture to be searched from among the posture regions WA1 to WA3 displayed in the display window W1. Then, the search unit 105 searches for a skeletal structure with a high similarity of feature amounts from among all the skeletal structures stored in the database 201 to be searched, using the skeletal structure designated by the user as the search query. The search unit 105 calculates the similarity between the feature amounts of the skeletal structure of the search query and the feature amounts of the skeletal structure to be searched, and extracts a skeletal structure whose calculated similarity is higher than a predetermined threshold. As the feature amounts of the skeletal structure of the search query, the pre-calculated feature amounts may be used, or the feature amounts obtained at the time of search may be used. Note that the search query may be input by moving each part of the skeletal structure according to a user operation, or the posture demonstrated by the user in front of the camera may be used as the search query.
[0057] In the present embodiment, similar to the classification method, by performing a search based on the feature amounts of the skeletal structure of a person, various search methods can be used. Note that the search method may be preset, or may be set arbitrarily by the user. For example, the search unit 105 performs a search by the following search method. Any one of the search methods may be used, or arbitrarily selected search methods may be combined. A plurality of search methods (search conditions) may be combined by a logical formula (for example, AND (logical product), OR (logical sum), NOT (negation)) for searching. For example, the search condition may be set to "(posture with the right hand raised) AND (posture with the left foot raised)" for searching.
[0058] (Search method 1) By performing a search using only the feature quantity in the height direction of the person to be searched, the influence of the lateral change of the person can be suppressed, and the robustness is improved against the change in the orientation and body shape of the person. For example, like the skeletal structures 501 to 503 in FIG. 13, even when the orientation and body shape of the person are different, the feature quantity in the height direction does not change significantly. Therefore, in the skeletal structures 501 to 503, it can be determined that they are in the same posture at the time of search (classification).
[0059] (Search method 2) When a part of the person's body is hidden in the partial search image, search is performed using only the information of the recognizable part. For example, like the skeletal structures 511 and 512 in FIG. 14, even when the left foot is hidden and the key points of the left foot cannot be detected, search can be performed using the feature quantities of the other detected key points. Therefore, in the skeletal structures 511 and 512, it can be determined that they are in the same posture at the time of search (classification). That is, classification and search can be performed using the feature quantities of some key points instead of all key points. In the examples of the skeletal structures 521 and 522 in FIG. 15, although the orientations of both feet are different, by using the feature quantities of the upper body key points (A1, A2, A31, A32, A41, A42, A51, A52) as the search query, it can be determined that they are in the same posture. Also, weights may be assigned to the parts (feature points) to be searched, or the threshold for similarity determination may be changed. When a part of the body is hidden, the hidden part may be ignored for search, or the hidden part may be taken into account for search. By including the hidden part in the search, postures in which the same part is hidden can be searched for.
[0060] (Search method 3) Search for the right and left sides of the searched person while ignoring the left - right of the skeletal structure, considering the skeletal structures with opposite right - left sides as the same skeletal structure. For example, as in the skeletal structures 531 and 532 in FIG. 16, the posture of raising the right hand and the posture of raising the left hand can be searched (classified) as the same posture. In the example of FIG. 16, although the positions of the key points A51 of the right hand, A41 of the right elbow, A52 of the left hand, and A42 of the left elbow in the skeletal structure 531 and the skeletal structure 532 are different, the positions of the other key points are the same. Among the key points A51 of the right hand and A41 of the right elbow of the skeletal structure 531 and the key points A52 of the left hand and A42 of the left elbow of the skeletal structure 532, when the key points of one skeletal structure are horizontally flipped, they will be in the same position as the key points of the other skeletal structure. Also, among the key points A52 of the left hand and A42 of the left elbow of the skeletal structure 531 and the key points A51 of the right hand and A41 of the right elbow of the skeletal structure 532, when the key points of one skeletal structure are horizontally flipped, they will be in the same position as the key points of the other skeletal structure. Therefore, they are judged to be the same posture.
[0061] (Search method 4) Search is first performed only based on the feature amount in the vertical direction (Y - axis direction) of the searched person according to the feature amounts in the vertical and horizontal directions, and then the obtained result is further searched using the feature amount in the horizontal direction (X - axis direction) of the person.
[0062] (Search method 5) Search based on the feature amounts of the skeletal structures in a plurality of images along the time series. For example, the feature amounts can be stacked in the time - series direction and searched based on the cumulative value. Further, search can also be performed based on the change (change amount) of the feature amounts of the skeletal structures in a plurality of consecutive images.
[0063] Furthermore, the search unit 105 displays the search results of the skeletal structure (S123). The search unit 105 acquires the necessary skeletal structures and images of people from the database 201, and displays the obtained skeletal structures and people as search results on the display unit 107. For example, when multiple search queries (search conditions) are specified, the search results are displayed for each search query. FIG. 17 shows a display example when searching with three search queries (postures). For example, as shown in FIG. 17, in the display window W2, the skeletal structures and people of the search queries Q10, Q20, and Q30 specified at the left end are displayed, and the skeletal structures and people of the search results Q11, Q21, and Q31 of each search query are arranged and displayed on the right side of the search queries Q10, Q20, and Q30.
[0064] The order in which the search results are arranged and displayed next to the search queries may be the order in which the corresponding skeletal structures are found, or the order of high similarity. When searching by weighting the partial search part (feature points), the display may be in the order of the similarity calculated with weighting. The display may be in the order of the similarity calculated only from the part (feature points) selected by the user. Also, around the image (frame) of the search result, the images (frames) before and after in the time series may be cut out for a certain period of time and displayed.
[0065] As described above, in the present embodiment, it is possible to detect the skeletal structure of a person from a 2D image and perform classification and search based on the feature amounts of the detected skeletal structure. As a result, it is possible to classify by similar postures with high similarity, and also to search for similar postures with high similarity to the search query (search key). By classifying and displaying similar postures from the image, the user can grasp the posture of the person in the image without specifying the posture or the like. Since the user can specify the posture of the search query from the classification results, it is possible to search for the desired posture even when the user has not grasped in detail the posture to be searched in advance. For example, since classification and search can be performed based on conditions such as the whole or a part of the skeletal structure of a person, flexible classification and search are possible.
[0066] (Embodiment 2) Next, Embodiment 2 will be described with reference to the drawings. In this embodiment, specific examples of feature amount calculation in Embodiment 1 will be described. In this embodiment, features are obtained by normalizing using the height of a person. Other aspects are the same as in Embodiment 1.
[0067] FIG. 18 shows the configuration of the image processing apparatus 100 according to this embodiment. As shown in FIG. 18, in addition to the configuration of Embodiment 1, the image processing apparatus 100 further includes a height calculation unit 108. Note that the feature amount calculation unit 103 and the height calculation unit 108 may be regarded as one processing unit.
[0068] The height calculation unit (height estimation unit) 108 calculates (estimates) the height (referred to as the number of height pixels) of a person in a 2D image when standing upright based on the 2D skeletal structure detected by the skeletal structure detection unit 102. The number of height pixels can also be said to be the height of the person in the 2D image (the length of the whole body of the person in the 2D image space). The height calculation unit 108 obtains the number of height pixels (number of pixels) from the length of each bone in the detected skeletal structure (the length in the 2D image space).
[0069] In the following examples, Specific Examples 1 to 3 are used as methods for obtaining the number of height pixels. Note that any one of Specific Examples 1 to 3 may be used, or a combination of arbitrarily selected multiple methods may be used. In Specific Example 1, the number of height pixels is obtained by summing the lengths of the bones from the head to the feet among the bones of the skeletal structure. When the skeletal structure detection unit 102 (skeletal estimation technology) does not output the top of the head and the feet, correction can be performed by multiplying by a constant as necessary. In Specific Example 2, the number of height pixels is calculated using a human body model showing the relationship between the length of each bone and the length of the whole body (height in the 2D image space). In Specific Example 3, the number of height pixels is calculated by fitting a 3D human body model to the 2D skeletal structure.
[0070] The feature quantity calculation unit 103 of the present embodiment is a normalization unit that normalizes the skeletal structure (skeletal information) of a person based on the calculated number of pixels of the person's height. The feature quantity calculation unit 103 stores the feature quantity (normalized value) of the normalized skeletal structure in the database 201. The feature quantity calculation unit 103 normalizes the height of each keypoint (feature point) included in the skeletal structure on the image by the number of pixels of the height. In the present embodiment, for example, the height direction is the vertical direction (Y-axis direction) in the two-dimensional coordinate (X-Y coordinate) space of the image. In this case, the height of the keypoint can be obtained from the value (number of pixels) of the Y coordinate of the keypoint. Alternatively, the height direction may be the direction of the vertical projection axis (vertical projection direction) obtained by projecting the direction of the vertical axis perpendicular to the ground (reference plane) in the three-dimensional coordinate space of the real world onto the two-dimensional coordinate space. In this case, the height of the keypoint can be obtained by obtaining the vertical projection axis obtained by projecting the axis perpendicular to the ground in the real world onto the two-dimensional coordinate space based on the camera parameters, and from the value (number of pixels) along this vertical projection axis. The camera parameters are the imaging parameters of the image. For example, the camera parameters are the posture, position, imaging angle, focal length, etc. of the camera 200. By imaging an object whose length and position are known in advance with the camera 200, the camera parameters can be obtained from the image. Distortion occurs at both ends of the captured image, and the vertical direction of the real world may not match the vertical direction of the image. In contrast, by using the parameters of the camera that captured the image, it is possible to know how much the vertical direction of the real world is tilted in the image. Therefore, by normalizing the value of the keypoint along the vertical projection axis projected into the image based on the camera parameters by the height, the keypoint can be quantified while considering the deviation between the real world and the image. The left-right direction (horizontal direction) is the left-right direction (X-axis direction) in the two-dimensional coordinate (X-Y coordinate) space of the image, or the direction obtained by projecting the direction parallel to the ground in the three-dimensional coordinate space of the real world onto the two-dimensional coordinate space.
[0071] Figs. 19 to 23 show the operations of the image processing apparatus 100 according to this embodiment. Fig. 19 shows the flow from image acquisition to search processing in the image processing apparatus 100. Figs. 20 to 22 show the flows of Specific Examples 1 to 3 of the height pixel number calculation process (S201) in Fig. 19, and Fig. 23 shows the flow of the normalization process (S202) in Fig. 19.
[0072] As shown in Fig. 19, in this embodiment, as the feature amount calculation process (S103) in Embodiment 1, a height pixel number calculation process (S201) and a normalization process (S202) are performed. Other aspects are the same as those in Embodiment 1.
[0073] Following image acquisition (S101) and skeleton structure detection (S102), the image processing apparatus 100 performs a height pixel number calculation process (S201) based on the detected skeleton structure. In this example, as shown in Fig. 24, the height of the skeleton structure of a person in an upright position in the image is defined as the height pixel number (h), and the height of each key point of the skeleton structure in the state of the person in the image is defined as the key point height (yi). Hereinafter, Specific Examples 1 to 3 of the height pixel number calculation process will be described.
[0074] <Specific Example 1> In Specific Example 1, the height pixel number is obtained using the length of the bones from the head to the feet. In Specific Example 1, as shown in Fig. 20, the height calculation unit 108 acquires the length of each bone (S211) and sums up the lengths of the acquired bones (S212).
[0075] The height calculation unit 108 acquires the lengths of the bones on the two-dimensional image from the head to the feet of a person, and obtains the number of height pixels. That is, from the image in which the skeletal structure is detected, among the bones in FIG. 24, the lengths (number of pixels) of bone B1 (length L1), bone B51 (length L21), bone B61 (length L31), and bone B71 (length L41), or bone B1 (length L1), bone B52 (length L22), bone B62 (length L32), and bone B72 (length L42) are acquired. The length of each bone can be obtained from the coordinates of each keypoint in the two-dimensional image. The value obtained by multiplying the sum of these, L1 + L21 + L31 + L41, or L1 + L22 + L32 + L42, by a correction constant is calculated as the number of height pixels (h). When both values can be calculated, for example, the longer value is used as the number of height pixels. That is, each bone has the longest length in the image when imaged from the front, and is displayed shorter when tilted in the depth direction with respect to the camera. Therefore, it is considered that the longer bone is more likely to be imaged from the front and is closer to the true value. For this reason, it is preferable to select the longer value.
[0076] In the example of FIG. 25, bone B1, bone B51 and bone B52, bone B61 and bone B62, bone B71 and bone B72 are detected without overlapping each other. The sum of these bones, L1 + L21 + L31 + L41, and L1 + L22 + L32 + L42 are obtained, and for example, the value obtained by multiplying the longer L1 + L22 + L32 + L42 on the left foot side where the detected bone length is long by a correction constant is used as the number of height pixels.
[0077] In the example of FIG. 26, bone B1, bone B51 and bone B52, bone B61 and bone B62, bone B71 and bone B72 are detected respectively, and the right foot bones B61 and B71 and the left foot bones B62 and B72 overlap. The sum of these bones, L1 + L21 + L31 + L41, and L1 + L22 + L32 + L42 are obtained, and for example, the value obtained by multiplying the longer L1 + L21 + L31 + L41 on the right foot side where the detected bone length is long by a correction constant is used as the number of height pixels.
[0078] In the example of FIG. 27, bones B1, B51 and B52, B61 and B62, B71 and B72 are detected respectively, and bones B61 and B71 of the right foot overlap with bones B62 and B72 of the left foot. The sums of these bones, L1+L21+L31+L41 and L1+L22+L32+L42, are obtained. For example, the value obtained by multiplying the longer L1+L22+L32+L42 on the left foot side where the detected bone length is long by a correction constant is used as the number of body height pixels.
[0079] In Specific Example 1, since the body height can be obtained by summing the lengths of the bones from the head to the feet, the number of body height pixels can be obtained by a simple method. Also, with the skeleton estimation technology using machine learning, as long as at least the skeleton from the head to the feet can be detected, the number of body height pixels can be accurately estimated even when the entire person is not necessarily shown in the image, such as in a crouched state.
[0080] <Specific Example 2> In Specific Example 2, the number of body height pixels is obtained using a two-dimensional skeleton model that shows the relationship between the length of the bones included in the two-dimensional skeleton structure and the length of the whole body of a person in the two-dimensional image space.
[0081] FIG. 28 is a human body model (two-dimensional skeleton model) 301 that shows the relationship between the length of each bone in the two-dimensional image space and the length of the whole body in the two-dimensional image space, which is used in Specific Example 2. As shown in FIG. 28, the relationship between the length of each bone of an average person and the length of the whole body (the ratio of the length of each bone to the length of the whole body) is associated with each bone of the human body model 301. For example, the length of the head bone B1 is 0.2 (20%) of the length of the whole body, the length of the right hand bone B41 is 0.15 (15%) of the length of the whole body, and the length of the right foot bone B71 is 0.25 (25%) of the length of the whole body. By storing the information of such a human body model 301 in the database 201, the average length of the whole body can be obtained from the length of each bone. In addition to the human body model of an average person, human body models may be prepared for each attribute of a person such as age, gender, nationality, etc. Thereby, the length of the whole body (body height) can be appropriately obtained according to the attribute of the person.
[0082] In Specific Example 2, as shown in FIG. 21, the height calculation unit 108 acquires the length of each bone (S221). The height calculation unit 108 acquires the length of all the bones (the length in the two-dimensional image space) in the detected bone structure. FIG. 29 is an example in which a person in a crouched state is imaged from the rear right diagonal, and the bone structure is detected. In this example, since the face and the left side of the person are not captured, the bones of the head, the left arm, and the left hand cannot be detected. Therefore, the lengths of the detected bones B21, B22, B31, B41, B51, B52, B61, B62, B71, and B72 are acquired.
[0083] Subsequently, as shown in FIG. 21, the height calculation unit 108 calculates the number of height pixels from the length of each bone based on the human body model (S222). The height calculation unit 108 refers to the human body model 301 showing the relationship between each bone and the length of the whole body as shown in FIG. 28, and obtains the number of height pixels from the length of each bone. For example, since the length of the right hand bone B41 is 0.15 times the length of the whole body, the number of height pixels based on the bone B41 is obtained by dividing the length of the bone B41 by 0.15. Also, since the length of the right foot bone B71 is 0.25 times the length of the whole body, the number of height pixels based on the bone B71 is obtained by dividing the length of the bone B71 by 0.25.
[0084] The human body model referred to at this time is, for example, a human body model of an average person, but the human body model may be selected according to the attributes of the person such as age, gender, and nationality. For example, when the face of a person is captured in the captured image, the attributes of the person are identified based on the face, and the human body model corresponding to the identified attributes is referred to. By referring to the information obtained by machine learning the face for each attribute, the attributes of the person can be recognized from the features of the face in the image. Also, when the attributes of the person cannot be identified from the image, an average human body model may be used.
[0085] Further, the number of body height pixels calculated from the bone length may be corrected using camera parameters. For example, when the camera is at a high position and the person is photographed looking down, in the two-dimensional bone structure, the horizontal length of bones such as the shoulder width is not affected by the depression angle of the camera, but the vertical length of bones such as the neck-waist bone becomes smaller as the depression angle of the camera increases. Then, the number of body height pixels calculated from the horizontal length of bones such as the shoulder width tends to be larger than the actual value. Therefore, by utilizing camera parameters, it is possible to know at what angle the person is being looked down upon by the camera, and thus correct the two-dimensional bone structure as if photographed from the front using this depression angle information. As a result, the number of body height pixels can be calculated more accurately.
[0086] Subsequently, as shown in FIG. 21, the body height calculation unit 108 calculates the optimum value of the number of body height pixels (S223). The body height calculation unit 108 calculates the optimum value of the number of body height pixels from the number of body height pixels obtained for each bone. For example, a histogram of the number of body height pixels obtained for each bone as shown in FIG. 30 is generated, and a large number of body height pixels are selected from it. That is, the number of body height pixels that is longer than the others is selected from among the plurality of body height pixels obtained based on a plurality of bones. For example, the top 30% is set as valid values, and in FIG. 30, the number of body height pixels based on bones B71, B61, and B51 is selected. The average of the selected number of body height pixels may be obtained as the optimum value, or the largest number of body height pixels may be used as the optimum value. To obtain the body height from the length of the bones in the two-dimensional image, when the bones are not formed from the front, that is, when the bones are imaged tilted in the depth direction as viewed from the camera, the length of the bones is shorter than when imaged from the front. Then, since the value with a large number of body height pixels is more likely to be imaged from the front than the value with a small number of body height pixels and is a more plausible value, a larger value is used as the optimum value.
[0087] In Specific Example 2, in order to obtain the number of body length pixels based on the bones of the detected skeletal structure using a human body model showing the relationship between the bones in the two-dimensional image space and the overall body length, even when all the skeletons from the head to the feet cannot be obtained, the number of body length pixels can be obtained from some of the bones. In particular, by adopting the larger value among the values obtained from a plurality of bones, the number of body length pixels can be accurately estimated.
[0088] <Specific Example 3> In Specific Example 3, the two-dimensional skeletal structure is fitted to a three-dimensional human body model (three-dimensional skeletal model), and the overall body skeletal vector is obtained using the number of body length pixels of the fitted three-dimensional human body model.
[0089] In Specific Example 3, as shown in FIG. 22, first, the body length calculation unit 108 calculates the camera parameters based on the image captured by the camera 200 (S231). The body length calculation unit 108 extracts an object whose length is known in advance from among a plurality of images captured by the camera 200, and obtains the camera parameters from the size (number of pixels) of the extracted object. Note that the camera parameters may be obtained in advance, and the obtained camera parameters may be acquired as needed.
[0090] Subsequently, the body length calculation unit 108 adjusts the arrangement and height of the three-dimensional human body model (S232). The body length calculation unit 108 prepares a three-dimensional human body model for calculating the number of body length pixels for the detected two-dimensional skeletal structure, and arranges it within the same two-dimensional image based on the camera parameters. Specifically, the "relative positional relationship between the camera and the person in the real world" is specified from the camera parameters and the two-dimensional skeletal structure. For example, assuming that the position of the camera is the coordinate (0, 0, 0), the coordinates (x, y, z) of the position where the person is standing (or sitting) are specified. Then, by assuming an image when the three-dimensional human body model is arranged and imaged at the same position (x, y, z) as the identified person, the two-dimensional skeletal structure and the three-dimensional human body model are superimposed.
[0091] FIG. 31 shows an example in which a person in a crouched position is imaged from the front left diagonal and a two-dimensional skeleton structure 401 is detected. The two-dimensional skeleton structure 401 has two-dimensional coordinate information. Although it is preferable to detect all bones, some bones may not be detected. For this two-dimensional skeleton structure 401, a three-dimensional human body model 402 as shown in FIG. 32 is prepared. The three-dimensional human body model (three-dimensional skeleton model) 402 has three-dimensional coordinate information and is a model of a skeleton having the same shape as the two-dimensional skeleton structure 401. Then, as shown in FIG. 33, the prepared three-dimensional human body model 402 is arranged and superimposed on the detected two-dimensional skeleton structure 401. Further, while superimposing, the height of the three-dimensional human body model 402 is adjusted to match the two-dimensional skeleton structure 401.
[0092] Note that the three-dimensional human body model 402 prepared at this time may be a model in a state close to the posture of the two-dimensional skeleton structure 401 or a model in an upright state, as shown in FIG. 33. For example, a three-dimensional human body model 402 in the estimated posture may be generated using a technique for estimating the posture in three-dimensional space from a two-dimensional image using machine learning. By learning the information of the joints in the two-dimensional image and the joints in the three-dimensional space, the posture in three-dimensional space can be estimated from the two-dimensional image.
[0093] Subsequently, as shown in FIG. 22, the height calculation unit 108 fits the three-dimensional human body model to the two-dimensional skeletal structure (S233). As shown in FIG. 34, the height calculation unit 108 deforms the three-dimensional human body model 402 so that the postures of the three-dimensional human body model 402 and the two-dimensional skeletal structure 401 match in a state where the three-dimensional human body model 402 is superimposed on the two-dimensional skeletal structure 401. That is, the height, body orientation, and joint angles of the three-dimensional human body model 402 are adjusted and optimized so that there is no difference from the two-dimensional skeletal structure 401. For example, the joints of the three-dimensional human body model 402 are rotated within the movable range of a person, and the entire three-dimensional human body model 402 is rotated or the overall size is adjusted. Note that the fitting of the three-dimensional human body model to the two-dimensional skeletal structure is performed in a two-dimensional space (two-dimensional coordinates). That is, the three-dimensional human body model is mapped onto the two-dimensional space, and the three-dimensional human body model is optimized in consideration of how the deformed three-dimensional human body model changes in the two-dimensional space (image).
[0094] Subsequently, as shown in FIG. 22, the height calculation unit 108 calculates the number of height pixels of the fitted three-dimensional human body model (S234). As shown in FIG. 35, when there is no difference between the three-dimensional human body model 402 and the two-dimensional skeletal structure 401 and the postures match, the height calculation unit 108 obtains the number of height pixels of the three-dimensional human body model 402 in that state. With the optimized three-dimensional human body model 402 in an upright state, the length of the whole body in the two-dimensional space is obtained based on the camera parameters. For example, the number of height pixels is calculated based on the length (number of pixels) of the bones from the head to the feet when the three-dimensional human body model 402 is upright. Similar to Specific Example 1, the lengths of the bones from the head to the feet of the three-dimensional human body model 402 may be summed up.
[0095] In Specific Example 3, by fitting the three-dimensional human body model to the two-dimensional skeletal structure based on the camera parameters and obtaining the number of height pixels based on the three-dimensional human body model, even when not all bones are shown in the front, that is, when there is a large error because all bones are shown obliquely, the number of height pixels can be accurately estimated.
[0096] <Normalization process> As shown in FIG. 19, after the height pixel number calculation process, the image processing apparatus 100 performs a normalization process (S202). In the normalization process, as shown in FIG. 23, the feature amount calculation unit 103 calculates the keypoint height (S241). The feature amount calculation unit 103 calculates the keypoint height (number of pixels) of all the keypoints included in the detected skeletal structure. The keypoint height is the length (number of pixels) in the height direction from the lowermost end of the skeletal structure (for example, a keypoint of either foot) to the keypoint. Here, as an example, the keypoint height is obtained from the Y coordinate of the keypoint in the image. Note that, as described above, the keypoint height may be obtained from the length in the direction along the vertical projection axis based on the camera parameters. For example, in the example of FIG. 24, the height (yi) of the keypoint A2 of the neck is the value obtained by subtracting the Y coordinate of the keypoint A81 of the right foot or the Y coordinate of the keypoint A82 of the left foot from the Y coordinate of the keypoint A2.
[0097] Subsequently, the feature amount calculation unit 103 specifies a reference point for normalization (S242). The reference point is a point serving as a reference for representing the relative height of the keypoints. The reference point may be set in advance or may be selectable by the user. The reference point is preferably at or higher than the center of the skeletal structure (above in the vertical direction of the image). For example, the coordinates of the keypoint of the neck are used as the reference point. Note that not only the neck but also the coordinates of the head or other keypoints may be used as the reference point. Not limited to keypoints, arbitrary coordinates (for example, the center coordinates of the skeletal structure, etc.) may be used as the reference point.
[0098] Subsequently, the feature quantity calculation unit 103 normalizes the keypoint height (yi) by the number of pixels of the height (S243). The feature quantity calculation unit 103 normalizes each keypoint using the keypoint height, reference point, and number of pixels of the height of each keypoint. Specifically, the feature quantity calculation unit 103 normalizes the relative height of the keypoint with respect to the reference point by the number of pixels of the height. Here, as an example of focusing only on the height direction, only the Y coordinate is extracted, and normalization is performed using the keypoint of the neck as the reference point. Specifically, with the Y coordinate of the reference point (keypoint of the neck) being (yc), the following formula (1) is used to obtain the feature quantity (normalized value). When using the vertical projection axis based on the camera parameters, (yi) and (yc) are converted into values in the direction along the vertical projection axis.
[0099]
Number
[0100] For example, when there are 18 keypoints, the coordinates (x0, y0), (x1, y1), ··· (x17, y17) of the 18 points of each keypoint are converted into an 18-dimensional feature quantity as follows using the above formula (1).
[0101]
Number
[0102] Figure 36 shows an example of the feature amounts of each keypoint obtained by the feature amount calculation unit 103. In this example, since the keypoint A2 of the neck is used as the reference point, the feature amount of the keypoint A2 is 0.0, and the feature amounts of the right shoulder keypoint A31 and the left shoulder keypoint A32 at the same height as the neck are also 0.0. The feature amount of the head keypoint A1 higher than the neck is -0.2. The feature amounts of the right hand keypoint A51 and the left hand keypoint A52 lower than the neck are 0.4, and the feature amounts of the right foot keypoint A81 and the left foot keypoint A82 are 0.9. When the person raises the left hand from this state, since the left hand becomes higher than the reference point as shown in Figure 37, the feature amount of the left hand keypoint A52 becomes -0.4. On the other hand, since normalization is performed using only the coordinates of the Y axis, as shown in Figure 38, the feature amounts do not change even if the width of the skeletal structure changes compared to Figure 36. That is, the feature amount (normalized value) of the present embodiment shows the feature in the height direction (Y direction) of the skeletal structure (keypoint) and is not affected by the change in the lateral direction (X direction) of the skeletal structure.
[0103] As described above, in the present embodiment, the skeletal structure of a person is detected from a two-dimensional image, and each keypoint of the skeletal structure is normalized using the number of pixels of the height obtained from the detected skeletal structure (the height in the upright state in the two-dimensional image space). By using this normalized feature amount, the robustness in classification, search, etc. can be improved. That is, since the feature amount of the present embodiment is not affected by the change in the lateral direction of the person as described above, it has high robustness against the change in the orientation of the person and the body shape of the person.
[0104] Furthermore, in the present embodiment, since it can be realized by detecting the skeletal structure of a person using a skeletal estimation technique such as OpenPose, it is not necessary to prepare learning data for learning the posture of the person. Also, by normalizing the keypoints of the skeletal structure and storing them in a database, classification and search of the posture of the person etc. become possible, so classification and search can be performed even for an unknown posture. Also, by normalizing the keypoints of the skeletal structure, a clear and easy-to-understand feature amount can be obtained, so different from a black box type algorithm such as machine learning, the user's acceptance of the processing result is high.
[0105] (Embodiment 3) Hereinafter, Embodiment 3 will be described with reference to the drawings. In this embodiment, a specific example of a process for searching for a video including a desired scene will be described.
[0106] FIG. 40 shows an example of a functional block diagram of the image processing apparatus 100 according to this embodiment. As shown in the figure, the image processing apparatus 100 includes a query acquisition unit 109, a query frame selection unit 112, a skeleton structure detection unit 102, a feature amount calculation unit 103, a change calculation unit 110, and a search unit 111. Note that the image processing apparatus 100 may further include other functional units described in Embodiments 1 and 2. An example of the hardware configuration of the image processing apparatus 100 according to this embodiment is the same as that in Embodiments 1 and 2.
[0107] The query acquisition unit 109 acquires a query video composed of a plurality of first frame images in time series. For example, the query acquisition unit 109 acquires a query video (video file) input / specified / selected by a user operation.
[0108] The query frame selection unit 112 selects at least a part of the plurality of first frame images as query frames. As shown in FIGS. 41 and 42, the query frame selection unit 112 can intermittently select query frames from among the plurality of first frame images in time series included in the query video. The number of first frame images between query frames may be constant or may be various. The query frame selection unit 112 can execute, for example, any one of the following selection processes 1 to 3.
[0109] - Selection process 1 - In selection process 1, the query frame selection unit 112 selects query frames based on user input. That is, the user makes an input designating at least a part of the plurality of first frame images as query frames. Then, the query frame selection unit 112 selects the first frame image designated by the user as a query frame.
[0110] - Selection Process 2 - In Selection Process 2, the query frame selection unit 112 selects query frames according to a predetermined rule.
[0111] Specifically, as shown in FIG. 41, the query frame selection unit 112 selects a plurality of query frames from among a plurality of first frame images at a predetermined fixed interval. That is, the query frame selection unit 112 selects a query frame every M frames. Examples of M include, but are not limited to, values from 2 to 10. M may be predetermined or may be selectable by the user.
[0112] - Selection Process 3 - In Selection Process 3, the query frame selection unit 112 selects query frames according to a predetermined rule.
[0113] Specifically, as shown in FIG. 42, after selecting one query frame, the query frame selection unit 112 calculates the similarity between that query frame and each of the first frame images in the time series order after that query frame. The similarity is the same concept as in Embodiments 1 and 2. Then, the query frame selection unit 112 selects, as a new query frame, the first frame image with a similarity that is less than or equal to a reference value and that is the earliest in the time series order.
[0114] Next, the query frame selection unit 112 calculates the similarity between the newly selected query frame and each of the first frame images in chronological order after that query frame. Then, the query frame selection unit 112 selects, as a new query frame, the first frame image whose similarity is equal to or less than the reference value and whose chronological order is the earliest. The query frame selection unit 112 repeats this process to select query frames. According to this process, the postures of the people included in adjacent query frames are somewhat different from each other. Therefore, it is possible to select a plurality of query frames showing characteristic postures of people while suppressing an increase in the number of query frames. The above reference value may be determined in advance, may be selectable by the user, or may be set by other means.
[0115] Returning to FIG. 40, the skeleton structure detection unit 102 detects the key points of the people (objects) included in each of the plurality of first frame images. The skeleton structure detection unit 102 may target only the query frames for the detection process, or may target all of the first frame images for the detection process. Since the configuration of the skeleton structure detection unit 102 is the same as that in Embodiments 1 and 2, detailed description thereof is omitted here.
[0116] The feature amount calculation unit 103 calculates, for each first frame image, the feature amount of the detected key points, that is, the feature amount of the detected two-dimensional skeleton structure. The feature amount calculation unit 103 may target only the query frames for the calculation process, or may target all of the first frame images for the calculation process. Since the configuration of the feature amount calculation unit 103 is the same as that in Embodiments 1 and 2, detailed description thereof is omitted here.
[0117] The change calculation unit 110 calculates the direction of change in the feature amount along the time axis of a plurality of time-series first frame images. The change calculation unit 110 calculates, for example, the direction of change in the feature amount between adjacent query frames. The feature amount is the feature amount calculated by the feature amount calculation unit 103. The feature amount is, for example, the height or area of the skeleton region and is expressed as a numerical value. The direction of change in the feature amount is divided into, for example, three types: "the direction in which the numerical value increases", "no change in the numerical value", and "the direction in which the numerical value decreases". "No change in the numerical value" may be the case where the absolute value of the change amount of the feature amount is 0, or may be the case where the absolute value of the change amount of the feature amount is less than or equal to the threshold value.
[0118] An example will be described with reference to FIG. 43. Comparing the images before and after the change shown in the figure, there is a difference in that the right arm that was lowered before the change is raised after the change. For example, the angle P formed by the keypoint A2, the keypoint A31, and the keypoint A41 is calculated as a feature amount. And in this case, the change calculation unit 110 determines the direction in which the numerical value increases as the direction of change in the feature amount along the time axis.
[0119] When three or more query frames are the processing targets, the change calculation unit 110 can calculate time-series data indicating the time-series change in the direction of change in the feature amount. The time-series data is, for example, like "the direction in which the numerical value increases" → "the direction in which the numerical value increases" → "the direction in which the numerical value increases" → "no change in the numerical value" → "no change in the numerical value" → "the direction in which the numerical value increases". Representing "the direction in which the numerical value increases" as, for example, "1", "no change in the numerical value" as, for example, "0", and "the direction in which the numerical value decreases" as, for example, "-1", the time-series data can be represented as a numerical sequence such as "111001".
[0120] When only two query frames are the processing targets, the change calculation unit 110 can calculate the direction of change in the feature amount that occurred between the two images.
[0121] Returning to FIG. 40, the search unit 111 searches for a video using, as a key, the direction of change in the feature amount calculated by the change calculation unit 110. Specifically, the search unit 111 searches for a DB video that matches the key from among the videos stored in the database 201 (hereinafter referred to as "DB videos"). The search unit 111 can execute, for example, either of the following video search processes 1 and 2.
[0122] - Video Search Process 1 - When using the time-series data of the direction of change in the feature amount as a key, the search unit 111 can search for a DB video whose similarity of the time-series data is equal to or higher than a reference value. The method for calculating the similarity of the time-series data is not particularly limited, and any technique can be adopted.
[0123] Note that, in advance, the above time-series data may be created for each of the DB videos stored in the database 201 by the same method as above and stored in the database. Alternatively, each time the search process is performed, the search unit 111 may process each of the DB videos stored in the database 201 by the same method as above to create the above time-series data for each DB video.
[0124] - Video Search Process 2 - When using, as a key, the direction of change in the feature amount that occurs between two query frames, the search unit 111 can search for a DB video that indicates the direction of change in the feature amount.
[0125] Note that, in advance, index data of the direction of change in the feature amount shown in each DB video may be created for each of the DB videos stored in the database 201 and stored in the database. Alternatively, each time the search process is performed, the search unit 111 may process each of the DB videos stored in the database 201 by the same method as above to create index data of the direction of change in the feature amount shown in each DB video for each DB video.
[0126] Next, an example of the processing flow of the image processing apparatus 100 will be described with reference to FIG. 44. Here, the purpose is to explain the processing flow. Since the details of each process have been described above, the description here will be omitted.
[0127] When the image processing apparatus 100 acquires a query video composed of a plurality of first frame images in time series (S400), it selects at least a part of the plurality of first frame images as query frames (S401).
[0128] Next, the image processing apparatus 100 detects the key points of the objects included in each of the plurality of first frame images (S402). Note that only the query frames selected in S401 may be the target of this process, or all the first frame images may be the target of this process.
[0129] Next, the image processing apparatus 100 calculates the feature amounts of the detected key points for each of the plurality of first frame images (S403). Note that only the query frames selected in S401 may be the target of this process, or all the first frame images may be the target of this process.
[0130] Next, the image processing apparatus 100 calculates the direction of change of the above feature amounts along the time axis of the plurality of first frame images in time series (S404). The image processing apparatus 100 calculates the direction of change of the feature amounts between adjacent query frames. The direction of change is divided into three types, for example, "the direction in which the numerical value increases", "no change in the numerical value", and "the direction in which the numerical value decreases".
[0131] When three or more query frames are the processing targets, the image processing apparatus 100 can calculate time series data indicating the time series change of the direction of change of the feature amounts. When only two query frames are the processing targets, the image processing apparatus 100 can calculate the direction of change of the feature amounts that occurred between the two images.
[0132] Next, the image processing apparatus 100 searches for the DB video using, as a key, the direction of change in the feature amount calculated in S404 (S405). Specifically, the image processing apparatus 100 searches for a DB video that matches the key from among the DB videos stored in the database 201. Then, the image processing apparatus 100 outputs the search result. The output of the search result can be realized by adopting any technique.
[0133] Here, a modification example of the present embodiment will be described. The image processing apparatus 100 of the present embodiment can be configured to adopt one or more of the following modification examples 1 to 7.
[0134] -Modification Example 1- As shown in the functional block diagram of FIG. 45, the image processing apparatus 100 may not have the query frame selection unit 112. In this case, the change calculation unit 110 can calculate the direction of change in the feature amount between adjacent first frame images. Then, when three or more first frame images are the processing targets, the change calculation unit 110 can calculate time-series data indicating the time-series change in the direction of change in the feature amount. When only two first frame images are the processing targets, the change calculation unit 110 can calculate the direction of change in the feature amount that occurred between the two images.
[0135] Next, an example of the processing flow of the image processing apparatus 100 in this modification example will be described with reference to FIG. 46. Here, it is for the purpose of explaining the processing flow. Since the details of each process have been described above, the description here will be omitted.
[0136] The image processing apparatus 100 acquires a query video composed of a plurality of first frame images in time series (S300). Next, the image processing apparatus 100 detects the key points of the objects included in each of the plurality of first frame images (S301). Next, the image processing apparatus 100 calculates the feature amounts of the detected key points for each of the plurality of first frame images (S302).
[0137] Next, the image processing apparatus 100 calculates the direction of change of the feature amount along the time axis of a plurality of first frame images in time series (S303). Specifically, the image processing apparatus 100 calculates the direction of change of the feature amount between adjacent first frame images.
[0138] Next, the image processing apparatus 100 searches for the DB video using the direction of change of the feature amount calculated in S303 as a key (S304). Specifically, the image processing apparatus 100 searches for a DB video that matches the key from among the DB videos stored in the database 201. Then, the image processing apparatus 100 outputs the search result. The output of the search result can be realized by adopting any technique.
[0139] -Modification Example 2- In the above embodiment, the image processing apparatus 100 detects the key points of a person's body and searches for the DB video using the direction of change thereof as a key. In Modification Example 2, the image processing apparatus 100 can detect the key points of an object other than a person and search for the DB video using the direction of change thereof as a key. The object is not particularly limited, and examples include animals, plants, natural substances, and artificial objects.
[0140] -Modification Example 3- In addition to the direction of change of the feature amount, the change calculation unit 110 can calculate the magnitude of the change of the feature amount. The change calculation unit 110 can calculate the magnitude of the change of the feature amount between adjacent query frames or between adjacent first frame images. The magnitude of the change of the feature amount can be represented, for example, by the absolute value of the difference between the numerical values indicating the feature amount. Alternatively, the magnitude of the change of the feature amount may be a value obtained by normalizing the absolute value.
[0141] When three or more images (query frames or first frame images) are the processing targets, in addition to the direction of change of the feature amount, the change calculation unit 110 can calculate time series data further indicating the time series change of the magnitude of the change.
[0142] When only two images (query frame or first frame image) are to be processed, the change calculation unit 110 can calculate the direction and magnitude of the change in feature amounts that occurred between the two images.
[0143] The search unit 111 searches the DB video using the direction and magnitude of the change calculated by the change calculation unit 110 as keys.
[0144] When using the time-series data of the direction and magnitude of the change in feature amounts as keys, the search unit 111 can search for DB videos whose similarity of the time-series data is equal to or higher than a reference value. The method for calculating the similarity of the time-series data is not particularly limited, and any technique can be adopted.
[0145] When using the direction and magnitude of the change in feature amounts that occurred between two images (query frame or first frame image) as keys, the search unit 111 can search for DB videos that show the direction and magnitude of the change in the feature amounts.
[0146] -Variant Example 4- In addition to the direction of the change in feature amounts, the change calculation unit 110 can calculate the speed of the change in feature amounts. This variant example is effective when query frames are selected at random intervals from the first frame image as shown in FIG. 42 and the direction of the change in feature amounts is calculated between adjacent query frames. In this case, it becomes possible to search for more similar DB videos by referring to the speed of the change in feature amounts between adjacent query frames.
[0147] The change calculation unit 110 can calculate the speed of the change in feature amounts between adjacent query frames. The speed can be calculated by dividing the magnitude of the change in feature amounts by a value indicating the magnitude of the time between adjacent query frames (such as the number of frames or a value converted to time based on the frame rate). The magnitude of the change in feature amounts can be represented, for example, by the absolute value of the difference in numerical values indicating the feature amounts. Alternatively, the magnitude of the change in feature amounts may be a value obtained by normalizing the absolute value.
[0148] When three or more query frames are to be processed, the change calculation unit 110 can calculate time-series data further indicating the speed of change in addition to the direction of change in feature amounts.
[0149] When only two query frames are to be processed, the change calculation unit 110 can calculate the direction and speed of change in feature amounts that occurred between the two images.
[0150] The search unit 111 searches the DB video using the direction of change and the speed of change calculated by the change calculation unit 110 as keys.
[0151] When using the time-series data of the direction and speed of change in feature amounts as keys, the search unit 111 can search for DB videos in which the similarity of the time-series data is equal to or higher than a reference value. The method for calculating the similarity of the time-series data is not particularly limited, and any technique can be adopted.
[0152] When using the direction and speed of change in feature amounts that occurred between two query frames as keys, the search unit 111 can search for DB videos indicating the direction and speed of change in the feature amounts.
[0153] - Modification Example 5 - So far, the search unit 111 has searched for DB videos that match the keys, but it may also search for DB videos that do not match the keys. That is, the search unit 111 may search for DB videos in which the similarity with the above time-series data as the key is less than the reference value. Further, the search unit 111 may search for DB videos that do not include the direction of change in feature amounts (which may include magnitude, speed, etc.) as the key.
[0154] Further, the search unit 111 may search for DB videos that match the search conditions in which a plurality of keys are connected by an arbitrary logical operator.
[0155] - Modification Example 6 - The search unit 111 uses the result calculated by the change calculation unit 110 (special SignsIn addition to the direction, magnitude, speed, etc. of the quantity change, the DB video can be retrieved by further using a representative image selected from the first frame image of the query video as a key. There may be one or more representative images. For example, the query frame may be used as the representative image, or a frame arbitrarily selected from the query frames may be used as the representative image, or the representative image may be selected from the first frame image by other means.
[0156] The search unit 111 can search for a DB video whose total similarity obtained by integrating the similarity with the query video calculated based on the representative image and the similarity with the query video calculated based on the result (especially the direction, magnitude, speed, etc. of the quantity change) calculated by the change calculation unit 110 is equal to or higher than a reference value from among the DB videos stored in the database 201. Signs Here, a method for calculating the similarity based on the representative image will be described. The search unit 111 can calculate the similarity between each DB video and the query video based on the following criteria.
[0157] · Increase the similarity of a DB video that includes a frame image whose similarity with the representative image is equal to or higher than a reference value.
[0158] · When there are multiple representative images, increase the similarity of a DB video that includes more frame images similar (similarity is equal to or higher than the reference value) to the multiple representative images. · When there are multiple representative images, the more similar the time series order of the multiple representative images is to the time series order of the frame images similar to each of the multiple representative images, the higher the similarity of the DB video.
[0159] The similarity between the representative image and the frame image is calculated based on the postures of the people included in each image. The more similar the postures are, the higher the similarity between the representative image and the frame image. The search unit 111 may calculate the similarity of the feature amounts of the skeletal structure described in the above embodiment as the similarity between the representative image and the frame image, or may calculate the similarity of the postures of people using other well-known techniques.
[0160] Next, the result calculated by the change calculation unit 110 (particularly Signs A method for calculating similarity based on time series data of the direction of feature change (which may further indicate the magnitude and speed of the feature change) will be described. When using time series data of the direction of feature change (which may further indicate the magnitude and speed of the feature change), the similarity of that time series data can be calculated as the similarity between each DB video and the query video.
[0161] When using the direction, magnitude, or speed of change in features between two query frames, the similarity of the DB video is increased if it shows the same direction of change as the query video and the magnitude and speed of change are similar to those shown in the query video.
[0162] The similarity based on the representative image and the result calculated by the change calculation unit 110 (particularly Signs There are various methods for integrating similarities based on the representative image and similarities based on the change calculation unit 110 (such as the direction, magnitude, and speed of the change in the amount). For example, each similarity may be normalized and added together. In this case, each similarity may be weighted. That is, the similarity based on the representative image or its normalized value multiplied by a predetermined weighting coefficient is integrated with the result calculated by the change calculation unit 110 (such as the direction, magnitude, and speed of the change in the amount). Signs Alternatively, the integration result may be calculated by adding a value obtained by multiplying a predetermined weighting coefficient to the similarity based on the amount of change (such as the direction, magnitude, or speed of the change in amount) or a standardized value thereof.
[0163] -Variation 7- As in the first and second embodiments, the image processing device 100 may constitute an image processing system 1 together with a camera 200 and a database 201 .
[0164] As described above, the image processing device 100 of this embodiment achieves the same effects as those of the first and second embodiments. Moreover, the image processing device 100 of this embodiment can search for a video using the direction of change in the posture of an object included in the image, the magnitude of the change, the speed of the change, and the like as keys. The image processing device 100 of this embodiment can accurately search for a video including a desired scene.
[0165] The embodiments of the present invention have been described above with reference to the drawings, but these are examples of the present invention, and various configurations other than the above can also be adopted.
[0166] Also, in the plurality of flowcharts used in the above description, a plurality of steps (processes) are described in order, but the execution order of the steps executed in each embodiment is not limited to the described order. In each embodiment, the order of the illustrated steps can be changed within a range that does not interfere with the content. Also, the above-described embodiments can be combined within a range where the contents do not conflict.
[0167] Some or all of the above embodiments can also be described as follows in the appended claims, but are not limited thereto. 1. Query acquisition means for acquiring a plurality of first frame images in time series, Skeleton structure detection means for detecting key points of an object included in each of the plurality of first frame images, Feature amount calculation means for calculating a feature amount of the detected key points for each of the first frame images, Change calculation means for calculating a direction of change of the feature amount along a time axis of the plurality of first frame images in time series, Search means for searching for a moving image using the calculated direction of change of the feature amount as a key, An image processing apparatus having the same. 2. The change calculation means further calculates a magnitude of the change, The search means searches for a moving image using the calculated magnitude of the change as a further key. The image processing apparatus according to 1. 3. The change calculation means further calculates a speed of the change, The search means searches for a moving image using the calculated speed of the change as a further key. The image processing apparatus according to 1 or 2. 4. The search means searches for a moving image using a representative image among the plurality of first frame images as a further key. The image processing apparatus according to any one of 1 to 3. 5. The image processing apparatus according to claim 4, wherein the search means searches for a moving image using the feature amount calculated from the representative image. 6. A computer A query acquisition step of acquiring a plurality of first frame images in time series, A skeleton structure detection step of detecting key points of an object included in each of the plurality of first frame images, A feature amount calculation step of calculating a feature amount of the detected key points for each of the first frame images, A change calculation step of calculating a direction of change of the feature amounts along the time axis of the plurality of first frame images in time series, A search step of searching for a moving image using the calculated direction of change of the feature amounts as a key, An image processing method for executing the above. 7. A computer Query acquisition means for acquiring a plurality of first frame images in time series, Skeleton structure detection means for detecting key points of an object included in each of the plurality of first frame images, Feature amount calculation means for calculating a feature amount of the detected key points for each of the first frame images, Change calculation means for calculating a direction of change of the feature amounts along the time axis of the plurality of first frame images in time series, and Search means for searching for a moving image using the calculated direction of change of the feature amounts as a key, A program for causing the computer to function as the above.
Explanation of Signs
[0168] 1 Image processing system 10 Image processing apparatus 11 Skeleton detection unit 12 Feature amount calculation unit 13 Recognition unit 100 Image processing apparatus 101 Image acquisition unit 102 Skeleton structure detection unit 103 Feature amount calculation unit 104 Classification unit 105 Search unit 106 Input section 107 Display section 108 Height calculation section 109 Query acquisition section 110 Change calculation section 111 Search section 112 Query frame selection section 200 Camera 201 Database 300, 301 Human body model 401 Two-dimensional bone structure
Claims
1. Query acquisition means for acquiring a plurality of first frame images in time series, Skeleton structure detection means for detecting key points of an object included in each of the plurality of first frame images, Feature amount calculation means for calculating a feature amount of the detected key points for each of the first frame images, Change calculation means for calculating a direction of change of the feature amount along the time axis of the plurality of first frame images in time series, Search means for searching for a video using the calculated direction of change of the feature amount as a key, which has, The feature amount of the key point is an image processing apparatus showing the feature of a human skeleton.
2. The change calculation means further calculates the magnitude of the change, The search means is the image processing apparatus according to claim 1, wherein the search means further uses the calculated magnitude of the change as a key to search for a video.
3. The change calculation means further calculates the speed of the change, The search means is the image processing apparatus according to claim 1 or 2, wherein the search means further uses the calculated speed of the change as a key to search for a video.
4. The search means is the image processing apparatus according to any one of claims 1 to 3, wherein the search means further uses a representative image among the plurality of first frame images as a key to search for a video.
5. The search means is the image processing apparatus according to claim 4, wherein the search means searches for a video using the feature amount calculated from the representative image.
6. The feature amount of the key point is The area of the skeleton region including the key point, The height of the skeleton region including the key point, or A value indicating the relative positional relationship in the height direction between a reference point and each of a plurality of key points, is the image processing apparatus according to any one of claims 1 to 5.
7. A computer A query acquisition step of acquiring a plurality of first frame images in time series, A skeleton structure detection step of detecting key points of an object included in each of the plurality of first frame images, A feature amount calculation step of calculating a feature amount of the detected key points for each of the first frame images, A change calculation step of calculating a direction of change of the feature amount along the time axis of the plurality of first frame images in time series, A search step of searching for a video using the calculated direction of change of the feature amount as a key, executes, The feature amount of the key point is an image processing method showing the feature of a human skeleton.
8. The computer In the change calculation step, the magnitude of the change is further calculated, The image processing method according to claim 7, wherein in the search step, a video is searched using the calculated magnitude of the change as a further key.
9. A computer, Query acquisition means for acquiring a plurality of first frame images in time series, Skeleton structure detection means for detecting key points of an object included in each of the plurality of first frame images, Feature amount calculation means for calculating a feature amount of the detected key points for each of the first frame images, Change calculation means for calculating a direction of change of the feature amounts along a time axis of the plurality of first frame images in time series, and Search means for searching a video using the calculated direction of change of the feature amounts as a key, functioning as, wherein the feature amount of the key point is a program showing features of a human skeleton.
10. The change calculation means further calculates the magnitude of the change, The search means is the program according to claim 9, wherein the search means searches a video using the calculated magnitude of the change as a further key.
Citation Information
Patent Citations
Methods for establishing an action recognition library, electronic devices, and storage media
CN109308438B
Image information updating system, image inputting device, image processing device, image updating device, image information updating method, image information updating program, and recording medium recording the program
JP2006260405A
Game device, control method for the game device, and program
JP2011194073A
Apparatus and method for retrieving object poses
JP2014522035A
Action analysis device and action analysis method
JP2020135747A