Search device, search method, and program
The search device enhances video similarity searches by extracting keyframes and analyzing human body postures and time intervals, improving the accuracy of finding matching movements.
Patent Information
- Application Number
- JP2023561978
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-17
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2041-11-17
AI Technical Summary
Existing technologies struggle to accurately search for videos that include human bodies moving similarly to the movement of a human body indicated in a query.
A search device and method that extracts keyframes from a query video, analyzing human body postures and time intervals between keyframes to find similar videos based on posture similarity and time interval matching.
Improves the accuracy of finding videos with human bodies moving similarly to the query by using keyframe extraction and posture/time interval analysis.
Smart Images

Figure 0007806807000001 
Figure 0007806807000002 
Figure 0007806807000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a search device, a search method, and a program. [Background technology]
[0002] Technologies related to the present invention are disclosed in Patent Document 1 and Non-Patent Document 1. Patent Document 1 discloses a technology for calculating feature amounts for each of multiple key points of a human body included in an image, and searching for still images including a human body with a posture similar to that of a human body indicated by a query based on the calculated feature amounts, or searching for videos including a human body with a movement similar to that of a human body indicated by a query. Non-Patent Document 1 also discloses a technology related to human skeleton estimation. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2021 / 084677 [Non-patent literature]
[0004] [Non-Patent Document 1] Zhe Cao, Tomas Simon, Shih-En Wei, Yaser Sheikh, "Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields", The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, P. 7291-7299 Summary of the Invention [Problem to be solved by the invention]
[0005] The present invention aims to improve the accuracy of searching for videos that include a human body that moves similarly to the movement of a human body indicated in a query. [Means for solving the problem]
[0006] According to the present invention, A keyframe extraction means for extracting a plurality of keyframes from the query video; a search means for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; A search device is provided, comprising:
[0007] Further, according to the present invention, The computer a keyframe extraction step of extracting a plurality of keyframes from the query video; a search process for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; A search method is provided that performs the following.
[0008] Further, according to the present invention, Computer, a keyframe extraction means for extracting a plurality of keyframes from the query video; a search means for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; A program is provided to function as a [Effects of the Invention]
[0009] According to the present invention, the accuracy of searching for videos that include a human body that moves similarly to the movement of a human body indicated in a query is improved. [Brief explanation of the drawings]
[0010] The above and other objects, features and advantages are described below. Suitable This will become more apparent from the following embodiments and the accompanying drawings.
[0011] [Figure 1] FIG. 10 is a diagram for explaining a process of extracting a key frame according to the present embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a hardware configuration of a search device according to the present embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a functional block diagram of the search device according to the present embodiment. [Figure 4] FIG. 10 is a diagram for explaining a process of extracting a key frame according to the present embodiment. [Figure 5] 3A to 3C are diagrams for explaining corresponding frames, time intervals between a plurality of key frames, and time intervals between a plurality of corresponding frames according to the present embodiment. [Figure 6] 10 is a flowchart illustrating an example of a processing flow of the search device of the present embodiment. [Figure 7] FIG. 2 is a diagram illustrating an example of a functional block diagram of the search device according to the present embodiment. [Figure 8] 1 is a diagram showing an example of a skeletal structure of a human body model detected by the search device of this embodiment. [Figure 9] FIG. 2 is a diagram showing an example of a skeletal structure of a human body model detected by the search device of this embodiment. [Figure 10] FIG. 2 is a diagram showing an example of a skeletal structure of a human body model detected by the search device of this embodiment. [Figure 11] FIG. 2 is a diagram showing an example of a skeletal structure of a human body model detected by the search device of this embodiment. [Figure 12] FIG. 10 is a diagram illustrating an example of feature amounts of key points calculated by the search device of the present embodiment. [Figure 13] FIG. 10 is a diagram illustrating an example of feature amounts of key points calculated by the search device of the present embodiment. [Figure 14] FIG. 10 is a diagram illustrating an example of feature amounts of key points calculated by the search device of the present embodiment. [Figure 15] 10 is a flowchart illustrating an example of a processing flow of the search device of the present embodiment. [Figure 16]10 is a flowchart illustrating an example of a processing flow of the search device of the present embodiment. [Figure 17] 10A and 10B are diagrams illustrating an example of a method in which a user specifies the weight of the similarity between the postures of the human body and the weight of the similarity between the time interval between key frames and the time interval between corresponding frames in this embodiment. [Figure 18] 10A and 10B are diagrams illustrating an example of a method in which a user specifies the weight of the similarity between the postures of the human body and the weight of the similarity between the time interval between key frames and the time interval between corresponding frames in this embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, like components are designated by like reference numerals, and the description thereof will be omitted as appropriate.
[0013] First Embodiment "overview" As shown in Figure 1, the search device of this embodiment extracts multiple key frames from a query video, and then searches for videos containing human bodies moving in a similar manner to the human body movements (time changes in the human body posture) shown in the query video based on the human body posture contained in each of the multiple key frames and the time intervals between the multiple key frames.
[0014] As described above, the search device of this embodiment has a feature of searching for videos based on two elements: the posture of the human body included in each of the plurality of key frames, and the time interval between the plurality of key frames.
[0015] "Hardware Configuration" Next, an example of the hardware configuration of a search device will be described. Each functional unit of the search device is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, programs loaded into the memory, a storage unit such as a hard disk that stores the programs (this can store programs that are pre-loaded when the device is shipped, as well as programs downloaded from storage media such as CDs (Compact Discs) or servers on the Internet), and a network connection interface. Those skilled in the art will understand that there are many variations in the realization methods and devices.
[0016] FIG. 2 is a block diagram illustrating an example of the hardware configuration of a search device. As shown in FIG. 2, the search device has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The search device does not necessarily have to have the peripheral circuit 4A. The search device may also be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices can have the above hardware configuration.
[0017] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data among them. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. Examples of input devices include a keyboard, a mouse, a microphone, physical buttons, a touch panel, etc. Examples of output devices include a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0018] "Function Configuration" 3 shows an example of a functional block diagram of the search device 10 of this embodiment. The search device 10 shown in the figure has a key frame extraction unit 11 and a search unit 12.
[0019] The key frame extraction unit 11 extracts a plurality of key frames from the query moving image.
[0020] A "query video" is a video that serves as a search query. The search device 10 searches for videos that include a human body that moves similarly to the movement of the human body indicated in the query video. A single video file may be specified as the query video, or a partial scene within a single video file may be specified as the query video. For example, the user specifies the query video. The specification of the query video can be realized using any technology.
[0021] A "key frame" is a frame that is part of a plurality of frames included in a query video. As shown in Figures 1 and 4, the key frame extraction unit 11 can intermittently extract key frames from a plurality of time-series frames included in a query video. The time intervals between key frames (the number of frames) may be constant or may vary. The key frame extraction unit 11 can execute, for example, any of the following extraction processes 1 to 3.
[0022] -Extraction process 1- In extraction process 1, the key frame extraction unit 11 extracts key frames based on user input. That is, the user inputs a part of a plurality of frames included in the query video to be designated as a key frame. Then, the key frame extraction unit 11 extracts the frame designated by the user as a key frame.
[0023] -Extraction process 2- In the extraction process 2, the key frame extraction unit 11 extracts key frames according to a predetermined rule.
[0024] Specifically, as shown in Fig. 1, the key frame extraction unit 11 extracts a plurality of key frames at predetermined intervals from a plurality of frames included in the query video. That is, the key frame extraction unit 11 extracts a key frame every M frames. M is an integer, and is, for example, between 2 and 10, but is not limited to this. M may be predetermined or may be selectable by the user.
[0025] -Extraction process 3- In the extraction process 3, the key frame extraction unit 11 extracts key frames according to a predetermined rule.
[0026] Specifically, as shown in FIG. 4, after extracting one key frame (for example, the first frame), the key frame extraction unit 11 calculates the similarity between that key frame and each of the frames following that key frame in chronological order. The similarity is the similarity between the postures of the human body included in each frame. There is no particular limitation on the method for calculating the posture similarity, but an example will be described in the following embodiment. Then, the key frame extraction unit 11 extracts, as a new key frame, the frame whose similarity is equal to or less than a reference value (design item) and which is the earliest in chronological order.
[0027] Next, the keyframe extraction unit 11 calculates the similarity between the newly extracted keyframe and each of the frames following it in chronological order. The keyframe extraction unit 11 then extracts the frame whose similarity is equal to or less than a reference value (design item) and which is the earliest in chronological order as a new keyframe. The keyframe extraction unit 11 repeats this process to extract multiple keyframes. According to this process, the postures of the human body included in adjacent keyframes differ to some extent. Therefore, it is possible to extract multiple keyframes showing characteristic postures of the human body while suppressing an increase in the number of keyframes. The reference value may be predetermined, may be user-selectable, or may be set by other means.
[0028] Returning to Fig. 3, the search unit 12 searches for videos similar to the query video based on the human body postures included in each of the multiple key frames extracted by the key frame extraction unit 11 and the time intervals between the multiple key frames. The search for videos by the search unit 12 may be to search for a scene similar to the query video from one video file, or to search for a video file including a scene similar to the query video from multiple video files, or may be something else.
[0029] Specifically, the search unit 12 searches for videos that satisfy the following conditions 1 and 2 as videos similar to the query video. Note that the search unit 12 may search for videos that further satisfy the following condition 3 in addition to the following conditions 1 and 2.
[0030] (Condition 1) A plurality of corresponding frames corresponding to each of a plurality of key frames are included. (Condition 2) The time intervals between a plurality of corresponding frames are similar to the time intervals between a plurality of key frames at a predetermined level or more. (Condition 3) The order in which multiple key frames appear in the query video matches the order in which multiple corresponding frames appear in the video.
[0031] Each condition is explained below.
[0032] -(Condition 1) Multiple corresponding frames are included, each corresponding to a multiple number of key frames. A corresponding frame is a frame containing a human body in a posture similar to the posture of the human body contained in the key frame at a predetermined level or more. The method for calculating posture similarity is not particularly limited, but an example will be described in the following embodiment. When Q (Q is an integer equal to or greater than 2) key frames are extracted from a query video, a video containing Q corresponding frames corresponding to each of the Q key frames will satisfy condition 1.
[0033] Figure 5 shows a query video consisting of 10 frames. In the figure, the first, fourth, sixth, eighth, and tenth frames marked with stars are extracted as key frames. Hereinafter, the Nth key frame in chronological order among multiple key frames will be referred to as the "Nth key frame." N is an integer greater than or equal to 1. In the example of Figure 5, the first frame will be referred to as the first key frame, the fourth frame as the second key frame, the sixth frame as the third key frame, the eighth frame as the fourth key frame, and the tenth frame as the fifth key frame.
[0034] In the example of FIG. 5, a video including five corresponding frames corresponding to the first to fifth key frames respectively satisfies condition 1. Incidentally, the video to be processed in FIG. 5 is a video that satisfies condition 1. The video to be processed is composed of 12 frames. In the figure, the first, third, seventh, eighth, and twelfth frames marked with a star are identified as corresponding frames. Hereinafter, the corresponding frame corresponding to the Nth key frame will be referred to as the "Nth corresponding frame." The first frame of the video to be processed is the first corresponding frame, the third frame is the second corresponding frame, the seventh frame is the third corresponding frame, the eighth frame is the fourth corresponding frame, and the twelfth frame is the fifth corresponding frame.
[0035] -(Condition 2) The time interval between multiple corresponding frames is similar to the time interval between multiple key frames at a predetermined level or more. First, the concepts of "time interval between a plurality of corresponding frames" and "time interval between a plurality of key frames" will be explained using FIG.
[0036] In the illustrated example, the time intervals between the plurality of corresponding frames are the time intervals between the first to fifth corresponding frames.
[0037] For example, the time intervals between multiple corresponding frames may be a concept that includes time intervals between temporally adjacent corresponding frames. In the example of Figure 5, the time intervals between temporally adjacent corresponding frames are the time interval between the first and second corresponding frames, the time interval between the second and third corresponding frames, the time interval between the third and fourth corresponding frames, and the time interval between the fourth and fifth corresponding frames.
[0038] Alternatively, the time interval between multiple corresponding frames may also include the time interval between the first and last corresponding frames in terms of time. In the example of Figure 5, the time interval between the first and last corresponding frames in terms of time is the time interval between the first and fifth corresponding frames.
[0039] Alternatively, the time intervals between multiple corresponding frames may be a concept that includes the time intervals between a reference corresponding frame determined by any method and each of the other corresponding frames. In the example of Figure 5, for example, if the first corresponding frame is the reference corresponding frame, the time intervals between the reference corresponding frame and each of the other corresponding frames are the time interval between the first and second corresponding frames, the time interval between the first and third corresponding frames, the time interval between the first and fourth corresponding frames, and the time interval between the first and fifth corresponding frames. Note that there may be one or more reference corresponding frames.
[0040] The "time interval between a plurality of corresponding frames" may be any one of the above-described types of time intervals, or may include a plurality of types. It is defined in advance which of the above-described types of time intervals will be the time interval between the plurality of corresponding frames. In the example of FIG. 5, the time interval between the plurality of corresponding frames is one or more of the following: the time interval between the first and second corresponding frames, the time interval between the second and third corresponding frames, the time interval between the third and fourth corresponding frames, and the time interval between the fourth and fifth corresponding frames (all of which are time intervals between temporally adjacent corresponding frames), the time interval between the first and fifth corresponding frames (all of which are time intervals between the first and last corresponding frames), the time interval between the first and second corresponding frames, the time interval between the first and third corresponding frames, the time interval between the first and fourth corresponding frames, and the time interval between the first and fifth corresponding frames (all of which are examples of time intervals between a reference corresponding frame and each of the other corresponding frames).
[0041] The concept of the time interval between multiple key frames is similar to the concept of the time interval between multiple corresponding frames described above.
[0042] The time interval between two frames may be indicated by the number of frames between the two frames, or may be indicated by the elapsed time between the two frames calculated based on the number of frames between the two frames and the frame rate.
[0043] Next, we will explain the concept of "the time interval between a plurality of corresponding frames is similar to the time interval between a plurality of key frames at a predetermined level or more." Here, we will explain separately the case where the time interval between a plurality of corresponding frames and the time interval between a plurality of key frames is one of the above-mentioned multiple types of time intervals, and the case where there are multiple types.
[0044] (When the time interval between multiple corresponding frames and the time interval between multiple key frames are one type of time interval) In this case, a state in which the difference between one type of time interval between multiple corresponding frames and one type of time interval between multiple key frames is equal to or less than a threshold is defined as a state in which the time interval between multiple corresponding frames is similar to the time interval between multiple key frames at a predetermined level or more. The threshold is a design factor and is set in advance. The "difference between time intervals" refers to the difference or rate of change.
[0045] As an example, the time interval between the first and last corresponding frames in time and the time interval between the first and last corresponding frames in time key For example, a state in which the difference between the time intervals between frames is equal to or less than a threshold value can be defined as a state in which the time intervals between multiple corresponding frames are similar to the time intervals between multiple key frames at a predetermined level or more. Note that, although the "time interval between multiple corresponding frames" is defined here as the "time interval between the first and last corresponding frames in terms of time" and the "time interval between multiple key frames" is defined as the "time interval between the first and last key frames in terms of time," this is merely an example and is not limiting.
[0046] (When the time intervals between multiple corresponding frames and the time intervals between multiple key frames include multiple types of time intervals) In this case, for each of the multiple types of time intervals, Multiple It is determined whether the difference between the time intervals between corresponding frames and the time intervals between multiple key frames is equal to or less than a threshold. The threshold is a design factor and is set in advance for each type of time interval. A state in which the difference is equal to or less than the threshold for a predetermined percentage or more of the multiple types of time intervals is defined as a state in which the time intervals between multiple corresponding frames are similar to the time intervals between multiple key frames at a predetermined level or more.
[0047] -(Condition 3) The order in which multiple key frames appear in the query video matches the order in which multiple corresponding frames appear in the video. Condition 3 is the first to Qth key frames extracted from the query video and the first to Qth corresponding frames in the video. Appearance A video in which the first to Qth corresponding frames appear in this order satisfies this condition, whereas a video in which the first to Qth corresponding frames do not appear in this order does not satisfy this condition.
[0048] Next, an example of the flow of processing by the search device 10 will be described with reference to the flowchart of FIG.
[0049] first, search The device 10 extracts a plurality of key frames from the query video (S10). search The device 10 searches for videos similar to the query video based on the posture of the human body included in each of the extracted key frames and the time interval between the extracted key frames (S11).
[0050] "Action and effect" As shown in Figure 1, the search device 10 of this embodiment extracts multiple key frames from a query video, and then searches for videos containing human bodies moving in a similar manner to the human body movements (time changes in the human body posture) shown in the query video based on the human body postures contained in each of the multiple key frames and the time intervals between the multiple key frames.
[0051] Specifically, the search device 10 searches for a video that includes multiple corresponding frames corresponding to multiple key frames, and the time intervals between the multiple corresponding frames are similar to the time intervals between the multiple key frames. The corresponding frames are frames that include a human body in a pose similar to that of the human body included in the key frames.
[0052] According to such a search device 10, videos are searched for that include human bodies in poses similar to each of the multiple poses of the human body shown in the query video and that have a similar speed of change in the poses (interval between key frames). For example, as shown in Fig. 1, if the query video shows a human body raising its right hand, videos are searched for that include a human body raising its right hand and that have a speed of the right hand raising movement similar to that shown in the query video.
[0053] According to the search device 10 of this embodiment, the accuracy of searching for videos that include a human body that moves similarly to the movement of a human body shown in a query video is improved.
[0054] <Second embodiment> The search device 10 of this embodiment embodies a method for calculating the similarity of human body postures. Fig. 7 shows an example of a functional block diagram of the search device 10 of this embodiment. As shown in the figure, the search device 10 has a keyframe extraction unit 11, a skeletal structure detection unit 13, a feature calculation unit 14, and a search unit 12.
[0055] The skeletal structure detection unit 13 performs processing to detect N (N is an integer equal to or greater than 2) key points of the human body included in the key frame. This processing by the skeletal structure detection unit 13 is realized using the technology disclosed in Patent Document 1. Although details are omitted, the technology disclosed in Patent Document 1 detects the skeletal structure using a skeletal estimation technology such as OpenPose disclosed in Non-Patent Document 1. The skeletal structure detected by this technology is composed of "key points," which are characteristic points such as joints, and "bones (bone links)," which indicate the links between the key points.
[0056] Fig. 8 shows the skeletal structure of a human body model 300 detected by the skeletal structure detection unit 13, and Figs. 9 to 11 show examples of detected skeletal structures. The skeletal structure detection unit 13 detects the skeletal structure of a human body model (two-dimensional skeletal model) 300 as shown in Fig. 8 from a two-dimensional image using a skeletal estimation technique such as OpenPose. The human body model 300 is a two-dimensional model made up of key points such as a person's joints and bones connecting each key point.
[0057] Skeletal structure detection unit 13 For example, feature points that can be key points are extracted from an image, and N key points of the human body are detected by referring to information obtained by machine learning of the image of the key points. The N key points to be detected are determined in advance. The number of key points to be detected (i.e., the number N) and which parts of the human body are to be used as key points to be detected vary widely, and any variation can be adopted.
[0058] In the example of FIG. 8, the following key points of a person are detected: head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right hip A61, left hip A62, right knee A71, left knee A72, right foot A81, and left foot A82. Furthermore, the bones of the person connected by these key points are detected as bones: bone B1 connecting head A1 and neck A2; bone B21 and bone B22 connecting neck A2 to right shoulder A31 and left shoulder A32, respectively; bone B31 and bone B32 connecting right shoulder A31 and left shoulder A32 to right elbow A41 and left elbow A42, respectively; bone B41 and bone B42 connecting right elbow A41 and left elbow A42 to right hand A51 and left hand A52, respectively; bone B51 and bone B52 connecting neck A2 to right hip A61 and left hip A62, respectively; bone B61 and bone B62 connecting right hip A61 and left hip A62 to right knee A71 and left knee A72, respectively; and bone B71 and bone B72 connecting right knee A71 and left knee A72 to right foot A81 and left foot A82, respectively.
[0059] Fig. 9 shows an example of detecting a person standing upright. In Fig. 9, an image of a person standing upright is captured from the front, and bones B1, B51 and B52, B61 and B62, and B71 and B72 are detected without overlapping, and bones B61 and B71 of the right foot are slightly more bent than bones B62 and B72 of the left foot.
[0060] Fig. 10 shows an example of detecting a person who is crouching. In Fig. 10, the image of the person crouching is captured from the right side, and bones B1, B51 and B52, B61 and B62, and B71 and B72 are detected as seen from the right side, with bones B61 and B71 of the right foot and bones B62 and B72 of the left foot being significantly bent and overlapping.
[0061] Fig. 11 shows an example of detecting a person who is lying down. In Fig. 11, the person lying down is imaged from the diagonal front left, and bones B1, B51 and B52, B61 and B62, and B71 and B72 as seen from the diagonal front left are detected, with bones B61 and B71 of the right foot and bones B62 and B72 of the left foot being bent and overlapping.
[0062] 7, the feature amount calculation unit 14 calculates the feature amount of the detected two-dimensional skeletal structure. For example, the feature amount calculation unit 14 calculates the feature amount of each of the detected key points.
[0063] Skeletal structure features indicate the characteristics of a person's skeleton and are used to search for a person's state (posture and movement) based on the person's skeleton. Typically, these features include multiple parameters. The features may be the features of the entire skeletal structure, the features of a portion of the skeletal structure, or multiple features for each part of the skeletal structure. The feature calculation method may be any method, such as machine learning or normalization, and normalization may involve finding a minimum or maximum value. Examples of feature values include features obtained by machine learning of the skeletal structure, the size of the skeletal structure from the head to the feet on the image, the relative positions of multiple key points in the vertical direction of the skeletal region containing the skeletal structure on the image, and the relative positions of multiple key points in the horizontal direction of the skeletal region. The size of the skeletal structure is the vertical height or area of the skeletal region containing the skeletal structure on the image. The vertical direction (height direction or vertical direction) is the up-down direction (Y-axis direction) in the image, for example, the direction perpendicular to the ground (reference plane). The left-right direction (horizontal direction) is the left-right direction (X-axis direction) in the image, and is, for example, a direction parallel to the ground.
[0064] In order to perform the search desired by the user, it is preferable to use features that are robust to the search process. For example, if the user desires a search that is not dependent on the person's orientation or body shape, features that are robust to the person's orientation and body shape may be used. By learning the skeletons of people facing in various directions in the same posture or the skeletons of people with various body shapes in the same posture, or by extracting features only in the up-down direction of the skeleton, features that are not dependent on the person's orientation or body shape can be obtained.
[0065] The above processing by the feature amount calculation unit 14 is realized using the technology disclosed in Patent Document 1.
[0066] 12 shows an example of the feature amounts of each of a plurality of key points calculated by the feature amount calculation unit 14. Note that the feature amounts of the key points illustrated here are merely examples and are not limited to these.
[0067] In this example, the feature values of key points indicate the relative positional relationships of multiple key points in the vertical direction of the skeletal region containing the skeletal structure on the image. Because the neck key point A2 is used as the reference point, the feature value of key point A2 is 0.0. The feature values of right shoulder key point A31 and left shoulder key point A32, which are at the same height as the neck, are also 0.0. The feature value of head key point A1, which is higher than the neck, is -0.2. The feature values of right hand key point A51 and left hand key point A52, which are lower than the neck, are 0.4, and the feature values of right foot key point A81 and left foot key point A82 are 0.9. If the person raises their left hand from this position, the left hand will be higher than the reference point as shown in Figure 13, and the feature value of left hand key point A52 will be -0.4. However, because normalization is performed using only the Y-axis coordinate, the feature values do not change even if the width of the skeletal structure changes, as shown in Figure 14, compared to Figure 7. That is, the feature amount (normalized value) in this example indicates the feature in the height direction (Y direction) of the skeletal structure (keypoint), and is not affected by changes in the lateral direction (X direction) of the skeletal structure.
[0068] The search unit 12 calculates the similarity of the posture of the human body based on the feature quantities of the key points described above, and searches for videos similar to the query video based on the calculation result. As a search method, the technology disclosed in Patent Document 1 can be adopted.
[0069] Other configurations of the search device 10 of this embodiment are the same as those of the first embodiment.
[0070] As described above, the search device 10 of this embodiment achieves the same effects as those of the first embodiment. Furthermore, the search device 10 of this embodiment can identify the posture of a human body based on the feature quantities of the two-dimensional skeletal structure of the human body. The search device 10 of this embodiment can accurately identify the posture of a human body. As a result, the accuracy of searching for videos that include a human body that moves similarly to the movement of the human body shown in the query video is improved.
[0071] <Third embodiment> This embodiment embodies the flow of processing by the search unit 12. The flowchart in Fig. 15 shows an example of the flow of processing by the search unit 12 of this embodiment.
[0072] In S20, the search unit 12 searches for a video including Q corresponding frames corresponding to the Q key frames, respectively. The Nth corresponding frame corresponding to the Nth key frame includes a human body in a pose whose similarity to the pose of the human body included in the Nth key frame is equal to or greater than a first threshold.
[0073] In S21, the search unit 12 searches for a moving image, from among the moving images searched in S20, for which the similarity between the time intervals between a plurality of corresponding frames and the time intervals between a plurality of key frames is equal to or greater than a second threshold. There are various methods for calculating the similarity between the time intervals between a plurality of corresponding frames and the time intervals between a plurality of key frames.
[0074] For example, if the time intervals between multiple corresponding frames and the time intervals between multiple key frames are one type of time interval, the difference between the time intervals is first calculated. The difference between the time intervals is a difference or a rate of change. This difference may be used as the similarity. Alternatively, the calculated difference may be normalized according to a predetermined rule to obtain a value that is used as the similarity.
[0075] On the other hand, if the time intervals between multiple corresponding frames and the time intervals between multiple key frames include multiple types of time intervals, first, the difference between the time intervals is calculated for each type of time interval. The difference between the time intervals is the difference or the rate of change. Then, a statistical value of the difference between the calculated time intervals for each type of time interval is calculated. Examples of the statistical value include, but are not limited to, the average, maximum, minimum, mode, and median. This statistical value may be used as the similarity. Alternatively, the calculated statistical value may be normalized according to a predetermined rule to be used as the similarity.
[0076] The concepts of "when the time interval between multiple corresponding frames and the time interval between multiple key frames is one type of time interval" and "when the time interval between multiple corresponding frames and the time interval between multiple key frames includes multiple types of time intervals" are as explained in the first embodiment.
[0077] The first threshold value referred to in S20 and the second threshold value referred to in S21 may be set in advance, and the search unit 12 may perform the above search process based on the preset first threshold value and second threshold value.
[0078] Alternatively, the user may be able to specify at least one of the first threshold and the second threshold, and the search unit 12 may determine at least one of the first threshold and the second threshold based on a user input, and perform the search process based on the determined first threshold and second threshold.
[0079] When the time intervals between a plurality of corresponding frames and the time intervals between a plurality of key frames include a plurality of types of time intervals as described in the first embodiment, a second threshold is set for each type of time interval.
[0080] Other configurations of the search device 10 of this embodiment are the same as those of the first and second embodiments.
[0081] The search device 10 of this embodiment achieves the same effects as those of the first and second embodiments. Furthermore, the search device 10 of this embodiment can determine whether the movements (posture changes) are similar and whether the speed of the movements (speed of posture changes) is similar, separately, in multiple stages, and can set criteria (first and second thresholds) for determining similarity for each stage. As a result, it becomes possible to search for similar videos using desired criteria.
[0082] <Fourth embodiment> In this embodiment, the processing flow by the search unit 12 is embodied. The processing flow by the search unit 12 in this embodiment is different from that described in the third embodiment. The flowchart in Fig. 16 shows an example of the processing flow by the search unit 12 in this embodiment.
[0083] In S30, the search unit 12 searches for a video including Q corresponding frames corresponding to the Q key frames, respectively. The Nth corresponding frame corresponding to the Nth key frame includes a human body in a posture whose similarity to the posture of the human body included in the Nth key frame is equal to or greater than a first threshold.
[0084] In S31, the search unit 12 calculates, for each video retrieved in S30, the similarity between the postures of the human body included in multiple corresponding frames and the postures of the human body included in multiple key frames (hereinafter referred to as "posture similarity"). There are various methods for calculating posture similarity. For example, for each pair of corresponding corresponding frames and key frames, the similarity between the postures of the human body included in each is calculated. The method disclosed in Patent Document 1 can be used as the method for calculating the similarity. Next, statistical values of the multiple similarities calculated for each pair are calculated. Examples of statistical values include, but are not limited to, the average, maximum, minimum, mode, and median. Then, the calculated statistical values are normalized according to a predetermined rule to calculate a value as the posture similarity. Note that the method for calculating posture similarity illustrated here is merely an example and is not limited thereto.
[0085] In S32, the search unit 12 calculates the similarity between the time intervals between multiple corresponding frames and the time intervals between multiple key frames for each moving image searched in S30 (hereinafter referred to as "time interval similarity"). There are various methods for calculating the time interval similarity.
[0086] For example, if the time intervals between multiple corresponding frames and the time intervals between multiple key frames are one type of time interval, the difference between those time intervals is first calculated. The difference between the time intervals is defined as a difference or a rate of change. The calculated difference is then normalized according to a predetermined rule to calculate the similarity.
[0087] On the other hand, if the time intervals between multiple corresponding frames and the time intervals between multiple key frames include multiple types of time intervals, first, the difference between the time intervals is calculated for each type of time interval. The difference between the time intervals is defined as a difference or a rate of change. Then, statistical values of the differences between the calculated time intervals for each type of time interval are calculated. Examples of statistical values include, but are not limited to, the average, maximum, minimum, mode, and median. Then, the calculated statistical values are normalized according to a predetermined rule to calculate the similarity between the time intervals.
[0088] The concepts of "when the time interval between multiple corresponding frames and the time interval between multiple key frames is one type of time interval" and "when the time interval between multiple corresponding frames and the time interval between multiple key frames includes multiple types of time intervals" are as explained in the first embodiment.
[0089] In S33, the search unit 12 calculates an integrated similarity for each moving image searched in S30 based on the pose similarity calculated in S31 and the time interval similarity calculated in S32.
[0090] For example, the search unit 12 may calculate the sum or product of the similarity of the posture and the similarity of the time interval as the integrated similarity.
[0091] Alternatively, the search unit 12 may calculate a statistical value of the posture similarity and the time interval similarity as the integrated similarity. Examples of the statistical value include, but are not limited to, an average value, a maximum value, a minimum value, a mode value, and a median value.
[0092] Alternatively, the search unit 12 may calculate a weighted average or weighted sum of the similarity of the posture and the similarity of the time interval as the integrated similarity.
[0093] In S34, the search unit 12 searches the videos searched for in S30 for videos whose integrated similarity calculated in S33 is equal to or greater than a third threshold.
[0094] In addition, when the weighted average or weighted sum of the posture similarity and the time interval similarity is calculated as the integrated similarity in S33, the weights of the posture similarity and the time interval similarity may be set in advance or may be specified by the user. When the weights are specified by the user, the user's specification may be accepted via, for example, sliders (user interface (UI) components) such as those shown in FIGS. 17 and 18. The slider shown in FIG. 17 is configured to specify a weight for each posture similarity and time interval similarity. The slider shown in FIG. 18 is configured to specify the ratio of importance between the posture similarity and the time interval similarity. Then, each weight is calculated based on the specified ratio of importance. Note that accepting user input via a slider is merely an example, and user input may be accepted by other methods.
[0095] The first threshold value referred to in S30 and the third threshold value referred to in S34 may be set in advance, and the search unit 12 may perform the above search process based on the first threshold value and the third threshold value set in advance.
[0096] Alternatively, the user may be able to specify at least one of the first threshold and the third threshold, and the search unit 12 may determine at least one of the first threshold and the third threshold based on a user input, and perform the search process based on the determined first threshold and the third threshold.
[0097] Other configurations of the search device 10 of this embodiment are the same as those of the first to third embodiments.
[0098] The search device 10 of this embodiment achieves the same effects as those of the first to third embodiments. Furthermore, the search device 10 of this embodiment can search for videos whose integrated similarity, which is an integration of the similarity of movement (similarity of posture) and the similarity of movement speed (similarity of time interval), satisfies a criterion. The search device 10 of this embodiment adjusts the weights of the similarity of posture and the similarity of time interval, making it possible to search for similar videos according to desired criteria.
[0099] <Fifth embodiment> The search device 10 of this embodiment has first and second search modes. The search device 10 searches for videos similar to a query video in a search mode specified by a user. The first search mode is a mode in which a search is performed using the method described in the third embodiment. The second search mode is a mode in which a search is performed using the method described in the fourth embodiment.
[0100] Other configurations of the search device 10 of this embodiment are the same as those of the first to fourth embodiments.
[0101] The search device 10 of this embodiment achieves the same effects as those of the first to fourth embodiments. Furthermore, the search device 10 of this embodiment is provided with a plurality of search modes, and can perform a search in a mode specified by the user. The search device 10 of this embodiment is preferable because it broadens the range of choices available to the user.
[0102] Sixth Embodiment In this embodiment, the user specifies a lower limit for the video length of the video to be searched for as a search condition. The search device 10 then searches for videos that satisfy the conditions of the first to fifth embodiments and whose video length is equal to or greater than the specified lower limit as videos similar to the query video. In this case, videos whose video length is less than the lower limit specified by the user are not searched for. As a result, videos that include a human body whose movements are similar to those of the human body shown in the query video but whose movement speed is faster than a predetermined level (videos whose video length is shorter than a predetermined level) are not searched for. This is explained in detail below.
[0103] The search unit 12 accepts a user input specifying a lower limit for the video length as a search condition. The search unit 12 may accept a user input specifying a lower limit for the video length based on the length of the query video. For example, the lower limit for the video length may be specified as "X times the length of the query video." In this case, the search unit 12 accepts a user input specifying X, where X is a number greater than 0 and less than or equal to 1.
[0104] Alternatively, the search unit 12 may accept a user input that directly specifies the lower limit of the moving image length using a numerical value or the like.
[0105] Next, a method for searching for a video whose video length satisfies the above search conditions will be described.
[0106] -Method 1- First, the search unit 12 determines a lower limit on the number of key frames to extract from the query video based on the lower limit of the video length specified by the user. The search unit 12 determines a lower limit on the number of key frames to extract from the query video so that the length of the video composed of the extracted key frames is the lower limit of the video length specified by the user.
[0107] For example, if the video length of the query video is "P frames" and the lower limit of the video length specified by the user is "0.5 times the video length of the query video," the search unit 12 determines 0.5 x P as the lower limit of the number of key frames to extract from the query video.
[0108] Also, if the video length of the query video is "R seconds" and the lower limit of the video length specified by the user is "0.5 times the video length of the query video," the search unit 12 determines 0.5 × R × F1 as the lower limit of the number of key frames to extract from the query video, where F1 is the frame rate.
[0109] Then, the key frame extraction unit 11 extracts key frames equal to or greater than the lower limit of the number of key frames determined by the search unit 12 from the query moving image.
[0110] For example, when extracting key frames in extraction process 1 described in the first embodiment, i.e., when extracting frames designated by the user as key frames, the condition for completing the user's designation process may be "designating as key frames a number equal to or greater than the lower limit of the number of key frames determined by the search unit 12." In other words, the user cannot end the process of designating key frames unless he or she designates as key frames a number equal to or greater than the lower limit of the number of key frames determined by the search unit 12.
[0111] Additionally, when extracting key frames in extraction process 2 described in the first embodiment, that is, when extracting key frames every M frames, the key frame extraction unit 11 can adjust the number of extracted key frames by adjusting the value of M. The key frame extraction unit 11 determines the value of M so that the number of extracted key frames is equal to or greater than the lower limit of the number of key frames determined by the search unit 12.
[0112] In addition, when extracting key frames in extraction process 3 described in the first embodiment, that is, when the pose similarity with the reference key frame is equal to or less than a reference value and the frame that is earliest in chronological order is extracted as a new key frame as shown in Fig. 4, key frame extraction unit 11 can adjust the number of extracted key frames by adjusting the reference value of similarity. Key frame extraction unit 11 determines the reference value of similarity so that the number of extracted key frames is equal to or greater than the lower limit of the number of key frames determined by search unit 12.
[0113] The search unit 12 searches for videos having multiple corresponding frames corresponding to each of the extracted multiple key frames. If the lower limit of the number of key frames to be extracted from the query video is determined so that the length of the video composed of the extracted key frames is the lower limit of the video length specified by the user, videos shorter than the lower limit of the video length specified by the user will inevitably not be searched for.
[0114] -Method 2- First, the search unit 12 identifies a lower limit for the video length based on user input. If the lower limit for the video length is specified as "X times the length of the query video," the search unit 12 identifies the product of the length of the query video and the X specified by the user as the lower limit for the video length. Alternatively, if the lower limit for the video length is directly specified as a number, etc., the search unit 12 identifies the number specified by the user as the lower limit for the video length.
[0115] Then, the search unit 12 searches for moving images in which the elapsed time between the first corresponding frame and the last corresponding frame is equal to or greater than the specified lower limit of the moving image length as moving images that satisfy the search condition of the lower limit of the moving image length.
[0116] Other configurations of the search device 10 of this embodiment are the same as those of the first to fifth embodiments.
[0117] The search device 10 of this embodiment achieves the same effects as those of the first to fifth embodiments. Furthermore, the search device 10 of this embodiment allows the user to specify the minimum length of the video, i.e., the time for performing the movement shown in the query video. Such a search device 10 According to this, videos that contain human bodies with movements similar to those of the human body shown in the query video but whose movement speed is faster than a predetermined level (videos whose length is shorter than a predetermined level) will not be searched. As a result, users can perform the search they want.
[0118] Although the embodiments of the present invention have been described above with reference to the drawings, these are merely examples of the present invention, and various other configurations may be adopted. The configurations of the above-described embodiments may be combined with each other, or some of the configurations may be replaced with other configurations. Furthermore, various modifications may be made to the configurations of the above-described embodiments without departing from the spirit of the invention. Furthermore, the configurations and processes disclosed in the above-described embodiments and modified examples may be combined with each other.
[0119] In addition, in the flowcharts used in the above description, multiple steps (processes) are described in order, but the order of execution of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the drawings can be changed to the extent that the content is not affected. Furthermore, the above-mentioned embodiments can be combined to the extent that the content is not contradictory.
[0120] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. 1. A keyframe extraction means for extracting multiple keyframes from a query video; a search means for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; A search device having the above configuration. 2. The search means The plurality of corresponding frames include a human body in a pose whose similarity to the pose of the human body included in each of the plurality of key frames is equal to or greater than a first threshold, and A search device as described in 1, having a first search mode that searches for videos in which the similarity between the time intervals between multiple key frames and the time intervals between multiple corresponding frames is greater than or equal to a second threshold as videos similar to the query video. 3. The search device according to 2, wherein the search means determines at least one of the first threshold value and the second threshold value based on a user input. 4. The search means For each video to be processed, identifying a plurality of corresponding frames corresponding to each of the plurality of key frames; calculating an integrated similarity based on a similarity between a posture of the human body included in each of the plurality of key frames and a posture of the human body included in each of the plurality of corresponding frames, and a similarity between a time interval between the key frames and a time interval between the corresponding frames; A search device according to any one of 1 to 3, having a second search mode in which the video to be processed, whose integrated similarity is equal to or greater than a third threshold, is searched for as a video similar to the query video. 5. The search device according to 4, wherein the time interval between the key frames includes at least one of the time interval between two temporally adjacent key frames and the time interval between the first and last key frames. 6. The search means 6. The search device according to claim 4 or 5, wherein the integrated similarity is calculated based on a weight of the similarity of the posture of the human body specified by the user and a weight of the similarity between the time interval between the key frames and the time interval between the corresponding frames. 7. The key frame extraction means 7. The search device according to any one of 1 to 6, wherein the search device extracts key frames that are equal to or greater than a lower limit of the key frames to be extracted, the lower limit being determined based on a lower limit of the length of the moving image designated by a user as a search condition. 8. The key frame extraction means 8. The search device according to claim 7, wherein the number of key frames to be extracted is determined so that the length of a video composed of the extracted multiple key frames is equal to or greater than a lower limit of the video length specified by the user. 9. The computer a keyframe extraction step of extracting a plurality of keyframes from the query video; a search process for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; The search method to perform. 10. Computer a keyframe extraction means for extracting a plurality of keyframes from the query video; a search means for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; A program that functions as a [Explanation of symbols]
[0121] 10 Search Device 11 Keyframe extraction section 12 Search section 13 Skeletal structure detection unit 14 Feature calculation unit 1A processor 2A Memory 3A input / output I / F 4A peripheral circuit 5A Bus
Claims
1. A keyframe extraction means for extracting a plurality of keyframes from the query video; a search means for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; and The search means The plurality of corresponding frames include a human body in a pose whose similarity to the pose of the human body included in each of the plurality of key frames is equal to or greater than a first threshold, and A search device having a first search mode that searches for videos similar to the query video for which the similarity between the time intervals between multiple key frames and the time intervals between multiple corresponding frames is greater than or equal to a second threshold.
2. The search device according to claim 1 , wherein the search means determines at least one of the first threshold value and the second threshold value based on a user input.
3. A keyframe extraction means for extracting a plurality of keyframes from the query video; a search means for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; and The search means For each video to be processed, identifying a plurality of corresponding frames corresponding to each of the plurality of key frames; calculating an integrated similarity based on a similarity between a posture of the human body included in each of the plurality of key frames and a posture of the human body included in each of the plurality of corresponding frames, and a similarity between a time interval between the key frames and a time interval between the corresponding frames; A search device having a second search mode that searches for the video to be processed whose integrated similarity is equal to or greater than a third threshold as a video similar to the query video.
4. 4. The search device according to claim 3, wherein the time interval between the key frames includes at least one of a time interval between two temporally adjacent key frames and a time interval between a first and a last key frame.
5. The search means 5. The search device according to claim 3, wherein the integrated similarity is calculated based on a weight of the similarity of the posture of the human body designated by the user and a weight of the similarity between the time interval between the key frames and the time interval between the corresponding frames.
6. The key frame extraction means The search device according to claim 1 , wherein the search device extracts key frames whose number is equal to or exceeds a lower limit of the number of key frames to be extracted, the lower limit being determined based on a lower limit of a moving image length designated by a user as a search condition.
7. The computer a keyframe extraction step of extracting a plurality of keyframes from the query video; a search process for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; Run The searching step includes: The plurality of corresponding frames include a human body in a pose whose similarity to the pose of the human body included in each of the plurality of key frames is equal to or greater than a first threshold, and A search method having a first search mode that searches for videos similar to the query video for which the similarity between the time intervals between multiple key frames and the time intervals between multiple corresponding frames is greater than or equal to a second threshold.
8. The computer a keyframe extraction step of extracting a plurality of keyframes from the query video; a search process for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; Run The searching step includes: For each video to be processed, identifying a plurality of corresponding frames corresponding to each of the plurality of key frames; calculating an integrated similarity based on a similarity between a posture of the human body included in each of the plurality of key frames and a posture of the human body included in each of the plurality of corresponding frames, and a similarity between a time interval between the key frames and a time interval between the corresponding frames; A search method having a second search mode in which the video to be processed, whose integrated similarity is equal to or greater than a third threshold, is searched for as a video similar to the query video.
9. Computer, a keyframe extraction means for extracting a plurality of keyframes from the query video; a search means for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; It functions as The search means The plurality of corresponding frames include a human body in a pose whose similarity to the pose of the human body included in each of the plurality of key frames is equal to or greater than a first threshold, and A program having a first search mode that searches for videos similar to the query video for which the similarity between the time intervals between multiple key frames and the time intervals between multiple corresponding frames is greater than or equal to a second threshold.
10. Computer, a keyframe extraction means for extracting a plurality of keyframes from the query video; a search means for searching for videos similar to the query video based on the posture of the human body included in each of the plurality of key frames and the time interval between the plurality of key frames; It functions as The search means For each video to be processed, identifying a plurality of corresponding frames corresponding to each of the plurality of key frames; calculating an integrated similarity based on a similarity between a posture of the human body included in each of the plurality of key frames and a posture of the human body included in each of the plurality of corresponding frames, and a similarity between a time interval between the key frames and a time interval between the corresponding frames; A program having a second search mode that searches for the video to be processed whose integrated similarity is equal to or greater than a third threshold as a video similar to the query video.
Citation Information
Patent Citations
Device and method for calculating similarity of moving image
JP2000339474A
Apparatus and method for retrieving object poses
JP2014522035A
Displaying video keyframes on online social networks
JP2019532422A
Image processing device, image processing method, and non-transitory computer-readable medium having image processing program stored thereon
WO2021084677A1