Image processing apparatus, image processing method, and program

The image processing apparatus improves template image preparation by using a UI screen with a reproduction and missing keypoint display area, facilitating the selection of high-quality sections for enhanced human body detection accuracy.

JP7697581B2Active Publication Date: 2025-06-24NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024505668
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-06-24
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

Existing technologies face challenges in improving the accuracy and workability of preparing template images for human body detection, particularly in ensuring the quality of registered images.

Method used

An image processing apparatus and method that generates a UI screen with a reproduction area for moving images and a missing keypoint display area to indicate undetected keypoints, allowing users to easily select and extract high-quality template images.

Benefits of technology

Enhances the workability of preparing template images by enabling users to intuitively identify and extract sections with good keypoint detection, improving the accuracy of human body detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007697581000001
    Figure 0007697581000001
  • Figure 0007697581000002
    Figure 0007697581000002
  • Figure 0007697581000003
    Figure 0007697581000003
Patent Text Reader

Abstract

The present invention provides an image processing device (10) comprising a screen generation unit (11) and an input reception unit (12). The screen generation unit (11) generates a screen including a playback region that plays back and displays a dynamic image including a plurality of frame images, and a missing key point display region that indicates a human body key point not detected in a human body included in a frame image displayed in the playback region. The input reception unit (12) receives an input specifying a section extracted from the dynamic image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a recording medium.

Background Art

[0002] Technologies related to the present invention are disclosed in Patent Document 1 and Non-Patent Document 1.

[0003] Patent Document 1 discloses a technique for calculating feature amounts of a plurality of key points of a human body included in an image, and searching for an image including a human body having a similar posture or a similar movement based on the calculated feature amounts, or classifying together those having a similar posture or movement. Further, Non-Patent Document 1 discloses a technique related to human skeleton estimation.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Non-Patent Documents

[0005]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] According to the technology disclosed in Patent Document 1 described above, by registering in advance an image including a human body in a desired posture or desired movement as a template image, it is possible to detect a human body in a desired posture or desired movement from the images to be processed. As a result of studying the technology disclosed in such Patent Document 1, the present inventor newly found that the detection accuracy deteriorates unless an image of a certain quality is registered as a template image, and that there is room for improvement in the workability of the work of preparing such a template image.

[0007] Since both Patent Document 1 and Non-Patent Document 1 described above do not disclose problems regarding template images and solutions therefor, there is a problem that the above problems cannot be solved.

[0008] An example of the object of the present invention is to provide an image processing apparatus, an image processing method, and a recording medium that solve the problem of workability of the work of preparing a template image of a certain quality in view of the above problems.

Means for Solving the Problems

[0009] According to one aspect of the present invention, screen generation means for generating a screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images, and a missing key point display area for indicating a key point of a human body that has not been detected in the human body included in the frame image displayed in the reproduction area, and causing the display unit to display the screen; input reception means for receiving an input for designating a section to be extracted from the moving image; An image processing apparatus having the above is provided.

[0010] Also, according to one aspect of the present invention, a computer generates a screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images, and a missing key point display area for indicating a key point of a human body that has not been detected in the human body included in the frame image displayed in the reproduction area, and causes the display unit to display the screen, Receiving an input for specifying a section to be extracted from the moving image An image processing method is provided.

[0011] Also, according to one aspect of the present invention, A computer A screen generation means for generating a screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images, and a missing keypoint display area for indicating a keypoint of a human body not detected in the human body included in the frame image displayed in the reproduction area, and causing the display unit to display the screen An input receiving means for receiving an input for specifying a section to be extracted from the moving image A recording medium recording a program for functioning as such is provided.

Advantages of the Invention

[0012] According to one aspect of the present invention, an image processing apparatus, an image processing method, and a recording medium for solving the problem of workability of the work of preparing a template image of a certain quality can be obtained.

Brief Description of the Drawings

[0013] The above-described object, and other objects, features, and advantages will become more apparent from the following Suitable described embodiments and the accompanying drawings below.

[0014]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Embodiments for Carrying Out the Invention

[0015] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description will be omitted as appropriate.

[0016] <The First Embodiment> FIG. 1 is a functional block diagram showing an overview of an image processing apparatus 10 according to a first embodiment. As shown in FIG. 1, the image processing apparatus 10 includes a screen generation unit 11 and an input reception unit 12. The screen generation unit 11 generates a screen including a reproduction area for displaying a moving image including a plurality of frame images, and a missing keypoint display area for indicating keypoints of a human body that have not been detected in the human body included in the frame image displayed in the reproduction area, and causes the display unit to display the screen. The input reception unit 12 receives an input for designating a section to be extracted from the moving image.

[0017] According to this image processing apparatus 10, it is possible to solve the problem of workability of the work of preparing a template image of a certain quality.

[0018] <Second Embodiment> "Overview" As shown in FIG. 2, for example, the image processing apparatus 10 generates a UI (User Interface) screen including a reproduction area for reproducing and displaying a moving image, and a missing keypoint display area for indicating keypoints of a human body that have not been detected in the human body included in the frame image displayed in the reproduction area, and causes the display unit to display the screen. Then, the image processing apparatus 10 can receive an input for designating a section to be extracted as a template image from the moving image via such a UI screen.

[0019] While referring to the reproduction area and the missing keypoint display area, the user can specify a location in the moving image including a human body having a desired posture or desired movement and a good keypoint detection state, and extract the specified location as a template image.

[0020] "Hardware Configuration" Next, an example of the hardware configuration of the image processing apparatus 10 will be described. Each functional unit of the image processing apparatus 10 is realized by an arbitrary combination of hardware and software centered around a CPU (Central Processing Unit) of an arbitrary computer, a memory, a program loaded into the memory, a storage unit such as a hard disk storing the program (in addition to the program stored in advance at the stage of shipping the apparatus, a program downloaded from a storage medium such as a CD (Compact Disc) or a server on the Internet can also be stored), and a network connection interface. And it is understood by those skilled in the art that there are various modifications to the realization method and apparatus.

[0021] FIG. 3 is a block diagram illustrating the hardware configuration of the image processing apparatus 10. As shown in FIG. 3, the image processing apparatus 10 includes a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The image processing apparatus 10 may not have the peripheral circuit 4A. Note that the image processing apparatus 10 may be composed of a plurality of physically and / or logically separated apparatuses. In this case, each of the plurality of apparatuses can include the above-described hardware configuration.

[0022] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to transmit and receive data from each other. The processor 1A is an arithmetic processing device such as a CPU or a GPU (Graphics Processing Unit). The memory 2A is a memory such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The input / output interface 3A includes an interface for acquiring information from an input device, an external device, an external server, an external sensor, a camera, etc., and an interface for outputting information to an output device, an external device, an external server, etc. The input device is, for example, a keyboard, a mouse, a microphone, a physical button, a touch panel, etc. The output device is, for example, a display, a speaker, a printer, a mailer, etc. The processor 1A can issue commands to each module and perform operations based on their operation results.

[0023] "Functional Configuration" FIG. 4 is a functional block diagram showing an overview of the image processing apparatus 10 according to the second embodiment. As shown in FIG. 4, the image processing apparatus 10 includes a screen generation unit 11, an input reception unit 12, a display unit 13, and a storage unit 14. Note that the image processing apparatus 10 may not have the storage unit 14. In this case, an external device configured to be communicable with the image processing apparatus 10 includes the storage unit 14. Also, the image processing apparatus 10 may not have the display unit 13. In this case, an external device configured to be communicable with the image processing apparatus 10 includes the display unit 13.

[0024] The storage unit 14 stores the results of the human body key point detection process performed on each of a plurality of frame images included in the moving image.

[0025] The "moving image" is an image that is the source of the template image. The template image is an image (a concept including still images and moving images) that is pre-registered in the technology disclosed in Patent Document 1 described above, and is an image including a human body in a desired posture or a desired movement (a posture or a movement that the user wants to detect).

[0026] The detection process of the key points of the human body is executed by the skeletal structure detection unit. The image processing apparatus 10 may include the skeletal structure detection unit, or another apparatus physically and / or logically separated from the image processing apparatus 10 may include the skeletal structure detection unit.

[0027] For each frame image, the skeletal structure detection unit detects N (N is an integer of 2 or more) key points of the human body included in each frame image. The process by the skeletal structure detection unit is realized using the technology disclosed in Patent Document 1. Although details are omitted, in the technology disclosed in Patent Document 1, the detection of the skeletal structure is performed using a skeletal estimation technology such as OpenPose disclosed in Non-Patent Document 1. The skeletal structure detected by the technology is composed of "key points" which are characteristic points such as joints, and "bones (bone links)" indicating the links between the key points.

[0028] FIG. 5 shows the skeletal structure of the human body model 300 detected by the skeletal structure detection unit, and FIGS. 6 to 8 show detection examples of the skeletal structure. The skeletal structure detection unit uses a skeletal estimation technology such as OpenPose to detect the skeletal structure of a human body model (2D skeletal model) 300 as shown in FIG. 5 from a 2D image. The human body model 300 is a 2D model composed of key points such as joints of a person and bones connecting the key points.

[0029] For example, the skeletal structure detection unit extracts feature points that can be key points from an image, and refers to information obtained by machine learning of the key point images to detect N key points of the human body. The N key points to be detected are predetermined. The number of key points to be detected (that is, the number of N) and which parts of the human body are to be detected as key points vary, and all variations can be adopted.

[0030] Hereinafter, as shown in FIG. 5, it is assumed that the head A1, neck A2, right shoulder A31, left shoulder A32, right elbow A41, left elbow A42, right hand A51, left hand A52, right waist A61, left waist A62, right knee A71, left knee A72, right foot A81, and left foot A82 are defined as N key points (N = 14) to be detected. In the human body model 300 shown in FIG. 5, as the bones of a person connecting these key points, bone B1 connecting the head A1 and the neck A2, bone B21 and bone B22 connecting the neck A2 with the right shoulder A31 and the left shoulder A32 respectively, bone B31 and bone B32 connecting the right shoulder A31 and the left shoulder A32 with the right elbow A41 and the left elbow A42 respectively, bone B41 and bone B42 connecting the right elbow A41 and the left elbow A42 with the right hand A51 and the left hand A52 respectively, bone B51 and bone B52 connecting the neck A2 with the right waist A61 and the left waist A62 respectively, bone B61 and bone B62 connecting the right waist A61 and the left waist A62 with the right knee A71 and the left knee A72 respectively, and bone B71 and bone B72 connecting the right knee A71 and the left knee A72 with the right foot A81 and the left foot A82 respectively are further defined.

[0031] FIG. 6 is an example of detecting a person in an upright state. In FIG. 6, an upright person is imaged from the front, and bone B1, bone B51 and bone B52, bone B61 and bone B62, bone B71 and bone B72 seen from the front are detected without overlapping each other, and the bones B61 and B71 of the right foot are slightly bent more than the bones B62 and B72 of the left foot.

[0032] FIG. 7 is an example of detecting a person in a crouched state. In FIG. 7, a crouched person is imaged from the right side, and bone B1, bone B51 and bone B52, bone B61 and bone B62, bone B71 and bone B72 seen from the right side are detected respectively, and the bones B61 and B71 of the right foot and the bones B62 and B72 of the left foot are greatly bent and overlapping.

[0033] FIG. 8 is an example of detecting a person in a lying-down state. In FIG. 8, a person in a lying-down state is imaged from the front left diagonal. Bones B1, B51, and B52, bones B61 and B62, and bones B71 and B72 seen from the front left diagonal are detected respectively. The bones B61 and B71 of the right foot and the bones B62 and B72 of the left foot are bent and overlapping.

[0034] FIG. 9 schematically shows an example of information stored in the storage unit 14. As shown in FIG. 9, in the storage unit 14, the detection results of the key points of the human body are stored for each frame image (for each frame image identification information). When a plurality of human bodies are included in one frame image, the detection results of the key points of each of the plurality of human bodies are stored in association with that frame image.

[0035] The storage unit 14 stores data capable of reproducing the human body model 300 in a predetermined posture as shown in FIGS. 6 to 8 as the detection results of the key points of the human body. In the detection results of the key points of the human body, it is shown which of the N key points of the detection target are detected and which key points are not detected. Further, the storage unit 14 may store data further indicating the positions of the detected key points of the human body within the frame image. In addition, the storage unit 14 may store attribute information regarding the moving image, such as the file name of the moving image, the shooting date and time, the shooting location, the identification information of the camera that shot, and the like.

[0036] Returning to FIG. 4, the screen generation unit 11 generates a UI screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images, and a missing key point display area for indicating the key points of the human body that are not detected in the human body included in the frame image displayed in the reproduction area, and causes the display unit 13 to display it.

[0037] FIG. 2 shows an example of the UI screen. The illustrated UI screen includes a reproduction area and a missing key point display area. Note that the layout of the reproduction area and the missing key point display area is not limited to the illustrated example.

[0038] In the playback area, a moving image is played and displayed. Although not shown, buttons for performing operations such as playback, pause, rewind, fast forward, slow playback, and stop may be displayed on the UI screen.

[0039] In the missing keypoint display area, information indicating the keypoints of the human body that were not detected in the human body included in the frame image displayed in the playback area is displayed. For example, as in the example shown in FIG. 2, a human body model that separately displays the detected keypoints and the undetected keypoints may be displayed. An object K1 outlined with a solid line corresponds to the detected keypoint, and an object K2 outlined with a broken line corresponds to the undetected keypoint. The method of separately displaying object K1 and object K2 is not limited to differentiating the mode of the outline, and the color, shape, size, brightness, etc. of the object may be made different, or other methods may be adopted. Also, an object as shown in FIG. 2 may be displayed corresponding to only one of the detected keypoints and the undetected keypoints, and the object corresponding to the other may be made non-displayed.

[0040] Note that the human body model displayed in the missing keypoint display area indicates the keypoints of the human body that were not detected, and does not indicate the posture of the human body. For this reason, the posture of the human body model displayed in the missing keypoint display area is always the same posture and does not change according to the posture of the human body included in the frame image displayed in the playback area. In the following embodiments, an example will be described in which the human body model displayed in the missing keypoint display area indicates the posture of the human body included in the frame image displayed in the playback area.

[0041] As another example of the information displayed in the missing keypoint display area, in addition to and / or instead of the human body model as shown in FIG. 2, at least one of "the number of undetected keypoints or the number of detected keypoints" and "the name of the undetected keypoint (head, neck, etc.) or the name of the detected keypoint" may be displayed in the missing keypoint display area.

[0042] Also, when a plurality of human bodies are included in the frame image displayed in the playback area, the screen generation unit 11 may select one human body from among the plurality of human bodies according to a predetermined rule, and display the key points of the human bodies that have not been detected in the selected human body in the missing key point display area. Examples of the rule for selecting one human body include, but are not limited to, "select the human body designated by the user" and "select the human body with the largest size within the frame image". In this case, the screen generation unit 11 may highlight the selected human body on the frame image displayed in the playback area. For example, the screen generation unit 11 may highlight the selected human body by superimposing a frame surrounding the human body, a mark corresponding to the human body, etc. on the frame image.

[0043] As a modification, when a plurality of human bodies are included in the frame image displayed in the playback area, the screen generation unit 11 may display the key points of the human bodies that have not been detected in each of the plurality of human bodies in the missing key point display area at once. For example, the screen generation unit 11 may, for each of the plurality of human bodies included in the frame image displayed in the playback area, display "the human body model displayed in the missing key point display area of FIG. 2", "the number of key points not detected, or the number of key points detected", and "the name of the key points not detected, or the name of the key points detected". In this case, it is preferable to display information indicating the correspondence between the plurality of human bodies included in the frame image displayed in the playback area and the detection results of the key points of the plurality of human bodies shown in the missing key point display area. For example, methods such as surrounding the "human body on the playback area" and the "detection result on the missing key point display area" corresponding to each other with a frame of the same color can be considered, but are not limited thereto.

[0044] In addition, while the moving image is being played in the playback area, the screen generation unit 11 may constantly display information as shown in FIG. 2 in the missing key point display area. In this case, the information displayed in the missing key point display area is also updated in accordance with the switching of the frame images displayed in the playback area. Additionally, the screen generation unit 11 may display, in the missing key point display area, the key points of the human body that were not detected in the human body included in the frame image displayed in the playback area at that time, only while the moving image on the playback area is paused.

[0045] The screen generation unit 11 can generate the above-described UI screen by using the "results of the human body key point detection process performed on each of the plurality of frame images included in the moving image" stored in the storage unit 14.

[0046] The display unit 13 that displays the UI screen may be a display or a projection device connected to the image processing apparatus 10. Additionally, a display or a projection device connected to an external device configured to be communicable with the image processing apparatus 10 may serve as the display unit 13 that displays the UI screen. In this case, the image processing apparatus 10 serves as a server and the external device serves as a client terminal. The external device is, for example, a personal computer, a smartphone, a smartwatch, a tablet terminal, a mobile phone, etc., but is not limited thereto.

[0047] Returning to FIG. 4, the input reception unit 12 receives an input for designating a section to be extracted as a template image from the moving image. The section is a partial time zone within the moving image having a time width. For example, the start position and the end position of the section are indicated by the elapsed time from the start of the moving image and the like.

[0048] The means for receiving the designation of the extraction section is not limited, and any technique can be adopted. In the case of the UI screen shown in FIG. 2, an operation of pressing a decision button corresponding to the start position of the extraction section while the frame image at the start position of the extraction section is displayed in the reproduction area, and an operation of pressing a decision button corresponding to the end position of the extraction section while the frame image at the end position of the extraction section is displayed in the reproduction area are performed, and an input for designating the extraction section is made.

[0049] In addition, as a means for receiving the designation of the extraction section, a slide bar indicating the playback time of the video, the elapsed time from the start, etc. may be displayed on the UI screen, and a means for receiving the designation of the start position and the end position of the extraction section on the slide bar may be adopted. In addition, as a means for receiving the designation of the extraction section, a means for automatically determining the position where the user starts playback as the start position of the extraction section and automatically determining the position where the user ends playback as the end position of the extraction section may be adopted. In addition, as a means for receiving the designation of the extraction section, a means for determining a position a predetermined number of frames before the reference position (reference frame) in the video designated by the user using the above-mentioned slide bar, etc. as the start position of the extraction section and determining a position a predetermined number of frames after the reference position as the end position of the extraction section may be adopted.

[0050] Next, an example of the processing flow of the image processing apparatus 10 will be described using the flowchart of FIG. 10.

[0051] The image processing apparatus 10 generates a UI screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images and a missing keypoint display area for indicating a keypoint of the human body that has not been detected in the human body included in the frame image displayed in the reproduction area, and causes the display unit 13 to display it (S10). Next, the image processing apparatus 10 receives an input for designating a section to be extracted from the moving image via the UI screen (S11).

[0052] In addition, when the image processing apparatus 10 receives an input for designating a section to be extracted from a moving image, it may cut out the section from the moving image, create another video file, and save it. Alternatively, when the image processing apparatus 10 receives an input for designating a section to be extracted from a moving image, it may store information indicating the designated section in the storage unit 14. For example, the file name of the moving image and information indicating the designated section (such as information indicating the start position and end position of the section) may be associated and stored in the storage unit 14.

[0053] "Operational Effects" According to the image processing apparatus 10 of the second embodiment, for example, as shown in FIG. 2, a UI screen including a reproduction area for reproducing and displaying a moving image and a missing keypoint display area for indicating a keypoint of a human body that has not been detected in the human body included in the frame image displayed in the reproduction area can be generated and displayed on the display unit 13. Then, the image processing apparatus 10 can receive an input for designating a section to be extracted as a template image from the moving image via such a UI screen.

[0054] The user can identify a location in the moving image that includes a human body in a desired posture or desired movement and has a good keypoint detection state while referring to the UI screen, and extract the identified location as a template image. According to this image processing apparatus 10, it is possible to solve the problem of workability in the task of preparing a template image of a certain quality.

[0055] Further, as shown in FIG. 2, the image processing apparatus 10 can display a UI screen that displays a human body model in which the detected keypoints and the undetected keypoints are separately displayed in the missing keypoint display area. Through such a human body model, the user can intuitively and easily grasp the undetected keypoints.

[0056] <Third Embodiment> In addition to the information (playback area, missing keypoint display area) described in the first and second embodiments, the image processing apparatus 10 according to the third embodiment is different from the image processing apparatuses 10 of the first and second embodiments in that it generates and displays a UI screen that further displays a human body model indicating the posture of the human body included in the frame image displayed in the playback area. This will be described in detail below.

[0057] The screen generation unit 11 generates a UI screen that further displays a human body model indicating the posture of the human body included in the frame image displayed in the playback area, in addition to the information (playback area, missing keypoint display area) described in the first and second embodiments, and causes the display unit 13 to display it. The human body model 300 shown in FIG. 5 in a predetermined posture as shown in FIGS. 6 to 8 is displayed on the UI screen. The screen generation unit 11 executes at least one of the first to third processes described below.

[0058] "First Process" In the first process, the screen generation unit 11 generates a UI screen that further includes a human body model display area, separate from the playback area and the missing keypoint display area. In the human body model display area, a human body model composed of keypoints detected in the human body included in the frame image displayed in the playback area and indicating the posture of the human body is displayed.

[0059] FIG. 11 shows an example of the UI screen. Although a human body model is displayed in both the human body model display area and the missing keypoint display area, the human body model displayed in the human body model display area indicates the posture of the human body, while the human body model displayed in the missing keypoint display area indicates the keypoints that were not detected, which is different.

[0060] In addition, when a plurality of human bodies are included in the frame image displayed in the playback area, the screen generation unit 11 may select one human body from among the plurality of human bodies according to a predetermined rule, and display a human body model indicating the posture of the selected human body in the human body model display area. Examples of the rule for selecting one human body include, but are not limited to, "select the human body designated by the user" and "select the human body with the largest size in the frame image". In this case, the screen generation unit 11 may highlight the selected human body on the frame image displayed in the playback area. For example, the screen generation unit 11 may highlight the selected human body by superimposing a frame surrounding the human body, a mark corresponding to the human body, or the like on the frame image.

[0061] As a modification, when a plurality of human bodies are included in the frame image displayed in the playback area, the screen generation unit 11 may display a plurality of human body models indicating the postures of each of the plurality of human bodies in the human body model display area. In this case, it is preferable to display information indicating the correspondence between the plurality of human bodies included in the frame image displayed in the playback area and the plurality of human body models displayed in the human body model display area. For example, methods such as surrounding the "human body on the playback area" and the "human body model on the human body model display area" that correspond to each other with frames of the same color can be considered, but are not limited thereto.

[0062] Also, the screen generation unit 11 may always display a human body model in the human body model display area while playing a moving image in the playback area. In this case, the posture of the human body model displayed in the human body model display area is also updated according to the switching of the frame images displayed in the playback area. In addition, the screen generation unit 11 may display a human body model indicating the posture of the human body included in the frame image displayed in the playback area at that time in the human body model display area only while the moving image on the playback area is paused.

[0063] "Second Process" In the second process, the screen generation unit 11 generates a UI screen in which a human body model indicating the posture of a human body is superimposed on the frame image displayed in the playback area. The human body model may be superimposed on the human body included in the frame image.

[0064] FIG. 12 shows an example of the UI screen. A human body model indicating the posture of the human body included in the frame image is superimposed on the frame image displayed in the playback area. The human body model is superimposed on the human body included in the frame image.

[0065] In addition, when the frame image displayed in the playback area includes a plurality of human bodies, the screen generation unit 11 may superimpose a plurality of human body models indicating the postures of the respective human bodies on the frame image. Each of the plurality of human body models is preferably superimposed on the corresponding human body.

[0066] Also, the screen generation unit 11 may constantly display a human body model on the frame image while playing a moving image in the playback area. In this case, according to the switching of the frame image displayed in the playback area, the posture and position of the human body model superimposed on the frame image are also updated. In addition, the screen generation unit 11 may superimpose a human body model indicating the posture of the human body included in the frame image displayed in the playback area at that time on the frame image only while the moving image on the playback area is paused.

[0067] "Third Process" In the third process, the screen generation unit 11 displays a human body model that indicates the key points of the human body that were not detected and that indicates the posture of the human body in the missing key point display area. In this case, the posture of the human body model displayed in the missing key point display area changes according to the posture of the human body included in the frame image displayed in the playback area. Specifically, the posture of the human body model displayed in the missing key point display area becomes the same posture as the posture of the human body included in the frame image displayed in the playback area.

[0068] FIG. 13 shows an example of the UI screen. The posture of the human model displayed in the missing key point display area is the same as the posture of the human body included in the frame image displayed in the playback area.

[0069] In addition, when a plurality of human bodies are included in the frame image displayed in the playback area, the screen generation unit 11 may select one human body from among the plurality of human bodies according to a predetermined rule, and display the detection result of the key points of the selected human body and the human model indicating the posture in the missing key point display area. Examples of the rule for selecting one human body include, but are not limited to, "select the human body designated by the user" and "select the human body with the largest size in the frame image". In this case, the screen generation unit 11 may highlight the selected human body on the frame image displayed in the playback area. For example, the screen generation unit 11 may highlight the selected human body by superimposing a frame surrounding the human body, a mark corresponding to the human body, etc. on the frame image.

[0070] As a modification, when a plurality of human bodies are included in the frame image displayed in the playback area, the screen generation unit 11 may display the detection results of the key points of each of the plurality of human bodies and a plurality of human models indicating the postures in the missing key point display area. In this case, it is preferable to display information indicating the correspondence relationship between the plurality of human bodies included in the frame image displayed in the playback area and the plurality of human models displayed in the missing key point display area. For example, methods such as surrounding the "human body on the playback area" and the "human model on the missing key point display area" corresponding to each other with frames of the same color can be considered, but it is not limited to this.

[0071] Further, the screen generation unit 11 may always display a human model in the missing key point display area while playing a moving image in the playback area. In this case, according to the switching of the frame image displayed in the playback area, the missing key points Display areaThe content of the human model (posture and detection results of key points) displayed in [the relevant area] is also updated. In addition, the screen generation unit 11 may display, in the missing key point display area, a human model indicating the posture of the human body and the detection results of the key points included in the frame image displayed in the reproduction area at that time, but only while the moving image on the reproduction area is paused.

[0072] Other configurations of the image processing apparatus 10 according to the third embodiment are the same as those of the image processing apparatus 10 according to the first and second embodiments.

[0073] According to the image processing apparatus 10 of the third embodiment, the same operational effects as those of the image processing apparatus 10 of the first and second embodiments are achieved. Further, according to the image processing apparatus 10 of the third embodiment, a UI screen for further displaying a human model indicating the posture of the human body included in the frame image displayed in the reproduction area can be generated and displayed.

[0074] While referring to the UI screen, the user can identify a location within the moving image that shows a desired posture or desired movement, has a good detection state of key points, and shows a correct posture and movement based on the detected key points (i.e., the key points are correctly detected), and extract the identified location as a template image. According to this image processing apparatus 10, the problem of workability in preparing a template image of a certain quality can be solved.

[0075] <Fourth Embodiment> The image processing apparatus 10 according to the fourth embodiment is different from the image processing apparatuses 10 according to the first to third embodiments in that, in addition to the information (reproduction area, missing key point display area) described in the first and second embodiments, it generates and displays a UI screen for further displaying a floor map indicating the installation position of the camera that captured the moving image. The UI screen generated by the image processing apparatus 10 according to the fourth embodiment may further display the information (human model indicating the posture of the human body included in the frame image displayed in the reproduction area) described in the third embodiment. This will be described in detail below.

[0076] In addition to the information (playback area, missing key point display area) described in the first and second embodiments, the screen generation unit 11 generates a UI screen that further displays a floor map indicating the installation position of the camera that captured the moving image, and causes the display unit 13 to display it. In addition to the above information, the screen generation unit 11 may generate a UI screen that further displays the information (human body model indicating the posture of the human body included in the frame image displayed in the playback area) described in the third embodiment, and cause the display unit 13 to display it. Hereinafter, some examples of the UI screen including the floor map will be shown.

[0077] "First Example" FIG. 14 shows an example of the UI screen generated by the screen generation unit 11. In the UI screen shown in FIG. 14, in addition to the playback area and the missing key point display area, a floor map is displayed. In this example, the camera is installed inside the bus. Therefore, the floor map is a map inside the bus. In the figure, the icon C1 indicates the installation position of the camera.

[0078] "Second Example" There may be cases where the same location is photographed by a plurality of cameras. The plurality of cameras are installed at different locations from each other. In this case, as in the example of FIG. 15, the screen generation unit 11 can generate a UI screen including a floor map indicating the installation positions of the plurality of cameras. In this example, three cameras are installed inside the bus. And in the floor map, icons C1 to C3 indicating the installation positions of each of the three cameras are shown.

[0079] In the case of this example, the input reception unit 12 can receive an input for designating one camera. Then, the screen generation unit 11 can reproduce and display the moving image captured by the designated camera among the plurality of cameras in the playback area. Note that, as shown in FIG. 15, the screen generation unit 11 may highlight the designated camera in the floor map. Further, the screen generation unit 11 may display information indicating the designated camera in the playback area. In the example shown in FIG. 15, character information identifying the designated camera "Camera C1" is superimposed and displayed on the moving image.

[0080] There are various means for the input reception unit 12 to receive an input specifying one camera. For example, the input reception unit 12 may receive an input for selecting an icon of one camera on the floor map, or it may be realized by other means.

[0081] Note that the input reception unit 12 may receive an input for changing the specified camera while a moving image is being played in the playback area. In this case, in response to the input for changing the specified camera, the moving image being played and displayed in the playback area is switched from the moving image captured by the camera specified before the change to the moving image captured by the camera specified after the change. At this time, the playback start position of the moving image captured by the camera specified after the change may be determined according to the playback end position of the moving image that was being played and displayed before the change. For example, a time stamp indicating the shooting date and time may be attached to the moving images captured by a plurality of cameras. Then, when the input reception unit 12 switches the moving image being played and displayed in the playback area in response to an input for changing the specified camera during playback of the moving image in the playback area, first, the shooting date and time of the playback end position of the moving image that was being played before the change may be specified. And the input reception unit 12 may play the moving image captured by the camera specified after the change from the portion shot at the specified shooting date and time.

[0082] "Third Example" There may be cases where the same location is photographed by a plurality of cameras. The plurality of cameras are installed at different locations from each other. In this case, as in the example of FIG. 16, the screen generation unit 11 can generate a UI screen including a floor map showing the installation positions of the plurality of cameras. In this example, three cameras are installed inside the bus. And on the floor map, icons C1 to C3 indicating the installation positions of each of the three cameras are shown.

[0083] In this example, the input reception unit 12 can receive an input for specifying one camera. Then, as shown in FIG. 16, the screen generation unit 11 can generate a UI screen that simultaneously reproduces and displays a plurality of moving images captured by each of the plurality of cameras in the reproduction area, and emphasizes and displays the moving image captured by the specified camera, and causes the display unit 13 to display it. In the illustrated example, the moving image captured by the specified camera is displayed in a larger screen than the moving images captured by other cameras, and further emphasized by superimposing character information of "being specified" on the moving image, but other methods may be used to realize the emphasis display.

[0084] Further, a time stamp indicating the shooting date and time may be attached to the moving images captured by the plurality of cameras. Then, the screen generation unit 11 may synchronize the reproduction timings and reproduction positions of the plurality of moving images so that the frame images captured at the same timing are simultaneously displayed in the reproduction area using the time stamp.

[0085] Note that, as shown in FIG. 16, the screen generation unit 11 may emphasize and display the specified camera on the floor map.

[0086] There are various means for the input reception unit 12 to receive an input for specifying one camera. For example, the input reception unit 12 may receive an input for selecting an icon of one camera on the floor map, or may receive an input for selecting a moving image captured by one camera on the reproduction area, or may be realized by other means.

[0087] Note that the input reception unit 12 may receive an input for changing the specified camera while a moving image is being reproduced in the reproduction area. In this case, the moving image emphasized and displayed in the reproduction area is switched according to the input for changing the specified camera.

[0088] In the case of the third example, information on the key points of the human body detected in the moving image captured by the designated camera among the plurality of moving images being played back and displayed in the playback area may be displayed in the missing key point display area. Also, when adopting the configuration of the third embodiment, a human body model showing the posture of the human body detected in the moving image captured by the designated camera among the plurality of moving images being played back and displayed in the playback area may be displayed on the UI screen.

[0089] Also, in the case of the third example, when the input reception unit 12 receives a user input designating one human body on one moving image displayed in the playback area, the screen generation unit 11 may highlight (such as surrounding with a frame) the human body appearing in other moving images. Identification of the same person appearing across a plurality of moving images is realized by face matching, appearance matching, position matching, etc.

[0090] "Fourth Example" The screen generation unit 11 may further indicate the position of the human body detected in the frame image displayed in the playback area on the floor map of the first to third examples. Also, the screen generation unit 11 may further indicate the position of the human body detected in the frame image captured by another camera at the same timing as the frame image displayed in the playback area on the floor map of the first to third examples.

[0091] FIG. 17 shows an example of the floor map displayed on the UI screen. The icon P indicates the position of the human body. The position of the human body can be specified by image analysis. For example, when the installation position and orientation of the camera are fixed, correspondence information indicating the correspondence between the positions in the frame images captured by each of the plurality of cameras and the positions in the floor map can be generated in advance. Then, using the said correspondence information, the position of the human body detected in the frame image can be converted into the position on the floor map.

[0092] Also, as shown in FIG. 20, information indicating the approximate shooting range of each camera may be displayed on the floor map. In the example shown in FIG. 20, the shooting range of each camera is indicated by a fan-shaped figure, but it is not limited to this. Also, in the example shown in FIG. 20, the shooting ranges of all cameras are displayed, but only the shooting range of the specified camera may be displayed. The shooting range of each camera may be automatically determined from the specifications of each camera (installation position, orientation, specifications (such as angle of view), etc.) or may be defined manually. Whether to include in the shooting range a position where a person is reflected in the camera but is far away and appears small, making it difficult to detect the skeleton, or a position where an obstacle obstructs the view is up to the definition of the shooting range and is freely determined.

[0093] Note that although an example of shooting inside a bus has been described here, the shooting location is not limited to this example.

[0094] Other configurations of the image processing apparatus 10 according to the fourth embodiment are the same as those of the image processing apparatus 10 according to the first to third embodiments.

[0095] According to the image processing apparatus 10 of the fourth embodiment, the same operational effects as those of the image processing apparatus 10 according to the first to third embodiments are realized. Also, according to the image processing apparatus 10 of the fourth embodiment, the user can identify the location to be extracted as a template image while checking the position of the camera that has taken the shot, switching and checking the moving images of the cameras taken simultaneously, comparing the moving images of the cameras taken simultaneously, and checking the positional relationship between the human body and the camera. According to this image processing apparatus 10, the problem of workability in preparing a template image of a certain quality can be solved.

[0096] <Fifth Embodiment> In the fifth embodiment, the camera is installed inside the moving body. And, in addition to the information (playback area, missing keypoint display area) described in the first and second embodiments, the image processing apparatus 10 of the fifth embodiment generates and displays a UI screen that further includes a moving body state display area indicating the state of the moving body at the timing when the frame image displayed in the playback area was captured. This is different from the image processing apparatus 10 of the first to fourth embodiments. The UI screen generated by the image processing apparatus 10 of the fifth embodiment may further display at least one of the information (human body model indicating the posture of the human body included in the frame image displayed in the playback area) described in the third embodiment and the information (floor map) described in the fourth embodiment. This will be described in detail below.

[0097] The screen generation unit 11 generates a UI screen that further includes a moving body state display area in addition to the information (playback area, missing keypoint display area) described in the first and second embodiments, and causes the display unit 13 to display it. In addition to the above information, the screen generation unit 11 may generate a UI screen that further displays at least one of the information (human body model indicating the posture of the human body included in the frame image displayed in the playback area) described in the third embodiment and the information (floor map) described in the fourth embodiment, and cause the display unit 13 to display it.

[0098] In the fifth embodiment, the camera is installed inside the moving body. The moving body is something that a person can ride on, and examples include buses, trains, airplanes, ships, vehicles, and the like. Information indicating the state of the moving body at the timing when the frame image displayed in the playback area was captured is displayed in the moving body state display area.

[0099] FIG. 18 shows an example of the UI screen generated by the screen generation unit 11. The moving body state display area is displayed on the UI screen shown in FIG. 18. And character information "stopped" is displayed in the area as the state of the moving body at the timing when the frame image displayed in the playback area was captured.

[0100] The state of the moving body is a state that can be identified by sensors installed on the moving body. Various states can be defined as states to be displayed in the moving body state display area. For example, when stopped, stationary, traveling, moving, going straight at less than X1 km / h, going straight at X1 km / h or more, turning right, turning left, turning right, turning left, ascending, descending, etc. are exemplified, but not limited thereto.

[0101] Based on the information acquired by various sensors installed on the moving body, moving body state information indicating the state of the moving body at each timing as shown in FIG. 19 can be generated and stored in the storage unit 14. The screen generation unit 11 can identify the state of the moving body at the timing when the frame image displayed in the reproduction area was taken based on the moving body state information, and display information indicating the identified state in the moving body state display area.

[0102] Other configurations of the image processing apparatus 10 according to the fifth embodiment are the same as those of the image processing apparatus 10 according to the first to fourth embodiments.

[0103] According to the image processing apparatus 10 of the fifth embodiment, the same operational effects as those of the image processing apparatus 10 of the first to fourth embodiments are realized. Further, according to the image processing apparatus 10 of the fifth embodiment, the user can identify the location to be extracted as the template image while confirming the state of the moving body at the timing of shooting. According to this image processing apparatus 10, the problem of workability in preparing a template image of a certain quality can be solved.

[0104] <Modification> "First Modification" In the above embodiment, image analysis processing such as processing for detecting key points in advance for a moving image was performed, the result was stored in the storage unit 14, and a characteristic UI screen was generated using the stored data. As a modification, when the moving image is reproduced and displayed in the reproduction area, image analysis processing such as processing for detecting key points for the moving image at that timing may be performed, and a UI screen may be generated using the result.

[0105] "Second Modification Example" Using image analysis techniques such as person tracking, the same person appearing across multiple frame images in a moving image may be identified. Then, when the user designates one human body shown in a certain frame image, the screen generation unit 11 identifies other frame images in which a human body that is the same person as the designated human body and has a better keypoint detection result than the designated human body appears, and may display the identified frame images on the UI screen as other candidates.

[0106] In addition, the screen generation unit 11 may identify other frame images in which a human body that is the same person as the designated human body, has a better keypoint detection result than the designated human body, and has a posture that is the same as or has a similarity degree to the posture of the designated human body equal to or greater than a threshold value appears, and display the identified frame images on the UI screen as other candidates.

[0107] Note that the frame images from a frame image a predetermined number of frames before the frame image in which the designated human body appears to a frame image a predetermined number of frames after may be narrowed down as the target for searching for the above other candidates.

[0108] "A human body with a better keypoint detection result than the designated human body" is a human body with a larger number of detected keypoints than the designated human body, etc. The similarity degree of the postures can be calculated using the method disclosed in Patent Document 1.

[0109] "Designation of one human body shown in a certain frame image" may be realized, for example, by an operation of designating one from among the human bodies shown in the frame image displayed in the reproduction area in a state where the moving image displayed in the reproduction area is paused.

[0110] As described above, the embodiments of the present invention have been described with reference to the drawings, but these are examples of the present invention, and various configurations other than the above may also be adopted.

[0111] In addition, in the plurality of flowcharts used in the above description, a plurality of steps (processes) are described in order. However, the execution order of the steps executed in each embodiment is not limited to the order of the description. In each embodiment, the order of the steps shown can be changed within a range that does not substantially affect the content. Also, the above-described embodiments can be combined within a range where the contents do not conflict with each other.

[0112] Some or all of the above embodiments can be described as follows in the appended claims, but are not limited thereto. 1. A screen generation means for generating a screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images, and a missing keypoint display area for indicating a keypoint of a human body not detected in the human body included in the frame image displayed in the reproduction area, and causing the display unit to display the screen; An input receiving means for receiving an input for designating a section to be extracted from the moving image; An image processing apparatus having the above. 2. The image processing apparatus according to 1, wherein the screen generation means generates the screen for further displaying a human body model indicating a posture of a human body included in the frame image displayed in the reproduction area. 3. The image processing apparatus according to 2, wherein the screen generation means further includes a human body model display area configured by the keypoints detected in the human body included in the frame image displayed in the reproduction area and displaying a human body model indicating the posture of the human body, and generates the screen. 4. The image processing apparatus according to 2, wherein the screen generation means generates the screen in which a human body model configured by the keypoints detected in the human body included in the frame image displayed in the reproduction area is superimposed and displayed on the frame image displayed in the reproduction area. 5. The image processing apparatus according to 2, wherein the screen generation means generates the screen in which, in the missing keypoint display area, the keypoints detected in the human body included in the frame image displayed in the reproduction area and the keypoints not detected are separately displayed, and a human body model indicating the posture of the human body is displayed. 6. The screen generation means generates the screen including a floor map indicating installation positions of a plurality of cameras, The input reception means receives an input designating one of the cameras, The screen generation means is the image processing apparatus according to any one of 1 to 5, which reproduces and displays the moving image captured by the designated camera in the reproduction area. 7. The image processing apparatus according to 6, wherein the screen generation means generates the screen in which the designated camera is highlighted on the floor map. 8. The screen generation means further includes a floor map indicating installation positions of a plurality of cameras, and generates the screen in which a plurality of moving images captured by each of the plurality of cameras are simultaneously reproduced and displayed in the reproduction area, The input reception means receives an input designating one of the moving images in the reproduction area, The screen generation means is the image processing apparatus according to any one of 1 to 5, which generates the screen in which the camera that captured the designated moving image is highlighted on the floor map. 9. The image processing apparatus according to any one of 6 to 8, wherein the floor map further indicates positions of human bodies detected within the frame image displayed in the reproduction area. 10. The image processing apparatus according to 9, wherein the floor map further indicates positions of human bodies detected within the frame image captured by another camera at the same timing as the frame image displayed in the reproduction area. 11. The moving image shows the state inside a moving body, The screen generation means is the image processing apparatus according to any one of 1 to 10, which generates the screen further including a moving body state display area indicating a state of the moving body at the timing when the frame image displayed in the reproduction area was captured. 12. A computer generates a screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images, and a missing keypoint display area indicating keypoints of a human body not detected in the human body included in the frame image displayed in the reproduction area, and causes the display unit to display the screen. Receiving an input for specifying a section to be extracted from the moving image, An image processing method. 13. A computer, generating a screen including a playback area for playing and displaying a moving image including a plurality of frame images, and a missing keypoint display area for indicating a keypoint of a human body not detected in the human body included in the frame image displayed in the playback area, and causing the display unit to display the screen; a screen generation means, input receiving means for receiving an input for specifying a section to be extracted from the moving image, A recording medium having recorded thereon a program for functioning as such.

Explanation of Signs

[0113] 10 Image processing apparatus 11 Screen generation unit 12 Input reception unit 13 Display unit 14 Storage unit 1A Processor 2A Memory 3A Input / output I / F 4A Peripheral circuit 5A Bus

Claims

1. Screen generation means for generating a screen including a playback area for playback-displaying a moving image including a plurality of frame images, and a missing keypoint display area for indicating a keypoint of the human body that has not been detected in the human body included in the frame image displayed in the playback area, and causing the display unit to display the screen; Input reception means for receiving an input for designating a section to be extracted from the moving image; An image processing apparatus having the above.

2. The image processing apparatus according to claim 1, wherein the screen generation means generates the screen for further displaying a human body model indicating the posture of the human body included in the frame image displayed in the playback area.

3. The screen generation means The screen further including a human body model display area for displaying a human body model composed of the keypoints detected in the human body included in the frame image displayed in the playback area and indicating the posture of the human body; The screen in which a human body model composed of the keypoints detected in the human body included in the frame image displayed in the playback area is superimposed and displayed on the frame image displayed in the playback area, or The screen in which, in the missing keypoint display area, the keypoints detected in the human body included in the frame image displayed in the playback area and the keypoints not detected are separately displayed, and a human body model indicating the posture of the human body is displayed; The image processing apparatus according to claim 2, which generates the above.

4. The screen generation means generates the screen including a floor map indicating the installation positions of a plurality of cameras; The input reception means receives an input for designating one of the cameras; The image processing apparatus according to any one of claims 1 to 3, wherein the screen generation means plays back and displays the moving image captured by the designated camera in the playback area.

5. The image processing apparatus according to claim 4, wherein the screen generation means generates the screen in which the designated camera is highlighted on the floor map.

6. The screen generation means further includes a floor map indicating the installation positions of a plurality of cameras, and generates the screen in which a plurality of moving images captured by each of the plurality of cameras are simultaneously played back and displayed in the playback area; The input reception means receives an input for designating one of the moving images in the playback area; The image processing apparatus according to any one of claims 1 to 3, wherein the screen generation means generates a screen in which the camera that has captured the specified moving image is highlighted on the floor map.

7. The image processing apparatus according to any one of claims 4 to 6, wherein the floor map further shows the positions of human bodies detected in the frame images captured by other cameras at the same timing as the frame images displayed in the reproduction area.

8. The moving image shows the interior of a moving body, The image processing apparatus according to any one of claims 1 to 7, wherein the screen generation means further generates a screen including a moving body state display area indicating the state of the moving body at the timing when the frame image displayed in the reproduction area was captured.

9. A computer, generates a screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images and a missing keypoint display area indicating keypoints of a human body that were not detected in the human body included in the frame images displayed in the reproduction area, and causes the screen to be displayed on a display unit, receives an input for designating a section to be extracted from the moving image, An image processing method.

10. A computer, a screen generation means for generating a screen including a reproduction area for reproducing and displaying a moving image including a plurality of frame images and a missing keypoint display area indicating keypoints of a human body that were not detected in the human body included in the frame images displayed in the reproduction area, and causing the screen to be displayed on a display unit, an input receiving means for receiving an input for designating a section to be extracted from the moving image, A program that functions as.

Citation Information

Patent Citations

  • Image retrieving apparatus, image retrieving method, and setting screen used therefor

    JP2019091138A

  • Image processing device, image processing method, and non-transitory computer-readable medium having image processing program stored thereon

    WO2021084677A1