Image processing device, image processing method, and program

An AI-driven image processing system addresses the inefficiencies in isolating specific members in video by using facial and skeletal recognition to track and correct errors, ensuring accurate extraction of desired images.

WO2025182601A1PCT designated stage Publication Date: 2025-09-04SONY MUSIC ENTERTAINMENT (JAPAN) INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/004869
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-14
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing methods for isolating specific members in video footage, such as those involving idol groups, are costly and inefficient, especially when members face away or are partially obscured, making it difficult to track and extract their images accurately.

Method used

An image processing system utilizing AI-based facial and skeletal recognition to identify and track specific individuals in video frames, assigning unique IDs, and correcting errors in real-time to ensure accurate extraction of the desired member's image.

Benefits of technology

Enables efficient and cost-effective generation of videos focusing on specific individuals by accurately tracking and correcting errors in facial recognition, ensuring continuous extraction of the intended member's image despite occasional misrecognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025004869_04092025_PF_FP_ABST
    Figure JP2025004869_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an image processing device, an image processing method, and a program that make it possible to easily generate a video in which a specific person is captured. The image processing device according to one aspect of the present technology acquires face information and skeleton information of each object captured in each frame constituting an entire video, calculates the coordinates of a specific object in each frame on the basis of the face information and the skeleton information, and extracts the video of the specific object from the entire video on the basis of the calculated coordinates. The present technology can be applied to a device for creating a tracking video of an individual member from a concert video of a group such as an idol group and a band composed of a plurality of members.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing method, and program

[0001] The present technology relates to an image processing device, an image processing method, and a program, and in particular to an image processing device, an image processing method, and a program that enable easy generation of an image in which a specific person appears.

[0002] Many people engage in activities called oshikatsu, which involve supporting their favorite idols, anime characters, etc. Fans who support real groups consisting of multiple members, such as idol groups or bands, usually have a specific favorite member, and they support their favorite member as the main focus of their activities.

[0003] For example, when watching a live performance, fans who have a favorite member may want to watch only the scenes in which that member appears, rather than scenes in which all members appear. To meet this need, it is common to prepare a camera for each member, or to edit footage shot with one camera and then cut out from each frame only the areas in which each member appears.

[0004] Because preparing a camera for each member and manually editing is costly, one proposed technology uses facial recognition to track the areas in which specific members appear and automatically cut them out from a single high-resolution video.

[0005] International Publication No. 2020-100664

[0006] When producing videos using facial recognition, it can be difficult to keep track of a specific member if they are facing away or to the side and their face is not visible.

[0007] The present technology has been developed in light of these circumstances, and makes it possible to easily generate an image in which a specific person appears.

[0008] An image processing device according to one aspect of the present technology includes an acquisition unit that acquires facial information and skeletal information of each object that appears in each frame that constitutes an overall video, and an image processing unit that calculates the coordinates of a specific object in each frame based on the facial information and the skeletal information, and cuts out an image of the specific object from the overall video based on the calculated coordinates.

[0009] In one aspect of the present technology, facial information and skeletal information of each object appearing in each frame that constitutes the overall video is obtained, the coordinates of a specific object in each frame are calculated based on the facial information and the skeletal information, and an image of the specific object is cut out from the overall video based on the calculated coordinates.

[0010] 12. FIG. 13 is a diagram showing a flow of video production according to an embodiment of the present technology. FIG. 14 is a diagram showing details of processing using AI. FIG. 15 is a diagram showing an example of person identification based on the trajectory of skeletal movement. FIG. 16 is a diagram showing the flow of the automatic PID numbering and face registration phase. FIG. 17 is a diagram showing an example of a face image. FIG. 18 is a diagram showing the flow of the tracking phase. FIG. 19 is a diagram showing an example of an error in face recognition AI. FIG. 20 is a diagram showing another example of an error in face recognition AI. FIG. 21 is a diagram showing an example of correction of a tracking target. FIG. 22 is a block diagram showing an example of the functional configuration of an image processing device. FIG. 23 is a block diagram showing an example of the configuration of a tracking video generation unit. FIG. 24 is a flowchart illustrating a series of processes for generating a tracking video. FIG. 25 is a flowchart illustrating the automatic PID numbering and face registration processing performed in step S1 of FIG. 12. FIG. 26 is a flowchart illustrating the tracking processing performed in step S2 of FIG. 25. FIG. 26 is a diagram showing an example of correction when face recognition fails midway. FIG. 27 is a flowchart illustrating cutout coordinate calculation processing 1. FIG. 28 is a diagram showing an example of correction when a new TID appears. FIG. 29 is a flowchart illustrating cutout coordinate calculation processing 2. FIG. 29 is a diagram showing an example of correction when erroneous face recognition occurs for a certain period of time. FIG. 29 is a flowchart illustrating cutout coordinate calculation processing 3. FIG. 29 is a diagram showing an example of correction when erroneous face recognition and a change of TID occur simultaneously. FIG. 10 is a diagram illustrating detection of a suspicious recognition result and detection of an erroneous recognition. FIG. 11 is a flowchart illustrating cutout coordinate calculation processing 4. FIG. 12 is a diagram illustrating a first example of the flow of the tracking phase. FIG. 13 is a diagram illustrating a second example of the flow of the tracking phase. FIG. 14 is a diagram illustrating a third example of the flow of the tracking phase. FIG. 15 is a diagram illustrating an example of a selection screen for a tracking target. FIG. 16 is a diagram illustrating an example of an editing screen. FIG. 17 is a diagram illustrating an enlarged view of a seek bar. FIG. 18 is a diagram illustrating an example of an editing screen. FIG. 19 is a diagram illustrating an example of correction of a tracking video. FIG. 20 is a diagram illustrating an example of correction of cutout coordinates. FIG. 21 is a diagram illustrating an example of decoding in GOP units. FIG. 22 is a diagram illustrating an example of correction of cutout coordinates. FIG. 23 is a diagram illustrating an example of the configuration of an image processing system. FIG. 24 is a block diagram illustrating an example of the configuration of a computer. FIG. 25 is a diagram illustrating an example of display of a tracking video.

[0011] Hereinafter, embodiments of the present technology will be described. The description will be made in the following order: 1. Overview of the present technology 2. Configuration of image processing device 3. Operation of image processing device 4. Flow of tracking phase 5. Editing of tracking target 6. Modified example

[0012] <<Outline of the Present Technology>> <Regarding AI Used for Image Processing> FIG. 1 is a diagram showing the flow of video production according to an embodiment of the present technology.

[0013] This technology detects areas in which specific people appear in each frame of a video (movie) that contains multiple people, and produces the video by cutting out only the areas in which the specific people appear. AI (Artificial Intelligence) is used to detect areas in which specific people appear. Note that using AI means performing processing using an inference model such as a neural network generated by machine learning. Hereinafter, the inference model generated by machine learning will be referred to simply as AI, where appropriate.

[0014] In the example of FIG. 1, live video of an idol group consisting of five members is prepared as the video to be processed. The video to be processed is video with a high resolution, such as 8K. For example, the video to be processed is video taken from a bird's-eye view by a camera installed so that the entire stage where the live performance is being performed is included in the angle of view. The video to be processed is a full video that shows multiple members.

[0015] Each frame of the entire video to be processed is used as input to the AI, as indicated by arrow A1. Each time a frame is input, the coordinates of a specific person are obtained as the AI's output, as indicated by arrow A2. The coordinates obtained as the AI's output indicate the position of the specific person to be tracked in the frame used for input.

[0016] As indicated by arrow A3, image processing is performed using the same image as the entire image used as input for the AI ​​as the cutout source, and as indicated by arrow A4, an image showing a specific person is generated. The image showing a specific person is generated by cutting out a partial area specified by coordinates obtained using AI from the cutout source image.

[0017] By performing the above process on each frame of the original video, a tracking video is generated that continuously tracks a specific person. In the example of Figure 1, a tracking video is generated that shows only one member who is in the center of the five members. It is possible to specify one or more arbitrary members as tracking targets and generate a tracking video that shows only the arbitrary members.

[0018] FIG. 2 shows details of the process using AI.

[0019] As shown in Figure 2, a skeleton recognition AI, which is a first inference model for skeleton recognition, and a face recognition AI, which is a second inference model for face recognition, are used. Each frame of the entire video is input to the skeleton recognition AI, as indicated by arrow A11.

[0020] In skeletal recognition AI, the skeletal coordinates of each person in the input frame are recognized (inferred) as shown in speech bubble #1. For example, the coordinates of multiple skeletal points such as the head, waist, shoulders of both arms, hands, and the tips of both feet are recognized for each person.

[0021] Additionally, a TrackID (TID) is assigned to each person whose skeletal coordinates have been recognized. The TID is an ID used to track a person using skeletal coordinates. In skeletal recognition AI, as shown in Figure 3, people appearing in each frame are identified based on the trajectory of the movement of the skeletal points. The same TID is assigned to each identified person in each frame. TID assignment is performed without using facial information.

[0022] In this way, the skeleton recognition AI is an inference model that takes each frame of the entire video as input and outputs skeletal coordinates, which are coordinate information of the skeletal points of the people appearing in each frame, and TID, which is skeletal information indicating the results of person identification based on the skeletal points. The output of the skeleton recognition AI is input to the face recognition AI, as shown by arrow A12 in Figure 2. In addition, the face image of each person identified by the skeletal coordinates is input to the face recognition AI.

[0023] During face registration, the face recognition AI assigns a Person ID (PID) to each person assigned a TID, as shown in speech bubble #2. For example, PID=1 is assigned to the person with TID=1 (the person to whom TID=1 is assigned), PID=2 is assigned to the person with TID=2, and so on, in the order in which TIDs were assigned.

[0024] Furthermore, face images are registered in association with PIDs. For example, the face image of a person with PID=1 is registered in association with PID=1, and the face image of a person with PID=2 is registered in association with PID=2. After the face images are registered, it becomes possible to identify people appearing in each frame based on their faces. People are identified by comparing the faces of people appearing in each frame with registered face images.

[0025] As indicated by arrow A13, the face recognition AI outputs the skeletal coordinates of each person and a combination of PID and TID. In the example of FIG. 2, the information for person A is output as a combination of skeletal coordinates (x, y) = (0.1, 0.4), PID = 1, and TID = 1, and the information for person B is output as a combination of skeletal coordinates (x, y) = (0.3, 0.3), PID = 2, and TID = 2. Similarly, skeletal coordinates and combinations of PID and TID are output for person C, person D, and person E. Although the example of FIG. 2 shows only one coordinate as the skeletal coordinate for each person, in reality, the coordinates of multiple skeletal points are obtained.

[0026] In this way, face recognition AI is an inference model that takes as input the facial image, skeletal coordinates, and TID of the person appearing in each frame, and outputs the skeletal coordinates and a combination of PID and TID. When identifying a person, the PID output by face recognition AI is facial information that indicates the results of person identification based on the face. For example, a PID of -1 is assigned to people who could not be identified (people who cannot be recognized).

[0027] The two AIs described above, the skeletal recognition AI and face recognition AI, are used to generate tracking video. A single inference model having the functions of skeletal recognition AI and face recognition AI may be prepared and used to generate tracking video. The series of processes for calculating the cut-out coordinates used to generate tracking video includes the automatic PID numbering and face registration phase and the tracking phase. The cut-out coordinates are coordinates that specify the area in the frame of the entire video that is the source of the cut-out, in which the person to be tracked is captured.

[0028] <Automatic PID Numbering & Face Registration Phase> Fig. 4 is a diagram showing the flow of the automatic PID numbering & face registration phase. Below, explanations that overlap with the above explanations will be omitted as appropriate.

[0029] As shown in the upper left of Figure 4, the automatic PID numbering and face registration phase is processed sequentially for each frame generated by decoding the video data of the entire video. Frames showing multiple members are read one by one and used as input for the skeleton recognition AI, as indicated by arrow A21. The skeleton recognition AI outputs the skeleton coordinates and TIDs of the people appearing in the input frames, as indicated by arrow A22. In the example of Figure 4, starting with the people appearing on the left side of the frame, TIDs 1 to 5 are assigned, and the skeleton coordinates of each person are output.

[0030] The output of the skeleton recognition AI is input to the face recognition AI along with the face image, as indicated by arrow A23. As shown in speech bubble #11, the face recognition AI assigns a PID to each person in the order in which the TIDs were assigned. The face image is also registered in association with the PID.

[0031] 4, as indicated by arrow A24, PID=1 is assigned to the person with TID=1, and PID=2 is assigned to the person with TID=2. PIDs=3 to 5 are also assigned to the people with TID=3 to 5, respectively.

[0032] FIG. 5 is a diagram showing an example of a face image.

[0033] When five people are assigned PIDs 1 to 5, facial images are associated with each PID and registered, as shown in FIG.

[0034] The above-described automatic PID numbering and face registration phase processing is repeated until the processing of the final frame of the entire video is completed. If the frame of the entire video used in the automatic PID numbering and face registration phase processing is designated as the first frame, in the processing of each first frame, the face image with the higher reliability is linked to the PID, and the face image is updated. The reliability of each face image is calculated based on the face direction, face size, etc., and is used to update the face image. After the automatic PID numbering and face registration phase processing, the tracking phase processing is performed.

[0035] <Tracking Phase> FIG. 6 is a diagram showing the flow of the tracking phase.

[0036] As shown in the upper left of Figure 6, the tracking phase is performed sequentially on each frame generated by decoding video data of the same image as the overall image used in the automatic PID numbering and face registration phase. Frames showing multiple members are read one by one and used as input for the skeleton recognition AI, as indicated by arrow A31. The skeleton recognition AI outputs the skeleton coordinates and TID of the person appearing in the input frame, as indicated by arrow A32.

[0037] The output of the skeleton recognition AI, along with the facial image, is input to the face recognition AI, as indicated by arrow A33. As indicated by speech bubble #21, the face recognition AI performs person identification for each person assigned a TID. Person identification is performed by comparing the face of the person shown in the input frame with each registered facial image. For example, a person whose similarity to a facial image registered in association with PID=1 is equal to or exceeds a threshold value is identified as the person with PID=1.

[0038] In the example of Figure 6, as indicated by arrow A34, the person with TID = 1 is identified as the person with PID = 1, and the person with TID = 2 is identified as the person with PID = 2. The people with TID = 3 to 5 are also identified as the people with PID = 3 to 5, respectively. The face recognition AI outputs the skeleton coordinates of each person appearing in the input frame, and the combination of PID and TID.

[0039] When a person to be tracked is specified, the cut-out coordinates are calculated based on the person's skeletal coordinates and are used to cut out the person from the entire video. The above tracking phase processing is repeated until the processing of the last frame of the entire video is completed. For example, when a person with PID=1 is specified as the tracking target, the coordinates of the area in which the person with PID=1 appears in each frame of the entire video are calculated, and the tracking video is generated.

[0040] If the frames of the entire image used in the tracking phase processing are referred to as second frames, face recognition in the tracking phase is performed by identifying which registered face image each person appears in each second frame. If multiple people appear in an input frame, the same PID is not assigned to multiple people, but each person is assigned a different PID.

[0041] <Regarding misrecognition by facial recognition AI> The person to be tracked is specified by PID. When continuing to track a specific person during the tracking phase processing, basically, it is sufficient to continue tracking the person with the specified PID based on the output of the facial recognition AI. However, the accuracy of facial recognition by facial recognition AI is not 100%, and errors such as misrecognition may occur.

[0042] 7 and 8 are diagrams showing examples of errors in face recognition AI.

[0043] Fig. 7 shows an example of a case where recognition was not possible (failed recognition), and Fig. 8 shows an example of a case where a person was recognized as a different person (misrecognition). The arrangement of dots in Fig. 7 and Fig. 8 indicates frames, and the numbers above the dots indicate frame numbers. The PID and TID below the dots are the PID and TID obtained as the recognition results for a person appearing in each frame.

[0044] - When facial recognition fails midway (Case 1) Figure 7A shows the recognition results for one person. PID=1 and TID=2 are obtained as the recognition results for the first and second frames. This recognition result shows that the person assigned TID=2 by the skeleton recognition AI was recognized by the facial recognition AI as the person with PID=1.

[0045] Furthermore, PID = -1 and TID = 2 are acquired as the recognition results for the third and fourth frames, and PID = 1 and TID = 2 are acquired as the recognition results for the fifth and sixth frames. Figure 7A shows a case in which a person who was recognized (identified) as PID = 1 up until the second frame becomes unrecognizable in the third frame, and then becomes recognizable again in the fifth frame. As shown in the box, the section from the third frame to the fourth frame, where recognition is unsuccessful, is the error section. Here, a small number of frames are used for explanation, but in reality, the error section will be a section of many more frames.

[0046] In facial recognition of live video footage, it may happen that a person's face cannot be recognized temporarily because they are facing away from the camera. Even in this case, it is expected that the face of that person will be recognized again after some time has passed, for example, when they turn to face the camera.

[0047] ・When a new TID appears (Case 2) Figure 7B also shows the recognition results for one person. The person to be tracked is not visible in the first or second frame, so no recognition results have been obtained from the skeleton recognition AI or face recognition AI. The recognition results for the third and fourth frames are PID = -1 and TID = 33. This recognition result indicates that the skeleton recognition AI has not yet been able to recognize the face of the person assigned TID = 33.

[0048] Furthermore, PID = 1 and TID = 33 are acquired as the recognition results for the fifth and sixth frames. Figure 7B shows a case in which face recognition of a person who appears in the third frame is possible by the fifth frame. As shown by the box, the section from the third frame to the fourth frame, where recognition is incorrect, is an error section.

[0049] As such, it can take time for a person to be identified. Typically, it takes several frames for a facial recognition AI to be able to recognize a face.

[0050] - When facial recognition errors occur for a certain period of time (Case 3) The upper row of A in Figure 8 shows the recognition results for person A (Mr. A). Person A is assigned PID=1. For person A, PID=1 and TID=3 are obtained as the recognition results for the first and second frames, and PID=2 and TID=3 are obtained as the recognition results for the third and fourth frames. Furthermore, PID=1 and TID=3 are obtained as the recognition results for the fifth and sixth frames, the same as the recognition results for the first two frames. The skeleton recognition result for person A is constant at TID=3.

[0051] The lower row of A in Fig. 8 shows the recognition results for person B (Mr. B), who is different from person A. Person B is assigned PID=2. For person B, PID=1 and TID=4 are obtained as the recognition results for the third and fourth frames, and PID=2 and TID=4 are obtained as the recognition results for the fifth frame and beyond. The skeleton recognition result is constant at TID=4.

[0052] 8A shows a case in which the facial recognition results for person A and person B are switched between the third and fourth frames, and then switched back to the fifth frame. As shown by the box, the section from the third frame to the fourth frame, in which the face is incorrectly recognized, is the error section.

[0053] - When a face recognition error and a TID change occur simultaneously (Case 4) The upper row of B in Figure 8 shows the recognition results for person A. For person A, PID = 1 and TID = 3 are obtained as the recognition results for frames 1 and 2, and PID = 2 and TID = 4 are obtained as the recognition results for frames 3 and 4. Furthermore, for frames 5 and 6, the face recognition result obtained is PID = 1, the same as the recognition results for frames 1 through 2, and the skeleton recognition result obtained is TID = 4, the same as frames 3 and 4.

[0054] The lower row of B in Fig. 8 shows the recognition results for person B. For person B, PID = 2 and TID = 4 are obtained as the recognition results for the second frame, and PID = 1 and TID = 3 are obtained as the recognition results for the third and fourth frames. Furthermore, as the recognition results for the fifth and sixth frames, PID = 2 is obtained as the face recognition result, the same as the recognition results for the first two frames, and TID = 3 is obtained as the skeleton recognition result, the same as the third and fourth frames.

[0055] 8B shows a case where the PID and TID of person A and person B are swapped between frames 3 and 4, and then the facial recognition results are restored in frame 5. As shown by the box, the section from frame 3 to frame 4, where the face is incorrectly recognized, is the error section.

[0056] In face recognition and skeletal recognition for live video footage, the recognition results may be swapped when a person overlaps or approaches. Face recognition is expected to return to normal (become correct) after some time has passed, such as when the person faces forward. However, skeletal recognition will not return to normal, as it is a person identification system based on the trajectory of the skeleton.

[0057] In this way, when performing face recognition on live video footage or the like, errors such as those in cases 1 to 4 may occur. With this technology, even if an error occurs in face recognition, the tracking target is corrected in order to continue tracking the same person, on the premise that correct face recognition will be possible after a few frames.

[0058] Correction of Tracking Target FIG. 9 is a diagram showing an example of correction of a tracking target.

[0059] The top row of Figure 9 shows the recognition results for people with PID=1 to 3. The horizontal direction is the frame direction. As shown by the box, a face misrecognition occurred in frames 3 to 6, with the person with TID=2 being recognized as the person with PID=1. Focusing on the person with PID=1, as shown by the colored area, the combination of PID and TID changes between frames 3 and 7. If the crop coordinates are calculated using the person with PID=1 as the tracking target, the person appearing in the tracking image will switch between frames 3 and 7.

[0060] In this technology, in order to be able to extract tracking video in which the same person appears, a suspicious recognition result (a recognition result that may be erroneous) is detected at the timing of the third frame when the combination of PID and TID changes, as shown in speech bubble #31. Also, at the timing of the seventh frame when the combination of PID and TID changes, erroneous recognition of the recognition results of the preceding few frames is detected (it is detected that the recognition result is erroneous), as shown in speech bubble #32.

[0061] After the misrecognition is detected, the tracking target is corrected so that it tracks the person who is believed to be correct, as shown in speech bubble #33. In the example of Fig. 9, in the section from the third frame to the sixth frame, the tracking target is corrected so that it tracks the person with PID = 2.

[0062] As described above, even if an error occurs, facial recognition by the facial recognition AI will recover after several frames. In the example of Figure 9, it can be considered that facial recognition by the facial recognition AI recovered at the timing of the seventh frame, which is the timing of the second change in the PID and TID combination. Therefore, based on the fact that the TID of the person with PID = 1 in the seventh frame is TID = 3, the tracking target is corrected so that the person with PID = 2, who is assigned TID = 3 in the section from the third frame to the sixth frame where the incorrect recognition was detected, is tracked. The person recognized as PID = 1 in the seventh frame and the person recognized as PID = 2 in the section from the third frame to the sixth frame are determined to be the same person based on the fact that their TIDs are the same.

[0063] Not only PID and TID combinations that have changed, but also PID and TID at the timing when two people overlap are detected as suspicious recognition results. In other words, suspicious recognition results are detected based on at least one of PID, TID, and skeleton coordinates.

[0064] As described above, processing using a tracking algorithm, including detection of suspicious recognition results, detection of erroneous recognition, and correction of the tracking target, is performed based on the recognition results of each person appearing in each frame.

[0065] The time between the detection of a suspicious recognition result and the detection of a false recognition is set in advance based on the expected time for the facial recognition AI to recover. The time it takes for facial recognition to recover varies depending on the characteristics of the song and group being played in the overall video. For example, in the case of a song in which the performers spend a long time facing forward while playing or a video of a group with little human movement, the time it takes for the facial recognition AI to recover will be shorter. The user will set the desired time depending on the characteristics of the song or group. The period between the detection of a suspicious recognition result and the detection of a false recognition may also be set using the number of frames rather than time.

[0066] If a misrecognition is detected within the threshold number of frames, the tracking target is corrected, but if no misrecognition is detected, the tracking target is not corrected. In this case, the recognition result by the facial recognition AI is considered correct, the tracking target is selected, and the cropping coordinates are calculated.

[0067] The series of processes for appropriately correcting the tracking target and calculating the extraction coordinates as described above will be described later with reference to a flowchart.

[0068] <<Configuration of Image Processing Device>> Fig. 10 is a block diagram showing an example of the functional configuration of the image processing device 1. The image processing device 1 is configured by a computer such as a PC. Other types of devices such as a tablet terminal or a smartphone may also be used as the image processing device 1. For example, the image processing device 1 is used by staff who create tracking videos of each member of an idol group.

[0069] The image processing device 1 is composed of an image acquisition unit 11, a tracking video generation unit 12, a tracking video recording unit 13, and a tracking video playback unit 14. At least some of the functional units shown in Fig. 10 are realized by the computer of the image processing device 1 executing a predetermined program.

[0070] The image acquisition unit 11 acquires a file of the entire video to be processed. For example, a file in a predetermined format such as XAVC is input to the image processing device 1 and acquired by the image acquisition unit 11. The image acquisition unit 11 decodes the video data of the acquired file and performs preprocessing such as color conversion and resolution conversion on each frame obtained by decoding. Each preprocessed frame is supplied to the tracking video generation unit 12.

[0071] The tracking video generation unit 12 detects an area in each frame supplied from the image acquisition unit 11 in which a specific person designated as a tracking target appears. The tracking video generation unit 12 cuts out the detected area from each frame and generates a tracking video of the specific person. The tracking video generation unit 12 outputs the tracking video data to the tracking video recording unit 13 to record it. For example, data of the entire video from which the tracking video was cut out is also supplied to the tracking video recording unit 13 and recorded.

[0072] The tracking video recording unit 13 is configured by a recording medium such as a computer HDD or SSD, etc. The tracking video recording unit 13 records the data of the tracking video generated by the tracking video generating unit 12 together with the data of the entire video.

[0073] The tracking video playback unit 14 reads and acquires the tracking video data from the tracking video recording unit 13 and plays it back. The tracking video played back by the tracking video playback unit 14 is displayed on a display, which is a display unit (not shown) of the image processing device 1. The image processing device 1, which is configured by a PC or the like, is provided with a display such as an LCD used to display various screens. An external display connected to the PC may be used to display the screen instead of an internal display built into the PC. The tracking video playback unit 14 functions as a display control unit that controls the display of the display unit.

[0074] FIG. 11 is a block diagram showing an example of the configuration of the tracking video generation unit 12.

[0075] The tracking video generation unit 12 is made up of a skeleton recognition unit 21, a face recognition unit 22, a cutout coordinate calculation unit 23, and a cutout processing unit 24. Data of each frame output from the image acquisition unit 11 is input to the skeleton recognition unit 21.

[0076] The skeleton recognition unit 21 has a skeleton recognition AI. Using the skeleton recognition AI, the skeleton recognition unit 21 recognizes the skeleton of each person appearing in the input frame. That is, the skeleton recognition unit 21 inputs the input frame to the skeleton recognition AI and acquires the skeleton coordinates and TID output from the skeleton recognition AI. Skeleton recognition by the skeleton recognition unit 21 is performed in the processing of the automatic PID numbering & face registration phase and the tracking phase. Information on the skeleton coordinates and TID, which are the results of skeleton recognition by the skeleton recognition unit 21, is supplied to the face recognition unit 22 together with the face image.

[0077] The face recognition unit 22 has a face recognition AI. The face recognition unit 22 uses the face recognition AI to recognize the face of each person. That is, during processing of the automatic PID numbering and face registration phase, the face recognition unit 22 inputs the skeleton coordinates and TID of each person appearing in the input frame to the face recognition AI and assigns a PID to each person. The face recognition unit 22 also registers the face image of each person by linking it to the PID.

[0078] Furthermore, during processing of the tracking phase, the face recognition unit 22 inputs the skeletal coordinates and TID of each person appearing in the input frames to the face recognition AI, and performs person identification for each person. Information on the skeletal coordinates and the combination of PID and TID, including the results of face recognition by the face recognition unit 22, is supplied to the cutout coordinate calculation unit 23. The skeleton recognition unit 21 and the face recognition unit 22 implement an acquisition unit that acquires the PID as face information and the TID as skeletal information of each person appearing in each frame constituting the entire video.

[0079] The cutout coordinate calculation unit 23 acquires the result of face recognition by the face recognition unit 22 and calculates cutout coordinates that specify the area in which the person to be tracked appears. When calculating the cutout coordinates, the tracking target is appropriately corrected using a tracking algorithm. Information on the cutout coordinates calculated by the cutout coordinate calculation unit 23 is supplied to the cutout processing unit 24.

[0080] The cut-out processing unit 24 cuts out an area specified by the cut-out coordinates from each frame of the entire video supplied from the image acquisition unit 11, and generates a tracking video. The tracking video generated by the cut-out processing unit 24 is output to and recorded in the tracking video recording unit 13. The cut-out coordinate calculation unit 23 and the cut-out processing unit 24 realize a video processing unit that calculates the coordinates of a specific person in each frame based on the PID and TID, and cuts out the video of the specific person from the entire video based on the calculated coordinates.

[0081] <<Operation of Image Processing Device>> Here, we will explain the operation of the image processing device 1 having the above configuration. The processing of each step shown in the following flowchart is not only performed in the order shown in the figure, but also in parallel with the processing of other steps or out of order as appropriate.

[0082] <Tracking Image Generation Processing> FIG. 12 is a flowchart illustrating a series of processing steps for generating a tracking image.

[0083] In step S1, automatic PID numbering and face registration processing is performed by the tracking video generation unit 12. The automatic PID numbering and face registration processing will be described later with reference to the flowchart of FIG.

[0084] In step S2, tracking processing is performed by the tracking image generating unit 12. The tracking processing will be described later with reference to the flowchart of Fig. 14. By the tracking processing, clipping coordinates of the tracking target are calculated.

[0085] In step S3, the cutout processing unit 24 of the tracking video generation unit 12 cuts out the area specified by the cutout coordinates from the frame of the entire video, and generates a tracking video.

[0086] <Automatic PID Numbering & Face Registration Processing> FIG. 13 is a flowchart illustrating the automatic PID numbering & face registration processing performed in step S1 of FIG.

[0087] In step S11, the image acquisition unit 11 plays back video content such as live footage and acquires an entire image to be processed. The image acquisition unit 11 performs preprocessing on the entire image and sequentially outputs frames obtained by the preprocessing to the tracking image generation unit 12.

[0088] In step S12, the skeleton recognition unit 21 of the tracking video generation unit 12 reads one frame of the entire video.

[0089] In step S13, the skeleton recognition unit 21 performs skeleton recognition using a skeleton recognition AI, and acquires the skeleton coordinates and TID of each person.

[0090] In step S14, the face recognition unit 22 performs face recognition using a face recognition AI and assigns a PID to each person.

[0091] In step S15, the face recognition unit 22 registers the face image of each person in association with the PID.

[0092] In step S16, the image acquisition unit 11 determines whether or not the frame is the last. If it is determined in step S16 that the frame is not the last, the process returns to step S11, and the same process is repeated. If it is determined in step S16 that the frame is the last, the process returns to step S1 in FIG. 12, and the subsequent processes are performed.

[0093] <Tracking Process> FIG. 14 is a flowchart illustrating the tracking process performed in step S2 of FIG.

[0094] In step S21, the image acquisition unit 11 plays back the video content and acquires the entire video to be processed.

[0095] In step S22, the skeleton recognition unit 21 of the tracking video generation unit 12 reads one frame of the entire video.

[0096] In step S23, the skeleton recognition unit 21 performs skeleton recognition using a skeleton recognition AI, and acquires the skeleton coordinates and TID of each person.

[0097] In step S24, the face recognition unit 22 performs face recognition using a face recognition AI, identifies each person, and assigns a PID to each person.

[0098] In step S25, the cutout coordinate calculation unit 23 performs a cutout coordinate calculation process, which is a process of appropriately correcting the tracking target by using a tracking algorithm and calculating the cutout coordinates.

[0099] The cutout coordinate calculation unit 23 detects a suspicious recognition result based on a change in the combination of PID and TID or based on the overlap of two people. If it detects that the suspicious recognition result is actually an erroneous recognition, the cutout coordinate calculation unit 23 corrects the tracking target based on the combination of PID and TID so that the correct person continues to be tracked. If necessary, the tracking target is corrected by correcting the PID detected as an erroneous recognition with the correct PID. The cutout coordinate calculation process will be described in detail later.

[0100] In step S26, the image acquisition unit 11 determines whether or not the frame is the last. If it is determined in step S26 that the frame is not the last, the process returns to step S21, and the same process is repeated. If it is determined in step S26 that the frame is the last, the process returns to step S2 in FIG. 12, and the subsequent processes are performed.

[0101] <Cut-out coordinate calculation process> Here, we will explain the details of the process using the tracking algorithm. As mentioned above, there are four types of errors in face recognition AI, cases 1 to 4. Processes to correct the four types of errors are performed on each frame.

[0102] When face recognition fails midway (Case 1) Fig. 15 is a diagram showing an example of correction when face recognition fails midway. The error shown in B of Fig. 15 is the error described with reference to A of Fig. 7.

[0103] In this case, the following information is saved and used to modify the tracking target (A in FIG. 15): Tracking TID

[0104] The tracking TID is the TID for the PID of the tracking target (the TID assigned to the person with the PID of the tracking target).

[0105] The correction for each frame is basically performed as follows: (1) Save the tracking TID. (2) If the TID for the PID changes, update the tracking TID. (3) If there is no PID to track (PID = -1), select the person with the saved tracking TID as the tracking target.

[0106] When PID=1 is designated as the tracking target, TID=2 is saved as the tracking TID at the timing of the first and second frames of B in FIG. 15 (processing (1) above).

[0107] Furthermore, since there is no PID of the tracking target at the timing of the third and fourth frames, the person with TID=2 is selected as the tracking target based on the saved tracking TID (processing (3) above).

[0108] In the fifth and sixth frames, when the recognition results for PID=1 and TID=2 are obtained, the person with PID=1 is used as the tracking target.

[0109] The combination of PID = -1 and TID = 2, which is the recognition result for the third frame, may be detected as a suspicious recognition result, and correction may be made in response to the detection of an erroneous recognition. In this case, in response to the detection of the combination of PID = 1 and TID = 2 in the fifth frame, an erroneous recognition of the recognition results for the third and fourth frames ahead is detected, and the tracking target is corrected by making the person with TID = 2 the tracking target in the third and fourth frames. If the PID of the person with TID = 2 changes from PID = -1 to PID = 1, the tracking target is corrected by changing the PID = -1 in the recognition results for the third and fourth frames to PID = 1.

[0110] 16 is a flowchart illustrating the cutout coordinate calculation process 1. The cutout coordinate calculation process 1 is a process that includes the above-described correction when face recognition fails midway.

[0111] In step S51, the cutout coordinate calculation unit 23 reads the AI ​​output of the current frame.

[0112] In step S52, the cutout coordinate calculation unit 23 determines whether or not a person with the tracking target PID is present.

[0113] If it is determined in step S52 that the person of the tracking target PID does not exist, the cutout coordinate calculation unit 23 determines in step S53 whether or not the tracking TID is stored.

[0114] If it is determined in step S53 that the tracking TID is stored, the cutout coordinate calculation unit 23 determines in step S54 whether or not a person in the tracking TID exists.

[0115] If it is determined in step S54 that a person with the tracking TID exists, in step S55, the cutout coordinate calculation unit 23 selects the person with the tracking TID as a tracking target. The cutout coordinate calculation unit 23 appropriately modifies the PID of the selected person using the tracking target PID.

[0116] If it is determined in step S53 that the tracking TID is not stored, or if it is determined in step S54 that the person with the tracking TID does not exist, the cropping coordinate calculation unit 23 determines in step S56 that the error cannot be handled. In this case, for example, a countermeasure is taken using another correction algorithm described later. The cropping coordinates of the previous frame may be used, or, as described later, a frame of the entire video may be used.

[0117] On the other hand, if it is determined in step S52 that a person with the tracking target PID exists, in step S57, the cutout coordinate calculation unit 23 uses the AI ​​output of the current frame and saves the output TID as the tracking TID. If the value of the output TID has changed, the tracking TID is updated.

[0118] In step S55, after selecting the person of the tracking TID as the tracking target, or after the processing of step S57, in step S58, the cutout coordinate calculation unit 23 calculates cutout coordinates that specify the area in which the tracking target person is captured and stores them in a buffer. After that, the process returns to step S25 in Fig. 14, and the subsequent processing is performed.

[0119] Assume that PID=1 is specified as the tracking target. When the current frame is the first frame of B in FIG. 15, after it is determined in step S52 that a person with the tracking target PID is present, in step S57, the AI ​​outputs PID=1 and TID=2 are used as is. Furthermore, TID=2 is saved as the tracking TID. In step S58, the extraction coordinates of the person with PID=1 are calculated. The same process is performed when the current frame is the second frame.

[0120] When the current frame is the third frame, after it is determined in step S52 that a person with the tracking target PID does not exist, TID=2 is saved as the tracking TID, and so it is determined in step S53 that the tracking TID is saved. Also, it is determined in step S54 that a person with the tracking TID exists, and after selecting the person with TID=2 as the tracking target in step S55, the clipping coordinates of the person with TID=2 are calculated in step S58. The same process is performed when the current frame is the fourth frame.

[0121] After selecting the person with TID=2 as the tracking target in the third and fourth frames, PID=-1 may be corrected to PID=1.

[0122] When a new TID appears (Case 2) Fig. 17 is a diagram showing an example of correction when a new TID appears. The error shown in Fig. 17B is the error described with reference to Fig. 7B.

[0123] In this case, the following information is saved and used to correct the tracking target (A in Figure 17): ・TID of the observed target and the frame number at the time of first observation ・AI output after detecting suspicious recognition results (for each frame)

[0124] The target TID is a newly appeared TID. The frame number at the time of first observation is the number of the frame in which the target TID was recognized.

[0125] Corrections for each frame are basically performed as follows: (1) When a combination of a new TID and PID = -1 appears (when the face recognition result of the person assigned the new TID is PID = -1), the observed TID, the frame number at the time of first observation, and the AI ​​output are saved, and the tracking target is put on hold. (2) When a PID other than PID = -1 is assigned to the person with the observed TID, the PID of the previous frame in which PID = -1 was recognized is corrected to the same PID as the newly assigned PID.

[0126] When PID=1 is designated as the tracking target, a new TID combination of TID=33 and PID=-1 appears at the timing of the third frame in B of Figure 17, so TID=33 is saved as the TID to be observed together with the frame number (processing (1) above).

[0127] Furthermore, at the timing of the fifth frame when PID=1 is assigned to the person with TID=33, PID=-1 in the third and fourth frames is corrected to PID=1 (process (2) above).

[0128] In this example, the combination of PID=-1 and TID=33, which is the recognition result for the third frame, is detected as a suspicious recognition result. That is, TID=33 is newly acquired, and PID=-1 is acquired as the PID of the person with TID=33, so a suspicious recognition result is detected.

[0129] Furthermore, in response to detecting the combination of PID=1, TID=33 in the fifth frame, erroneous recognition of the recognition results of the third and fourth frames ahead is detected, and the tracking target is corrected by changing the PID=-1 in the recognition results of the third and fourth frames to PID=1. When the PID of the person with TID=33 changes from PID=-1 to PID=1, the tracking target is corrected by changing the PID=-1 in the recognition results of the third and fourth frames to PID=1.

[0130] 18 is a flowchart illustrating the extraction coordinate calculation process 2. The extraction coordinate calculation process 2 is a process that includes the above-mentioned correction when a new TID appears.

[0131] In step S61, the cutout coordinate calculation unit 23 reads the AI ​​output of the current frame.

[0132] In step S62, the extraction coordinate calculation unit 23 determines whether a new TID has been detected or whether an observation target TID exists.

[0133] If it is determined in step S62 that a new TID has been detected or that a target TID exists, then in step S63, the extraction coordinate calculation unit 23 saves the first observed TID as the target TID and saves the current frame number.

[0134] In step S64, the extraction coordinate calculation unit 23 determines whether the number of frames in the information storage section is within a specified number. The information storage section is the section after the frame where storage of the observation target TID, etc., begins. The section from the detection of the suspicious recognition result to the detection of the erroneous recognition corresponds to the information storage section.

[0135] If it is determined in step S64 that the number of frames in the information storage section is within the specified number, then in step S65, the cutout coordinate calculation unit 23 stores the AI ​​output of the current frame.

[0136] In step S66, the cutout coordinate calculation unit 23 determines whether a PID has been assigned to the person of the observation target TID.

[0137] If it is determined in step S66 that a PID has been assigned to the person of the observation target TID, then in step S67 the cutout coordinate calculation unit 23 corrects PID=-1 in the information storage section with the assigned PID.

[0138] On the other hand, if it is determined in step S64 that the number of frames in the information storage section is not within the specified number, then in step S68, the cutout coordinate calculation unit 23 cancels storage using each AI output of the frames in the information storage section.

[0139] If it is determined in step S62 that a new TID has not been detected and that no observation target TID exists, then in step S69, the extraction coordinate calculation unit 23 uses the AI ​​output of the current frame.

[0140] After the processing of step S67, step S67, or step S69, in step S70, the cutout coordinate calculation unit 23 calculates cutout coordinates that designate the area in which the person to be tracked appears, and stores the calculated coordinates in the buffer.

[0141] On the other hand, if it is determined in step S66 that a PID has not been assigned to the person of the observation target TID, the extraction coordinate calculation unit 23 switches to the next frame in step S71 while keeping the information stored.

[0142] After the process of step S70 or step S71, the process returns to step S25 in FIG. 14, and the subsequent processes are carried out.

[0143] Assume that PID=1 is specified as the tracking target. If the current frame is the third frame in B of FIG. 17, it is determined in step S62 that a new TID has been detected. In step S63, TID=33, which is the first observed TID, is saved as the observation target TID. The same process is performed when the current frame is the fourth frame.

[0144] 17B, ​​it is determined in step S62 that the observed TID exists, and the processes from step S63 onward are carried out. After PID=1 and TID=33, which are the AI ​​outputs for the fifth frame, are saved in step S65, it is determined in step S66 that a PID has been assigned to the person with the observed TID, and in step S67, the PID=-1 for the person with TID=33 in the third and fourth frames is corrected using PID=1, which is the AI ​​output for the fifth frame.

[0145] In the third and fourth frames, the corrected cutout coordinates of the person with PID=1 are calculated. In the fifth and sixth frames, the AI ​​output is used as is to calculate the cutout coordinates of the person with PID=1.

[0146] - When a face recognition error occurs for a certain period of time (Case 3) Fig. 19 is a diagram showing an example of correction when a face recognition error occurs for a certain period of time. The error shown in B of Fig. 19 is the error described with reference to A of Fig. 8.

[0147] In this case, the following information is saved and used to correct the tracking target (A in Figure 19): The PID and TID combination one frame before the frame in which the combination changed, and the frame number of the frame in which the combination changed The PID and TID combination two steps before, and the frame number when that combination was observed AI output after detecting a suspicious recognition result (every frame)

[0148] The combination of PID and TID of the previous frame is the combination of PID and TID of the frame immediately preceding the current frame when the combination of PID and TID, which is the recognition result of the current frame input as the processing target, has changed.

[0149] The PID and TID combination two steps back is the PID and TID combination two steps back from the current combination when the PID and TID combination has been changed multiple times. The PID and TID combination two steps back is saved along with the frame number of the frame for which the recognition result for that combination was obtained.

[0150] Corrections for each input frame are basically performed as follows: (1) The AI ​​output and frame number of the frame one frame before (one step before) the frame in which the PID-TID combination changed are saved. (2) If the PID-TID combination changes again, the PID-TID combination saved as the previous step is treated as the combination two steps before, and the current combination (the changed combination) is compared with the combination two steps before. (3) If the two compared combinations are the same, it is determined that the PID for the section between the frame in which the recognition result for the combination two steps before was obtained and the current frame is incorrect, and the person assigned the same value as the current TID value is selected as the tracking target. (4) If the two compared combinations are different, it is determined that the PID for the section between the frame in which the recognition result for the combination two steps before was obtained and the current frame are correct, and the AI ​​output is used as is to select the tracking target.

[0151] When PID=1 is specified as the tracking target, the combination of PID and TID changes at the timing of the third frame in B of Figure 19, so the AI ​​output from one frame before, PID=1, TID=3, is saved (processing (1) above).

[0152] Furthermore, since the combination of PID and TID has changed again at the timing of the fifth frame, the current combination of PID=1, TID=3 is compared with the combination two steps earlier, PID=1, TID=3 in the second frame (process (2) above). If the recognition result of the fifth frame is used as the reference, the combination of PID=2, TID=3, which is the recognition result of the third and fourth frames, is the combination one step earlier, and the combination of PID=1, TID=3, which is the recognition result of the second frame, is the combination two steps earlier.

[0153] Because the combination in the fifth frame is the same as the combination in the second frame, it is determined that the PID recognition results for the third and fourth frames are incorrect. As the tracking target for the third and fourth frames, the person with PID=2, which is assigned the same value as the current TID value, TID=3, is selected (process (3) above).

[0154] In this example, the combination of PID=2 and TID=3, which is the recognition result for the third frame, is detected as a suspicious recognition result. In other words, when the TID of a person assigned PID=1 remains unchanged at TID=3, a suspicious recognition result is detected in response to a change in the combination of PID and TID.

[0155] Furthermore, if the combination of PID=1 and TID=3 is detected in the fifth frame and determined to be the same as the combination from two frames earlier, an erroneous recognition of the recognition results in the preceding third and fourth frames is detected. The person with PID=2, who is assigned TID=3, is selected as the tracking target for the third and fourth frames, and the tracking target is corrected. Taking advantage of the fact that the TID does not change when face recognition returns (only the PID changes) and that the PID and TID combination returns to the combination before the error period after the error period, if the PID of the person with TID=3 changes to PID=1, the tracking target is corrected by changing PID=2 in the recognition results for the third and fourth frames to PID=1.

[0156] 20 is a flowchart illustrating the cutout coordinate calculation process 3. The cutout coordinate calculation process 3 is a process that includes the above-described correction when erroneous face recognition occurs in a certain section.

[0157] In step S81, the cutout coordinate calculation unit 23 reads the AI ​​output of the current frame.

[0158] In step S82, the cutout coordinate calculation unit 23 determines whether or not the combination of PID with TID of the previous frame has changed.

[0159] If it is determined in step S82 that the combination of PID and TID of the previous frame has changed, then in step S83 the cutout coordinate calculation unit 23 saves the combination of PID and TID of the previous frame and the current frame number.

[0160] In step S84, the cutout coordinate calculation unit 23 determines whether the number of frames in the information storage section is within a specified number.

[0161] If it is determined in step S84 that the number of frames in the information storage section is within the specified number, then in step S85, the cutout coordinate calculation unit 23 stores the AI ​​output of the current frame.

[0162] In step S86, the cutout coordinate calculation unit 23 determines whether or not a combination from two stages before that includes the same TID is stored.

[0163] If it is determined in step S86 that a combination from two steps ago containing the same TID has been saved, then in step S87 the cutout coordinate calculation unit 23 determines whether the combination of PID and TID from two steps ago and the current combination are the same.

[0164] If it is determined in step S87 that the combination of PID and TID two steps ago and the current one is the same, in step S88, the cutout coordinate calculation unit 23 selects the person with the same TID as the current TID as the tracking target for the information storage section.

[0165] If it is determined in step S84 that the number of frames in the information storage section is not within the specified number, or if it is determined in step S87 that the combination of PID and TID two steps ago and the current combination is not the same, then in step S89 the extraction coordinate calculation unit 23 uses each AI output of the frames in the information storage section and cancels storage.

[0166] If it is determined in step S82 that the combination of PIDs for TIDs in the previous frame has not changed, then in step S90 the cropping coordinate calculation unit 23 determines whether a combination two steps prior that includes the same TID has been saved, and if it is determined that a combination two steps prior that includes the same TID has not been saved, then in step S91 the AI ​​output of the current frame is used. If it is determined in step S90 that a combination two steps prior that includes the same TID has been saved, then the process proceeds to step S84, and the subsequent processes are performed.

[0167] After the processing of step S88, step S89, or step S91, in step S92, the cutout coordinate calculation unit 23 calculates cutout coordinates that designate the area in which the person to be tracked appears, and stores the calculated coordinates in a buffer.

[0168] If it is determined in step S86 that a combination containing the same TID from two steps prior has not been saved, the cropping coordinate calculation unit 23 switches to the next frame in step S93 while keeping the information saved. Here, the PID and TID combination from one frame prior and its frame number are changed to the information from two steps prior. In addition, the current AI output (after the combination has changed) is saved as the information from one frame prior.

[0169] After the process of step S92 or step S93, the process returns to step S25 in Fig. 14, and the subsequent processes are carried out. Through the above process, the error shown in B in Fig. 19 is corrected and the tracking target is selected.

[0170] - When a face recognition error and a TID change occur simultaneously (Case 4) Fig. 21 is a diagram showing an example of correction when a face recognition error and a TID change occur simultaneously. The error shown in B of Fig. 21 is the error described with reference to B of Fig. 8.

[0171] In this case, the following information is saved and used to correct the tracking target (A in Figure 21): PID and TID combination from the previous frame Overlap information AI output after overlap detection (every frame)

[0172] The overlap information includes the PID and TID combination of the two overlapping people and the frame number when the overlap was detected. The overlap is detected for each frame based on the difference in skeletal coordinates. For example, two people whose skeletal coordinate difference is below a threshold are determined to be overlapping people.

[0173] Corrections for each frame are basically performed as follows: (1) When overlapping people are detected, overlap information including the combination of the PID and TID of the two overlapping people is saved. (2) It is checked whether the PIDs of the two overlapping people will be swapped after a few frames. (3) If the PIDs have been swapped, the person assigned the same value as the current TID value is selected as the tracking target for the section between the frame where the overlapping was detected and the current frame. (4) If the PIDs have not been swapped, it is determined that the AI ​​output for the section between the frame where the overlapping was detected and the current frame is correct, and the AI ​​output is used as is to select the tracking target.

[0174] When PID=1 is designated as the tracking target, overlap between person A and person B is detected at the timing of the second frame of B in FIG. 21, and therefore the combination of PID=2, TID=4, which is the AI ​​output of person A, and the combination of PID=1, TID=3, which is the AI ​​output of person B, are saved (processing (1) above).

[0175] Furthermore, at the timing of the fifth frame, it is confirmed that the PID of person A and the PID of person B have been swapped while the TID remains the same (process (2) above).

[0176] For the third and fourth frames, which are sandwiched between the second frame where overlap was detected and the current frame, the fifth frame, the person with PID=2, which is assigned the same value as the current TID value, TID=4, is selected as the tracking target (processing (3) above).

[0177] In this example, the recognition result in the second frame, in which two people overlap, is detected as suspicious. That is, when the difference in coordinates is below the threshold and the PID and TID combinations of the two people are swapped, a suspicious recognition result is detected.

[0178] Furthermore, since it is determined in the fifth frame that the PIDs of the two people have been swapped, erroneous recognition of the recognition results in the preceding third and fourth frames is detected. The detection of suspicious recognition results and erroneous recognition is shown in FIG. 22 . The tracking target is corrected by selecting the person with PID=2, who is assigned TID=4, as the tracking target in the third and fourth frames. If the PID-TID combination of one person and the PID-TID combination of the other person are swapped, and only the PIDs of the two people are swapped again, the tracking target is corrected by changing PID=2 in the recognition results in the third and fourth frames to PID=1, and PID=1 to PID=2.

[0179] 23 is a flowchart illustrating the cutout coordinate calculation process 4. The cutout coordinate calculation process 4 is a process that includes the above-mentioned correction when a face recognition error and a TID change occur simultaneously.

[0180] In step S101, the cutout coordinate calculation unit 23 reads the AI ​​output of the current frame.

[0181] In step S102, the cutout coordinate calculation unit 23 determines whether an overlap between two people has been detected or whether overlap information has been saved.

[0182] If it is determined in step S102 that an overlap between two people has been detected or that overlap information has been saved, then in step S103, when an overlap between two people is newly detected, the cutout coordinate calculation unit 23 saves the current frame number and the combination of the PID and TID of the two overlapping people as overlap information.

[0183] In step S104, the cutout coordinate calculation unit 23 determines whether the number of frames in the information storage section is within a specified number.

[0184] If it is determined in step S104 that the number of frames in the information storage section is within the specified number, then in step S105, the clipping coordinate calculation unit 23 stores the AI ​​output of the current frame.

[0185] In step S106, the cutout coordinate calculation unit 23 determines whether or not there is a PID whose associated TID has been changed among the PIDs stored as overlap information.

[0186] If it is determined in step S106 that there is a PID whose TID has changed, then in step S107, the cutout coordinate calculation unit 23 determines whether the stored TID of one person is equal to the current TID of the other person.

[0187] If it is determined in step S107 that the TIDs of the two people are the same, in step S108, the extraction coordinate calculation unit 23 swaps the PIDs of the two people in the information storage section and selects the tracking target using the recognition results of the overlapping PIDs of the two people.

[0188] In step S109, the cutout coordinate calculation unit 23 erases the overlap information and cancels the storage of the AI ​​output.

[0189] On the other hand, if it is determined in step S104 that the number of frames in the information storage section is not within the specified number, or if it is determined in step S107 that the TIDs of the two people are not equal, then in step S110 the cutout coordinate calculation unit 23 cancels storage using each AI output of the frames in the information storage section.

[0190] If it is determined in step S102 that overlapping of two people has not been detected and that overlapping information has not been saved, then in step S111, the cutout coordinate calculation unit 23 uses the AI ​​output of the current frame.

[0191] After the processing of step S109, step S110, or step S111, in step S112, the cutout coordinate calculation unit 23 calculates cutout coordinates that designate an area in which the person to be tracked appears, and stores the calculated coordinates in a buffer.

[0192] On the other hand, if it is determined in step S106 that there is no PID whose associated TID has been changed among the PIDs stored as overlap information, then in step S113, the cropping coordinate calculation unit 23 switches to the next frame while retaining the information.

[0193] After the process of step S112 or step S113, the process returns to step S25 in Fig. 14, and the subsequent processes are performed. Through the above process, the error shown in B in Fig. 21 is corrected and the tracking target is selected.

[0194] The above-described cutout coordinate calculation processes 1 to 4 are performed in step S25 of FIG. 14, and the person being tracked is corrected as appropriate to calculate the cutout coordinates. By detecting suspicious recognition results based on a change in the combination of PID and TID or based on the overlap of two people, and correcting the tracking target based on the combination of PID and TID, it becomes possible to generate tracking video that continues to track the correct person. Because there is no need to manually select the correct person and correct the tracking target, the user can easily generate tracking video that continues to track the correct person.

[0195] <<Flow of the Tracking Phase>> FIG. 24 is a diagram showing a first example of the flow of the tracking phase.

[0196] In the example of Fig. 24, generation of a tracking video is performed after calculation of the cutout coordinates for all frames of the entire video has been completed. The upper part of Fig. 24 shows the flow of calculation of the cutout coordinates for all frames of the entire video, and the lower part shows the flow of generation of the tracking video.

[0197] In calculating the cutout coordinates of the tracking target, preprocessing is performed on the entire image obtained by decoding the video data, and the preprocessed frame is used as input to the AI, as shown in the upper part of Fig. 24. Based on the recognition result output from the AI, processing including correction by the tracking algorithm described above is performed, and the cutout coordinates are calculated.

[0198] Since it is necessary to use information from future frames, the tracking algorithm performs corrections while accumulating the recognition results for a certain number of frames. This process of calculating the cutout coordinates of the tracking target is performed for all frames of the entire video.

[0199] On the other hand, in generating a tracking video, the tracking video is cut out from the entire video obtained by decoding the video data, as shown in the lower part of Fig. 24. The tracking video is cut out based on the cutout coordinates of each frame.

[0200] When multiple people are specified as tracking targets, multiple tracking videos are extracted and the data of each tracking video is encoded to generate a file for each tracking video. In the example at the bottom of Fig. 24, an MP4 format file is generated as the tracking video file.

[0201] 25 is a diagram showing a second example of the flow of the tracking phase. Duplicate descriptions will be omitted as appropriate.

[0202] In the example of Figure 25, the calculation of the cutout coordinates and the generation of the tracking image are performed in parallel. That is, preprocessing is performed on the entire image obtained by decoding the video data, and the preprocessed frame is used as input to the AI. Based on the recognition result output from the AI, the cutout coordinate calculation process, including correction by the tracking algorithm described above, is performed.

[0203] Furthermore, until the calculation of the cutout coordinates is completed, frames of the entire video to be cut out are accumulated in a buffer (BUF). For example, the frame at time t1 is accumulated in the buffer at least until the cutout coordinates of the frame at time t1 are calculated, and the frame at time t2 is accumulated in the buffer at least until the cutout coordinates of the frame at time t2 are calculated. After the cutout coordinates are calculated, the entire video to be cut out is read out from the buffer frame by frame and used to cut out the tracking video.

[0204] This makes it possible to reduce the waiting time for the user compared to when a tracking video is generated after calculation of the cutout coordinates for all frames of the entire video has been completed.

[0205] FIG. 26 is a diagram illustrating a third example of the flow of the tracking phase.

[0206] In the example of Fig. 26, calculation of the cutout coordinates and generation of the tracking video are performed in parallel to reduce the waiting time for the user. In the example of Fig. 25, frames of the whole video from which the cutout is made are stored in a buffer, but in the example of Fig. 26, decoding of the whole video from which the cutout is made is delayed.

[0207] That is, the entire video used to calculate the cutout coordinates and the entire video from which the cutout is made are decoded separately. For example, the decoding of the entire video from which the cutout is made starts with a delay corresponding to the time it takes to calculate the cutout coordinates for the same frame.

[0208] The full video from which the clipping is performed is a high-resolution video, and a large buffer is required to store it. However, by delaying the decoding of the full video from which the clipping is performed, such a large buffer becomes unnecessary, as shown in FIG.

[0209] <<Editing a Tracking Target>> <UI Example> Fig. 27 is a diagram showing an example of a selection screen for a tracking target. The selection screen shown in Fig. 27 is displayed by, for example, the tracking video playback unit 14 after processing in the automatic PID numbering and face registration phase.

[0210] In the example of Fig. 27, face images of five people registered by the processing of the automatic PID numbering and face registration phase are displayed. Using such a screen, the user selects a person to be tracked. After the person to be tracked is selected, the processing of the tracking phase begins. It is also possible to automatically generate a tracking video of all people appearing in the overall video without selecting a person to be tracked.

[0211] The tracking video generated by the processing of the tracking phase is played back in the image processing device 1 in response to a user operation. The tracking target can be manually edited while viewing the tracking video. The states of the tracking target before and after correction by the tracking algorithm are visualized and presented to the user.

[0212] FIG. 28 is a diagram showing an example of the editing screen.

[0213] An original video area 101 is formed in approximately the center of the editing screen. The original video area 101 is a display area for the entire video from which the tracking video was cut out. In the example of Fig. 28, an entire video showing five people, people A to E, is displayed.

[0214] When a tracking video is generated with each of persons A to E as the tracking target, rectangular frame images 111-1 to 111-5 surrounding each person are displayed in the overall video. The frame images 111-1 to 111-5 indicate clipped areas of the tracking video with persons A to E as the tracking target, respectively. For example, the frame images 111-1 to 111-5 are displayed using the colors assigned to each person. By displaying the frame images 111-1 to 111-5 in different colors, the user can easily confirm which area of ​​each person was tracked.

[0215] Face icons 112-1 to 112-5, each consisting of a small face image, are displayed above frame images 111-1 to 111-5. A face icon corresponding to each person is displayed at the position of each person appearing in the full video. Face icons 112-1 to 112-5 are used, for example, to edit the cut-out area. When the user swaps face icon 112-1 and face icon 112-2, for example, by dragging and dropping, the cut-out coordinates are corrected so that the tracking targets for the swapped sections are swapped.

[0216] Preview area 102 on the left side of the editing screen is a display area for the tracking video of a person selected on the entire video. When, for example, the area of ​​frame image 111-1 is selected by a click operation or the like, the tracking video of person A is displayed in preview area 102 as shown in FIG.

[0217] A seek bar 103 is displayed below the original image area 101 and the preview area 102. The seek bar 103 indicates sections where the recognition accuracy of the facial recognition AI is low, i.e., sections where the PID recognition result may be incorrect. For example, sections detected as error sections in processing using a tracking algorithm, and sections where the similarity in comparison with registered facial images is lower than a threshold, are indicated as sections where the accuracy of facial recognition is low.

[0218] FIG. 29 is an enlarged view of the seek bar 103. As shown in FIG.

[0219] The entire seek bar 103 indicates the entire section of the content. Sections indicated by color are sections where the accuracy of facial recognition is low. The color of each section represents a person with low facial recognition accuracy, such as a section where person A's facial recognition accuracy is low is displayed in the color assigned to person A, and a section where person B's facial recognition accuracy is low is displayed in the color assigned to person B.

[0220] In the example of Fig. 29, sections 113-1, 113-2, and 113-4 are sections where the accuracy of face recognition is low for person A. Sections 113-3, 113-5, and 113-6 are sections where the accuracy of face recognition is low for person B, person C, and person D, respectively. The user can correct the tracking target by focusing on sections where the accuracy of face recognition is low for a specific person.

[0221] Returning to the explanation of Fig. 28, an edit item selection area 104 is formed on the right side of the editing screen, which is a display area for information used to select an editing target. In the example of Fig. 28, face images of persons A to E and an "All" button are displayed in the edit item selection area 104, and the "All" button is selected from among them. By selecting the "All" button, it becomes possible to edit the tracking targets of all people. The editing screen shown in Fig. 28 is a screen used to edit the tracking targets of all people.

[0222] 30 is a diagram showing an example of the display of the editing screen when person A is selected as the editing target. When person A is selected as the editing target using the display in the editing item selection area 104, the editing screen of FIG. 30 is displayed.

[0223] When person A is the tracking target corrected by the tracking algorithm, a solid-line frame image is displayed as frame image 111-1 surrounding person A, as shown in original video area 101 in Fig. 30. In this case, for example, a dashed-line frame image is displayed as frame image 111-2 surrounding person B, who is the tracking target before correction. The dashed-line frame image 111-2 indicates that person B is a candidate for editing the tracking target to another person. Frame images 111-3 to 111-5 of an inconspicuous color such as gray are displayed for the other people.

[0224] When person A is selected as the editing target, the seek bar 103 indicates a section where the accuracy of face recognition is low as well as a section where all the people are shown from above.

[0225] <Tracking Target Editing Process> By using the editing screen, the user can manually edit the tracking target that was automatically selected by the tracking phase process. When an editing operation is performed, the tracking video is corrected as shown in Fig. 31. The arrangement of people shown in Fig. 31 represents the people that appear in each frame of the tracking video. In the example of Fig. 31, correction is made to a section in which a different person was selected as the tracking target in the tracking video before correction.

[0226] FIG. 32 is a diagram showing an example of exchanging the cutout coordinates.

[0227] A in Fig. 32 shows the cutout coordinates of each frame of the person with PID = 1. B in Fig. 32 shows the cutout coordinates of each frame of the person with PID = 2.

[0228] Assume that a facial recognition error occurs in the sections indicated by the double-headed arrows in Figures 32A and 32B. A person who should be assigned PID=1 is recognized as a person with PID=2, and a person who should be assigned PID=2 is recognized as a person with PID=1.

[0229] When an operation to swap the person with PID=1 and the person with PID=2 is performed on the editing screen, with the section indicated by the bidirectional arrow as the editing section, the cut-out coordinates of the edit section of the person with PID=1 are overwritten with the cut-out coordinates of the edit section of the person with PID=2, as shown in Fig. 33. The overwriting of the cut-out coordinates is performed by the cut-out coordinate calculation unit 23. Although Fig. 33 only shows the editing content of the cut-out coordinates of the person with PID=1, the cut-out coordinates of the person with PID=2 are also overwritten with the cut-out coordinates of the person with PID=1 in a similar manner.

[0230] When the cutout coordinates are corrected in this way, for example, the entire video is decoded again from the first frame, and the tracking video is cut out using the corrected cutout coordinates, which makes it possible to correct the tracking video so that it continues to track the correct person.

[0231] Instead of decoding again from the first frame of the entire video, only the editing section may be decoded and the tracking video may be extracted. In this case, a section including the editing section, for example, a GOP (Group of Pictures) unit, is decoded. The entire video and tracking video are encoded using a coding method that uses GOPs.

[0232] FIG. 34 is a diagram showing an example of decoding in units of GOPs.

[0233] The pre-modification video shown in the upper part of Fig. 34 is the tracking video before modification. The section indicated by the bidirectional arrow is the editing section. In this case, a plurality of GOPs including the editing section in the tracking video before modification are decoded. The GOPs indicated by the dark frame are the GOPs to be decoded. Also, as shown in the middle part of Fig. 34, a plurality of GOPs including the editing section in the entire video are decoded. In the example of Fig. 34, the pre-modification tracking video is decoded so that a greater number of GOPs than the GOPs of the entire video are decoded.

[0234] The tracking video is cut out from each frame of the decoded entire video using the corrected cut-out coordinates. The newly cut-out frames of the tracking video are added as frames of the editing section and re-encoded to generate the corrected tracking video as shown in the lower part of Fig. 34. Each frame of the editing section of the corrected tracking video shown in the lower part of Fig. 34 is a frame composed of video cut out based on the corrected cut-out coordinates as shown in Fig. 35.

[0235] By decoding in GOP units and editing the tracking video, it is possible to recreate the tracking video in a shorter time than when the entire video is decoded again from the first frame and edited.

[0236] <<Modifications>> <System Configuration> FIG. 36 is a diagram showing an example of the configuration of an image processing system.

[0237] The image processing device 301 shown in Fig. 36 is a server on the Internet. In the image processing device 301, a predetermined program is executed to realize an image processing unit 301A as shown in the balloon. The image processing unit 301A has the same configuration as the image processing device 1 described with reference to Fig. 10.

[0238] The image processing device 301 receives the entire video uploaded via the Internet from a user terminal 302 such as a smartphone, and generates a tracking video as described above. The tracking video data generated by the image processing device 301 is transmitted to the user terminal 302 and displayed on the user terminal 302. In this way, the tracking video generation function may be provided on a server on the Internet. An editing screen may be displayed on the screen of the user terminal 302, and the tracking video may be edited using the user terminal 302.

[0239] 37 is a block diagram showing an example of the hardware configuration of the image processing device 301. The image processing device 1 is also configured by a computer having a similar configuration.

[0240] A CPU (Central Processing Unit) 1001 , a ROM (Read Only Memory) 1002 , and a RAM (Random Access Memory) 1003 are interconnected by a bus 1004 .

[0241] An input / output interface 1005 is also connected to the bus 1004. An input unit 1006 including a keyboard, a mouse, etc., and an output unit 1007 including a display, a speaker, etc. are connected to the input / output interface 1005. In addition, a storage unit 1008 including a hard disk, a nonvolatile memory, etc., a communication unit 1009 including a network interface, etc., and a drive 1010 that drives removable media 1011 are also connected to the input / output interface 1005.

[0242] In a computer configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.

[0243] <Other Examples> FIG. 38 is a diagram showing a display example of a tracking video.

[0244] There may be cases where the person to be tracked does not appear anywhere in the overall video. In sections where the person to be tracked does not appear anywhere in the overall video, frames of the overall video may be inserted to generate a tracking video. When the tracking video is played back, in sections where the person to be tracked does not appear, instead of the video showing the person to be tracked, a video showing a bird's-eye view of people other than the person to be tracked is displayed, as shown in FIG.

[0245] Although the tracking video is generated by cutting out an area that shows the entire body of the person to be tracked, the tracking video may also be generated by cutting out an area that shows only a specific part of the person to be tracked, such as the face, hands, or feet. This makes it possible to generate a tracking video that focuses on a specific part of a specific person. The part of the tracking video to be generated is selected, for example, by the user.

[0246] Although the tracking target is assumed to be a person, the tracking video may be generated for any moving object such as an animal, character, etc. Skeleton recognition and face recognition are performed for the arbitrary object, and the cutout coordinates are calculated as described above.

[0247] Example of Computer Configuration The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the program constituting the software is installed in a computer having the configuration shown in FIG.

[0248] The program executed by the CPU 1001 is installed in the storage unit 1008 by being recorded on, for example, a removable medium 1011 or provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital broadcasting.

[0249] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.

[0250] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0251] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0252] The embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible without departing from the spirit of the present technology.

[0253] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.

[0254] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.

[0255] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.

[0256] <Examples of Combinations of Configurations> The present technology can also have the following configurations.

[0257] (1) An image processing device comprising: an acquisition unit that acquires face information and skeletal information of each object appearing in each frame constituting an overall video; and an image processing unit that calculates coordinates of a specific object in each frame based on the face information and the skeletal information, and cuts out a video of the specific object from the overall video based on the calculated coordinates. (2) The image processing device described in (1), wherein the acquisition unit acquires, as the face information, information indicating an identification result of the object based on a face. (3) The image processing device described in (2), wherein the acquisition unit further acquires coordinate information indicating the coordinates of skeletal points of the object. (4) The image processing device described in (3), wherein the acquisition unit acquires, as the skeletal information, information indicating an identification result of the object based on a trajectory of coordinates of skeletal points in a plurality of frames. (5) The image processing device according to (3) or (4), wherein the video processing unit detects the face information of the specific object that may be erroneous based on at least one of the face information, the skeletal information, and the coordinate information, and detects an error in the face information of the specific object based on a combination of the face information and the skeletal information acquired during a period up to a predetermined threshold number of frames after the detection of the face information that may be erroneous, and corrects the face information to be correct. (6) The image processing device according to (5), wherein the video processing unit detects the face information that may be erroneous based on a combination of the face information and the skeletal information, or based on the fact that a difference between coordinates indicated by the coordinate information of two or more of the objects is equal to or less than a threshold. (7) The image processing device according to (6), wherein the video processing unit detects the face information that may be erroneous in response to a change in the face information from information assigned to the specific object to information indicating unrecognizable, when there is no change in the skeletal information of the specific object. (8) The image processing device according to (7), wherein, when the information indicating unrecognizable is changed to information assigned to the specific object, the video processing unit detects an error in the face information and corrects the information indicating unrecognizable using the information assigned to the specific object.(9) The image processing device according to any one of (6) to (8), wherein the video processing unit detects the face information that may be erroneous in response to information indicating unrecognizable being acquired as the face information of the object whose skeleton information has been newly acquired. (10) The image processing device according to (9), wherein the video processing unit detects an error in the face information when the information indicating unrecognizable has changed to information assigned to the specific object, and corrects the information indicating unrecognizable using the information assigned to the specific object. (11) The image processing device according to any one of (6) to (10), wherein the video processing unit detects the face information that may be erroneous in response to a change in the combination of the face information and the skeleton information when the skeleton information of the specific object remains unchanged. (12) The image processing device according to (11), wherein the video processing unit detects an error in the face information when the face information has changed to information assigned to the specific object, and corrects the error in the face information using the information assigned to the specific object. (13) The image processing device according to any of (6) to (12), wherein the video processing unit detects the face information that may be erroneous when a combination of the face information and the skeletal information of two or more of the objects, for which a difference in coordinates indicated by the coordinate information is equal to or less than a threshold, is swapped. (14) The image processing device according to (13), wherein the video processing unit detects an error in the face information when the face information in the swapped combination is swapped again, and corrects the error in the face information using the swapped face information. (15) The image processing device according to any of (2) to (14), wherein the acquisition unit, when registering face images of the objects, acquires the face images of each of the objects shown in a first frame, and registers the acquired face images in association with the face information assigned to each of the objects, and, when identifying the faces of the objects, assigns the face information associated with the face image to an object shown in a second frame whose face has a high similarity to the face image.(16) The image processing device according to any of (1) to (15), further comprising a display control unit that displays the overall video on an external or internal display unit and displays a corresponding icon at the position of each of the objects appearing in the overall video, wherein the video processing unit corrects coordinates associated with the face information based on a user operation on the icon. (17) The image processing device according to any of (1) to (16), wherein the display control unit displays information indicating a section in which the face information of the object is likely to be incorrect. (18) The image processing device according to any of (1) to (17), wherein, if the specific object is not shown in a frame of the overall video from which the video is cut out, the video of the specific object is generated using the frame of the overall video from which the video is cut out. (19) An image processing method, wherein an image processing device acquires face information and skeletal information of each object shown in each frame constituting the overall video, calculates coordinates of the specific object in each frame based on the face information and the skeletal information, and cuts out a video of the specific object from the overall video based on the calculated coordinates. (20) A program that causes a computer to execute a process of acquiring face information and skeletal information of each object that appears in each frame that constitutes an entire video, calculating coordinates of a specific object in each frame based on the face information and the skeletal information, and cutting out a video of the specific object from the entire video based on the calculated coordinates. (21) The image processing device described in (3), wherein the acquisition unit acquires the coordinate information that indicates coordinates of a plurality of skeletal points of the object, and the video processing unit cuts out a video of the specific part of the specific object from the entire video based on the coordinates of the skeletal point that corresponds to a specific part selected by a user.(22) The image processing device according to (3), wherein the acquisition unit acquires the skeletal information and the coordinate information of each of the objects appearing in the input frames based on the output of a first inference model that receives as input frames constituting the entire video, and acquires the face information of each of the objects based on the output of a second inference model that receives as input the skeletal information and the coordinate information of each of the objects. (23) The image processing device according to (1), wherein the video processing unit starts decoding the entire video from which the clipping is performed with a delay after decoding the entire video used to calculate the coordinates of the specific object.

[0258] REFERENCE SIGNS LIST 1 image processing device, 11 image acquisition unit, 12 tracking video generation unit, 13 tracking video recording unit, 14 tracking video playback unit, 21 skeleton recognition unit, 22 face recognition unit, 23 cut-out coordinate calculation unit, 24 cut-out processing unit

Claims

1. An image processing device comprising: an acquisition unit that acquires facial information and skeletal information of each object that appears in each frame that constitutes an overall image; and an image processing unit that calculates the coordinates of a specific object in each frame based on the facial information and the skeletal information, and cuts out an image of the specific object from the overall image based on the calculated coordinates.

2. The image processing device according to claim 1, wherein the acquisition unit acquires, as the face information, information indicating a result of identifying the object based on a face.

3. The image processing device according to claim 2, wherein the acquisition unit further acquires coordinate information indicating coordinates of skeleton points of the object.

4. The image processing device according to claim 3, wherein the acquisition unit acquires, as the skeleton information, information indicating the result of identification of the object based on a trajectory of coordinates of skeleton points in a plurality of frames.

5. The image processing device of claim 4, wherein the image processing unit detects the facial information of the specific object that may be erroneous based on at least one of the facial information, the skeletal information, and the coordinate information, and detects errors in the facial information of the specific object based on a combination of the facial information and the skeletal information obtained between the detection of the facial information that may be erroneous and a frame within a predetermined threshold, and corrects the facial information to the correct one.

6. The image processing device according to claim 5, wherein the image processing unit detects the face information that may be erroneous based on a combination of the face information and the skeletal information, or based on the difference between the coordinates indicated by the coordinate information of two or more of the objects being less than a threshold value.

7. The image processing device according to claim 6, wherein the image processing unit detects potentially erroneous face information in response to a change in the face information from information assigned to the specific object to information indicating unrecognizable when there is no change in the skeletal information of the specific object.

8. The image processing device according to claim 7, wherein the video processing unit, when the information indicating unrecognizable changes to information assigned to the specific object, detects an error in the face information and corrects the information indicating unrecognizable using the information assigned to the specific object.

9. The image processing device according to claim 6, wherein the image processing unit detects the face information that may be erroneous in response to the acquisition of information indicating that the skeleton information is unrecognizable as the face information of the newly acquired object.

10. The image processing device according to claim 9, wherein the video processing unit, when the information indicating unrecognizable changes to information assigned to the specific object, detects an error in the face information and corrects the information indicating unrecognizable using the information assigned to the specific object.

11. The image processing device according to claim 6, wherein the image processing unit detects potentially erroneous face information in response to a change in the combination of the face information and the skeletal information when there is no change in the skeletal information of the specific object.

12. The image processing device according to claim 11, wherein the video processing unit detects an error in the face information when the face information changes to information assigned to the specific object, and corrects the error in the face information using the information assigned to the specific object.

13. The image processing device according to claim 6, wherein the image processing unit detects potentially erroneous face information when the combination of face information and skeletal information of two or more objects whose coordinate difference indicated by the coordinate information is less than a threshold is swapped.

14. The image processing device according to claim 13, wherein when the facial information in the swapped combination is swapped again, the video processing unit detects an error in the facial information and corrects the error in the facial information using the swapped facial information again.

15. The image processing device described in claim 2, wherein the acquisition unit, when registering a facial image of the object, acquires the facial image of each of the objects shown in the first frame, and registers the acquired facial image by linking it to the facial information assigned to each of the objects, and when identifying the face of the object, assigns the facial information linked to the facial image to the object shown in the second frame whose face is highly similar to the facial image.

16. An image processing device as described in claim 1, further comprising a display control unit that displays the overall image on an external or internal display unit and displays a corresponding icon at the position of each of the objects shown in the overall image, wherein the image processing unit modifies the coordinates associated with the face information based on a user's operation on the icon.

17. The image processing device according to claim 16, wherein the display control unit displays information indicating an area in which the facial information of the object may be incorrect.

18. The image processing device according to claim 1, wherein, if the specific object is not shown in the frame of the entire image from which the image was cut out, the image processing unit generates an image of the specific object using the frame of the entire image from which the image was cut out.

19. An image processing method in which an image processing device acquires face information and skeletal information of each object appearing in each frame that constitutes an entire image, calculates the coordinates of a specific object in each frame based on the face information and the skeletal information, and cuts out an image of the specific object from the entire image based on the calculated coordinates.

20. A program that causes a computer to execute the following process: acquire facial information and skeletal information of each object that appears in each frame that makes up the entire video; calculate the coordinates of a specific object in each frame based on the facial information and skeletal information; and cut out an image of the specific object from the entire video based on the calculated coordinates.

Citation Information

Patent Citations

  • Correlated display of biometric id, feedback and user interaction state

    JP2017504849A

  • Automatic switching device, automatic switching method, and program

    JP2022091640A

  • Display control system and display control method

    WO2018142494A1

  • Video processing system, video processing method, and non-transitory computer-readable medium

    WO2023281620A1