Information processing apparatus, information processing method, and program

The information processing apparatus addresses the challenge of creating varied and balanced combined videos by classifying and displaying shooting information ratios, allowing for optimal video selection and combination.

JP2025113784APending Publication Date: 2025-08-04CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024008110
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-23
Publication Date
2025-08-04

AI Technical Summary

Technical Problem

Existing technologies struggle to generate combined moving images that are rich in variations and well-balanced, as it is difficult to confirm and select videos with diverse shooting conditions for optimal combination.

Method used

An information processing apparatus that recognizes shooting information for each video, classifies videos into groups with similar shooting information, and displays the ratios of these groups to the user, facilitating the selection of videos for well-balanced combination.

Benefits of technology

Enables the generation of combined videos with varied and visually appealing content by ensuring a balanced distribution of shooting conditions, such as subjects, camera work, and frame rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025113784000001_ABST
    Figure 2025113784000001_ABST
Patent Text Reader

Abstract

To facilitate creation of a coupled moving image obtained by coupling moving images having a lot of changes with good balance.SOLUTION: An information processing apparatus has: recognition means that recognizes photographic information for every moving image of a plurality of moving images; classification and summing means that classifies the plurality of moving images into moving images having similar photographic information on the basis of the photographic information for every moving image, and sums up the ratios of the moving images having the similar photographic information to the plurality of moving images; and display means that displays, to a user, the ratios summed up by the classification and summing means.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to information processing technology for acquiring and editing moving images and the like.

Background Art

[0002] As one of the technologies for processing moving images, there is a technology for combining a plurality of moving images to generate one combined moving image. Further, Patent Document 1 discloses a shooting assist method for determining whether the number of shots or the ratio of a specific subject based on reference data is within a predetermined range in still image shooting and assisting shooting based on the determination result.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] A moving image obtained by combining a plurality of moving image data may sometimes become a monotonous moving image in which similar images continue. On the other hand, if moving images rich in changes can be combined well-balancedly, it is considered that a combined moving image with good appearance can be generated. For this reason, in order to generate a combined image with good appearance, it is desirable to prepare in advance a plurality of moving images rich in changes by shooting or the like. However, it is not easy to confirm whether a plurality of moving images rich in changes are obtained. Further, in order to generate a combined moving image with good appearance, it is desired to combine moving images rich in changes well-balancedly, but it is not easy to select well-balancedly moving images rich in changes from among a plurality of moving images. Note that the technology of Patent Document 1 is a technology for still images, and is also a technology for assisting whether a specific subject is sufficiently shot in still image shooting, and thus cannot be applied to the generation of combined moving images.

[0005] Therefore, an object of the present invention is to facilitate the generation of a combined video in which videos rich in changes are combined in a well-balanced manner.

Means for Solving the Problems

[0006] The information processing apparatus of the present invention includes: a recognition unit that recognizes shooting information for each of a plurality of videos; a classification aggregation unit that classifies the plurality of videos into groups of videos having similar shooting information based on the shooting information for each of the videos, and aggregates the ratios of the videos having the similar shooting information with respect to the plurality of videos; and a display unit that displays the ratios aggregated by the classification aggregation unit to the user.

Effects of the Invention

[0007] According to the present invention, it becomes easy to generate a combined video in which videos rich in changes are combined in a well-balanced manner.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Modes for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The embodiments described below do not limit the present invention, and not all combinations of the features described in this embodiment are essential for the solution means of the present invention. The configuration of the embodiment can be appropriately modified or changed according to the specifications of the device to which the present invention is applied and various conditions (usage conditions, usage environment, etc.). Also, in the following embodiments, duplicate descriptions of the same or similar configurations and processing steps will be omitted.

[0010] FIG. 1 is a diagram showing each functional unit of the information processing apparatus 1 according to this embodiment. That is, the information processing apparatus 1 of this embodiment executes each processing step represented by each functional unit shown in FIG. 1. Also, FIG. 2 is a diagram showing a configuration example of the video file 2 generated and held by the information processing apparatus 1 of this embodiment. The information processing apparatus 1 has an imaging unit 11 capable of capturing a video for generating video data 21 in the video file 2. The video captured by the imaging unit 11 is sent to the file generation unit 12. Note that the imaging unit 11 may be provided as an external device of the information processing apparatus 1. In this case, the video captured by the imaging unit 11 is transmitted to the information processing apparatus 1 by wireless communication or wired communication, and the information processing apparatus 1 sends the received video to the file generation unit 12.

[0011] The file generation unit 12 includes a data generation unit 121 and a shooting information generation unit 122. The data generation unit 121 converts the video input from the imaging unit 11 into video data 21 in a predetermined format that can be handled within the information processing apparatus 1. The shooting information generation unit 122 recognizes the shooting conditions at the time of shooting the video based on the video captured by the imaging unit 11, and generates shooting information 22 according to the recognition result. In the case of this embodiment, the shooting information generation unit 122 recognizes, as the shooting conditions at the time of shooting the video, the date at the time of shooting the video, the length of the video, the subject shown in the video, the camera work, and the frame rate. Details of the shooting information 22 according to the result of recognizing those shooting conditions will be described later.

[0012] Then, the file generation unit 12 generates a video file 2 by attaching the shooting information 22 generated by the shooting information generation unit 122 to the video data 21 converted by the data generation unit 121, and sends it to the file recording unit 13. That is, the shooting information 22 is sent to the file recording unit 13 as a part of the video file 2 of the video data 21. The file recording unit 13 records the plurality of video files 2 respectively generated by the file generation unit 12 from the plurality of videos shot by the imaging unit 11.

[0013] Note that the file generation unit 12 may be provided in an external device (not shown) different from the information processing device 1. In this case, the video file 2 generated by the file generation unit 12 of the external device is transmitted to the information processing device 1 by wired communication or wireless communication, and the file recording unit 13 of the information processing device 1 records the received video file 2.

[0014] FIG. 2 is a diagram showing a data configuration example of the video file 2 including the video data 21 and the shooting information 22. As shown in FIG. 2, the shooting information 22 includes date information 221, video length information 222, subject recognition information 223, camera work information 224, and frame rate information 225. The date information 221 is information indicating the date when the video of the video data 21 was shot. The video length information is information indicating the time length of the video of the video data 21. The subject recognition information 223 is information indicating the recognition result of the subject shown in the video of the video data 21. The recognition of the subject and specific examples of the subject will be described later. The camera work information 224 is information indicating the camera work. The details and specific examples of the camera work will be described later. The frame rate information 225 is information indicating how many images the video is composed of per second. Specific examples of the frame rate will be described later. The information included in the shooting information 22 is not limited to the above-described each information, and other information other than these each information may be included. Also, the shooting information 22 may have an information configuration that does not include any one or more of the above-described each information.

[0015] FIG. 3 is a diagram showing the detailed functional units of the imaging information generation unit 122. The imaging information generation unit 122 executes the processing steps represented by the respective functional units shown in FIG. 3. The imaging information generation unit 122 includes a subject recognition unit 122a, a camera work recognition unit 122b, and a frame rate recognition unit 122c as recognition units that respectively recognize at least a subject, camera work, and frame rate. Note that the imaging information generation unit 122 may be configured not to include any one or more of the above-described three recognition units. In addition to these three recognition units, the imaging information generation unit 122 may include a recognition unit that recognizes the date when the video was shot and the length of the video, or may further include a recognition unit that recognizes other shooting conditions.

[0016] The subject recognition unit 122a recognizes a subject appearing in the video shot by the imaging unit 11, and sets the recognition result as the subject recognition information 223 of the imaging information 22. Examples of the subject recognition process include a human body recognition process for recognizing a person as a specific subject in the video. For example, when performing the human body recognition process, the subject recognition unit 122a holds dictionary information capable of identifying the human body, compares the dictionary information with the subject in the video, and when a subject with characteristics similar to the human body is detected, recognizes that the subject is a person. On the other hand, videos that do not include people are often videos of landscapes or the like. Therefore, for example, the subject recognition unit 122 can determine whether a person is included in the video using the above-described human body recognition process, and when no person is included, recognize that the video is a video of a landscape or the like. In this way, the subject recognition unit 122a can distinguish to some extent whether the video shot by the imaging unit 11 is a video of a person or a video of a landscape.

[0017] Note that the subject recognition process may be an animal recognition process that recognizes an animal other than a person as a specific subject. Further, the subject recognition process may include a face recognition process that recognizes a person's face as a specific subject, or a recognition process that more specifically identifies an individual or an entity among people and animals. Furthermore, the subject recognition process may be a process of recognizing two or more specific subjects such as those people, animals, people's faces, individuals, or entities.

[0018] Also, the subject recognition unit 122a may perform a process of including the subject recognition information of all recognized subjects in the shooting information 22. Further, for example, the subject recognition unit 122a may not include all the recognized subject recognition information in the shooting information 22, but may perform a process of determining whether to include it in the shooting information 22 based on a predetermined determination condition. For example, the subject recognition unit 122a may use, as a predetermined determination condition, whether a subject has been recognized for a predetermined time or more, and may include only the subject recognition information of the subjects recognized for a predetermined time or more in the shooting information 22. In the video file 2 shown in FIG. 2, an example is shown in which the result of recognizing a person as a subject is included in the shooting information 22, and an example is shown in which subject recognition information 223 including information "person present", which means that a person has been recognized in the subject recognition process, is recorded.

[0019] The camera work recognition unit 122b recognizes the camera work when a video is shot by the imaging unit 11, and sets the recognition result as the camera work information 224 of the shooting information 22. FIGS. 4(a) to 4(d) are schematic diagrams for explaining examples of camera work recognized by the camera work recognition unit 122b, and represent an example of a video 3 in which a person 31 and a tree 32 that is part of the background are shown as subjects. Arrows 31a and 32a in the figure represent the movements of the person 31 and the tree 32 in the video, respectively. In the figure, the subjects to which the arrows 31a and 32a are attached are shown to be moving in the direction of the arrows in the video 3, and the subjects to which the arrows 31a and 32a are not attached are shown to be stationary in the video 3.

[0020] Figure 4(a) shows the case where both the person 31 and the tree 32 are stationary in Video 3. Since the tree 32 is an object that does not move in the real space, when the tree 32 is stationary in Video 3, it is considered that the imaging unit 11 (the information processing apparatus 1 equipped with the imaging unit 11) that is shooting this Video 3 is also in a stationary state. Therefore, it can be said that Figure 4(a) represents a state where both the person 31 and the information processing apparatus 1 are stationary when the imaging unit 11 of the information processing apparatus 1 shoots Video 3. Figure 4(b) shows the case where the person 31 is moving in Video 3 but the tree 32 is stationary. That is, Figure 4(b) represents a state where the person 31 is moving but the imaging unit 11 (the information processing apparatus 1 equipped with the imaging unit 11) is stationary when the imaging unit 11 shoots Video 3. Figure 4(c) shows the case where the person 31 is stationary in Video 3 but the tree 32 is moving. That is, Figure 4(c) represents a state where the person 31 and the imaging unit 11 (the information processing apparatus 1 equipped with the imaging unit 11) are moving in parallel at substantially the same speed when the imaging unit 11 shoots Video 3. Figure 4(d) shows the case where both the person 31 and the tree 32 are moving at substantially the same speed in Video 3. That is, Figure 4(d) represents a state where the person 31 is stationary and the imaging unit 11 (the information processing apparatus 1 equipped with the imaging unit 11) is moving when the imaging unit 11 shoots Video 3. As in the examples of Figures 4(a) to 4(d), even if it is Video 3 that shoots the same person 31 and the tree 32 which is a part of the background, various patterns of Video 3 can be obtained depending on the movements of the person 31 and the imaging unit 11 (the information processing apparatus 1 equipped with the imaging unit 11) during the shooting.

[0021] The camera work recognition unit 122b of the present embodiment recognizes "camera work" based on a plurality of patterns of the movement of the subject and the background of Video 3 as illustrated in Figures 4(a) to 4(d). Hereinafter, the recognition process of camera work performed by the camera work recognition unit 122b will be described in detail.

[0022] The camera work recognition unit 122b performs the same human body recognition process as the subject recognition unit 122a to recognize whether a person 31 is included in the moving image 3. When the person 31 is recognized, it further recognizes whether the person 31 is moving in the moving image 3. Whether the person 31 is moving in the moving image 3 can be recognized, for example, by determining whether the position information of the person 31 recognized by the human body recognition process in the moving image 3 has changed by a predetermined amount or more.

[0023] Moreover, the information processing apparatus 1 of the present embodiment also has a movement detection unit 123. The movement detection unit 123 is composed of a sensor such as an acceleration sensor that can detect the movement of the information processing apparatus 1, and can detect how the imaging unit 11 of the information processing apparatus 1 moves in the real space. The movement information detected by the movement detection unit 123 is sent to the shooting information generation unit 122 and used for the recognition of the camera work by the camera work recognition unit 122b. When the imaging unit 11 is an external device of the information processing apparatus 1, the movement detection unit 123 may be included in the imaging unit 11. In this case, the movement detection unit 123 detects how the imaging unit 11 itself, which is an external device, moves in the real space.

[0024] The camera work recognition unit 122b can detect, by the above-described processing, whether the person 31 is moving in the video and whether the imaging unit 11 (the information processing apparatus 1 including the imaging unit 11) is moving. Then, by using these two pieces of information, the camera work recognition unit 122b can recognize which of the four camera works shown in FIGS. 4(a) to 4(d) the video was shot with. For example, if the person 31 is not moving in video 3 and the movement of the imaging unit 11 (the information processing apparatus 1 including the imaging unit 11) is not detected either, the camera work recognition unit 122b recognizes that the video was shot with the camera work shown in FIG. 4(a). Also, for example, if the person 31 is moving in video 3 but the movement of the imaging unit 11 (the information processing apparatus 1 including the imaging unit 11) is not detected, the camera work recognition unit 122b recognizes that the video was shot with the camera work shown in FIG. 4(b). Also, for example, if the person 31 is not moving in video 3 and the movement of the imaging unit 11 (the information processing apparatus 1 including the imaging unit 11) is detected, the camera work recognition unit 122b recognizes that the video was shot with the camera work shown in FIG. 4(c). Further, for example, if the person 31 is moving in video 3 and the movement of the imaging unit 11 (the information processing apparatus 1 including the imaging unit 11) is also detected, the camera work recognition unit 122b recognizes that the video was shot with the camera work shown in FIG. 4(d).

[0025] The camera work recognition unit 122b recognizes, by the processing as described above, with which camera work each video was shot, and sends the recognition result to the file recording unit 13. The file recording unit 13 records the recognition result of the camera work as the camera work information 224 of the shooting information 22. The shooting information 22 of the video file 2 shown in FIG. 2 shows the case where the camera work of FIG. 4(b) is recognized, and (b) indicating that the camera work information 224 is the camera work of FIG. 4(b) is recorded.

[0026] In addition, in this embodiment, an example in which four camera works shown in FIGS. 4(a) to 4(d) are recognized has been given. However, the present invention is not limited to this, and the camera work recognition unit 122b may be configured to recognize other camera works. Also, in this embodiment, an example in which the detection output of the movement detection unit 123 is used for the movement detection of the imaging unit 11 (the information processing apparatus 1 including the imaging unit 11) has been given. However, other methods may be used. For example, similar to the detection of the movement of the person 31, when the movement of the tree 32, which is a part of the background in the video, is detected, the configuration may be such that the movement of the imaging unit 11 (the information processing apparatus 1 including the imaging unit 11) is detected.

[0027] The frame rate recognition unit 122c recognizes the frame rate indicating how many images the video 3 is composed of per second. As the unit of this frame rate, fps is generally used. For example, 120fps indicates that the video is composed of 120 images per second. This frame rate recognition process may be performed based on the settings of the imaging unit 11 at the time of video shooting, or may be performed based on the video data 21. That is, the frame rate recognition unit 122 may recognize the frame rate from the value of the frame rate set in the imaging unit 11 at the time of video shooting, or may measure how many frames the video data 21 is composed of per second to recognize the frame rate.

[0028] The information on the frame rate recognized by the frame rate recognition unit 122c is sent to the file recording unit 13 and recorded as the frame rate information 225 of the shooting information 22. The shooting information 22 of the video file 2 shown in FIG. 2 shows a state in which the frame rate of the video data 21 is recorded as 120fps.

[0029] The shooting information recognition unit 14 of the information processing apparatus 1 shown in FIG. 1 recognizes the shooting information 22 associated with the video data 21 of each video file 2 recorded in the file recording unit 13. In the case of this embodiment, the shooting information recognition unit 14 reads the shooting information 22 given to the video data 21 of the video file 2 recorded in the file recording unit 13. Note that the shooting information recognition unit 14 may have a configuration including recognition units such as a subject recognition unit 122a, a camera work recognition unit 122b, and a frame rate recognition unit 122c as shown in FIG. 3. In this case, the shooting information recognition unit 14 reads the video data 21 of the video file 2 from the file recording unit 13, and each recognition unit recognizes the shooting information 22 based on the video of the video data 21. Then, the shooting information recognition unit 14 sends the shooting information 22 to the classification aggregation unit 15 or the classification extraction unit 17.

[0030] Based on the shooting information 22 recognized by the shooting information recognition unit 14, the classification aggregation unit 15 classifies each video file 2 for each video file 2 having similar shooting information 22. Then, the classification aggregation unit 15 calculates the ratio of the video files 2 having similar shooting information 22 among all the video files 2 to be subjected to classification aggregation.

[0031] FIG. 5 is a diagram showing, in a table, an example of the result of the classification aggregation unit 15 reading the shooting information 22 of each video file 2 recorded in the file recording unit 13 and classifying and aggregating the shooting information 22 of those video files 2. In the case of the example of FIG. 5, a total of 10 video files represented by A to J are recorded in the file recording unit 13. In FIG. 5, the row of subject recognition information represents only the recognition result as to whether a person is included in the subject in the video. When a person is included, it is described as "person present", and when a person is not included, it is described as "person absent". (a) to (d) in the row of camera work information represent the recognition results of the camera work illustrated in FIGS. 4(a) to 4(d). For example, (a) in the row of camera work information of video file A indicates that the camera work shown in FIG. 4(a) has been recognized. The date information, video length information, and frame rate information are as shown in FIG. 5 respectively. For example, 3 / 1 of the date information indicates that it is March 1st, 10 seconds of the video length information indicates that the time length of the video is 10 seconds, and 60fps of the frame rate information indicates that the frame rate is 60fps.

[0032] In the case of the example shown in FIG. 5, the classification tabulation unit 15 targets all 10 video files from A to J for classification tabulation, and calculates the ratio of video files having similar shooting information. That is, the classification tabulation unit 15 calculates the ratio in which each of the subject recognition information, camera work information, and frame rate information is shot among the total 10 video files from A to J. For example, in the case of subject recognition information, there are 7 with "person present" and 3 with "person absent". Therefore, the classification tabulation unit 15 calculates that the ratio of video files in which a person is included as a subject among the entire 10 video files recorded in the file recording unit 13 is 70%, and the ratio of video files in which a person is not included as a subject is 30%. Also, in the case of camera work information, there are 4 for (a), 2 for (b), 1 for (c), and 3 for (d). Therefore, the classification tabulation unit 15 can calculate that the ratio of the camera work of FIG. 4(a) is 40%, the ratio of the camera work of FIG. 4(b) is 20%, the ratio of the camera work of FIG. 4(c) is 10%, and the ratio of the camera work of FIG. 4(d) is 30% in the entire 10 video files. Also, in the case of frame rate information, there are 5 at 30 fps, 3 at 60 fps, and 2 at 120 fps. Therefore, the classification tabulation unit 15 can calculate that the ratio of 30 fps is 50%, the ratio of 60 fps is 30%, and the ratio of 120 fps is 20% in the entire 10 video files.

[0033] In this embodiment, an example of calculating the ratio based on the number of video files 2 recorded in the file recording unit 13 (10 in the example of FIG. 5) has been given, but it is not limited to this. The classification tabulation unit 15 may calculate the ratio based on, for example, the time of the video length. For example, in the case of FIG. 5, the total video length of the entire 10 video files is 300 seconds, while the total video length of the 7 video files with "person present" is 150 seconds. Therefore, when the total video length of the entire 10 video files is based on 300 seconds, the ratio of video files with subject recognition information being "person present" is 50%. In this way, the classification tabulation unit 15 may calculate the ratio based on the length of the video. In addition, the classification tabulation unit 15 may calculate the ratio by combining the number of video files and the length of the video.

[0034] As yet another example, the classification and aggregation unit 15 may limit the video files in the video file 2 recorded in the file recording unit 13 that are the targets when calculating the ratio. For example, when the shooting information includes date information as in the present embodiment, the classification and aggregation unit 15 may calculate the ratio only for the videos of the video files having the shooting information of a specific date. For example, as shown in FIG. 5, when calculating the ratio only for the video files with the date information being 4 / 1 (April 1st), there are 4 video files with the date information being 4 / 1, and among them, 2 video files have the subject recognition information being "with people". Therefore, the classification and aggregation unit 15 can calculate the ratio of the video files with the subject recognition information being "with people" in the video files with the date information being 4 / 1 as 50%. In this way, the classification and aggregation unit 15 can calculate the ratio only for the videos of the video files having specific shooting information such as a specific date.

[0035] The display unit 16 of the information processing apparatus 1 shown in FIG. 1 is composed of a liquid crystal display or the like, and performs various displays for transmitting information to the user of this apparatus. In the case of the present embodiment, the display unit 16 can perform a display based on the ratio calculated by the classification and aggregation unit 15 as described above.

[0036] FIG. 6 is a diagram showing an example of a display based on the calculation results of the respective ratios performed by the classification and aggregation unit 15 based on the video file illustrated in FIG. 5. FIG. 6 shows an example in which the ratios calculated by the classification and aggregation unit 15 corresponding to the subject recognition information, the camera work information, and the frame rate information are displayed as the bar graphs 161, 162, and 163. In the case of the example of FIG. 5 described above, the ratio in which a person is included in the subject calculated based on the subject recognition information is 70%, and the ratio in which a person is not included is 30%. Therefore, in the bar graph 161 corresponding to the subject recognition information, the ratio in which a person is included is represented by the length of the band 161a, and the ratio in which a person is not included is represented by the length of the band 161b. Also, in the case of the camera work information, as described above, the camera work of FIG. 4(a) is 40%, the camera work of FIG. 4(b) is 20%, the camera work of FIG. 4(c) is 10%, and the camera work of FIG. 4(d) is 30%. Therefore, the bar graph 162 corresponding to the camera work information is represented by the lengths of the bands 162a to 162d according to the ratios of the respective camera works. Also, in the case of the frame rate information, as described above, 30 fps is 50%, 60 fps is 30%, and 120 fps is 20%. Therefore, the bar graph 163 corresponding to the frame rate information is represented by the lengths of the bands 163a to 163c according to the ratios of the respective frame rates.

[0037] Note that, in the example of FIG. 6, an example in which the ratio calculated by the classification and aggregation unit 15 is displayed as a bar graph is given, but the present invention is not limited thereto, and the display unit 16 can perform display in various ways, such as other display forms such as a pie chart or direct display of the ratio numbers.

[0038] The combining unit 18 of the information processing apparatus shown in FIG. 1 appropriately combines the video data 21 of the plurality of video files 2 recorded in the file recording unit 13 to generate one combined video. Here, when creating a combined video by combining the video data 21 of a plurality of video files 2, it is more desirable to combine videos having a plurality of different shooting information than to combine videos with similar shooting information such as the subject, camera work, frame rate, etc. This is because combining similar videos tends to result in a monotonous combined video, while on the other hand, when combining videos having a plurality of different shooting information, it is considered that a combined video with a rich and good-looking appearance can be obtained.

[0039] For example, a combined video in which a video showing a scenery is inserted between videos showing people is considered to be a combined video with a sharp and good-looking appearance rather than a combined video that always shows only people or only scenery. Similarly, for camera work, a combined video composed of various camera works is more likely to create a combined video with a sharp and good-looking appearance than a combined video composed of only the same camera work. Also, in the case of the frame rate, for example, when shooting at a high frame rate such as 120fps, it can be played back in slow motion by playing it back at 30fps during playback. By inserting some slow-motion videos into the combined video, a combined video with a sharp and good-looking appearance can be created.

[0040] That is, in order to enable the generation of a combined video with a rich and good-looking appearance, it is important to shoot videos with various subjects, various frameworks, and various frame rates in a well-balanced manner when shooting the videos to be used as materials for the combined video. For example, it is important to shoot videos showing people and videos showing scenery without people in a well-balanced manner, shoot with various camera works, and further shoot videos with a frame rate set at 120fps or the like.

[0041] Therefore, in the present embodiment, by displaying the ratio for each piece of shooting information as shown in FIG. 6 on the display unit 16, the user of the information processing apparatus 1 can grasp at a glance whether various pictures are shot in a well-balanced manner. That is, by looking at the display in FIG. 6, the user can grasp at a glance what ratio of videos of various subjects, camera works, and frame rates are shot for each video file as a video. In other words, the user can grasp which shooting information lacks video shooting, and can obtain an opportunity to perform shooting under shooting conditions that match the shooting information. Then, when shooting of videos with insufficient shooting information is performed, the file recording unit 13 is in a state where videos of various subjects, various camera works, and various frame rates are recorded. By combining the videos of a plurality of video files recorded in this file recording unit 13, it becomes possible to create a richly changing combined video by combining video data having a plurality of different shooting information.

[0042] For example, by checking the band graph 161 in FIG. 6 from the display on the display unit 16, the user can check at a glance what balance the videos with people and the videos without people are shot in. For example, in the case of the band graph 161 shown in FIG. 6, the user can recognize that the ratio of the videos without people is less than that of the videos with people. The videos without people are often landscape videos. By noticing that the ratio of this video is small, the user can notice that it is better to shoot a little more landscape videos. Note that there are individual differences in the optimal values for the ratio between the videos with people and the videos without people, and since the user can judge whether the ratio is appropriate while checking the band graph 161, the user can perform shooting at a shooting ratio suitable for the user. By recording the video files obtained by performing such shooting in the file recording unit 13, it becomes possible to create a richly changing combined video in which the videos with people and the landscape videos are combined in a well-balanced manner.

[0043] For example, the user can quickly check, by looking at the bar graph 162 shown in FIG. 6 on the display unit 16, the balance of camera work in each video file. For example, in the case of the bar graph 162 shown in FIG. 6, the user can recognize that the camera work shown in FIG. 4(c) is the least, and thus can notice that it is better to increase the shooting with the camera work shown in FIG. 4(c). Note that there are individual differences in the optimal values for the ratio of camera work, and since the user can judge whether the ratio is appropriate while checking the bar graph 162, shooting can be performed at a shooting ratio suitable for the user. By recording the video files obtained through such shooting in the file recording unit 13, it becomes possible to create a richly varying combined video in which videos shot with various camera works are combined in a well-balanced manner.

[0044] For example, the user can quickly check, by looking at the bar graph 163 shown in FIG. 6 on the display unit 16, the ratio at which the frame rate can be shot in each video file. For example, in the case of the bar graph 163 shown in FIG. 6, the user can recognize that the ratio of videos shot at a high frame rate of 120 fps is small, and thus can notice that it is better to shoot more videos at a frame rate of 120 fps. Note that there are individual differences in the optimal values for the ratio of the frame rate, and since the user can judge whether the ratio is appropriate while checking the bar graph 163, shooting can be performed at a shooting ratio suitable for the user. By recording the video files obtained through such shooting in the file recording unit 13, it becomes possible to create a richly varying combined video in which videos shot with various camera works are combined in a well-balanced manner. For example, it is possible to create a richly varying combined video in which videos that can be played back in slow motion are combined in a well-balanced manner.

[0045] The display shown in FIG. 6 is preferably performed at a timing such as when the video shooting ends, when the video is being confirmed, when the operation of the information processing apparatus 1 ends, or before combining videos to generate a combined video. For example, by performing the display when the video shooting ends, it is possible to give the user a hint as to what kind of shooting should be done next. Also, by performing such a display when confirming the video during video playback or when the operation of the information processing apparatus 1 ends, the user can know the shooting information that is still insufficient in a series of shootings. Further, by displaying before combining videos to generate a combined video, the user can confirm before video combination whether the videos corresponding to various shooting information are included in a well-balanced manner.

[0046] Also, in the information processing apparatus 1 of the present embodiment, in the video classified and aggregated by the classification aggregation unit 15, when the ratio of the shooting information under specific shooting conditions is less than a predetermined ratio, a display prompting to satisfy the predetermined ratio may be performed on the display unit 16. For example, when the predetermined ratios for videos with and without a person as a subject are each set to 40% or more, in the examples of FIGS. 5 and 6, the videos without a person are only 30%, which is less than 40% of the predetermined ratio. In such a case, the display unit 16 may perform a display recommending shooting videos without a person, such as "Let's shoot more landscapes." Thereby, it is possible to assist a user who does not know at what ratio shooting should be done in creating a good-looking combined video.

[0047] The classification extraction unit 17 of the information processing apparatus 1 shown in FIG. 1 calculates the ratios for each of the subject recognition information, camera work information, and frame rate information based on the shooting information of the video files obtained from the shooting information recognition unit 14, in the same manner as performed by the classification aggregation unit 15. Note that the classification extraction unit 17 may acquire information indicating the ratios in each of the subject recognition information, camera work information, and frame rate information from the classification aggregation unit 15. Then, the classification extraction unit 17 extracts videos from the video files recorded in the file recording unit 13 based on extraction conditions such that the ratio of videos with similar shooting information falls within a predetermined range of ratios.

[0048] For example, when extraction conditions are set such that the ratio of videos including a person as a subject is in the range of 40% to 50% and the ratio of videos not including a person is also in the range of 40% to 50%, the classification extraction unit 17 extracts videos from the file recording unit 13 based on the extraction conditions. For example, when each video file as shown in FIG. 5 is recorded in the file recording unit 13, the classification extraction unit 17 extracts videos from, for example, each of the video files A, B, C, D, G, and J based on the extraction conditions. As described above, the video files A, B, and C have subject recognition information of "with person", and the video files D, G, and J have subject recognition information of "without person". Thus, the videos extracted by the classification extraction unit 17 have 50% of videos including a person and 50% of videos not including a person, satisfying the set extraction condition of the ratio in the range of 40% to 50%.

[0049] Then, each video data of the six video files A, B, C, D, G, and J extracted by the classification extraction unit 17 is sent to the combining unit 18. Therefore, in the combining unit 18, each video data of the six video files A, B, C, D, G, and J is combined to generate one combined video. That is, the combined video becomes a combined video that well-balancedly includes a video with a person and a video of a landscape or the like without a person, and becomes an attractive and good-looking video with a sense of rhythm.

[0050] In the foregoing description, an example was given of extracting a video based on extraction conditions set for subject recognition information. However, videos may also be extracted based on the extraction conditions set for other shooting information such as camera work information and frame rate information, respectively. Of course, the extraction conditions are not limited to being limited to only one of the subject recognition information, camera work information, and frame rate information, and may be extraction conditions combining two or more of them.

[0051] Also, in the case of the classification extraction unit 17 as well, when calculating the ratio of videos, it is not limited to calculating based on the number of video files. For example, it may be calculated based on the time of the video length, or they may be combined for calculation. Also, in the foregoing example, an example was given in which all the video files recorded in the file recording unit 13 were the targets of extraction. However, it is not limited to this, and among all the video files recorded in the file recording unit 13, the video files to be extracted may be limited. For example, only the videos of the video files having shooting information on a specific date may be the targets of extraction, and the videos may be extracted such that specific shooting information satisfies a predetermined ratio for those target video files. Thereby, for example, it becomes possible to extract a video for generating a combined video corresponding to the conditions of a specific date such as a certain day.

[0052] Also, before combining the videos in the combining unit 18, the information processing apparatus 1 of the present embodiment may calculate the ratio of the videos of each shooting information in the classification aggregation unit 15 for the videos to be combined, and display the ratio calculation result on the display unit 16. In the case of this example, before generating the combined video, the user can confirm the ratio of the videos of each shooting information, and can notice in advance if the ratio of a certain video is insufficient. Also, the videos combined in the combining unit 18 are not limited to the videos extracted by the classification extraction unit 17, and may be arbitrarily selected videos by the user.

[0053] FIG. 7 is a diagram showing an example of a hardware configuration to which the information processing apparatus 1 of the present embodiment can be applied. The information processing apparatus 700 includes a CPU 701, a ROM 702, a RAM 703, a large-capacity memory 704, a network interface 706, an input device 707, a display device 708, a camera 709, a sensor 710, etc.

[0054] The CPU 701 controls the information processing apparatus 700 in an overall manner. The ROM 702 stores a control program for the CPU 701 to control the information processing apparatus 700 and an information processing program for executing the processing steps related to the functional units of the information processing apparatus shown in FIGS. 1 and 3 described above. The RAM 703 is a memory in which the program read from the ROM 702 is expanded and executed by the CPU 701. Also, the RAM 703 is used as a temporary storage area for temporarily storing data to be processed in various processes.

[0055] The camera 709 is an imaging device corresponding to the imaging unit 11 described above. The sensor 710 corresponds to the movement detection unit 123 described above and includes an acceleration sensor or the like. The network interface 706 includes a communication circuit or the like for performing communication via a network (not shown). The CPU 701 can exchange images and various information with external devices or the like via the network interface 706. The large-capacity memory 704 is an HDD, an SSD, or the like, and can store videos or the like photographed by the camera 709. The file recording unit 13 described above can record the video file 2 in the large-capacity memory 704. Then, the CPU 701 can also read out the shooting information of the video file stored in the large-capacity memory 704 and perform the processing described above.

[0056] The display device 708 is a display device on which the display is performed by the display unit 16 described above, and can display videos, still images, text, etc. The display content of FIG. 6 described above is displayed on the screen of the display device 708. The input device 707 is a device equipped with at least one of a keyboard for input, a pointing device for a screen display on the display device 708, a mouse, a touch panel, etc. A user of the information processing device 700 can input various instructions, information, etc. via the input device 707, and the CPU 701 performs processing in accordance with the information input by the user.

[0057] The hardware configuration of the information processing device 700 has components similar to those of hardware components installed in, for example, a smartphone, a tablet terminal, a personal computer, and even a digital camera or a digital video camera. Therefore, various functions realized by the information processing device 700 can be implemented as software (programs) executed by the CPU 701. That is, the CPU 701 can realize the processing steps related to each functional unit shown in FIGS. 1 and 3 by executing the information processing program according to this embodiment. Of course, each functional unit shown in FIGS. 1 and 3 may also be realized as a circuit configuration.

[0058] The present invention can also be realized by supplying a program that realizes one or more of the functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program.The present invention can also be realized by a circuit (e.g., ASIC) that realizes one or more of the functions. The above-described embodiments are merely examples of specific implementations of the present invention, and the technical scope of the present invention should not be construed as being limited by these embodiments. In other words, the present invention can be implemented in various forms without departing from its technical concept or main features.

[0059] The disclosure of this embodiment includes the following configuration, method, and program. (Configuration 1) recognition means for recognizing shooting information for each of a plurality of moving images; Based on the shooting information for each of the videos, classify the plurality of videos into groups of videos having similar shooting information, and a classification and aggregation means for aggregating the ratio of the videos having the similar shooting information with respect to the plurality of videos; A display means for displaying the ratio aggregated by the classification and aggregation means to the user; An information processing apparatus characterized by comprising the above. (Configuration 2) A recognition means for recognizing the shooting information for each of the plurality of videos; Based on the shooting information for each of the videos, classify the plurality of videos into groups of videos having similar shooting information, and classification and extraction means for extracting videos from the plurality of videos so that the ratio of the videos having the similar shooting information with respect to the plurality of videos satisfies a predetermined ratio; A combining means for combining the videos extracted by the classification and extraction means; An information processing apparatus characterized by comprising the above. (Configuration 3) Classification and extraction means for extracting videos from the plurality of videos so that the ratio of the videos having the similar shooting information with respect to the plurality of videos satisfies a predetermined ratio; A combining means for combining the videos extracted by the classification and extraction means; The information processing apparatus according to Configuration 1, characterized by comprising the above. (Configuration 4) The classification and aggregation means calculates the ratio of the videos having the similar shooting information with respect to the plurality of videos based on at least one of the number of the plurality of videos and the length of the videos. The information processing apparatus according to Configuration 1 or 3. (Configuration 5) The classification and extraction means calculates the ratio of the videos having the similar shooting information with respect to the plurality of videos based on at least one of the number of the plurality of videos and the length of the videos. The information processing apparatus according to Configuration 2 or 3. (Configuration 6) The classification and aggregation means calculates the ratio of videos having similar shooting information to videos having specific shooting information among the plurality of videos, and is the information processing apparatus according to any one of Configurations 1, 3, and 4. (Configuration 7) The classification and extraction means calculates the ratio of videos having similar shooting information to videos having specific shooting information among the plurality of videos, and is the information processing apparatus according to any one of Configurations 2, 3, and 5. (Configuration 8) The shooting information includes at least any one of information on a specific subject reflected in the video, information on camera work when the video is shot, and information on the frame rate of the video, and is the information processing apparatus according to any one of Configurations 1 to 7. (Configuration 9) The recognition means recognizes at least any one of a person, an animal, a human face, a specific individual, and a specific entity as the specific subject from the video, and includes information on the recognized specific subject in the shooting information, and is the information processing apparatus according to Configuration 8. (Configuration 10) The recognition means includes only information on a specific subject recognized from the video for a predetermined time or longer as the information on the specific subject included in the shooting information, and is the information processing apparatus according to Configuration 9. (Configuration 11) The recognition means recognizes camera work when the video is shot based on at least any one of the movement of the subject reflected in the video and the movement of the device that shot the video, and is the information processing apparatus according to Configuration 8. (Configuration 12) The display means performs display of the ratio at least at any one of the timing when shooting of the video ends, when the video is confirmed, when the operation of the information processing apparatus ends, and before generation of a combined video to which the video is combined, and is the information processing apparatus according to Configuration 3. (Configuration 13) When the ratio of videos having the specific photographed information aggregated by the classification aggregation means is less than a predetermined ratio, the display means performs a display prompting to satisfy the predetermined ratio. The information processing apparatus according to any one of Configurations 1, 3, 4, 6, and 12. (Method 1) A recognition step of recognizing the photographed information for each of the plurality of videos; A classification aggregation step of classifying the plurality of videos into groups of videos having similar photographed information based on the photographed information for each of the videos, and aggregating the ratio of videos having the similar photographed information with respect to the plurality of videos; A display step of displaying the ratio aggregated in the classification aggregation step to the user; An information processing method characterized by comprising: (Method 2) A recognition step of recognizing the photographed information for each of the plurality of videos; A classification extraction step of classifying the plurality of videos into groups of videos having similar photographed information based on the photographed information for each of the videos, and extracting videos from the plurality of videos so that the ratio of videos having the similar photographed information satisfies a predetermined ratio; A combining step of combining the videos extracted in the classification extraction step; An information processing method characterized by comprising: (Program 1) A program for causing a computer to function as the information processing apparatus according to any one of Configurations 1 to 13.

Explanation of Reference Numerals

[0060] 1: Information processing apparatus, 12: File generation unit, 13: File recording unit, 14: Photographed information recognition unit, 15: Classification aggregation unit, 16: Display unit, 17: Classification extraction unit, 18: Combining unit, 121: Data generation unit, 122: Photographed information generation unit, 123: Movement detection unit

Claims

1. recognition means for recognizing shooting information for each of the plurality of videos; classification aggregation means for classifying the plurality of videos into groups of videos having similar shooting information based on the shooting information for each of the videos, and aggregating the ratio of the videos having the similar shooting information with respect to the plurality of videos; display means for displaying the ratio aggregated by the classification aggregation means to the user; An information processing apparatus characterized by comprising:

2. recognition means for recognizing shooting information for each of the plurality of videos; classification extraction means for classifying the plurality of videos into groups of videos having similar shooting information based on the shooting information for each of the videos, and extracting videos from among the plurality of videos so that the ratio of the videos having the similar shooting information with respect to the plurality of videos satisfies a predetermined ratio; combining means for combining the videos extracted by the classification extraction means; An information processing apparatus characterized by comprising:

3. classification extraction means for extracting videos from among the plurality of videos so that the ratio of the videos having the similar shooting information with respect to the plurality of videos satisfies a predetermined ratio; combining means for combining the videos extracted by the classification extraction means; The information processing apparatus according to claim 1, characterized by comprising:

4. The classification aggregation means calculates the ratio of the videos having the similar shooting information with respect to the plurality of videos based on at least one of the number of the plurality of videos and the length of the videos. The information processing apparatus according to claim 1 or 3.

5. The classification extraction means calculates the ratio of the videos having the similar shooting information with respect to the plurality of videos based on at least one of the number of the plurality of videos and the length of the videos. The information processing apparatus according to claim 2 or 3.

6. The classification aggregation means calculates the ratio of the videos having the similar shooting information with respect to the videos having specific shooting information among the plurality of videos. The information processing apparatus according to claim 1 or 3.

7. The classification extraction means calculates the ratio of the videos having the similar shooting information with respect to the videos having specific shooting information among the plurality of videos. The information processing apparatus according to claim 2 or 3.

8. The information processing apparatus according to any one of claims 1 to 3, wherein the shooting information includes at least one of information on a specific subject shown in the moving image, information on camera work when the moving image is shot, and information on the frame rate of the moving image.

9. The information processing apparatus according to claim 8, wherein the recognition means recognizes at least one of a person, an animal, a human face, a specific individual, and a specific entity as the specific subject from the moving image, and includes information on the recognized specific subject in the shooting information.

10. The information processing apparatus according to claim 9, wherein the recognition means includes only information on a specific subject recognized from the moving image for a predetermined time or longer as the information on the specific subject in the shooting information.

11. The information processing apparatus according to claim 8, wherein the recognition means recognizes camera work when the moving image is shot based on at least one of the movement of the subject shown in the moving image and the movement of the apparatus that shot the moving image.

12. The information processing apparatus according to claim 3, wherein the display means performs the display of the ratio at at least one of the timing of the end of shooting of the moving image, the timing of confirmation of the moving image, the timing of the end of operation of the information processing apparatus, and the timing before generation of a combined moving image to which the moving image is combined.

13. The information processing apparatus according to claim 1 or 3, wherein when the ratio of the moving images having the specific shooting information aggregated by the classification aggregation means is less than a predetermined ratio, the display means performs a display prompting satisfaction of the predetermined ratio.

14. A recognition step of recognizing shooting information for each of a plurality of moving images; A classification aggregation step of classifying the plurality of moving images into groups having similar shooting information based on the shooting information for each of the moving images, and aggregating the ratio of the moving images having the similar shooting information with respect to the plurality of moving images; A display step of displaying the ratio aggregated in the classification aggregation step to a user; An information processing method characterized by including:

15. A recognition step of recognizing shooting information for each of a plurality of moving images; Based on the shooting information for each of the plurality of videos, classify the plurality of videos into groups of videos having similar shooting information, and extract videos from the plurality of videos so that the ratio of videos having the similar shooting information to the plurality of videos satisfies a predetermined ratio, which is a classification and extraction step; A combining step of combining the videos extracted in the classification and extraction step; An information processing method, characterized by comprising the above.

16. A computer, A recognition means for recognizing the shooting information for each of the plurality of videos; Based on the shooting information for each of the videos, classify the plurality of videos into groups of videos having similar shooting information, and a classification and totaling means for totaling the ratio of videos having the similar shooting information to the plurality of videos; A display means for displaying the ratio totaled by the classification and totaling means to the user; A program for causing the computer to function as an information processing apparatus having the above.

17. A computer, A recognition means for recognizing the shooting information for each of the plurality of videos; Based on the shooting information for each of the videos, classify the plurality of videos into groups of videos having similar shooting information, and a classification and extraction means for extracting videos from the plurality of videos so that the ratio of videos having the similar shooting information to the plurality of videos satisfies a predetermined ratio; A combining means for combining the videos extracted by the classification and extraction means; A program for causing the computer to function as an information processing apparatus having the above.

Citation Information

Patent Citations

  • Shooting assist method, program thereof, recording medium, recording medium, shooting device, and shooting system

    JP2011211695A