Video generation device, method for operating video generation device, and program for operating video generation device
The video generation device addresses the challenge of unnatural transitions by analyzing still images to control variable elements like subject movement and camera work, resulting in more realistic video output.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2026-03-05
AI Technical Summary
Existing video generation technologies struggle to produce videos with natural changes, lacking control over variable elements such as subject movement, background changes, and camera work, resulting in unnatural transitions.
A video generation device that includes a processor to analyze still images for variable elements, control the speed of these elements based on image analysis and user input, and generate videos with more natural changes by incorporating subject attributes, background variations, and camera work adjustments.
The device generates videos with more natural image changes by controlling the speed of variable elements, enhancing the realism and variety of video content.
Smart Images

Figure JP2025028527_05032026_PF_FP_ABST
Abstract
Description
Video creation device, video creation device operation method, and video creation device operation program
[0001] The technology of the present disclosure relates to a moving image generating device, an operating method for a moving image generating device, and an operating program for a moving image generating device.
[0002] JP 2021-033961 A describes a technology in which a computer stores video generation models, which have been trained to generate and output video based on input still images, in different categories, receives content including text and still images, determines a category based on the text or still images contained in the received content, selects the video generation model corresponding to the determined category, and inputs still images from the content into the selected video generation model to generate video.
[0003] The technology according to the present disclosure provides a moving image generation device, an operating method for a moving image generation device, and an operating program for a moving image generation device that can generate moving images with more natural changes than conventional devices.
[0004] The video generation device according to the disclosed technique includes a processor, which, when generating a video based on a still image, acquires information about variable elements of the still image and controls the speed at which the variable elements change in the video based on the information about the variable elements.
[0005] A variable element may be an element that can be changed in the animation.
[0006] The variable elements may include at least one of a subject that can be given movement in the video, a background that can be changed in the video, and camera work that can be changed in the video.
[0007] The information about the variable element may include information obtained by image analysis of a still image.
[0008] The information about the variable element may include at least one of an attribute and a state of the subject recognized from the still image.
[0009] The processor may further use the conditions under which the still image was taken to control the speed.
[0010] In the case where the still image is included in one of a plurality of still images taken in continuous shooting, the processor may further use the plurality of still images taken in continuous shooting for speed control.
[0011] The information about the variable element may include information obtained through a user operation.
[0012] The information about the variable elements may include information about camera work that can be changed in the video.
[0013] The processor may be capable of accepting a user's wishes regarding the presentation of the video.
[0014] The processor may output the generated video so that it can be previewed, and when a user request for the video is input, may regenerate the video in response to the input request.
[0015] A method for operating a moving image generating device according to the technology of the present disclosure is a method for operating a moving image generating device equipped with a processor, in which, when generating a moving image based on a still image, the processor acquires information regarding variable elements of the still image and controls the speed at which the variable elements change in the moving image based on the information regarding the variable elements.
[0016] The operating program of the video generation device relating to the technology of the present disclosure is an operating program of a video generation device equipped with a processor, which, when generating a video based on a still image, causes the processor to perform processing including obtaining information about variable elements of the still image and controlling the speed at which the variable elements change in the video based on the information about the variable elements.
[0017] According to the technology of the present disclosure, it is possible to generate moving images with more natural image changes than before.
[0018] 1 is a diagram illustrating an overview of a video generation device. FIG. 2 is a diagram illustrating an example of a hardware configuration constituting a video generation device. FIG. 3 is a diagram illustrating an overview of processing of a video generation device. FIG. 4 is a diagram illustrating an overview of generation condition acquisition processing. FIG. 5 is a table illustrating specific examples of variable elements. FIG. 6 is a table illustrating the relationship between specific examples of subject attributes and movement speed. FIG. 7 is a diagram illustrating an example of controlling movement speed according to the type of subject. FIG. 8 is a diagram illustrating another example of controlling movement speed according to the type of subject. FIG. 9 is a diagram illustrating an example of a still image including multiple subjects. FIG. 10 is a diagram illustrating specific examples of the state of the subject. FIG. 11 is a diagram illustrating an example of using shooting conditions for speed control. FIG. 12 is a diagram illustrating an example of including a background as a variable element. FIG. 13 is a diagram illustrating an example of including camera work as a variable element. FIG. 14 is a diagram illustrating an example of a method for specifying camera work. A flowchart illustrating an example of the procedure for video generation processing. FIG. 15 is a diagram illustrating an example of accepting a user's wishes regarding the direction of a video. An example of using multiple still images captured in rapid succession to generate a video.
[0019] As shown in FIG. 1, the video production device 10 is a device that generates a video MV based on a still image SI. The still image SI is, for example, an arbitrary still image designated by a user 11. The still image SI is, for example, a photograph taken with a camera, but may also be an image other than a photograph, such as an illustration. Based on the data of such a still image SI, the video production device 10 generates a video MV in which the subject or the like depicted in the still image SI moves. The video production device 10 is, for example, configured by a personal computer.
[0020] 2, the computer constituting the video generating device 10 includes a display unit 21, an input unit 22, a processor 23, a memory 24, a communication I / F (Interface) 25, and a storage 27. These components are interconnected via a bus line 26.
[0021] The display unit 21 is a display that displays still images SI, videos MV, and operation screens. The input unit 22 is a keyboard, mouse, and the like for inputting operation instructions. The display unit 21 is configured as a touch panel display, for example, and also functions as an input unit for inputting operation instructions. The storage 27 is a data storage device such as a flash memory, a hard disk drive, or a solid state drive. The memory 24 is a work memory for the processor 23 to execute processing, and is configured by a RAM (Random Access Memory), etc.
[0022] The storage 27 stores control programs such as an operating system, various application programs, and various data associated with these programs. The processor 23 is, for example, a CPU (Central Processing Unit). The processor 23 loads programs stored in the storage 27 into the memory 24 and executes processing in accordance with the programs, thereby providing overall control of each part of the computer. The communication I / F 25 is a network interface for communication via a communication network and performs transmission control. Still images SI and videos MV are communicated with external devices such as cameras via the communication I / F 25.
[0023] 3, a program 41 is stored in the storage 27 of the video production device 10. The program 41 is an example of an operating program for the video production device according to the technique of the present disclosure. By executing the program 41, the processor 23 functions as a processing unit that executes each process of the video production device.
[0024] The program 41 includes a video generation model ML, which is a machine learning model that generates a video MV from a still image SI. The video generation model ML is also called generative AI (artificial intelligence). The processor 23 executes a process for generating a video MV using the video generation model ML. Conceptually, the process executed by the processor 23 is composed of a generation condition acquisition process and a video generation process. The generation conditions are conditions for generating a video MV, and the generation condition acquisition process is executed based on, for example, the input still image SI and an input request from the user 11.
[0025] The generation conditions include information about variable elements of the still image SI. The processor 23 controls the speed at which the variable elements change in the video MV based on the information about the variable elements. The variable elements are elements that can be changed in the video MV. More specifically, the variable elements include at least one of a subject that can be given movement in the video MV, a background that can be changed in the video MV, and camerawork that can be changed in the video MV.
[0026] In FIG. 3, the still image SI is exemplified as an image of a running person as a subject. The video MV generated from this still image SI is exemplified as a video of a running person gradually approaching from a distance. The symbol F indicates a frame included in the video MV. Although three frames F are shown in the video MV, in reality, three or more frames F are included. More specifically, the number of frames F included corresponds to the frame rate and playback time of the video MV. The symbol T indicates time.
[0027] In the example shown in Figure 3, a person, which is an example of a subject that can be given movement in a video MV, is used as a variable element of the still image SI. Furthermore, the shadow of a running person, which is an example of a background that can be changed in a video MV, is used as a variable element of the still image SI. In the video MV, the shadow also shifts as the person moves. The video MV shown in Figure 3 is intended to be shot with a fixed camera angle of view and shooting position, but as will be described later, it is also possible to add camerawork such as zooming, translation, or panning.
[0028] Figure 4 shows the generation condition acquisition process in more detail. Because video generation is performed by the video generation model ML, it is difficult to clearly distinguish the individual processes, but the concept is as shown in Figure 4. That is, the generation condition acquisition process includes processes such as image analysis, user request acceptance, determination of variable elements to be adopted, and speed estimation. The video generation model ML is trained to perform these processes.
[0029] First, the subject of the still image SI is detected by image analysis. For example, the image analysis includes object recognition such as semantic segmentation using a convolutional neural network, or pattern matching. The processor 23 acquires information about the subject of the still image SI, such as attributes including the type of the subject and the state of the subject, through image analysis. The attributes of the subject detected through image recognition include, for example, the type of person, the age group of young people, and the gender of male. Furthermore, the image recognition also detects the shadow of a person in the background.
[0030] In accepting user input, user requests such as information about camerawork such as zoom, information about the subject's movement speed, and information about the subject's movement direction, such as moving the subject from left to right, are input. In this way, information about variables may be acquired through user operations. User input can be performed by gesture input using a finger on a touch panel display, or by inputting a natural language prompt as text or voice. Voice input requires a microphone as hardware.
[0031] The movement speed instruction may not be a numerical value, but may be an instruction such as "move the person slowly" or "move the person quickly." The movement direction instruction may be an instruction such as "move the person from right to left" or "move the person from back to front."
[0032] As will be described later, the movement speed of a subject is basically determined based on the results of image analysis of the type of subject, etc. (see FIG. 6, etc.), but may also be obtained by user input. The direction of movement can also be determined based on the posture of the subject in the still image SI, etc. The direction of movement may also be obtained based on the results of image analysis, or may also be obtained by user input. The example video MV shown in FIG. 3 does not include camerawork.
[0033] In this way, through image analysis or user input, the processor 23 acquires information about variable elements such as the subject, background, and camerawork. The processor 23 then determines which variable elements to use as generation conditions, all or part of the acquired information about the variable elements. Information about which variable elements to prioritize in a video MV can be learned in advance, for example, in a video generation model. Alternatively, the variable elements to prioritize can be specified by user input. Furthermore, candidate variable elements to be adopted by the processor 23 may be presented to the user 11 via the display unit 21, with the user 11 making the final decision.
[0034] Furthermore, the processor 23 controls the speed at which the variable elements change in the video MV based on information about the variable elements employed. The speed at which the variable elements change is the speed at which the subject displaces, including the subject's movement speed and the speed at which the subject's posture changes. The speed at which the variable elements change also includes the speed at which the background, such as shadows, changes in the video MV, as well as the speed of camerawork, such as zooming and panning. The processor 23 determines the speed at which the variable elements change in the video MV based on the results of image analysis and information input by the user.
[0035] An example of information related to variable elements is shown in Figure 5. The information related to variable elements includes the variable elements themselves and information that forms the basis for estimating the speed at which the variable elements change. In Figure 5, the variable elements are broadly categorized as the subject, background, and camerawork, which can change the movement of a video.
[0036] Of these, the subject is classified into medium-level categories, such as subject attributes and subject state. The subject attributes include sub-categories such as subject type, such as people, animals, vehicles, and other moving elements. The subject's age and gender are included in the category of people. The animals include dogs, cats, and birds. The vehicles include cars, trains, and airplanes. The other moving elements include windmills, clouds, and rivers. The subject state includes sub-categories such as the subject's posture, including its direction; the subject's distance, which is reflected in its size in the image; and the appearance, including the degree of blur. The video generation model ML is trained to recognize subjects through image analysis according to the classification shown in FIG. 5. Furthermore, if the image analysis results are inappropriate, the user can input information about the subject, such as the subject's attributes.
[0037] In the example shown in FIG. 5 , anything that changes in brightness or color is classified as background. The background includes sub-items such as sunlight, weather, and lighting. Shadows are an element that changes depending on the direction of the sun, so as an example, they are included in sunlight. Weather conditions include weather changes such as sunny, cloudy, and rainy. Lighting includes neon signs and lighthouse lights (searchlights, etc.). The video generation model ML is trained by image analysis so that it can recognize backgrounds according to the classification shown in FIG. 5 . It is also possible for the user to input information about the background.
[0038] Furthermore, camerawork relates to changes in composition and includes zooming, which changes the angle of view, as well as panning, tilting, rolling, and translation, which change the shooting direction. Camerawork is input, for example, by a user. Of course, even in the absence of user input, the processor 23 can automatically add camerawork as a variable element to the video MV. For example, if image analysis of a still image SI reveals that the subject is a running person, the processor 23 may automatically decide to add zoom to create a more impactful image.
[0039] Furthermore, priority information may be set for each variable element as shown in Fig. 5. Specifically, the moving image generation model ML is trained so that variable elements are adopted according to such priorities.
[0040] Furthermore, in video generation, the speed of displacement, such as movement and posture change, is controlled according to the attributes of the subject, as conceptually shown in FIG. 6 . Specifically, the video generation model ML is trained so that the speed of displacement changes according to the type of subject. As shown in FIG. 6 , the speed of displacement of the subject differs according to the attributes of major categories such as living things, vehicles, play equipment, and other elements. Furthermore, for people, the speed of displacement differs according to attributes such as age and gender. For example, babies cannot walk fast. Infants and children aged 0 to 11 years old can generally walk and run faster than babies, but their walking and running speeds are slower than those of older young people (teens to 30s).
[0041] For example, in Fig. 7, Fig. 7(A) shows an example of generating a video MV based on a still image SI of a mother and a girl facing each other, and Fig. 7(B) shows an example of generating a video MV based on a still image SI of a baby. In Fig. 7(A), a video MV of a girl walking toward her mother is generated, and in Fig. 7(B), a video MV of a baby walking is generated. In this case, if the girl and the baby move at the same speed, it would give an unnatural impression, so the speed of movement is determined depending on the type of subject, such as the girl or the baby.
[0042] Returning to Figure 6, middle-aged people (in their 40s to 60s) move at slower speeds than young people. Elderly people (over 70s) move at even slower speeds. In addition, the way they move differs depending on their age. Furthermore, while there is little difference between boys and girls in terms of the speed of change in babies, there may be differences between boys and girls depending on their age.
[0043] Similarly, the speed at which animals move varies greatly between dogs and cats and birds. Similarly, the speed at which vehicles move varies between automobiles, trains, and airplanes. Furthermore, the speed at which vehicles move varies depending on the subcategory that is further subdivided into intermediate categories, such as passenger cars and racing cars for automobiles, local trains and bullet trains for trains, and airplanes and helicopters for airplanes.
[0044] For example, in Fig. 8, Fig. 8(A) is an example in which a moving image MV is generated based on a still image SI of a passenger car as a subject, and Fig. 8(B) is an example in which a moving image MV is generated based on a still image SI of a racing vehicle as a subject. In this case, if the moving speeds of the passenger car and the racing vehicle were the same, it would give an unnatural impression, so the moving speed is determined depending on the type of subject, i.e., passenger car or racing vehicle.
[0045] 6, the speed of playground equipment also differs depending on the type of medium-sized category, such as balls, swings, kites, and balloons. Similarly, the speed of balls also differs depending on the type of small-sized category, such as soccer balls, volleyballs, and baseballs. Other objects, such as windmills, clouds, rivers, and plants swaying in the wind, also differ in speed.
[0046] The moving image generation model ML is trained so that the speed at which these subjects move in the moving image MV becomes an appropriate speed according to their attributes.
[0047] Furthermore, as shown in Fig. 9, a still image SI may contain multiple subjects. In this case, the attributes of each subject are naturally determined, and the displacement speed is determined according to the determined attributes. In Fig. 9, the subjects include two boys, one girl, a car, a river, and grass swaying in the wind.
[0048] In this way, when a still image SI contains multiple subjects, it may be possible to specify the subject to be displaced in the video MV. One possible method of specification is, for example, as shown in FIG. 9 , a method in which the user specifies the subject with a finger FG in the still image SI displayed on a touch panel display. FIG. 9 shows an example in which a river is specified by a tap operation with a finger FG. In general, compared to moving objects such as people, cars, and airplanes, rivers can be difficult to detect using image recognition. The method of specifying a variable element through user operation is particularly effective in such cases.
[0049] 9 depicts a scene in which a boy and a girl are chasing each other. In such a case, the user 11 may specify the direction of movement of each of the boy and the girl. The direction of movement is specified, for example, by a swipe operation using a finger FG. In this way, it is possible to naturally depict a scene of chasing, for example, a boy running toward another boy or girl, in the generated video MV.
[0050] Another method of specifying the subject to be displaced may be for the processor 23 to display an identification frame (not shown) such as a bounding box for multiple subjects recognized by image analysis, and for the user 11 to specify the subject to be displaced from within that frame.
[0051] Next, the state of the subject shown in FIG. 5 , which is an example of information related to variable elements, will be described using the still image SI shown in FIG. 9 as an example. The state of the subject is organized as shown in FIG. 10 as an example. As shown in FIG. 10 , the state of the subject includes, for example, the subject distance, the subject's behavior, posture, and appearance. The still image SI shown in FIG. 9 is an image of a boy and girl running. Running is an example of the subject's behavior. Furthermore, the posture of the subject, in the case of a running subject, includes the degree to which the subject leans forward and the extent to which the legs are spread apart. Furthermore, the appearance of the subject includes the degree of blur.
[0052] Such a subject's state can be used as information that affects the subject's movement speed. For example, even if the actual movement speed is the same, the closer the subject distance (for example, whether it is in front or behind), the faster the movement speed in the video MV is considered to be. Furthermore, regarding the subject's posture, the more forward the subject leans, the faster the running speed is considered to be, so the more forward the subject leans, the faster the movement speed is considered to be. Furthermore, regarding how the subject is captured in a still image SI, if the shutter speed is the same, the faster the movement speed, the greater the degree of blur. In other words, the greater the degree of blur relative to the shutter speed, the faster the movement speed.
[0053] 11 , the processor 23 acquires the shutter speed from the shooting conditions included in the metadata attached to the still image SI. The metadata is, for example, additional information included in the image file, such as an Exchangeable Image File Format (Exif) tag. The processor 23 uses the state of the subject acquired by image analysis and the shutter speed acquired from the metadata to control the speed of the subject in the moving image MV.
[0054] In this way, the moving image generation model ML is trained so that the speed at which these subjects move in the moving image MV becomes an appropriate speed according to the state of the subjects.
[0055] The example shown in FIG. 12 illustrates an example in which the background shown in FIG. 5 changes. FIG. 12(A) illustrates an example in which a video MV is generated based on a still image SI in which the subject is an airplane and the shadow of the airplane appears in the background. That is, in the example shown in FIG. 12(A), the variable elements are the airplane as the subject and the shadow cast by sunlight. The speed at which the shadow moves is affected not only by the speed at which the airplane is moving but also by the direction of the shadow (i.e., corresponding to the direction of the sun). The video generation model ML is trained to control the speed at which the shadow moves, taking these elements into consideration.
[0056] 12B, a still image SI has a lighthouse as its subject, and the lighthouse searchlight is the variable element. That is, the variable element is lighting. The moving image generation model ML is trained so that the speed at which the lighting state changes is controlled according to the type of lighting, such as the lighthouse searchlight.
[0057] The example shown in Fig. 13 is an example in which the camerawork shown in Fig. 5 is added as a variable element. In Fig. 13(A), the variable elements are the subject and camerawork, and in the video MV, not only does the subject, the girl, move closer to her mother, but, unlike Fig. 7(A), zooming is added as camerawork. The zooming changes the angle of view from one that includes both the mother and the girl to one that shows a close-up of the girl's face.
[0058] The example shown in Figure 13(B) is an example in which a video MV is generated based on still images SI of two men and women as subjects. In Figure 13(B), the only variable element used in the video MV is camerawork that combines zoom and pan to sequentially capture close-ups of the faces of the two men and women, and the subjects do not move. In other words, the video MV generated in Figure 13(B) is a so-called photo movie in which the angle of view or cropping range changes within the shooting range of the still images SI. The generated video MV may be such a photo movie.
[0059] The speed of the camerawork as shown in Fig. 13 is input, for example, through a user operation as shown in Fig. 14. As shown as an example in Fig. 14, the speed is input by the user 11 swiping with a finger FG on a touch panel display on which a still image SI is displayed. For example, if the swiping speed is increased, the speed of the camerawork such as zooming or panning will increase, and if the swiping speed is decreased, the speed will decrease.
[0060] In addition, camerawork specifications other than speed are also input through user operations such as touch gestures. For example, the direction of camerawork is controlled by the position and direction of swiping the finger FG. Zooming can also be instructed by touch gestures such as pinching out and pinching in. Furthermore, camerawork specifications may be made by inputting a natural language prompt, such as "zoom so that the person's face is enlarged" or "pan from right to left." Furthermore, user operations can also specify not only speed but also the amount of change in variables such as the amount of movement of the subject and the amount of change in camerawork.
[0061] In the changing background example shown in FIG. 12, the rate at which the daylight or lighting conditions change may be input by user interaction, such as touch gestures or natural language prompts as shown in FIG.
[0062] The moving image generation procedure in moving image generation device 10 will be described below with reference to the flowchart shown in Fig. 15. In step ST1000, processor 23 of moving image generation device 10 waits for user 11 to designate a still image SI and input a moving image generation instruction. If a moving image generation instruction is input (Y in step ST1000), processor 23 proceeds to step ST1100 and acquires the generation conditions for moving image MV through moving image generation model ML. The generation conditions are acquired by image analysis and user input as shown in Figs. 4 and 11.
[0063] In step ST1200, if there is a user input regarding the generation conditions, the input user request is added. If there is a user request regarding variable elements including camerawork and the speed of the variable elements, in step ST1300, processor 23 adds the user request as a generation condition.
[0064] In step ST1400, the processor 23 generates a moving image in accordance with the acquired generation conditions using the moving image generation model ML. In generating the moving image, the speed of the variable elements is controlled based on information about the variable elements.
[0065] In step ST1500, the processor 23 outputs the generated moving image MV to the display unit 21 so that the moving image MV can be previewed.
[0066] In step ST1600, the user 11 checks the previewed moving image MV and inputs a modification request as necessary. Modification requests include, for example, a request to add a variable element to be moved, a request to modify the speed of a variable element by speeding it up or slowing it down, and a request to modify the direction of movement. Such modification requests are input by touch gestures such as those described with reference to FIG. 14, for example. Of course, they may also be input by a prompt in natural language. The prompt may be input by text or voice.
[0067] If there is a request for revision, processor 23 returns to step ST1300 and repeats the subsequent processes to regenerate the video MV. If there is no request for revision, processor 23 proceeds to step ST1700 and stores the generated video MV in storage 27.
[0068] In step ST1800, if another moving image is to be generated, the processor 23 returns to step ST1000, and if another moving image is not to be generated, the processor 23 ends the moving image generation process.
[0069] As described above, the video generating device 10 according to the technology of the present disclosure includes a processor 23. When generating a video MV based on a still image SI, the processor 23 acquires information about variable elements of the still image SI and controls the speed at which the variable elements change in the video MV based on the information about the variable elements. This allows for the generation of videos with more natural image changes than in the past. For example, by controlling the speed at which a subject moves according to the attributes of the subject, a video MV is generated in which the subject moves at a speed according to the attributes.
[0070] Furthermore, in the above embodiment, the variable elements are elements that can be changed in the video MV, so that a more natural-looking video MV can be generated.
[0071] In the above embodiment, the variable elements include at least one of a subject that can be animated in the video music video, a background that can be changed in the video music video, and camerawork that can be changed in the video music video, thereby enabling the creation of a variety of video music videos.
[0072] In the above embodiment, the information on the variable elements includes information obtained by image analysis of the still images SI, which makes it easier to generate a video MV that matches the content of the still images SI.
[0073] In the above embodiment, the information about the variable elements includes at least one of the attributes and state of the subject recognized from the still image SI. Since the movement of the subject is an impressive element in a video MV, controlling the speed according to the attributes and / or state of the subject makes it possible to generate a video MV that looks more natural.
[0074] In the above embodiment, the processor 23 controls the speed based on the shooting conditions of the still image SI. For example, as shown in Fig. 11, the moving speed of the subject is controlled based on the degree of blur of the subject in the still image SI and the shutter speed, thereby more accurately reflecting the actual movement of the subject in the still image SI.
[0075] In the above embodiment, the cause of blurring is explained as being due to the movement of the subject, but camera shake can also be considered as a cause of blurring. In that case, the cause of blurring may be identified using information from a camera shake detection sensor or the like. If camera shake is the cause of blurring, it may be possible to use the degree of blurring without using it to estimate the subject's speed.
[0076] Furthermore, in the above embodiment, the information about the variable elements includes information obtained through user operation. This makes it easier to reflect the intentions of the user 11 in the video music video. As described above, the operation of the user 11 includes any of gesture input, text input, and voice input. Such operation by the user 11 makes it easy to input information such as the speed, direction, and amount of movement of the variable elements. Here, the information about the variable elements includes information about camerawork that can be changed in the video music video. Camerawork is an important element in the direction of the video music video. By making it possible to specify camerawork through user operation, it is easy to specify a variety of camerawork. This makes it possible to produce a wide variety of different direction in the video music video.
[0077] (Variation 1) Furthermore, as shown in FIG. 16 , the processor 23 may be able to accept a user's wishes regarding the direction of the video music video. The user's wishes regarding the direction of the video music video may be, for example, "I want a dynamic video," "I want a powerful video," or "I want a video with a calm impression." These may be input, for example, by a natural language prompt. When the processor 23 accepts such a user's wishes, it generates a video music video that incorporates direction that reflects the wishes. This makes it easier for the processor 23 to generate a video music video that reflects the user's intentions.
[0078] 17 , when a still image SI is included in one of a plurality of still images captured in continuous shooting, the processor 23 may further use the plurality of still images SI captured in continuous shooting for speed control. For example, if there are a plurality of still images SI captured in continuous shooting that show a moving subject, the moving speed of the subject can be estimated based on the interval between continuous shooting of the plurality of still images SI and the amount of movement during that interval. This allows the processor 23 to accurately reflect the moving speed of the subject in the moving image MV.
[0079] Furthermore, multiple still images SI captured in continuous shooting may not only be used to estimate speed, but also may be input to a moving image generation model ML to output a moving image MV. In other words, multiple still images SI may be used as the basis for a moving image MV that generates the multiple still images SI.
[0080] In the above embodiment, an example was shown in which the video generation model ML is used from generation condition acquisition to video generation, but instead of processing everything using a machine learning model, for example, a machine learning model may be used for image analysis and other processing may be performed using rule-based processing. For example, the determination of variables to be used in the video MV and speed control may be processed using table data in the form conceptually shown in Figures 5, 6, and 10. Furthermore, a rule-based method such as pattern matching may be used for part of the image analysis.
[0081] Furthermore, in the above embodiment, the video production device 10 has been described as being configured as a stand-alone personal computer, but the video production device 10 may also be configured as a server that provides services to multiple client terminals via a communication network. The video production device 10 may also be configured as a mobile terminal such as a smartphone or tablet terminal. By configuring the video production device 10 as a mobile terminal such as a smartphone, users can easily enjoy videos MV generated based on captured still images SI. The functions of the video production device 10 may also be incorporated into a digital camera.
[0082] The above description can be used to understand the technologies described in the following supplementary items. [Supplementary Item 1] A moving image generation device including a processor, wherein when generating a moving image based on a still image, the processor acquires information about variable elements of the still image, and controls the speed at which the variable elements change in the moving image based on the information about the variable elements. [Supplementary Item 2] The moving image generation device according to Supplementary Item 1, in which the variable elements are elements that can be changed in the moving image. [Supplementary Item 3] The moving image generation device according to Supplementary Item 2, in which the variable elements include at least one of a subject that can be given movement in the moving image, a background that can be changed in the moving image, and camerawork that can be changed in the moving image. [Supplementary Item 4] The moving image generation device according to any one of Supplementary Item 1 to Supplementary Item 3, in which the information about the variable elements includes information acquired by image analysis of the still image. [Supplementary Item 5] The moving image generation device according to Supplementary Item 4, in which the information about the variable elements includes at least one of an attribute and a state of the subject recognized from the still image. [Supplementary Item 6] The moving image generation device according to Supplementary Item 5, wherein the processor further uses shooting conditions of the still image for speed control. [Supplementary Item 7] The moving image generation device according to Supplementary Item 4, wherein, when the still image is included in one of a plurality of still images shot in continuous shooting, the processor further uses the plurality of still images shot in continuous shooting for speed control. [Supplementary Item 8] The moving image generation device according to any one of Supplementary Item 1 to Supplementary Item 7, wherein the information on variable elements includes information acquired through user operation. [Supplementary Item 9] The moving image generation device according to Supplementary Item 8, wherein the information on variable elements includes information on camerawork that can be changed in the moving image. [Supplementary Item 10] The moving image generation device according to any one of Supplementary Item 1 to Supplementary Item 9, wherein the processor is able to accept a user's request for the presentation of the moving image. [Supplementary Item 11] The moving image generation device according to Supplementary Item 1, wherein the processor outputs the generated moving image in a manner that allows a preview to be displayed, and when a user request for the moving image is input, regenerates the moving image in accordance with the input request.[Supplementary Item 12] A method for operating a moving image generation device having a processor, wherein the processor, when generating a moving image based on a still image, acquires information about variable elements of the still image, and controls the speed at which the variable elements change in the moving image based on the information about the variable elements. [Supplementary Item 13] An operating program for a moving image generation device having a processor, wherein, when generating a moving image based on a still image, the program causes the processor to execute processes including acquiring information about variable elements of the still image, and controlling the speed at which the variable elements change in the moving image based on the information about the variable elements.
[0083] In the above embodiment, the video generation process is executed by an arbitrary computer. Furthermore, the arbitrary computer may execute these processes by a processor as hardware, a program as software, or a combination thereof. In this case, the processor is configured to execute various processes in this embodiment in cooperation with the program, and may function as each unit or each means in this embodiment. Furthermore, the order in which the processes are executed by the processor is not limited to the order described above and may be changed as appropriate.
[0084] The given computer may be a general-purpose computer, a computer for specific applications, a workstation, or any other system capable of executing each process. The processor may be configured with one or more pieces of hardware, and the type of hardware is not limited. For example, the processor may be configured with hardware such as a central processing unit (CPU), a micro processing unit (MPU), a programmable logic device such as a field programmable gate array (FPGA), a dedicated circuit for executing specific processes such as an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or a neural processing unit (NPU). The type of hardware may also be a combination of different types of hardware. When multiple pieces of hardware are configured to execute one or more processes of a certain processor, the multiple pieces of hardware may be located in devices physically separated from each other, or may be located in the same device. Furthermore, in any embodiment, the order of the processes performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is configured by an electric circuit (circuitry) or the like that combines circuit elements such as semiconductor elements.
[0085] Furthermore, the program may be software, such as firmware or microcode. The program may also be, for example, a group of program modules, each function of which may be implemented by a processor configured to perform the respective function. The program may be program code and / or multiple code segments stored in one or more non-transitory computer-readable media (e.g., storage media and / or other storages). The program may be stored across multiple non-transitory computer-readable media that reside in physically separate devices. The program code or code segment may represent a procedure, a function, a subprogram, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. The program code or code segment may be connected to another code segment or a hardware circuit by sending or receiving information, data, arguments, parameters, or memory contents.
[0086] The technology of the present disclosure can also be appropriately combined with the various embodiments and / or various modified examples described above. Furthermore, the technology is not limited to the above embodiments, and various configurations can be adopted without departing from the spirit of the present disclosure. Furthermore, the technology of the present disclosure also covers, in addition to programs, storage media that non-temporarily store programs. The storage medium is, for example, a computer-readable non-temporary storage medium such as a USB (Universal Serial Bus) memory, a flexible disk, or a CD-ROM (Compact Disc Read Only Memory). The program may also be provided online via a network such as the Internet. The technology of the present disclosure also covers, in addition to programs, program products. A program product includes any type of product for providing a program. Like a program, a program product may be provided stored on a computer-readable non-temporary storage medium or provided online.
[0087] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0088] In this specification, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."
[0089] The disclosure of Japanese Patent Application No. 2024-150355, filed on August 30, 2024, is incorporated herein by reference in its entirety. In addition, all documents, patent applications, and technical standards described herein are incorporated herein by reference to the same extent as if each individual document, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.
Claims
1. A video generation device comprising a processor, which, when generating a video based on a still image, acquires information regarding variable elements of the still image, and controls the speed at which the variable elements change in the video based on the information regarding the variable elements.
2. The moving image generating device according to claim 1, wherein the variable element is an element that can be changed in the moving image.
3. The video generation device of claim 2, wherein the variable elements include at least one of a subject that can be given movement in the video, a background that can be changed in the video, and camera work that can be changed in the video.
4. The video generation device according to claim 1, wherein the information relating to the variable elements includes information obtained by performing image analysis on the still images.
5. The video generation device according to claim 4, wherein the information relating to the variable elements includes at least one of the attributes and state of the subject recognized from the still image.
6. The video generating device according to claim 5, wherein the processor further uses the shooting conditions of the still images to control the speed.
7. The video generation device according to claim 4, wherein, when the still image is included in one of a plurality of still images taken in continuous shooting, the processor further uses the plurality of still images taken in continuous shooting to control the speed.
8. The video generation device according to claim 1, wherein the information relating to the variable elements includes information acquired through user operations.
9. The video generation device according to claim 8, wherein the information about the variable elements includes information about camera work that can be changed in the video.
10. The video generating device according to claim 1, wherein the processor is capable of accepting a user's wishes regarding the presentation of the video.
11. The video generating device according to claim 1, wherein the processor outputs the generated video in a manner that allows it to be previewed, and when a user request for the video is input, regenerates the video in accordance with the input request.
12. A method for operating a moving image generating device equipped with a processor, wherein the processor, when generating a moving image based on a still image, acquires information regarding variable elements of the still image, and controls the speed at which the variable elements change in the moving image based on the information regarding the variable elements.
13. An operating program for a video generation device equipped with a processor, which, when generating a video based on a still image, causes the processor to execute processing including: acquiring information about variable elements of the still image; and controlling the speed at which the variable elements change in the video based on the information about the variable elements.
Citation Information
Patent Citations
Method of presenting image
JP2000259144A
Imaging apparatus
JP2009284411A
Image creation device, image creation method, and program
JP2019204476A
Method and system for generating an animation from a static image
US20230058036A1