Video generation method and related equipment

By generating a second image containing more background information and generating a video of the mirror based on it, the problem of lack of video effects in the prior art is solved, and a higher interest and user experience is achieved.

CN120050480APending Publication Date: 2025-05-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510191957.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The video effects generated by the prior art have certain shortcomings and have not fully met the user's needs for fun.

Method used

By acquiring the first image, a second image containing more background information is generated, and a mirror video is generated based on the second image. The first frame image is enlarged, and the image size of the tail frame image is basically the same as the second image, so that the mirror effect with a zoom effect is achieved.

Benefits of technology

It improves the fun and user experience of the video, making the video generation richer and more vivid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050480A_ABST
    Figure CN120050480A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and related equipment. The method comprises the steps that a first image is acquired, and the first image comprises a target object and first background information; a second image is generated based on the first image, the second image comprises the target object and second background information, and the area of the second background information is larger than that of the first background information; a mirror moving video is generated based on the second image, the mirror moving video comprises multiple video frames, and the first frame of image in the multiple video frames is obtained after the second image is amplified by a preset multiple; the difference between the image size of the tail-frame image in the plurality of video frames and the image size of the second image is within a preset difference range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a video generation method and related devices. Background Art

[0002] With the development of computer technologies, the demand for interestingness in the video field is increasing day by day. Users hope to use video software to generate various interesting videos.

[0003] However, the inventors of the present disclosure have found that there are still certain deficiencies in the effects of videos generated by related technologies. Summary of the Invention

[0004] The present disclosure provides a video generation method and related devices to solve or partially solve the above problems.

[0005] In a first aspect of the present disclosure, there is provided a video generation method, including:

[0006] Obtaining a first image, where the first image includes a target object and first background information;

[0007] Generating a second image based on the first image, where the second image includes the target object and second background information, and the area of the second background information is larger than the area of the first background information;

[0008] Generating a panning video based on the second image, where the panning video includes multiple video frames, the first frame image in the multiple video frames is obtained by magnifying the second image by a preset multiple, and the difference between the image size of the last frame image in the multiple video frames and the image size of the second image is within a preset difference range.

[0009] In a second aspect of the present disclosure, there is provided a video generation device, including:

[0010] An obtaining module, configured to: obtain a first image, where the first image includes a target object and first background information;

[0011] A first generation module, configured to: generate a second image based on the first image, where the second image includes the target object and second background information, and the area of the second background information is larger than the area of the first background information;

[0012] A second generation module, configured to: generate a panning video based on the second image, where the panning video includes multiple video frames, the first frame image in the multiple video frames is obtained by magnifying the second image by a preset multiple, and the difference between the image size of the last frame image in the multiple video frames and the image size of the second image is within a preset difference range.

[0013] In a third aspect of the present disclosure, a computer device is provided, including one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the programs include instructions for executing the method according to the first aspect.

[0014] In a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium containing a computer program is provided. When the computer program is executed by one or more processors, the processors are caused to execute the method according to the first aspect.

[0015] In a fifth aspect of the present disclosure, a computer program product is provided, including a computer program, characterized in that when the computer program is executed by a processor, the steps of the method according to the first aspect are implemented.

[0016] The video generation method and related devices provided by the embodiments of the present disclosure can generate a second image containing more background information based on the first image, and then generate a panning video based on the second image. The first frame image of the panning video is enlarged, and the image size of the last frame image is substantially the same as that of the second image, so that a video with a zoom effect and a panning effect can be realized, improving the interestingness and user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only the embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 The schematic diagram of an exemplary system provided by the embodiments of the present disclosure is shown.

[0019] Figure 2A The flowchart of an exemplary method provided by the embodiments of the present disclosure is shown.

[0020] Figure 2B The flowchart of an exemplary method for generating a panning video according to the embodiments of the present disclosure is shown.

[0021] Figure 3A The schematic diagram of an exemplary first image according to the embodiments of the present disclosure is shown.

[0022] Figure 3B The schematic diagram of an exemplary second image according to the embodiments of the present disclosure is shown.

[0023] Figure 3C The schematic diagram of a video frame according to the embodiments of the present disclosure is shown.

[0024] Figure 3D A schematic diagram showing another video frame according to an embodiment of the present disclosure.

[0025] Figure 3E A schematic diagram showing a video frame after magnifying a second image by a preset multiple according to an embodiment of the present disclosure.

[0026] Figure 4 A schematic diagram showing a variable speed curve according to an embodiment of the present disclosure.

[0027] Figure 5 A schematic diagram showing the hardware structure of an exemplary computer device provided by an embodiment of the present disclosure.

[0028] Figure 6 A schematic diagram showing an exemplary device provided by an embodiment of the present disclosure. Detailed implementation manners

[0029] To make the objectives, technical solutions, and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0030] It should be noted that unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be the ordinary meanings understood by those of ordinary skill in the field to which the present disclosure belongs. The "first", "second", and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0031] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.

[0032] For example, when responding to an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure's technical solution based on the prompt message.

[0033] As an optional but non-limiting implementation manner, when responding to an active request from a user, the manner of sending a prompt message to the user can be, for example, in the form of a pop-up window. The prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0034] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0035] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 provided by an embodiment of the present disclosure.

[0036] As Figure 1 shown, the system 100 may include a terminal device 102, a server 106, and a database server 108. A medium (e.g., a network) for providing a communication link may be included between the terminal device 102 and the server 106 and the database server 108. The network may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0037] Various application programs (APPs) or software may be installed on the terminal device 102. For example, video generation application programs or software, video editing application programs or software, collaborative office application programs or software, image processing application programs or software, video conferencing application programs or software, reading application programs or software, video application programs or software, social application programs or software, payment application programs or software, web browsers, and instant messaging tools, etc. In some embodiments, these application programs or software can all be used to generate videos or edit videos.

[0038] The terminal device 102 here can be either hardware or software. When the terminal device 102 is hardware, it can be various electronic devices with a display screen, including but not limited to smartphones, tablet computers, e-book readers, MP3 players, laptop computers, and desktop computers, etc. When the terminal device 102 is software, it can be installed in the above-listed electronic devices. It can be implemented as multiple software or software modules (for example, to provide distributed services), or it can be implemented as a single software or software module. Specific limitations are not made here.

[0039] The server 106 can be a server that provides various services, such as a background server that supports various application programs or software displayed on the terminal device 102. The database server 108 can also be a database server that provides various services. It can be understood that when the relevant functions of the database server 108 can be implemented in the server 106, the database server 108 may not be set in the system 100.

[0040] The server 106 and the database server 108 here can also be either hardware or software. When they are hardware, they can be implemented as a distributed server cluster composed of multiple servers, or they can be implemented as a single server. When they are software, they can be implemented as multiple software or software modules (for example, to provide distributed services), or they can be implemented as a single software or software module. Specific limitations are not made here.

[0041] It should be noted that the method for editing a document provided by the embodiments of the present disclosure can be executed by the server 106. It should be understood that Figure 1 the numbers of the terminal devices, users, servers, and database servers in

[0042] In one embodiment, a video generation or video editing application or software can be installed in the terminal device 102, and the user 104 can use the application or software installed in the terminal device 102 to generate a video or edit a video.

[0043] As mentioned above, the inventors of the present disclosure found that there are still certain deficiencies in the effects of the videos generated by the related technologies.

[0044] In view of this, the embodiments of the present disclosure provide a video generation method to solve or partially solve the above problems.

[0045] Figure 2AThe flowchart of an exemplary method 200 provided by an embodiment of the present disclosure is shown. The method 200 can be used to generate a video and / or edit a video. Optionally, the method 200 can be implemented by the Figure 1 terminal device 102 or the server 106, or can be implemented by the Figure 1 system 100.

[0046] As Figure 2A shown, the method 200 may include the following steps.

[0047] In step 202, a first image 1022 is obtained.

[0048] There are many ways to obtain the first image 1022. For example, the user 104 can select an image locally on the terminal device 102 as the first image 1022, or the user 104 can receive an image from elsewhere using the terminal device 102 as the first image 1022, or the user 104 can download an image using the terminal device 102 as the first image 1022, and so on.

[0049] In some embodiments, the user 104 can upload the first image 1022 in a video generation or video editing application or software for subsequent processing, so that the application or software can be used to generate a video or edit a video, and so on. In some embodiments, when the server 106 executes the method 200, the terminal device 102 can also upload the first image 1022 to the server 106 so that it can obtain the first image 1022.

[0050] In some embodiments, the first image 1022 may include a target object and first background information. The target object can be any entity object, such as a person, an animal, and so on. Exemplarily, the target object can be an object (or “target”) that can be recognized based on an object detection algorithm, so that the target object can be recognized from the first image 1022 based on the object detection algorithm to distinguish it from the background information (first background information) other than the target object in the first image 1022.

[0051] Figure 3A The schematic diagram of an exemplary first image 1022 according to an embodiment of the present disclosure is shown.

[0052] Exemplarily, as Figure 3A shown, the first image 1022 may include a person, and the person can be recognized as the target object. The part other than the target object in the first image 1022 can be the first background information.

[0053] In step 204, a second image is generated based on the first image.

[0054] After acquiring the first image, illustratively, in this step, image expansion may be performed based on the first image to obtain a second image for generating a video in a subsequent step.

[0055] In some embodiments, the second image includes the target object and second background information, and the area of ​​the second background information is larger than the area of ​​the first background information, so that the second image is an image with a larger space relative to the first image 1022, and the second image can have richer content relative to the first image 1022.

[0056] Figure 3B A schematic diagram of an exemplary second image 300 according to an embodiment of the present disclosure is shown.

[0057] For example, Figure 3B As shown, the second image 300 can be generated based on the first image 1022. The second image 300 includes the target object (e.g., a person), and the size of the target object is smaller in the second image 300 than in the first image 1022. Meanwhile, the second image 300 also includes more background information than in the first image 1022, so that a larger background space and content appear in the second image 300.

[0058] Optionally, the second image 300 may include background information generated based on the first image 1022 so that the second image may include background information not included in the first image, thereby enriching the content of the second image.

[0059] Optionally, the amount of second background information of the second image 300 is greater than that of the first background information of the first image 1022, and the target object in the second image 300 is smaller than that in the first image, thereby achieving a larger background space and content in the second image.

[0060] In some embodiments, generating the second image based on the first image may be to first determine the target object in the first image, and then generate the second image based on the target object, so that the second image includes the target object and the second background information. In this way, under the premise of ensuring that the second image contains the target object, the background information in the second image can be added, so that the subsequently generated video is a video containing the target object, thereby improving the fun and user experience.

[0061] In some embodiments, reference Figure 3A and Figure 3BWhen generating the second image 300 based on the first image 1022, detailed information may be added to the target object (eg, expanding the torso, extending the arms, etc.), thereby making the expanded second image 300 more realistic.

[0062] In some embodiments, a pre-trained image generation model may be used, the first image 1022 may be input into the model, and then the second image 300 may be output. Optionally, the image generation model may be an artificial intelligence model based on machine learning, and the image generation model may be pre-trained using multiple original images, so that multiple desired second images may be generated based on the first image.

[0063] In step 206, a camera movement video is generated based on the second image.

[0064] Optionally, the camera movement video includes multiple video frames, a first frame image in the multiple video frames is obtained by magnifying the second image by a preset multiple, and a difference in image size between the last frame image in the multiple video frames and the second image is within a preset difference range.

[0065] In this way, a camera movement video is generated based on the second image, the first frame image of the camera movement video is enlarged, and the image size of the last frame image is basically the same as the second image, thereby realizing a video with a zoom effect and a camera movement effect, thereby improving fun and user experience.

[0066] like Figure 3E As shown, the first frame image can be obtained by enlarging the second image by a preset multiple. In this way, when multiple video frames are rendered, the first frame image can present an effect of a larger target object and a smaller background area after being drawn on the canvas 320, thereby obtaining the desired camera video effect.

[0067] In some embodiments, the image size of the last frame image is the same as the image size of the second image and / or the last frame image is the second image with the original image size, so as to obtain a more realistic camera movement effect.

[0068] In some embodiments, the multiple video frames all include the target object (e.g., a person), and the target object in the multiple video frames includes multiple orientations. Optionally, taking the target object as a person as an example, the orientation may refer to the facial orientation of the target object, or may refer to the body orientation of the target object, or may refer to the common orientation of the face and body of the target object. The orientation of the front of the person in each video frame may be different, so that the video generated based on the multiple video frames of the target object with multiple orientations may have a camera movement effect.

[0069] Figure 3C and Figure 3DSchematic diagrams of video frames 310A and 310B according to embodiments of the present disclosure are respectively shown. Figure 3E A schematic diagram of a video frame after magnifying a second image by a preset multiple according to an embodiment of the present disclosure is shown.

[0070] As Figure 3E shown, the orientation of the target object in this video frame is close to 0°. As Figure 3C shown, the orientation of the target object in the video frame is about 90°. As Figure 3D shown, the orientation close to the target object in the video frame is about 180°. In summary, the orientations of the target objects in multiple video frames can be different from each other, and the angles can change gradually. In this way, by changing the orientation of the target object, the 3D stereoscopic effect of the target object can be enhanced in the subsequent generated video, improving the video effect.

[0071] In some embodiments, the area sizes of the background information of the multiple video frames are different from each other. Exemplarily, as Figure 3C and Figure 3D shown, the area of the background information of video frame 310B is larger than that of video frame 310A, such that when changing from Figure 3C to Figure 3D , a visual effect of gradually increasing space can be presented. It can be understood that the areas of the background information of each video frame in multiple video frames can be different from each other, so that a video effect of gradually changing background space and content can be formed when generating a video. Correspondingly, when the background information gradually increases, the sizes of the target objects in the multiple video frames can be gradually reduced and different from each other.

[0072] In some embodiments, the angular range of the orientation can be 0 to 360°, so that a panning video can obtain a 360° panning effect, increasing the interest of the panning video. It can be understood that according to different requirements, the user can also obtain other panning effects by inputting an orientation range. For example, a 180° panning effect.

[0073] In some embodiments, the angle of the orientation has a negative correlation with the area size of the background information of the multiple video frames.

[0074] For example, as Figure 3C , 3D shown, when the orientation increases from 90° of video frame 310A to 180° of video frame 310B, the area of the background information also increases, thus realizing a negative correlation between the angle of the orientation and the area size of the background information. In this way, when the orientation angle changes, the background information also changes, so that a better panning effect can be achieved.

[0075] It can be understood that if the angle with the target object facing forward is 0°, then a change in the orientation to the left or right can both be an increase in the angle. At this time, the area of the background information can gradually increase (for example, from Figure 3C changing to Figure 3D ), so that the angle of the orientation is negatively correlated with the area size of the background information, and a better camera movement effect can be achieved.

[0076] In some embodiments, the size of the target object in the multiple video frames is negatively correlated with the area size of the background information of the multiple video frames. As shown in Figure 3C , 3D , when the target object gradually becomes smaller, the amount of background information gradually increases, thus achieving the effect of gradually increasing the background space and content, and further better achieving the camera movement effect.

[0077] In a more specific embodiment, the size, orientation angle of the target object and the amount of background information can change together, so as to achieve a more immersive 3D camera movement effect.

[0078] In some embodiments, generating a camera movement video based on the second image includes: arranging the multiple video frames in sequence based on the size of the target object, the angle of the orientation, and / or the area size of the background information of the multiple video frames to generate the camera movement video.

[0079] For example, if the sizes of the target objects in each video frame are different, the multiple video frames can be arranged according to the size of the target object to obtain an arrangement order. Correspondingly, the size of the target object in the generated camera movement video changes gradually, thus presenting an effect that the target object gradually moves away (the target object changes from large to small).

[0080] Also, for example, if the orientation angles of the target objects in each video frame are different, the multiple video frames can be arranged according to the size of the orientation angle to obtain an arrangement order. Correspondingly, the orientation of the target object in the generated camera movement video changes gradually, thus presenting an effect of rotation of the target object.

[0081] Furthermore, for example, if the area sizes of the background information in each video frame are different, the multiple video frames can be arranged according to the area size of the background information to obtain an arrangement order. Correspondingly, the area size of the background information in the generated camera movement video changes gradually, thus presenting an effect of gradually increasing the background content.

[0082] It can be understood that the above three parameters can be arranged and combined to implement the above arrangement operation, so that the generated first camera movement video can have various effects.

[0083] In a more specific embodiment, when the size and orientation of the target object in the multi-frame video frames gradually change and the area size of the background information gradually changes, the video generated based on the multi-frame video frames can obtain a better camera movement effect, enabling the viewer to have an immersive feeling.

[0084] In some embodiments, a pre-trained video generation model can be utilized to input the second image into the video generation model, and the camera movement video is output. Optionally, the video generation model can be an artificial intelligence model based on machine learning, and the video generation model is pre-trained based on the original camera movement video and the first-frame image of the multi-frame video frames of the original camera movement video.

[0085] Optionally, the original camera movement video can be a camera movement video generated based on the original image. The first-frame image of the multi-frame video frames of the original camera movement video is obtained by magnifying the original image by a preset multiple (for example, 1.5 times, 2 times, 3 times, etc.), and the difference between the image size of the last-frame image in the multi-frame video frames of the original camera movement video and the image size of the original image is within a preset difference range, for example, the image sizes are the same or not much different.

[0086] Next, the original camera movement video and its first-frame image are paired, and then the paired data is used as training data to train the initial model, thereby obtaining the video generation model. In this way, when the video generation model is input with the second image (the magnified image), the corresponding camera movement video can be generated, and the difference between the image size of its last-frame image and the image size of the second image can be within a preset difference range.

[0087] In some embodiments, generating the camera movement video based on the second image may further include: generating the second camera movement video based on the multi-frame video frames of the camera movement video in the reverse order of the arrangement order of the multi-frame video frames of the camera movement video, and then generating the target camera movement video based on the camera movement video and the second camera movement video. In this way, the second camera movement video can be the reverse-play video of the camera movement video, so that the effect of the target camera movement video generated based on the camera movement video and the second camera movement video can be more abundant.

[0088] In this step, generating the target camera movement video based on the camera movement video and the second camera movement video can obtain a better camera movement effect. Optionally, the camera movement video and the second camera movement video can be directly spliced to obtain the target camera movement video, which is easier to process.

[0089] In some embodiments, as Figure 2B shown, generating the camera movement video based on the second image may further include the following steps:

[0090] In step 212, input the second image into the video generation model to output an intermediate video.

[0091] In this step, the video generation model can generate an intermediate video with a camera movement effect based on the second image, and then the intermediate video can be processed by video editing.

[0092] Next, in step 214, video parameters can be obtained.

[0093] In this step, the video effect of the final camera movement video can be enhanced by setting video parameters for the intermediate video.

[0094] Optionally, the video parameters can be setting the trajectory of the video, the speed change curve, the zoom in or out of key frames, and so on.

[0095] In step 216, process the intermediate video based on the video parameters to obtain the camera movement video.

[0096] In this step, further processing the intermediate video based on the video parameters can add more effects to the camera movement video to match the camera movement, enhancing the interest and user experience.

[0097] In some embodiments, the video parameters include a speed change curve; processing the intermediate video based on the video parameters to obtain the camera movement video includes: performing speed change processing on the intermediate video based on the speed change curve to obtain the camera movement video.

[0098] Figure 4 FIG. shows a schematic diagram of a speed change curve 400 according to an embodiment of the present disclosure.

[0099] As Figure 4 shown, exemplarily, user 104 can use a video generation or video editing application or software of terminal device 102 to edit the speed change curve, so that the camera movement video can be speed-changed according to the speed change curve. Exemplarily, as Figure 4 shown, when the speed change curve is raised, the video frames in the corresponding time period of the curve part will be played at an accelerated speed. It can be understood that when the speed change curve is lowered, the video frames in the corresponding time period of the curve part will be played at a decelerated speed.

[0100] In this way, by setting the speed change curve, the camera movement video can have a speed change effect, thereby enhancing the video interest and user experience.

[0101] In some embodiments, the video parameters include the movement trajectory of the target object; processing the intermediate video based on the video parameters to obtain the camera movement video includes: moving the position of the target object in multiple video frames of the intermediate video based on the movement trajectory to obtain the camera movement video.

[0102] Exemplarily, user 104 can use a video generation or video editing application or software of the terminal device 102 to edit the movement trajectory of the target object, so that the target object in the camera movement video can move according to the movement trajectory. Therefore, based on the setting of the movement trajectory, the target object in the camera movement video can move according to a predetermined trajectory, increasing the controllability of the target object and the interest of the video at the same time.

[0103] As can be seen from the above embodiments, the video generation method provided by the embodiments of the present disclosure can generate a second image containing more background information based on the first image, and then generate a camera movement video based on the second image. The first frame image of the camera movement video is enlarged, and the image size of the last frame image is basically the same as that of the second image. Thus, a video with a zoom effect and a camera movement effect can be realized, enhancing the interest and user experience. By using this video generation method, any photo can be made cool, enhancing the interest of video generation.

[0104] In some embodiments, the size and orientation of the target object in multiple second images gradually change, and the area size of the background information gradually changes. Thus, in the case of inputting the first picture, a video with multiple angles and a large space can be obtained with one key. Moreover, the video can have a better 3D camera movement effect, enhancing the user experience.

[0105] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.

[0106] It should be noted that some embodiments of the present disclosure are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from that in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0107] The embodiments of the present disclosure also provide a computer device for implementing the above method 200. Figure 5 The hardware structure diagram of an exemplary computer device 500 provided by the embodiments of the present disclosure is shown. The computer device 500 can be used to implement Figure 1 the server 106, and can also be used to implement Figure 1The terminal device 102. In some scenarios, this computer device 500 can also be used to implement Figure 1 the database server 108.

[0108] Such as Figure 5 As shown, the computer device 500 may include: a processor 502, a memory 504, a network module 506, a peripheral interface 508, and a bus 510. Among them, the processor 502, the memory 504, the network module 506, and the peripheral interface 508 are communicatively connected to each other inside the computer device 500 through the bus 510.

[0109] The processor 502 may be a central processing unit (CPU), a graphics processor, a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. The processor 502 can be used to execute functions related to the technologies described in this disclosure. In some embodiments, the processor 502 may further include multiple processors integrated as a single logic component. For example, as Figure 5 shown, the processor 502 may include multiple processors 502a, 502b, and 502c.

[0110] The memory 504 can be configured to store data (e.g., instructions, computer code, etc.). As Figure 5 shown, the data stored in the memory 504 may include program instructions (e.g., program instructions for implementing the method 200 of the embodiments of this disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). The processor 502 can also access the program instructions and data stored in the memory 504, and execute the program instructions to operate on the data to be processed. The memory 504 may include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 504 may include a random access memory (RAM), a read-only memory (ROM), an optical disc, a magnetic disk, a hard disk, a solid state drive (SSD), a flash memory, a memory stick, etc.

[0111] The network interface 506 can be configured to provide communication with other external devices to the computer device 500 via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC), etc.), a cellular network, the Internet, or a combination of the above. It can be understood that the type of the network is not limited to the above specific examples.

[0112] The peripheral interface 508 can be configured to connect the computer device 500 to one or more peripheral devices to enable information input and output. For example, the peripheral devices may include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, various sensors, etc., and output devices such as a display, a speaker, a vibrator, an indicator light, etc.

[0113] The bus 510 can be configured to transfer information between various components of the computer device 500 (such as the processor 502, the memory 504, the network interface 506, and the peripheral interface 508), such as an internal bus (e.g., the processor - memory bus), an external bus (USB port, PCI - E bus), etc.

[0114] It should be noted that although the architecture of the computer device 500 shown above only shows the processor 502, the memory 504, the network interface 506, the peripheral interface 508, and the bus 510, in the specific implementation process, the architecture of the computer device 500 may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the architecture of the computer device 500 above may also only include the components necessary to implement the solution of the embodiments of the present disclosure, and do not necessarily include all the components shown in the figure.

[0115] The embodiments of the present disclosure also provide a video generation device. Figure 6 The schematic diagram of the exemplary device 600 provided by the embodiments of the present disclosure is shown. As Figure 6 shown, the device 600 can be used to implement the method 200 and may further include the following modules.

[0116] An acquisition module 602, configured to: acquire a first image, where the first image includes a target object and first background information;

[0117] A first generation module 604, configured to: generate a second image based on the first image, where the second image includes the target object and second background information, and the area of the second background information is larger than the area of the first background information;

[0118] A second generation module 606, configured to: generate a panning video based on the second image, where the panning video includes multiple video frames, the first frame image in the multiple video frames is obtained by magnifying the second image by a preset multiple, and the difference between the image size of the last frame image in the multiple video frames and the image size of the second image is within a preset difference range.

[0119] In some embodiments, the first generation module 604 is configured to:

[0120] Determine the target object in the first image;

[0121] Generate the second image based on the target object, where the second image includes the target object and the second background information.

[0122] In some embodiments, each of the multiple video frames includes the target object, and the target object in the multiple video frames has multiple orientations;

[0123] The area sizes of the background information of the multiple video frames are different from each other; and / or,

[0124] The sizes of the target object in the multiple video frames are different from each other.

[0125] In some embodiments, the angle of the orientation is negatively correlated with the area size of the background information of the multiple video frames, and / or, the size of the target object in the multiple video frames is negatively correlated with the area size of the background information of the multiple video frames.

[0126] In some embodiments, the second generation module 606 is configured to:

[0127] Arrange the multiple video frames in sequence based on the size of the target object, the angle of the orientation, and / or the area size of the background information of the multiple video frames to generate the panning video.

[0128] In some embodiments, the angle range of the orientation is 0 to 360°, and / or, the image size of the tail frame image is the same as the image size of the second image, and / or, the tail frame image is the second image with the original image size.

[0129] In some embodiments, the second generation module 606 is configured to:

[0130] Input the second image into a video generation model to output the panning video;

[0131] Wherein, the video generation model is pre-trained based on the original panning video and the first frame image of the multiple video frames of the original panning video.

[0132] In some embodiments, the third generation module 608 is configured to:

[0133] Input the second image into a video generation model to output an intermediate video;

[0134] Obtain video parameters;

[0135] Process the intermediate video based on the video parameters to obtain the panning video.

[0136] In some embodiments, the video parameters include a variable speed curve;

[0137] The third generation module 608 is configured to: perform variable speed processing on the intermediate video based on the variable speed curve to obtain the panning video.

[0138] In some embodiments, the video parameters include the movement trajectory of the target object;

[0139] The third generation module 608 is configured to: move the position of the target object in multiple video frames of the intermediate video based on the movement trajectory to obtain the panning video.

[0140] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0141] The device in the above embodiments is used to implement the corresponding method 200 in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0142] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the method 200 in any of the foregoing embodiments.

[0143] The computer-readable medium in this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device.

[0144] The computer instructions stored in the storage medium in the above embodiments are used to cause the computer to execute the method 200 in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated herein.

[0145] Based on the same inventive concept, corresponding to the method 200 in any of the above embodiments, the present disclosure also provides a computer program product, which includes a computer program. In some embodiments, the computer program is executable by one or more processors to cause the processors to execute the method 200. Corresponding to the execution subjects of the respective steps in the embodiments of the method 200, the processors executing the corresponding steps may belong to the corresponding execution subjects.

[0146] In some embodiments, the computer program product includes a program module for resolving operation conflicts, and the program module can be compiled into a binary instruction set (e.g., wasm) based on a stack virtual machine and deployed in a terminal device and / or compiled into a static library and deployed in a server.

[0147] The computer program product of the above embodiments is used to cause a processor to execute the method 200 described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0148] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, and they are not provided in detail for the sake of brevity.

[0149] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order not to make the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (e.g., circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0150] Although the present disclosure has been described in connection with specific embodiments of the present disclosure, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0151] Embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A video generation method, comprising: Acquire a first image, wherein the first image includes a target object and first background information; generating a second image based on the first image, wherein the second image includes the target object and second background information, and an area of ​​the second background information is larger than an area of ​​the first background information; A camera movement video is generated based on the second image, the camera movement video includes multiple video frames, the first frame image in the multiple video frames is obtained by enlarging the second image by a preset multiple, and the image size of the last frame image in the multiple video frames is different from the image size of the second image within a preset difference range.

2. The method of claim 1, wherein: Generating a second image based on the first image includes: determining the target object in the first image; The second image is generated based on the target object.

3. The method of claim 1, wherein: The multiple video frames all include the target object, and the target object in the multiple video frames includes multiple directions; The area sizes of the background information of the multiple video frames are different; and / or, The sizes of the target objects in the multiple video frames are different.

4. The method of claim 3, wherein: The orientation angle is negatively correlated with the area size of the background information of the multiple video frames, and / or the size of the target object in the multiple video frames is negatively correlated with the area size of the background information of the multiple video frames.

5. The method of claim 4, wherein: Generating a camera movement video based on the second image includes: Based on the size of the target object, the angle of the orientation and / or the area size of the background information of the multiple video frames, the multiple video frames are arranged in sequence to generate the camera movement video.

6. The method of claim 3, wherein: The angle range of the orientation is 0 to 360 degrees, and / or the image size of the last frame image is the same as the image size of the second image, and / or the last frame image is the second image with an original image size.

7. The method of claim 1, wherein: Generating a camera movement video based on the second image includes: Inputting the second image into a video generation model, and outputting the camera movement video; The video generation model is pre-trained based on the original camera movement video and the first frame image of multiple video frames of the original camera movement video.

8. The method of claim 1, wherein: Generating a camera movement video based on the second image includes: Inputting the second image into a video generation model, and outputting an intermediate video; Get video parameters; The intermediate video is processed based on the video parameters to obtain the camera movement video.

9. The method of claim 8, wherein: The video parameters include a speed change curve; Processing the intermediate video based on the video parameters to obtain the camera movement video includes: The intermediate video is subjected to speed change processing based on the speed change curve to obtain the camera movement video.

10. The method of claim 8, wherein: The video parameters include the moving trajectory of the target object; Processing the intermediate video based on the video parameters to obtain the camera movement video includes: The position of the target object in the multiple video frames of the intermediate video is moved based on the movement trajectory to obtain the camera movement video.

11. A video generating device, comprising: An acquisition module is configured to: acquire a first image, wherein the first image includes a target object and first background information; A first generating module is configured to: generate a second image based on the first image, wherein the second image includes the target object and second background information, and an area of ​​the second background information is larger than an area of ​​the first background information; The second generating module is configured to generate a camera movement video based on the second image, wherein the camera movement video includes multiple video frames, wherein the first frame image in the multiple video frames is obtained by enlarging the second image by a preset multiple, and the image size of the last frame image in the multiple video frames differs from the image size of the second image within a preset difference range.

12. A computer device comprising one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the programs include instructions for executing the method according to any one of claims 1-10.

13. A non-volatile computer-readable storage medium containing a computer program, which, when executed by one or more processors, causes the processors to perform the method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.