Video generation method and electronic equipment

By generating lock screen videos, the problem that the digital image lock screen interface cannot achieve dynamic effects and individual elements changes are solved, dynamic effects and personalized settings are achieved, and user experience is enhanced.

CN120455755AActive Publication Date: 2025-08-08HONOR DEVICE CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202411251901.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-06
Publication Date
2025-08-08
Estimated Expiration
2044-09-06

AI Technical Summary

Technical Problem

The lock screen interface based on digital images cannot achieve dynamic effects and cannot achieve changes to individual elements in the lock screen interface, resulting in low playability and lack of personalized settings.

Method used

By obtaining video description information, use preset materials to select image animations of foreground image materials, background image materials and digital images, combine to generate target videos, realize the dynamic effect of digital images, and allow changes to individual elements.

Benefits of technology

It realizes the dynamic effects of digital images and personalized settings of the lock screen interface, enhancing playability and fun.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455755A_ABST
    Figure CN120455755A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of terminals, and provides a video generation method and electronic equipment. The video generation method comprises the following steps: acquiring video description information, wherein the video description information is used for describing a scene of a target video and / or an action of a digital image in the target video; based on the video description information, target scene materials and target animation materials are obtained in a preset material set, the target scene materials comprise scene image materials, the scene image materials comprise foreground image materials and / or background image materials, and the target animation materials comprise image animations of the digital images moving based on the motion data; and combining the target scene material and the target animation material to obtain a target video. According to the technical scheme, the problems that a screen locking interface based on a digital image cannot achieve a dynamic effect and cannot achieve change of a single element in the screen locking interface can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal technology, and in particular to a video generation method and electronic device. Background Art

[0002] With the development of science and technology, the popularity of electronic devices has increased, and the importance of electronic device security has also continued to increase. The lock screen is an important function that can improve the security of electronic devices and can be used to prevent unauthorized access. Specifically, when an electronic device is in the lock screen state, the lock screen interface is displayed on the screen of the electronic device. The lock screen interface includes a lock screen image and basic information such as the date and time. Users can verify permissions by entering a password, pattern, fingerprint, or facial recognition. Only if the permission verification is passed can the user access the content or functions in the electronic device; otherwise, the screen will always remain in the lock screen interface.

[0003] The lock screen interface can be a natural scenery lock screen interface, an abstract art lock screen interface, a digital image lock screen interface, etc. The digital image lock screen interface includes a cartoon or anime style digital image, which is usually colorful and cute, and can provide a pleasant visual experience.

[0004] However, avatar-based lock screens typically display basic information like the date and time over a static image containing the avatar. Static images result in fixed lock screen content, preventing dynamic effects and resulting in poor display quality. Furthermore, when editing elements within a cartoon-based lock screen, only the entire static image can be replaced; individual elements cannot be modified, hindering personalized lock screen settings. Summary of the Invention

[0005] The present application provides a video generation method and electronic device to solve the problem that a lock screen interface based on a digital image cannot achieve dynamic effects and cannot change a single element in the lock screen interface.

[0006] To achieve the above objectives, in a first aspect, an embodiment of the present application provides a video generation method, the method comprising:

[0007] Obtain video description information, which is used to describe the scene of the target video and / or the action of the digital image in the target video; based on the video description information, obtain target scene material and target animation material from a preset material set, wherein the first target scene material includes scene image material, the scene image material includes foreground image material and / or background image material, and the target animation material includes image animation of the digital image based on motion data; combine the target scene material and the target animation material to obtain the target video.

[0008] The video generation method shown in the present application first obtains the video description information used to describe the lock screen video, and then, based on the video description information, selects the foreground image material, the background image material and the image animation containing the digital image from the pre-set material set containing multiple materials, and combines them to generate the target video. Since the video generated by the present application uses the image animation of the digital image movement, it can achieve a dynamic effect based on the digital image. In addition, the present application realizes the decoupling and componentization of each material in the video, and obtains the target video by combining multiple materials. Therefore, when it is necessary to change an element in the lock screen interface, it is only necessary to change the material corresponding to the element. By customizing any material, it is beneficial to the personalized setting of the lock screen interface, and enhances playability and fun.

[0009] In one implementation, before obtaining target scene materials and target animation materials from a preset material set, the method further includes: constructing a material set, wherein the materials in the material set include at least one image animation and a scene image material. Using this implementation, multiple different materials can be generated in advance to construct a material set. When combining the materials to create a video, one only needs to select the pre-generated materials from the material set.

[0010] In one implementation, constructing a material collection includes obtaining at least one material and a corresponding material description tag, wherein the material description tag identifies the scene or action corresponding to the material; and constructing the material collection based on the correspondence between the material and the material description tag. In this implementation, the pre-constructed material collection includes not only the material but also the correspondence between the material and the description tag. Subsequently, the corresponding material can be found in the material collection based on the description tag, thereby ensuring a higher degree of match between the video generated from the material and the video description information.

[0011] In one implementation, obtaining at least one material and a material description tag corresponding to the material includes: generating at least one scene image material based on a preset scene tag, and using the scene tag as the material description tag corresponding to the scene image material, wherein the scene tag is used to identify the scene corresponding to the scene image material; generating at least one image animation based on a preset animation tag, and using the animation tag as the material description tag corresponding to the image animation, wherein the animation tag is used to identify the action or the scene in which the action occurs in the image animation. Using this implementation, the scene image material is generated using the scene tag, and the image animation is generated using the animation tag. Since the material is generated based on the tag, the correspondence between the tag and the material can be guaranteed. Therefore, based on the scene tag, the scene image material whose scene meets the requirements can be found; based on the animation tag, the image animation whose action or the scene in which the action occurs can be found, which can ensure that the scene and action of the combined target video meet the requirements of the video description information.

[0012] In one implementation, at least one scene image material is generated based on a preset scene label, including: inputting at least one scene description text corresponding to the scene label into a semantic model to obtain a scene description feature corresponding to the scene description text, wherein the scene description text is used to describe the scene identified by the scene label; and inputting the scene description feature into an image generation model to obtain a scene image material corresponding to the scene description feature. Using this implementation, foreground image material and / or background image material are generated using a semantic model and an image generation model. Through the above-mentioned generative artificial intelligence model, a large amount of foreground image material and / or background image material can be quickly obtained, which is conducive to improving the production efficiency of material resources, while also reducing the demand for designers and lowering costs.

[0013] In one implementation, before inputting at least one scene description text corresponding to a scene label into the semantic model, the method further includes: setting at least one scene label for each scene category; and determining at least one scene description text for each scene label. Using this implementation, multiple scene labels can be set for each scene category. Since scene labels are used to identify scenes, each scene category can correspond to multiple different scenes. The scene image material thus obtained can cover multiple different scenes and provide users with a wider range of choices. In addition, the present implementation determines multiple scene description texts for each scene label, which can describe the scene corresponding to the scene label more comprehensively and meticulously, thereby improving the degree of match between the generated foreground image material and / or background image material and the scene label. Furthermore, setting multiple scene description texts is also conducive to generating a larger number of foreground image materials and / or background image materials.

[0014] In one implementation, after generating at least one scene image material based on a preset scene label, the method further includes: if the scene image material does not match a preset screening rule, eliminating the scene image material, where the screening rule includes: the scene image material matches the scene label and / or the naturalness of the scene image material meets a preset condition. This implementation eliminates poorly rendered material after the scene image material is generated, thereby ensuring the quality of the generated target video.

[0015] In one implementation, at least one image animation is generated based on a preset animation tag, including: designing an action sequence based on at least one animation description text corresponding to the animation tag, wherein the animation description text is used to describe the action or action scenario identified by the animation tag; capturing motion data generated by the user during movement in accordance with the action sequence; redirecting the motion data to a digital image, and adjusting the redirected motion data to obtain an image animation. Using this implementation, the user utilizes motion capture technology to capture the motion data generated by the user's movement and redirects it to the digital image to obtain an image animation. In the image animation thus generated, the digital image's movements are more natural and realistic, which helps improve the viewing experience of video viewers. Furthermore, generating image animation using motion capture technology eliminates the need for frame-by-frame drawing and modeling, thereby improving the efficiency of image animation generation.

[0016] In one implementation, before designing an action sequence based on at least one animation description text corresponding to an animation tag, the method further includes: setting at least one animation tag corresponding to an action category, wherein the animation tag corresponding to the action category is used to identify the action in the avatar animation; and / or setting at least one animation tag corresponding to an action scene category, wherein the animation tag corresponding to the action scene category is used to identify the action scene corresponding to the avatar animation; and determining at least one animation description text for each animation tag. In this implementation, action categories and animation scene categories are designed for the avatar animation, and multiple animation tags are set for each action category and animation scene category. Since the animation tag for the action scene category is used to identify the action scene, the action scene category can correspond to multiple different action scenes. The resulting avatar animation can cover multiple different action scenes, providing users with a wider range of choices. Since the animation tag for the action category is used to identify the action of the digital avatar, the action category can correspond to multiple different actions. The resulting avatar animation can cover multiple different actions, providing users with a wider range of choices. Furthermore, this implementation method specifies multiple animation description texts for each animation tag, enabling a more comprehensive and detailed description of the action or scene corresponding to the animation tag, thereby improving the match between the generated image animation and the animation tag. Furthermore, setting multiple animation description texts also facilitates the generation of a larger number of image animations.

[0017] In one implementation, determining at least one animation description text corresponding to an animation tag includes: determining the character corresponding to the animation tag, where the character can be a human or an animal; if there are multiple characters, determining the interactive actions between the multiple characters, and determining the animation description text corresponding to the animation tag based on the interactive actions. In this implementation, the action description text can describe the interactive actions between the multiple characters. By designing the interactive actions, an animated image of multiple characters interacting can be generated, ultimately resulting in a target video of multiple digital characters interacting. Furthermore, the characters are not limited to human characters; they can also be animal characters, increasing the fun and playability.

[0018] In one implementation, after constructing a resource set, the method further includes adding resource information for each resource to the resource set. The resource information includes at least one of a resource description tag, resource description text, an index, a default resource identifier, extended information, and a resource version number. The resource description text for scene image resources is referred to as scene description text, while the resource description text for image animations is referred to as animation description text. This implementation adds detailed resource information to each resource, facilitating resource management. Furthermore, the index improves retrieval efficiency, and the resource version number enables version control and tracking. The provision of resource information facilitates better utilization of resource resources.

[0019] In one implementation, after constructing a material collection based on the correspondence between materials and material description tags, the method further includes: constructing at least one material subset based on the scene image materials and image animations in the material collection; and constructing a mapping relationship table based on the correspondence between each material subset and the material description tag. Using this implementation, multiple material subsets are constructed, and a mapping relationship is determined between each material subset and the material description tag. Therefore, based on the mapping relationship table, the corresponding material subset can be found using the material description tag, and the target video can be generated using the materials in the material subset.

[0020] In one implementation, a mapping relationship table is constructed based on the correspondence between each material subset and the material description tag, including: for each material subset, determining the correspondence between the material subset and the material description tag based on the correspondence between each material in the material subset and the material description tag; and constructing a mapping relationship table based on the correspondence between the material subset and the material description tag. Using this implementation, the correspondence between the material and the material description tag is converted into the correspondence between the material subset and the material description tag. Therefore, when searching for target scene materials and target animation materials, it is only necessary to determine the correspondence between the material subset and the material description tag, without having to determine the correspondence between each material and the material description tag separately. Therefore, this implementation can simplify the material search process, quickly find the required material, and apply it.

[0021] In one implementation, obtaining target scene and animation materials from a preset material collection includes: determining material description tags corresponding to video description information; and obtaining the target scene and animation materials corresponding to the material description tags from the material collection based on a pre-built mapping relationship table. This implementation method, through table lookup, is simple and quick to obtain the target scene and animation materials from the material collection, thereby improving the efficiency of generating the target video.

[0022] In one implementation, obtaining target scene materials and target animation materials corresponding to material description tags from a material set based on a mapping relationship table includes: determining, based on the mapping relationship table, a material subset corresponding to the material description tags as a target material subset; and determining target scene materials and target animation materials within the target material subset. This implementation ensures that the target scene materials and target animation materials are materials corresponding to the material description tags by determining the target scene materials and target animation materials within the target material subset corresponding to the material description tags.

[0023] In one implementation, determining target scene materials and target animation materials within a target material subset includes: obtaining a first target index and a second target index, wherein the first target index is used to distinguish different scene image materials corresponding to the same scene label, and the second target index is used to distinguish different image animations corresponding to the same animation label; and determining, within the material subset, the scene image material corresponding to the first target index as the target scene material, and the image animation corresponding to the second target index as the target animation material. This implementation improves the efficiency of determining target scene materials and target animation materials by using the first target index and the second target index to search within the target material subset.

[0024] In one implementation, a material collection includes a scene image material collection and an animation material collection. Retrieving target scene materials and target animation materials from the preset material collection includes: retrieving target scene materials from the scene image material collection, and retrieving target animation materials from the animation material collection. With this implementation, the scene image materials and the animation characters are stored in the material collection, respectively, facilitating systematic material management, facilitating rapid location of target scene materials and target animation materials, and facilitating archiving and backup of materials.

[0025] In one implementation, combining target scene material and target animation material to obtain a target video includes: sequentially overlaying the target scene material and target animation material in a preset order to obtain the target video; and rendering each frame of the target video based on lighting change information and / or camera trajectory information to obtain the rendered target video. This implementation method, sequentially overlaying the target scene material and target animation material, facilitates obtaining a video with a sense of depth, making the target video more three-dimensional and realistic. After obtaining the stereoscopic video, rendering each frame of the video can further improve video quality and enhance the user's visual experience.

[0026] In one implementation, after obtaining the target video, the method further includes: setting the target video as the lock screen video of the electronic device. Using this implementation, a lock screen video containing a dynamic digital image can be provided to the user, thereby improving the user experience.

[0027] In one implementation, before acquiring the target scene and animation materials from a preset material set, the method further includes: acquiring a digital image in response to a video generation instruction; and editing the digital image in response to an editing instruction. This implementation method first acquires a default digital image and then edits it based on the user's editing instructions, thereby increasing user engagement and enhancing the playability of the target video. Furthermore, the user-customized digital image's appearance is more in line with the user's aesthetic taste.

[0028] In one implementation, obtaining video description information includes determining the video description information based on current scene information, where the current scene information includes at least one of current user status information, current date information, and user input information. In this implementation, determining the video description information based on the current scene allows the resulting target video to better match the current scene, providing an immersive user experience.

[0029] In a second aspect, the present application also provides an electronic device comprising a memory and a processor; the memory and the processor are coupled; wherein the memory is used to store computer program code, the computer program code comprises computer instructions, and when the processor executes the computer instructions, the electronic device executes the video generation method as described in the first aspect and any implementation method above.

[0030] In a third aspect, the present application also provides a chip system, which includes a processor; the processor is coupled to a memory, the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the video generation method in the first aspect and any implementation method mentioned above is executed.

[0031] In a fourth aspect, the present application also provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is run on a computer, the computer executes the video generation method as described in the first aspect and any implementation method above.

[0032] In a fifth aspect, the present application also provides a computer program product, which includes: a computer program or instructions, which, when the computer program or instructions are run on a computer, enables the computer to execute the video generation method in the first aspect and any implementation method described above.

[0033] It can be understood that the beneficial effects that can be achieved by the technical solutions provided in the second to fifth aspects mentioned above can be referred to the beneficial effects in the first aspect and any optional implementation thereof, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solution of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0035] Figure 1 It is a lock screen display effect image based on digital image;

[0036] Figure 2 It is a schematic diagram of changes to elements in the lock screen interface;

[0037] Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application;

[0038] Figure 4 is a schematic diagram of the software structure of the electronic device provided in an embodiment of the present application;

[0039] Figure 5 This is the first flow chart of the video generation method provided in the embodiment of the present application;

[0040] Figure 6 This is a second flow chart of the video generation method provided in an embodiment of the present application;

[0041] Figure 7 This is a schematic diagram of the process of generating scene image materials provided in an embodiment of the present application;

[0042] Figure 8 This is a schematic diagram of the process of generating an image animation provided by an embodiment of the present application;

[0043] Figure 9 This is the third flow chart of the video generation method provided in the embodiment of the present application;

[0044] Figure 10 This is the fourth flow chart of the video generation method provided in the embodiment of the present application;

[0045] Figure 11 This is the fifth flow chart of the video generation method provided in the embodiment of the present application;

[0046] Figure 12 This is the sixth flow chart of the video generation method provided in the embodiment of the present application;

[0047] Figure 13 Schematic diagram of the superposition of target scene material and target animation material provided in an embodiment of the present application;

[0048] Figure 14 This is a schematic diagram of a method of superimposing a target scene material and a target animation material to obtain a target video, provided by an embodiment of the present application;

[0049] Figure 15 This is the seventh flow chart of the video generation method provided in the embodiment of the present application;

[0050] Figure 16 This is the eighth flow chart of the video generation method provided in the embodiment of the present application;

[0051] Figure 17 This is a schematic diagram of a lock screen video setting provided by an embodiment of the present application;

[0052] Figure 18 This is a schematic diagram of a video generation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The following will clearly describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments of the present application, other embodiments obtained by ordinary technicians in this field without making any creative work are all within the scope of protection of this application.

[0054] Hereinafter, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature identified with "first," "second," etc., may explicitly or implicitly include one or more of the features. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0055] In addition, in this application, directional terms such as "upper", "lower", "inner" and "outer" are defined relative to the orientation of the components in the drawings. It should be understood that these directional terms are relative concepts. They are used for relative description and clarification, and they can change accordingly according to changes in the orientation of the components in the drawings.

[0056] The following explains the professional terms mentioned in the embodiments of the present application to facilitate understanding by those skilled in the art.

[0057] A digital avatar is a virtual character created using digital technology. A digital avatar can be two-dimensional (such as a cartoon character) or three-dimensional (such as a 3D animated character) and can be used in animation, games, virtual reality (VR), augmented reality (AR), and other environments. In an embodiment of the present application, a digital avatar can be used in a lock screen video. In this way, when the electronic device is in the lock screen state, a dynamic lock screen effect containing the digital avatar can be displayed.

[0058] The following first describes the application scenarios of the embodiments of the present application with reference to the accompanying drawings.

[0059] With the development of science and technology, the popularity of electronic devices has increased, and the importance of electronic device security has also continued to increase. The lock screen is an important function that can improve the security of electronic devices and can be used to block unauthorized access. Specifically, when an electronic device is in the lock screen state, the lock screen interface is displayed on the screen of the electronic device. The lock screen interface includes a lock screen image and basic information such as date and time. Users can verify permissions by entering a password, pattern, fingerprint or facial recognition. Only after the permission verification is passed can the user unlock the electronic device and access the content or functions in the electronic device. Otherwise, the screen will always remain in the lock screen interface.

[0060] The lock screen interface can be a natural scenery lock screen interface, an abstract art lock screen interface, a digital image lock screen interface, etc. The digital image lock screen interface includes a cartoon or anime style digital image, which is usually colorful and cute, and can provide a pleasant visual experience.

[0061] Figure 1 It is a lock screen display effect image based on digital image.

[0062] like Figure 1 As shown in (a) in FIG, in the lock screen state, the screen of the electronic device displays a lock screen interface 10 including a digital image. Figure 1 As shown in (b) of FIG, when receiving a user's touch, slide or other operation, the electronic device can display an unlocking interface 20, in which a numeric keyboard for inputting a lock screen password is displayed. Figure 1As shown in (c) in the figure, after the user enters the correct lock screen password through the numeric keypad, the verification is passed and the electronic device is unlocked. It should be understood that the embodiment of the present application only uses the unlocking method using a password as an example. In actual applications, it is also possible to use patterns or other methods to unlock, and the lock screen display effect image will also change accordingly.

[0063] However, this digital image-based lock screen interface has the following problems.

[0064] Currently, digital avatar-based lock screens typically display basic information like the date and time on a static image containing the digital avatar. Due to the limitations of static images, the lock screen can only display fixed content, with no dynamic effects. Consequently, when the screen is locked, users are left with a static, unchanging screen, resulting in a poor viewing experience.

[0065] Furthermore, because the digital image-based lock screen consists of a single static lock screen image and basic information like the date and time, editing a specific element of the lock screen requires replacing the entire static image. This method of replacing the entire image doesn't allow for modifying individual elements of the lock screen, hindering personalized lock screen settings.

[0066] Figure 2 It is a schematic diagram of changes to elements in the lock screen interface.

[0067] Figure 2 (a) is a schematic diagram of the lock screen interface before the elements are changed; Figure 2 (b) in the figure is a schematic diagram of the lock screen interface after the elements are changed.

[0068] like Figure 2 As shown in (a) in the figure, the lock screen image is a static image containing a digital image and a willow tree. The static image shows the digital image standing under the willow tree. In the lock screen state, the electronic device displays a lock screen interface. In addition to the digital image, the lock screen interface also contains the willow tree element. At this time, if the willow tree in the lock screen interface needs to be changed to a pine tree, the current lock screen image can only be replaced with a picture containing a cartoon digital image and a pine tree. Figure 2 As shown in (b) of the figure, the lock screen image is replaced with a static image containing a cartoon digital image and a pine tree. The static image shows the digital image standing under the pine tree. In the locked state, the electronic device displays another lock screen interface, which includes the pine tree in addition to the digital image.

[0069] In summary, the current lock screen interface based on digital images cannot achieve dynamic effects, nor can the elements in the lock screen interface be changed at will. Therefore, the playability is low, and it cannot meet the personalized needs of users, which affects the user experience and lacks appeal.

[0070] In order to solve the above problems, an embodiment of the present application provides a video generation method for generating a lock screen video based on a digital image.

[0071] The video generation method provided in the embodiment of the present application first obtains the video description information used to describe the lock screen video, and then, based on the video description information, selects the foreground image material, the background image material and the image animation containing the digital image from the pre-set material set containing multiple materials, and combines them to generate the target video. Since the video generated by the present application uses the image animation of the digital image movement, it can achieve a dynamic effect based on the digital image. In addition, since the video generated by the present application is obtained by combining multiple materials, when it is necessary to change an element in the lock screen interface, it is only necessary to change the material corresponding to the element. For example, for a lock screen interface containing a digital image and a willow tree, if it is necessary to change the willow tree in the lock screen interface to a pine tree, it is only necessary to replace the material corresponding to the willow tree with the material corresponding to the pine tree. Therefore, the embodiment of the present application is more conducive to the personalized setting of the lock screen interface, and enhances the playability and fun.

[0072] The video generation method provided in the embodiments of the present application can be applied to electronic devices. In some embodiments, the electronic device can be a mobile phone, a tablet computer, a handheld computer, a personal computer (PC), an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, an in-vehicle device, or other mobile terminal. The embodiments of the present application do not impose any particular restrictions on the specific type of the electronic device.

[0073] Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.

[0074] like Figure 3As shown, the electronic device 100 may include a processor 110, a memory 120, an antenna 1, an antenna 2, a mobile communication module 130, a wireless communication module 140, a sensor module 150, a display screen 160, etc. Among them, the sensor module 150 may include a pressure sensor 150A, a touch sensor 150B, a fingerprint sensor 150C, an ambient light sensor 150D, a temperature sensor 150E, a gyroscope sensor 150F, a proximity light sensor 150G, etc. In the embodiment of the present application, a video generation instruction from a user may be received through the pressure sensor 150A and the touch sensor 150B.

[0075] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0076] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.

[0077] The memory 120 can be used to store computer executable program code, and the executable program code includes instructions. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the foldable electronic device 100 (such as audio data, a phone book, etc.), etc. In addition, the memory 120 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the foldable electronic device 100 by running instructions stored in the memory 120 and / or instructions stored in a memory provided in the processor.

[0078] In the embodiment of the present application, the code for implementing the video generation method of the embodiment of the present application may be stored in a non-volatile memory. When generating a video, the electronic device 100 may load the executable code stored in the non-volatile memory into a random access memory.

[0079] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 130, the wireless communication module 14, the modem and the baseband processor.

[0080] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization.

[0081] The mobile communication module 130 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 130 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 130 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 130 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 130 can be set in the same device as at least some of the modules of the processor 110.

[0082] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs the sound signal through the audio device or displays an image or video through the display screen 160. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 130 or other functional modules.

[0083] Video codecs are used to compress or decompress digital video. Terminal 100 may support one or more video codecs. This allows terminal 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4. In the embodiment of the present application, the video codec may be used to generate a video and display the generated video on the display screen 160 of electronic device 100.

[0084] The wireless communication module 140 can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 140 can be one or more devices that integrate at least one communication processing module. The wireless communication module 140 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 140 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0085] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 130 , and antenna 2 is coupled to wireless communication module 140 , so that electronic device 100 can communicate with a network and other devices via wireless communication technology.

[0086] Electronic device 100 implements display functionality through a GPU, display screen 160, and an application processor. A GPU is a microprocessor for image processing that connects display screen 160 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.

[0087] In the embodiment of the present application, the electronic device 100 implements the video generation method provided in the embodiment of the present application, which mainly relies on the video codec and the image calculation and processing capabilities provided by the GPU.

[0088] Display screen 160 is used to display images, videos, and the like. Display screen 160 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniLED, a microLED, a micro-oLED, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or N display screens 160, where N is a positive integer greater than 1.

[0089] In the embodiment of the present application, the ability of the electronic device 100 to display the lock screen video depends on the display function provided by the above-mentioned GPU, display screen 160, and application processor.

[0090] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is only for illustrative purposes and does not constitute a structural limitation on the electronic device. In other embodiments of the present application, the electronic device may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0091] Generally speaking, the implementation of the functions of the electronic device 100 requires not only hardware support but also software cooperation. The software system of the electronic device 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a microservice architecture, or a cloud architecture. The embodiment of the present application takes the Android system with a layered architecture as an example to illustrate the software structure of the electronic device 100.

[0092] Figure 4This is a schematic diagram of the software structure of the electronic device provided in an embodiment of the present application.

[0093] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0094] The application layer can include a series of application packages.

[0095] like Figure 4 As shown, the application package may include camera, calendar, map, WLAN, music, short message, gallery, call, navigation, video, lock screen and other applications.

[0096] Among them, a lock screen application is an application that can provide a lock screen interface. When an electronic device is locked by a lock screen application, the screen of the electronic device displays the lock screen interface. The user must unlock the screen by entering a password, pattern, fingerprint recognition or facial recognition before accessing the content and functions on the electronic device.

[0097] The lock screen application can provide a lock screen resource download path, allowing the user to download the corresponding lock screen image or lock screen video through this path and use the lock screen image or lock screen video to set the lock screen interface of the electronic device. In addition, the lock screen application can also provide the function of generating a lock screen image or lock screen video. In an embodiment of the present application, the lock screen application can use the materials in the material collection based on the video description information to generate a video that matches the video description information as the lock screen video of the electronic device.

[0098] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0099] like Figure 4 As shown, the application framework layer may include a window manager, a content provider, a phone manager, a resource manager, a notification manager, a view system, a lock screen management service, and the like.

[0100] Content providers are used to store and retrieve data and make it accessible to applications. This data can include videos, images, audio, calls made and received, browsing history and bookmarks, and phone books.

[0101] The lock screen management service is used to provide a download path for lock screen resources, and is also used to manage and set lock screen resources. It may include lock screen image gallery, lock screen video gallery, personalized recommendations, download and settings, community and user upload functions, etc.

[0102] Android Runtime includes core libraries and a virtual machine. Android Runtime is responsible for scheduling and management of the Android system.

[0103] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.

[0104] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.

[0105] The system library can include multiple functional modules, such as a surface manager, a 3D graphics processing library (such as OpenGL ES (OpenGL for Embedded Systems)), a 2D graphics engine (such as SGL (Simple Graphics Library)), and media libraries.

[0106] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.

[0107] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0108] A 2D graphics engine is a drawing engine for 2D drawings.

[0109] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0110] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.

[0111] The following is an example of the workflow of the software and hardware of the electronic device 100, combined with the lock screen video setting scenario.

[0112] First, the lock screen application calls the lock screen management service in the application framework layer and launches the lock screen application through the lock screen management service. The lock screen application then calls the 2D graphics engine and / or 3D graphics processing library in the system library to generate a target video based on the digital image. After the lock screen application sets the target video as the lock screen video for the electronic device, it calls the media library in the system library to play the lock screen video while the screen is locked.

[0113] Figure 5 This is a flowchart of an embodiment of the video generation method provided in an embodiment of the present application.

[0114] like Figure 5 As shown, this embodiment includes the following steps S11-S15.

[0115] S11: Build a material collection.

[0116] The materials in the material set include at least one image animation and scene image material.

[0117] Electronic devices can construct a material collection using multiple different types of materials. It will be understood that the target video is generated by combining image animation and scene image materials. Therefore, the materials in the material collection include image animation and at least one of foreground image material and background image material to ensure that the material collection contains sufficient materials for selection. It should be noted that the material collection can be constructed by the electronic device, a server, or other device, and this application does not limit this.

[0118] The embodiment of the present application pre-constructs a material set including multiple materials, so that when generating a target video, pre-generated materials can be selected from the material set, thereby improving the efficiency of video generation.

[0119] S12: Building a mapping relationship table based on the materials in the material set.

[0120] The electronic device can construct a mapping relationship table so that it can search the mapping relationship table in subsequent steps to find the material corresponding to the video description information. For example, each material can be used as a value by constructing a key-value pair, and the description of the material or the material ID can be used as the corresponding key. For example, for a scene image material containing an image of a water cup, the key of the material can be set to "cup". Thereafter, the scene image material can be found by "cup". It should also be noted that the mapping relationship table can be constructed by an electronic device, or by a server or other device, and this application does not limit this.

[0121] The embodiment of the present application constructs a mapping relationship table for the material, through which the material in the material collection can be found, thereby improving the efficiency of material search. In addition, the structure of the mapping table is usually flexible, so it is easy to add or modify mapping relationships when the material collection is updated or expanded.

[0122] S13: Obtain video description information.

[0123] An electronic device can obtain video description information used to describe a target video, so as to facilitate the generation of a target video corresponding to the video description information in subsequent steps. The video description information can be used to describe the scene of the target video and / or the actions of a digital avatar in the target video. The video description information can be information input into the electronic device by a user via text or voice, or it can be determined by the electronic device based on the user's current state, the current environmental state, etc. It is understood that the video description information can be used to describe the scene of the target video to be generated and the actions of the digital avatar in the target video. For example, the video description information "Celebrating a Birthday" and "Mid-Autumn Festival" describe the scene of the target video, while the video description information "Listening to Music" describes the actions of the digital avatar in the target video. It is understood that the video description information can also describe both the scene and the actions of the digital avatar. For example, "Eating Mooncakes on Mid-Autumn Festival" describes the scene of the target video as Mid-Autumn Festival and also describes the actions of the digital avatar in the target video as eating mooncakes.

[0124] S14: Based on the video description information and the mapping relationship table, a target scene material and a target animation material are obtained from a preset material set.

[0125] The electronic device obtains target scene material and target animation material that match the content described in the video description information. For example, a material set containing multiple materials can be pre-set, and then the target scene material and target animation material can be selected from the material set by querying the mapping relationship table.

[0126] The target scene material includes scene image material, which includes foreground image material and / or background image material. The foreground image material and background image material are used to form different layers in the target video. The foreground image material is the layer in the video closest to the viewer, usually located in the front of the video, and can be elements such as objects and plants. The background image material is the layer in the video farther away from the viewer, usually located in the back of the video, and can be primary colors such as walls and sky.

[0127] The target animation material includes an image animation in which a digital character moves based on motion data. In the image animation, the digital character performs the specified actions according to the motion data. The digital character in the image animation is a pre-set character, perhaps a cartoon or anime-style character. The actions corresponding to the motion data in the image animation match the video description information. For example, if the video description information is "celebrating a birthday," the motion data may correspond to actions such as blowing out candles and eating cake.

[0128] Furthermore, it is understood that in the target video, the digital image must complete a specified action, so the image animation containing the digital image is dynamic. However, the elements in the scene image material can be dynamic (such as twinkling stars, floating clouds, etc.) or static (such as a stationary table, etc.). Therefore, the scene image material can be a dynamic video or a static image.

[0129] Optionally, in one implementation, the target scene material and target animation material may also include other types of material, such as audio material. For example, for a birthday scene, the target scene material may include a background image material showing a birthday party background, a foreground image material showing elements such as a birthday cake, and an audio material serving as background music for playing a "Happy Birthday" song.

[0130] The embodiment of the present application determines the target scene material and target animation material based on the video description information, which can ensure that the generated target video matches the content described in the video description information. In addition, the use of a mapping relationship table to determine the target scene material and target animation material can simplify the material search logic, thereby improving query efficiency.

[0131] S15: Combine the target scene material and the target animation material to obtain a target video.

[0132] The electronic device combines the target scene material and the target animation material obtained in the above steps to obtain a target video. For example, the target scene material and the target animation material may be superimposed in a certain order. For example, the target video may be obtained by superimposing the background image material, the image animation, and the foreground image material in order from far to near the viewer.

[0133] The embodiment of the present application generates a target video by combining target scene material and target animation material, thereby achieving decoupling between different materials in the target video. When the user needs to modify the content in the target video, he only needs to modify the material corresponding to the content.

[0134] In one implementation, step S15 may be followed by the following steps:

[0135] S16: Setting the target video as the lock screen video of the electronic device.

[0136] The electronic device sets the target video as the lock screen video of the electronic device, so that the electronic device can play the target video in the lock screen state, providing a personalized visual experience for the user.

[0137] Optionally, in one implementation, after step S15, the target video can be saved in the local lock screen resource library of the electronic device so that it can be retrieved from the lock screen resource library at any time as a lock screen video. It can also be uploaded to the lock screen resource community of the lock screen management service, allowing users to communicate and share their own videos with others, increasing social interaction. If you are not satisfied with the generated target video, you can exit directly without saving the target video.

[0138] The following combination Figure 6-Figure 8 , further introduce the technical solutions involved in step S1.

[0139] Figure 6 This is a flowchart of another embodiment of the video generation method provided in an embodiment of the present application.

[0140] like Figure 6 As shown, in this embodiment, the aforementioned step S11 may include the following steps S111-S113:

[0141] S111: Obtain at least one material and a material description tag corresponding to the material.

[0142] The material description tag is used to identify the scene or action corresponding to the material.

[0143] The electronic device determines the material description tag corresponding to the material, and then can use the material and the corresponding material description tag to construct a material set in a subsequent step. Thereafter, the corresponding material can be found in the material set by using the material description tag.

[0144] For the same scene or action, multiple different materials can be pre-set, and the material description tag can identify the scene or action corresponding to the material. Therefore, different materials can correspond to the same scene or action, that is, to the same material description tag. For example, a foreground image material containing a full moon can have a corresponding material description tag of "Mid-Autumn Festival"; a foreground image material containing mooncakes can also have a corresponding material description tag of "Mid-Autumn Festival". For example, an image animation of a digital character waving its right hand can have a corresponding material description tag of "greeting"; an image animation of a digital character waving its left hand can also have a corresponding material description tag of "greeting".

[0145] The embodiment of the present application obtains the material and its corresponding material description tag so that the required material can be found based on the material description tag in subsequent steps.

[0146] In one implementation, step S111 may include the following steps:

[0147] Step 1: Generate at least one scene image material based on a preset scene tag. In addition, the scene tag can also be used as a material description tag corresponding to the scene image material.

[0148] The scene tag is used to identify the scene corresponding to the scene image material.

[0149] Taking foreground image material as an example, we first generate it based on pre-set scene tags. Because scene tags identify the scene to which the foreground image material corresponds, the scene image material generated based on these tags includes elements from the specified scene. For example, for a birthday scene, the scene tag might be "birthday." The foreground image material generated based on the birthday scene tag might include elements typically seen during birthday celebrations, such as birthday cakes.

[0150] It can be understood that since the foreground image material is generated based on a preset scene tag, the elements displayed in the foreground image material match the scene tag. Based on this, the scene tag can be used as the material description tag corresponding to the foreground image material. For example, in this step, the foreground image material generated based on the birthday scene tag displays elements that match the birthday scene, such as a birthday cake. Therefore, the birthday scene tag can be used as the material description tag corresponding to the foreground image material to establish a correspondence between the birthday scene tag and the foreground image material. Thereafter, the foreground image material can be found in the material library using the birthday scene tag.

[0151] Furthermore, to expand the library's resource pool and provide a wider selection of content for the same target video description, multiple foreground image assets can be generated for the same scene tag. When generating a target video, users can choose foreground image assets based on their preferences, facilitating personalized design. Alternatively, they can randomly select from eligible foreground image assets or follow specific rules, providing a sense of freshness and avoiding visual fatigue from viewing the same target video.

[0152] The embodiment of the present application generates multiple scene image materials so that in the subsequent steps, the target scene material can be selected from the generated scene image materials. In addition, since the scene image materials are generated using scene tags, it is only necessary to limit the scene tags to obtain the scene image materials that meet the requirements.

[0153] For background image materials, the generation method and scene label determination method are similar to those for foreground image materials and will not be repeated here.

[0154] The following combination Figure 7, further introducing the generation process of scene image materials.

[0155] Figure 7 It is a schematic diagram of the process of generating scene image materials provided in an embodiment of the present application.

[0156] like Figure 7 As shown, first, for each scene category, multiple scene tags are defined, and each scene tag corresponds to a scene. In this way, the scene definition is realized. Then, several scene description texts are defined for each scene. In some application scenarios, the scene description text is input into an artificial intelligence model including a semantic model and an image generation model, and the scene image material output by the artificial intelligence model is obtained. In other application scenarios, scene image materials are produced based on the scene description text in an artificially produced manner. Finally, a foreground image material library is constructed using the foreground image material, and a background image material library is constructed using the background image material. The foreground image material library and background image material library obtained in this way respectively store a number of foreground image materials and background image materials, which can be used as part of a material set.

[0157] Below Figure 7 The process is further described in the following. In one implementation, the first step may include the following steps:

[0158] For each scene category, set at least one scene label. For each scene label, determine at least one scene description text. Input the at least one scene description text corresponding to the scene label into the semantic model to obtain scene description features corresponding to the scene description text. Input the scene description features into the image generation model to obtain scene image materials corresponding to the scene description features.

[0159] In some embodiments, multiple scene categories may be predefined, and a scene label may be set for each scene category, where the scene label is a brief description of the scene.

[0160] Exemplarily, the following scenario categories may be defined: personal specific scenario category, activity operation scenario category, personal status scenario category, natural scenario category, and daily general scenario category.

[0161] For personal-specific scenarios, you can set the following scenario tags: Birthday, Anniversary, etc. For event-related scenarios, you can set the following scenario tags: Holiday, Event, Spring, Summer, Autumn, Winter, etc. For personal status scenarios, you can set the following scenario tags: Life Service, Music, Sports, Work, Rest, etc. For natural scenarios, you can set the following scenario tags: Weather, Sunrise, Sunset, etc.

[0162] It should be noted that the embodiments of the present application only take the aforementioned scene categories and scene labels as examples. In some implementations, other scene categories and scene labels may also be defined, and the present application does not limit this.

[0163] The embodiments of the present application have designed a variety of different scene categories and scene tags, which can cover a variety of scenes and meet the personalized needs of different customers.

[0164] In some embodiments, when determining the scene description text for each scene tag, a relatively detailed scene description may be designed for the scene identified by the scene tag and used as the scene description text.

[0165] For example, for a birthday scene tag, the following scene description text can be set: birthday cake, birthday candles, balloons, gifts, etc. For a spring scene tag, the following scene description text can be set: flowers, grass, trees, picnic basket, outdoor barbecue grill, kite, etc. For a music listening tag, the following scene description text can be set: headphones, record player, tape, music poster, instrument model, sheet music, etc.

[0166] It should be noted that the embodiments of the present application only take the aforementioned scene labels and scene description texts as examples. In some implementations, other scene description texts may also be defined.

[0167] It should be noted that the aforementioned embodiment only provides a method for determining scene labels and scene description texts. In some implementations, other methods may also be used to determine scene labels and scene description texts. Exemplarily, models such as GPT (Generative Pre-trained Transformer) may be used to generate scene labels and scene description texts; it is also possible to collect historical lock screen interfaces of multiple users, count the historical lock screen interfaces and cluster them to obtain scene labels and scene description texts. It will be understood that the above is only an exemplary description of the method for determining scene labels and scene description texts. In actual applications, other methods may also be used, and this application does not limit this.

[0168] The embodiment of the present application designs a variety of different scene description texts for each scene label, which can improve the diversity of scene image materials. In some embodiments, taking the foreground image material as an example, when using the semantic model to obtain the scene description features corresponding to the scene description text, the scene description text can be input into a semantic model such as BERT (Bidirectional Encoder Representations from Transformers) or GPT that can recognize the meaning expressed by the text, and the corresponding scene description features can be extracted from the semantic model scene description text. The scene description features can be used to generate corresponding pictures or videos as foreground image materials.

[0169] It can be understood that for a scene label, the contents described by different scene description texts can be mutually exclusive or can appear at the same time. For example, for the aforementioned birthday scene label, the contents described by scene description texts such as birthday cake, birthday candles, balloons, and gifts can appear at the same time. Therefore, when generating foreground image materials, scene description texts such as birthday cake, birthday candles, balloons, and gifts can be input into the semantic model at the same time. In this way, the semantic description features simultaneously include information about birthday cake, birthday candles, balloons, and gifts. Therefore, the final foreground image material can simultaneously include elements such as birthday cake, birthday candles, balloons, and gifts.

[0170] Furthermore, scene description texts for birthday cakes, birthday candles, balloons, and gifts can also be separately input into the semantic model. This ultimately yields multiple foreground image materials, each containing one of the elements: birthday cake, birthday candles, balloons, and gifts. It will be appreciated that by generating separate foreground image materials for each element, any number of foreground image materials that meet the requirements can be used when generating the target video, yielding a wider variety of combinations to choose from. For example, a foreground image material containing a birthday cake and a foreground image material containing birthday candles can be used to generate a target video containing both the birthday cake and the birthday candles; alternatively, only a foreground image material containing the birthday cake can be used to generate a target video containing only the birthday cake.

[0171] For background image materials, the method for extracting scene description features is similar to that for foreground image materials and will not be repeated here.

[0172] The embodiment of the present application utilizes a semantic model to extract scene description features corresponding to the scene description text, and can accurately identify the information expressed by the scene description text, thereby helping to obtain scene image materials that better match the scene description information.

[0173] In some embodiments, taking foreground image material as an example, when using an image generation model to obtain scene image material corresponding to scene description features, the scene description features are input into image generation models such as GAN (Generative Adversarial Network) and cGAN (Conditional Generative Adversarial Network), and the image generation model is used to generate corresponding foreground image material based on the scene description features.

[0174] It should be noted that the embodiments of this application only use the aforementioned image generation model to generate foreground image material as an example. In some implementations, other models can also be used to generate foreground image material. For example, a pre-trained artificial intelligence model (such as a Stable Diffusion model) can be used to generate corresponding pictures or videos based on scene description text as foreground image material.

[0175] In particular, for some special and detailed scenes, foreground image materials can also be generated artificially.

[0176] For background image materials, the generation method is similar to that of foreground image materials and will not be repeated here.

[0177] The embodiment of the present application utilizes an image generation model to quickly generate a large amount of scene image materials, thereby improving material production efficiency and reducing material production costs.

[0178] In some implementations, after obtaining the scene image material, the following steps may be further included:

[0179] In the case that the scene image material does not match the preset screening rules, the scene image material is eliminated, wherein the screening rules include: the scene image material matches the scene label, and / or the naturalness of the drawing of the scene image material meets the preset conditions.

[0180] Considering the possibility of inconsistencies between the scene image material generated by the model and the model input, the scene image material is compared with the model input, namely the scene label, to determine whether the two match. If they match, the scene image material is considered to represent the scene identified by the scene label and can be retained. If they do not match, the scene image material is considered to represent the scene identified by the scene label and is therefore discarded.

[0181] Furthermore, given the possibility that the quality of the scene image materials generated by the model may vary, lower-quality scene image materials may be discarded. For example, scene image materials with unnatural graphics may be discarded, while those with natural graphics may be retained to obtain a more realistic target video.

[0182] It should be noted that the present embodiment only illustrates the method of screening scene image materials based on the naturalness of the drawing. In the actual screening process, other screening conditions can also be set. For example, the realism, style, and detail richness of the scene image materials can be used as screening conditions, and scene image materials that do not meet the screening conditions can be eliminated accordingly.

[0183] The embodiment of the present application can improve the quality of scene image materials in the material set by screening scene image materials and eliminating scene image materials with poor effects, so as to obtain a target video with better effects.

[0184] Step 2: Generate at least one image animation according to the preset animation tag, and use the animation tag as the material description tag corresponding to the image animation.

[0185] The animation tag is used to identify an action or a scene where an action occurs in an image animation.

[0186] The electronic device generates an image animation and determines its corresponding material description tag. Optionally, the image animation is first generated according to a pre-set animation tag. Since the animation tag can identify the action or the scene where the action occurs in the image animation, the image animation generated according to the animation tag contains the specified action or the scene where the action occurs. Exemplarily, the greeting animation tag can identify the greeting action in the image animation. The image animation generated according to the greeting animation tag may include a greeting action, such as waving. Exemplarily, the work animation tag can identify the work scene in the image animation. The image animation generated according to the work animation tag may include actions that occur in a work scene, such as operating a computer, making a phone call, etc.

[0187] It can be understood that since the image animation is generated based on a preset animation tag, the action or scene in which the action occurs displayed in the image animation matches the animation tag, so the animation tag can be used as the material description tag corresponding to the image animation. Exemplarily, in this step, the image animation generated based on the greeting animation tag displays an action that matches waving and greeting, so the greeting animation tag can be used as the material description tag corresponding to the image animation to establish a corresponding relationship between the greeting animation tag and the image animation. Thereafter, the image animation can be found in the material library through the greeting animation tag. Exemplarily, in this step, the image animation generated based on the work animation tag displays an action that matches the work scene, so the work animation tag can be used as the material description tag corresponding to the image animation to establish a corresponding relationship between the work animation tag and the image animation. Thereafter, the animation can be found in the material library through the work animation tag.

[0188] In addition, to increase the amount of material in the material library and provide a wider range of choices for the same target video description information, multiple image animations can be generated for the same animation tag. For example, for the greeting animation tag, multiple image animations can be generated, each showing a waving, nodding, shaking hands, and other actions.

[0189] The following combination Figure 8 , further introduces the generation process of image animation.

[0190] Figure 8 It is a flowchart of generating image animation provided by an embodiment of the present application.

[0191] like Figure 8 As shown, first, multiple animation tags are defined for action categories and action scene categories. Each animation tag corresponds to an action or action scene. In this way, action definition and action scene definition are achieved. Then, several animation description texts are defined for each action or action scene. If the animation description text includes a pet character, when designing the action sequence, it is necessary to design the interactive actions between the human character and the pet character. Otherwise, it is only necessary to design the action sequence of the human character separately. After the action sequence design is completed, the motion data generated when the action sequence moves is captured, that is, the motion capture data. The motion capture data is redirected to the digital image to obtain an image animation of the digital image based on the motion data. Finally, the image animation is used to build an image animation library. In this way, the image animation library obtained stores several image animations and can be used as part of the material collection.

[0192] Below Figure 8 The process is further described in the following. In one implementation, the second step may include the following steps:

[0193] At least one animation tag corresponding to an action category is set, and / or at least one animation tag corresponding to a category of a scene in which the action occurs is set. For each animation tag, at least one animation description text is determined. An action sequence is designed based on the at least one animation description text corresponding to the animation tag, wherein the animation description text is used to describe the action identified by the animation tag or the scene in which the action occurs. Motion data generated by a user performing movements in accordance with the action sequence is captured. The motion data is redirected to a digital avatar, and the redirected motion data is adjusted to produce an avatar animation.

[0194] In some embodiments, multiple animation categories can be predefined, and animation tags can be assigned to each category. The animation tags are brief descriptions of the animation. Unlike the multiple scene categories defined for scene image materials in the previous embodiment, this embodiment defines two animation categories for image animation: action category and action scene category.

[0195] For example, for the action category, the following animation tags can be defined to identify actions in the avatar animation: greeting animation tag, heart animation tag, fist bump animation tag, hands on hips animation tag, stretching animation tag, etc. For the action scene category, the following animation tags can be defined to identify the action scene corresponding to the avatar animation: work animation tag, rest animation tag, listening to music animation tag, running animation tag, walking animation tag, cycling animation tag, etc.

[0196] In particular, it is understood that the scene tags in the personal status scene category typically describe scenes involving the actions of the digital avatar. Meanwhile, the animation tags in the action scene category typically describe scenes in which the digital avatar's actions occur. Therefore, when defining tags, the scene tags in the personal status category and the animation tags in the action scene category can be set to be highly consistent. For example, the work scene tags in the personal status category and the work animation tags in the action scene category correspond to each other.

[0197] It should be noted that the embodiments of the present application only take the aforementioned animation categories and animation tags as examples. In some implementations, other animation categories and animation tags may also be defined.

[0198] The embodiment of the present application designs a variety of different animation categories and animation tags, which can cover a variety of sports scenes and meet the personalized needs of different customers.

[0199] In some embodiments, when determining the animation description text for each animation tag, a relatively detailed description of the action or action occurrence scene identified by the animation tag may be designed and used as the animation description text.

[0200] For example, for a greeting animation tag, the following animation description texts may be set: waving left hand, waving right hand, nodding, shaking hands, etc. For a work animation tag, the following animation description texts may be set: typing, making a phone call, reading documents, carrying goods, etc.

[0201] It should be noted that the embodiments of the present application only take the aforementioned animation tags and animation description texts as examples. In some implementations, other animation description texts may also be defined.

[0202] The embodiment of the present application designs a variety of different animation description texts for each animation tag, which can improve the diversity of image animation.

[0203] In one implementation, the aforementioned “determining at least one animation description text for each animation tag” may include the following steps:

[0204] Determine the character corresponding to the animation tag. If there are multiple characters, determine the interactive actions between the multiple characters. Based on the interactive actions, determine the animation description text corresponding to the animation tag.

[0205] When designing animation description text for an animation label, electronic devices can consider the situation of multiple characters and design interactive actions for multiple characters to increase the fun of the image animation. The multiple characters can include human characters or animal characters.

[0206] For example, a running animation tag might correspond to a human character and a pet character. When designing the animation description, in addition to describing the running movements of the human character and the pet character separately, you can also design interactive actions between the human character and the pet character. For example, if the human character makes a command gesture, the pet character can then perform actions such as jumping or stopping based on the command gesture.

[0207] Alternatively, the same animation tag can be assigned to different characters to generate multiple character-based animations. For example, a running animation tag can be assigned to a single character to generate multiple running animations; and a running animation tag can be assigned to both a character and a pet character to generate multiple running animations of both the character and the pet together. Using both sets of animations as the running animation tag increases the diversity of the generated target videos.

[0208] Alternatively, for multiple characters, separate actions can be designed for each character, without designing interactive actions between the characters. This can also generate an image animation containing multiple characters, and the target video generated using this image animation contains multiple characters, with the actions of each character being independent of each other.

[0209] Alternatively, a separate image animation can be generated for each character. When generating a target video, several image animations are selected and combined to obtain the target video. The target video thus obtained also contains multiple characters, and the actions of each character are independent of each other.

[0210] The embodiment of the present application improves the richness of the content in the target video through the interactive action design of multiple characters. In addition, the addition of animal characters is more popular with pet users and can increase the appeal to users who like animals.

[0211] In some embodiments, when designing an action sequence based on an animation description text, the action described in the animation description text may be first divided into multiple steps. Alternatively, based on the scenario in which the action described in the animation description text occurs, the action corresponding to the scenario may be determined, and then the action may be divided into multiple steps.

[0212] It is understood that the animation description text is a more detailed description of the image animation than the animation label, which can describe the specific action or the scene in which the action occurs. For example, if the animation label is a greeting, the corresponding animation description text can be a wave, a handshake, a nod, a hug, etc.

[0213] An action sequence consists of a series of sequential steps that achieve the content described in the animation description. For example, for the animation description of waving, the following action sequence can be designed: [Body facing the viewer of the animated image; raise right hand; swing wrist up and down; lower arm]. By executing each step in the action sequence, the greeting action can be completed. The animated greeting image based on this action sequence can include a waving animation.

[0214] Optionally, for an animation tag, the contents described by different animation description texts may be mutually exclusive or may appear at the same time. For example, for the aforementioned greeting animation tag, the contents described by animation description texts such as waving and hugging may appear at the same time. Therefore, the two animation description texts of waving and hugging may be combined to design the following action sequence: [body facing the viewer of the image animation; raise the right hand; gently swing the wrist up and down; walk towards the viewer; open arms; release the hug; return to standing position]. The greeting action can be completed by executing each step in the action sequence in sequence. The image animation of greeting obtained based on this action sequence includes a waving animation and a hugging animation.

[0215] In some embodiments, when capturing motion data generated by a user moving according to an action sequence, the user can first perform the action corresponding to the animation tag by following the action sequence designed in the aforementioned steps. In some implementations, the user performing the action sequence can be a professional motion capture actor. Because motion capture actors can accurately perform complex movements and expressions, the accuracy of the motion data can be guaranteed, resulting in a more expressive animated image. In other implementations, the user performing the action sequence can also be the user of the target video, resulting in a more personalized animated image to meet the user's personalized needs.

[0216] As the user moves through the motion sequence, motion capture technology can be used to capture motion data. Motion capture technology (also known as motion capture technology) is a technology that can track and record the motion trajectory of a human body or other object in real three-dimensional space. In embodiments of the present application, motion capture technology can be used to capture the user's motion trajectory and convert it into motion data.

[0217] For example, optical motion capture technology can be used to capture motion data. Specifically, reflective or actively luminous markers can be added to specific locations on the user's body, such as joints, head, and hands. During the user's movements, a camera can be used to capture the reflected light from the markers, generating the motion trajectory of specific locations such as joints, head, and hands. This motion trajectory can then be converted into motion data in the form of three-dimensional coordinates.

[0218] It should be noted that the embodiments of the present application only take optical motion capture technology as an example. In some implementations, other motion capture technologies may also be used. For example, inertial motion capture technology may be used, and the user wears a gyroscope, and the user's motion data is calculated based on the rotation information of the gyroscope during the user's movement. For example, visual capture technology may also be used to record the user's movements through a camera, and use deep learning and other algorithms to identify the user's joint information to obtain motion data. In actual application, appropriate motion capture technology can be selected based on the application scenario. For example, for professional motion capture actors, optical motion capture technology and inertial motion capture technology can be used to capture their motion data; for users of electronic devices, visual motion capture technology can be selected to capture their motion data.

[0219] In some embodiments, when redirecting motion data to a digital avatar, the motion data corresponding to each specific position can first be bound to the corresponding position in the digital avatar. For example, if the motion data includes motion trajectory data for each joint, the motion trajectory data for the wrist joint can be bound to the wrist joint of the digital avatar. The motion trajectory of the wrist joint can then be mapped to the wrist joint of the digital avatar, achieving motion data redirection. At this point, the aforementioned motion data is applied to the digital avatar, and the digital avatar can perform corresponding actions based on the motion data, resulting in an animation effect.

[0220] Optionally, considering the unnatural motion of the action sequence and the collection deviation, the generated image animation will not be realistic enough. Therefore, the redirected motion data can be fine-tuned to achieve a more realistic animation effect.

[0221] Optionally, if the motion data is not fully suitable for the cartoon style, the redirected motion data can be fine-tuned. For example, by adjusting the motion data, the amplitude and rhythm of the movement can be adjusted, so that the digital image performs the corresponding movement based on the fine-tuned motion data, which is more consistent with the cartoon style.

[0222] By capturing the user's motion data and redirecting it to a digital avatar, the present embodiment can produce an avatar animation that matches the user's actual movements, which helps to improve the naturalness and realism of the avatar animation. Furthermore, compared to directly designing the avatar animation, the present embodiment can also improve the efficiency of avatar animation production.

[0223] S112: Constructing a material set based on the correspondence between the materials and the material description tags.

[0224] A material collection is constructed based on the correspondence between materials and material description tags. After that, the corresponding materials can be quickly found in the material collection based on the material description tags.

[0225] In some implementations, after constructing the material set, the following step S113 is further included:

[0226] S113: Add the material information of each material to the material set.

[0227] When storing materials in a material collection, corresponding material information is added to each material. The material information includes at least one of a material description tag, a material description text, an index, a default material identifier, extended information, and a material version number corresponding to the material.

[0228] In some implementations, the material information is as shown in Table 1 below:

[0229] Table 1 Material information table

[0230] Field type meaning classification string Scene Tags scene string Animation Tags description string Material description text index int index defaultflag int Default material identification extend json-string Extended Information version string Material version number

[0231] In Table 1, the material description tags corresponding to scene image materials and avatar animations are respectively the scene tag and animation tag. For a description of scene tags, refer to the first step of step S111 above. These tags can be personal, holiday operation, personal status, nature, or general daily use. For a description of animation tags, refer to the second step of step S111 above. These tags can be birthday, holiday, listening to music, exercise, work, rest, sunrise, sunset, weather-sunny, weather-rainy, weather-snow, weather-fog, weather-haze, weather-wind, or general animation tags. These are not detailed here. Material description text briefly describes the material. The material description text corresponding to scene image materials is the scene description text, while the material description text corresponding to avatar animations is the animation description text. The index is a unique identifier for the material, such as the file code corresponding to the material, and can be used to distinguish different materials corresponding to the same tag. The default material identifier is used to indicate whether the material can be used as the default material. When generating the target video, a default material can be randomly selected based on the default material identifier. Optionally, the default material is the material corresponding to the general daily scene tag. Extended information includes the material's effective time and validity time, which can be determined based on the time zone. During the time periods defined by the effective and effective times, a material is in the effective state. During all other time periods, the material is in the effective state. By setting the effective and effective states, you can make a material available during specific times and unavailable during other times. For example, some materials can be available only during the day, while others can be available only at night. A material version number is an identifier added when a material is modified or updated. It can be used to record different historical versions of the same material and avoid confusion between different versions.

[0232] The embodiment of the present application records the material information of each material and adds it to the material set, which facilitates the management of the materials and is particularly suitable for situations where there are a large number of materials.

[0233] The following combination Figure 9 , further introduces the technical solution involved in step S12.

[0234] Figure 9 This is a flowchart of another embodiment of the video generation method provided in an embodiment of the present application.

[0235] like Figure 9 As shown, in this embodiment, the aforementioned step S12 may include the following steps S121-S122:

[0236] S121: Construct at least one material subset based on the scene image material and the image animation in the material set.

[0237] The material set includes multiple scene image materials and image animations, and the material subset is obtained by selecting several scene image materials and image animations from the material set. For example, in the material set {F, B, A}, F = {f1, f2, ..., f n}, is a set of multiple foreground image materials; B = {b1, b2, ..., b m}, is a collection of multiple background image materials; A={a1,a2,…,a k}, is a collection of multiple image animations. You can use several materials in the material set to construct a material subset {F i ,B i ,A i}. Among them, F i is a subset of F, B i A is a subset of B. i is a subset of A. Therefore, the material subset {F i ,B i ,A i} is a subset of the material set {F,B,A}.

[0238] Optionally, the materials in the same material subset may be materials with an associated relationship. For example, the image animation corresponding to the running action, the scene image material corresponding to the treadmill, and the scene image material corresponding to the playground may be placed in the same material subset.

[0239] S122: Construct a mapping relationship table based on the corresponding relationship between each material subset and the material description tag.

[0240] It is understood that a material subset includes several materials, and there is a correspondence between the material description tags and the material subset. Based on this, the correspondence between the material subset and the material description tags can be obtained to construct a mapping relationship table. Afterwards, the material description tags can be queried in the mapping relationship table to obtain the corresponding material subset.

[0241] For example, the aforementioned material subset including the image animation corresponding to the running action, the scene image material corresponding to the treadmill, and the scene image material corresponding to the playground can correspond to the material description tag "running".

[0242] The mapping relationship table constructed in the embodiment of the present application contains the correspondence between the material subsets and the material description tags. Therefore, when querying the material description tags, all the materials in the material subset can be directly obtained, which is conducive to improving the search efficiency.

[0243] In one implementation, step S122 may include the following steps:

[0244] For each material subset, the corresponding relationship between the material subset and the material description tag is determined based on the corresponding relationship between each material in the material subset and the material description tag. Based on the corresponding relationship between the material subset and the material description tag, a mapping relationship table is constructed.

[0245] Considering that a material subset is composed of multiple materials, the correspondence between the material subset and the material description tags can be determined based on the correspondence between the materials in the material subset and the material description tags, thereby constructing a mapping relationship table. For example, the mapping relationship table can be in the form of key-value pairs, where the key and value are the material subset and its corresponding material description tag, respectively.

[0246] Optionally, in the material subset, the material description tag corresponding to each material corresponds to the material subset. i ,B i ,A i}, F i The material in the corresponding material description tag l fi , B i The material in the corresponding material description tag l bi , A i The material in the corresponding material description tag l ai , then the material subset {F i ,B i ,A i}Also corresponds to label l fi 、l bi and l ai That is, there is a one-to-many relationship between material subsets and material description tags.

[0247] Optionally, in the material subset, the material description tags corresponding to all materials form a composite tag, and the material subset corresponds to the composite tag. i ,B i ,A i}, F i The material in the corresponding material description tag l fi , B i The material in the corresponding material description tag l bi , A i The material in the corresponding material description tag l ai , then the material subset {F i ,B i ,A i} corresponds to the compound tag {l fi ,l bi ,l ai}.

[0248] Optionally, after determining the material description tag corresponding to the material subset, it can also be determined as the material description tag corresponding to each material in the material subset. i ,B i ,A i}Also corresponds to label l fi 、l bi and l ai , then we can determine F i The material in the tag also corresponds to l fi 、l bi and l ai , B i The material in the tag also corresponds to l fi 、l bi and l ai , A i The material in the tag also corresponds to l fi 、l bi and l ai That is, a material corresponds to both a scene tag and an animation tag. For example, the material subset {F u ,B i ,A i} corresponds to the compound tag {l fi ,l bi ,l ai}, then we can determine F i 、B i and A i The materials in the corresponding composite tag {l fi ,l bi ,l ai}.

[0249] The material subset in the embodiment of the present application includes both scene image material and image animation. Therefore, two different types of materials, scene image material and image animation, can be found at the same time based on a material description tag corresponding to the material subset.

[0250] The following combination Figure 10 , further introduces the technical solution involved in step S13.

[0251] Figure 10 This is a flowchart of another embodiment of the video generation method provided in an embodiment of the present application.

[0252] like Figure 10 As shown, in this embodiment, the aforementioned step S13 may include the following step S131:

[0253] Step S131: Determine video description information based on current scene information.

[0254] The electronic device determines video description information based on the current scene information to obtain a target video that matches the current scene. The current scene information may include current user status information, current date information, user input information, etc.

[0255] The current user status information can be used to identify the user's exercise status and can be obtained through a variety of different channels. For example, the current user status information can be determined using the built-in sensors of the electronic device. For example, if the built-in sensors of the mobile phone are used to detect that the user is running, the current user status can be determined to be running. For example, the status of the application loaded on the electronic device can be monitored, and the current user status information can be determined based on the application status. For example, if it is detected that the music program is playing music, the current user status information can be determined to be listening to music. For example, the current user status information can be determined based on information sent by other electronic devices. For example, if cycling information sent by a sports watch is received, the current user status can be determined to be cycling.

[0256] The current date information can be used to determine seasons, holidays, anniversaries, etc., and can be obtained based on a calendar application installed on the electronic device, etc. For example, if the current date is obtained as January 1, it can be determined that the current date is New Year's Day.

[0257] User input information can be information input by the user through text or voice, etc., and is input by the user according to actual needs.

[0258] It is understood that the video description information can be determined by a single current scene information or by a combination of multiple current scene information. For example, if the current date information is New Year's Day, the video description information can be determined to be New Year's Day, and a target video related to New Year's Day can be generated. For example, if the current user status is running and the current date is winter, the video description information can be determined to be {winter, running}, and a target video of running in winter can be generated.

[0259] It can be understood that the embodiments of the present application only take the aforementioned current scene information as an example. In some implementations, other current scene information may also be used. Exemplarily, the current scene information may include current time information, which is used to determine day and night and daily scheduled activities, etc. For example, if the current time is determined to be 12:00 noon, it can be determined that the current time is lunch time, and "eating lunch" is used as the current scene information to generate a target video related to "eating lunch" in subsequent steps. Exemplarily, if the current weather is detected to be light rain based on a weather application, "light rain" can be used as the current scene information to generate a target video related to "light rain" in subsequent steps.

[0260] The embodiments of the present application design a variety of methods for determining video description information, which can obtain a target video that is more in line with the current actual scene or more in line with user preferences.

[0261] The following combination Figure 11 , further introduces the technical solution involved in step S14.

[0262] Figure 11 This is a flowchart of another embodiment of the video generation method provided in an embodiment of the present application.

[0263] like Figure 11 As shown, in this embodiment, the aforementioned step S14 may include the following steps S141-S142:

[0264] S141: Determine a material description tag corresponding to the video description information.

[0265] The electronic device determines the material description tag corresponding to the video description information, and uses the material description tag to determine the corresponding target scene material and target animation material. Optionally, semantic features of the video description information may be extracted, and then a material description tag matching the semantic features may be found among multiple predefined material description tags. For example, for the video description information "birthday," the material description tag "birthday" is found among multiple material description tags to semantically match the word "birthday." In this case, the material description tag corresponding to the video description information "birthday" is "birthday." Optionally, in the technical solutions involved in steps S13 and S131, the video description information is directly generated using words in the material description tag. In the technical solution involved in step S141, the video description information may be directly used as the corresponding material description tag. For example, after detecting that the current date is the user's birthday, the video description information "birthday" is directly generated using the word "birthday" in the material description tag. The video description information thus generated may be directly used as the corresponding material description tag.

[0266] It is understood that video description information can correspond to one material description tag or multiple material description tags. For example, the aforementioned video description information "running" can be determined to correspond to the "running" animation tag. For example, the aforementioned video description information {winter, running} can be determined to correspond to the "winter" scene tag and the "running" animation tag. In other words, the video description information {winter, running} corresponds to the compound tag "winter running."

[0267] The embodiment of the present application obtains a material description tag based on complex video description information that does not have a unified format. Since the material description tag is recorded in the mapping relationship table, the embodiment of the present application establishes a connection between the video description information and the mapping relationship table through the material description tag, so that in subsequent steps, the mapping relationship table can be queried to obtain the material that matches the video description information.

[0268] S142: Based on the pre-built mapping relationship table, the target scene material and the target animation material corresponding to the material description tag corresponding to the video description information are obtained from the material set.

[0269] The electronic device determines the target scene material and target animation material based on the material description tags corresponding to the video description information. Since each material has a direct correspondence with a material description tag, the electronic device can use a table lookup to find the material corresponding to the material description tag corresponding to the video description information in the mapping relationship table, and use this as the target scene material and target animation material. The mapping relationship table is a pre-constructed table that records the correspondence between materials and material description tags. For an explanation and construction method of this table, please refer to the technical solution involved in step S12 above and will not be repeated here.

[0270] The embodiment of the present application determines the target scene material and target animation material by querying the mapping relationship table. Since the above steps have established the connection between the video description information and the mapping relationship table, the embodiment of the present application can ensure that the material queried based on the mapping relationship table conforms to the content described in the video description information.

[0271] In some implementations, step S142 may include the following steps:

[0272] Based on the mapping relationship table, the material subset corresponding to the material description tag corresponding to the video description information is determined as the target material subset. In the target material subset, the target scene material and the target animation material are determined.

[0273] A material subset includes several materials within a material set. In the aforementioned embodiment, a mapping table has been constructed based on the correspondence between each material subset and a material description tag. Therefore, the mapping table can be used to search for the material description tag corresponding to the target video information and obtain the corresponding material subset as the target material subset.

[0274] Optionally, if the target video corresponds to a material description tag, or corresponds to multiple material description tags, and multiple material description tags correspond to the same material subset, then the material subset can be used as the target material subset, and the target scene material and target animation material can be determined from it. For example, if the target video corresponds to a material description tag "running", and a material subset {F1, B1, A1} corresponds to both the scene tag "summer" and the animation tag "running", then this material subset {F1, B1, A1} can be determined as the target material subset. Therefore, the target scene material can be determined in F1 and B1, and the target animation material can be determined in A1.

[0275] Optionally, if the target video corresponds to multiple material description tags, and the multiple material description tags correspond to different material subsets, then the corresponding materials can be selected from each material subset based on the material description tags corresponding to the target video. For example, the target video corresponds to two material description tags "winter" and "running", and none of the existing material subsets correspond to these two material description tags at the same time. Only the material subset {F1, B1, A1} corresponds to the scene tag "summer" and the animation tag "running" at the same time, and the material subset {F2, B2, A2} corresponds to the scene tag "winter" and the animation tag "cycling" at the same time. At this time, it can be determined that the material subsets {F1, B1, A1} and {F2, B2, A2} are both target material subsets, and the target animation material is determined in A1, and the target scene material is determined in F2 and B2.

[0276] It can be understood that since the target material subset is pre-established, compared with the method of directly searching for the target material, the embodiment of the present application first searches for the target material subset and then determines the target material in the target material subset, which can effectively improve the search efficiency.

[0277] In one implementation, “determining target scene material and target animation material in the target material subset” may include the following steps:

[0278] A first target index and a second target index are obtained. In the material subset, the scene image material corresponding to the first target index is determined as the target scene material, and the image animation corresponding to the second target index is determined as the target animation material.

[0279] Considering that the target material set includes multiple materials, and the index can distinguish different materials in the target material set, the target scene material and the target animation material can be uniquely specified in the target material set according to the target index.

[0280] The target index includes a first target index and a second target index. The first target index is used to distinguish different scene image materials corresponding to the same scene tag, and the second target index is used to distinguish different image animations corresponding to the same animation tag. For example, the target material set includes scene image materials f1, f2, b1, b2 and image animations a1, a2, and their corresponding indexes are f01, f02, b01, b02, a01, and a02, respectively. If the first target indexes are f01 and b01 and the second target index is a02, the target scene materials f1, b1 and the target animation material a2 can be uniquely determined.

[0281] Optionally, the target index can be input by the user in the form of text, voice, or a selection box. For example, the materials in the target material set can be displayed on a display screen. The user can click on any scene image material to set it as the target scene material, and click on any image animation to set it as the target animation material.

[0282] Optionally, the target index can be randomly generated. When the target material set includes multiple scene image materials and multiple animation images, a first target index and a second target index can be randomly generated, and then the target scene material and the target animation material can be determined using the randomly generated target index.

[0283] The embodiment of the present application utilizes target indexes to uniquely and quickly determine target scene materials and target animation materials, thereby improving video generation efficiency and accuracy.

[0284] It will be understood that the aforementioned embodiment merely uses a method for determining target scene materials and target animation materials using a target index as an example. In some implementations, other methods may also be used to determine target scene materials and target animation materials in a target material set. Exemplarily, the target material may be determined based on a default identification bit in the material information. For example, the default identification bit information of each material in the target material set may be queried. Materials whose default identification bits meet the conditions are designated as default materials, and the default materials may be determined to be the target scene materials or target animation materials.

[0285] The following combination Figures 12 to 14 , further introduces the technical solution involved in step S15.

[0286] Figure 12 This is a flowchart of another embodiment of the video generation method provided in an embodiment of the present application.

[0287] like Figure 12 As shown, in this embodiment, the aforementioned step S15 may include the following steps S151-S152:

[0288] S151: Superimposing target scene materials and target animation materials in sequence according to a preset order to obtain a target video.

[0289] The electronic device can sequentially overlay target scene materials and target animation materials in a certain order, and maintain a certain interval between adjacent materials to obtain a target video with a sense of depth.

[0290] It is understood that the target scene material may include any number of foreground image materials. Figure 13 , which describes two situations: the target scene material includes multiple foreground image materials and does not include any foreground image materials.

[0291] Figure 13This is a schematic diagram of the superposition of the target scene material and the target animation material provided in the embodiment of the present application.

[0292] like Figure 13 As shown in (a) in the figure, the target scene material includes the following two foreground image materials: a water cup image material and a grass image material. The background image material, the image animation, the water cup image material and the grass image material can be superimposed in order from far to near from the viewer to obtain the target video. Figure 13 As shown in (b), in the target video at this time, part of the water cup is blocked by the grass.

[0293] like Figure 13 As shown in (c) in the figure, the target scene material includes the following two foreground image materials: a water cup image material and a grass image material. The background image material, the image animation, the grass image material and the water cup image material can be superimposed in the order from far to near from the viewer to obtain the target video. Figure 13 As shown in (d) in the figure, in the target video at this time, part of the grass is blocked by the water cup.

[0294] like Figure 13 As shown in (e) in FIG, for example, the target scene material does not include foreground image material, but only background image material. The background image material and the image animation can be superimposed in the order from far to near from the viewer to obtain the target video. Figure 13 As shown in (f), in the target video at this time, there are only digital images and grass clouds in the background.

[0295] The embodiments of the present application can overlay multiple foreground image materials in different orders, thereby generating different videos using the same foreground image material. Furthermore, the embodiments of the present application can also omit the foreground image material, thereby generating a video without foreground elements. This design increases the diversity of the target videos.

[0296] It can be understood that the target scene material may include any number of background image materials, and the superposition method thereof is similar to the superposition example of any number of foreground image materials, which will not be described in detail here.

[0297] S152: Render each frame of the target video according to the illumination change information and / or the camera trajectory information to obtain a rendered target video.

[0298] It can be understood that in three-dimensional space, the position, intensity, and direction of the light source will affect the distribution of light and dark on the surface of the object, thereby creating a visual sense of three-dimensionality. Therefore, the target video can be rendered based on the lighting change information, and the three-dimensional sense can be enhanced through light and dark contrast and light and shadow effects. For example, a light source can be set for the target video. The relative position between the points at different positions in the target video and the light source is different, so the light received is also different, that is, the degree of brightness and darkness is different. Based on this, light and dark changes can be applied to different positions of the target video to simulate the shape and depth of each element in the target video in three-dimensional space.

[0299] It can be understood that in three-dimensional space, camera movement trajectories can be used to create visual layers and enhance the three-dimensionality of the image. Therefore, the target video can be rendered based on the camera movement information, enhancing the three-dimensional effect through the lens movement path. For example, a single camera can be set up for the target video. Points at different locations in the target video have different relative positions to the camera. Rendering the target video based on the perspective and overlapping occlusion effects at each location can enhance the spatial perception of the target video.

[0300] Optionally, the illumination change information and the camera movement trajectory information may be information pre-stored in the renderer, or may be custom information input by a user.

[0301] The embodiment of the present application takes into account the impact of lighting and camera movement on the video and renders the target video, which can increase the three-dimensional sense of objects in the target video and improve the realism of the target video.

[0302] Figure 14 This is a schematic diagram of an embodiment of the present application for superimposing target scene material and target animation material to obtain a target video.

[0303] like Figure 14 As shown, the background image material, the digital human image animation, and the foreground image material are sequentially superimposed, and each frame is rendered according to lighting changes and camera movement trajectory to obtain the final target video. For the material superposition and image rendering methods, please refer to the technical solutions involved in steps S151-S152 above. The target video includes grass and cloud elements in the background, the digital human greeting animation, and the water cup element in the foreground. Furthermore, the digital pet image animation can also be superimposed, which is the same as the digital human image animation and will not be further described here.

[0304] The following combination Figure 15 and Figure 16 , another implementation method is introduced.

[0305] Figure 15 This is a flowchart of another embodiment of the video generation method provided in an embodiment of the present application.

[0306] like Figure 15As shown, first, the video material resources are searched locally on the electronic device. The video material resources may include digital image resources, scene image materials, and image animations. If the electronic device has no local resources, a prompt message is generated to prompt the user to download the resources locally. If the electronic device has local resources, the user edits the digital image, and the edited digital image has an appearance that suits the user's preferences. Then, based on the edited digital image, a target video is generated. If the video generation fails, a corresponding prompt message is generated to prompt the user. Otherwise, the generated target video is saved to a specified location. The specified location can be a local location on the electronic device or a location on the cloud server, which is not limited here.

[0307] The following combination Figure 16 ,right Figure 15 The process is further introduced in the following.

[0308] Figure 16 This is a flowchart of another embodiment of the video generation method provided in an embodiment of the present application.

[0309] like Figure 16 As shown, this embodiment includes the following steps S21-S26.

[0310] S21: Responding to the video generation instruction, obtaining a digital image.

[0311] The electronic device obtains a default digital image in response to the video generation instruction, so as to perform personalized image editing based on the default image.

[0312] Optionally, the video generation instruction can be input by the user, such as a user can input a video generation instruction to actively generate a target video. The video generation instruction can also be input on a scheduled basis, such as generating a target video as a lock screen video every day, so that the user can watch a new lock screen video every day, increasing the user's sense of freshness. The video generation instruction can also be generated based on changes in current scene information. For example, if the current scene information changes from "sunny weather" to "rainy weather", a video generation instruction can be generated in response to this change to generate a target video corresponding to the new current scene information.

[0313] Optionally, a default digital image can be pre-stored locally on the electronic device. Therefore, before generating the digital image, it is determined whether the digital image is stored locally on the electronic device. If not, the digital image is downloaded from the cloud server to the local electronic device.

[0314] Optionally, in addition to determining whether the electronic device has a digital image stored locally, it may also be determined whether the electronic device has a material collection stored locally. If not, the material collection is downloaded from the cloud server to the electronic device. Alternatively, rather than determining whether the electronic device has a material collection stored locally in this step, in the subsequent step of acquiring the target scene material and target animation material, it may be determined directly whether the electronic device has materials matching the video description information locally. If not, the materials are downloaded from the cloud server to the electronic device.

[0315] S22: Edit the digital image in response to the editing instruction.

[0316] Based on the user's editing instructions, the electronic device modifies the appearance of the digital avatar to achieve customization of the digital avatar. For example, based on the default avatar in the technical solution involved in step S21, the digital avatar's hairstyle, hair color, skin color, etc. can be modified to obtain a digital avatar that meets the user's personalized needs, thereby increasing the user's freedom.

[0317] S23: Obtain video description information.

[0318] For the description of step S23 , please refer to the aforementioned step S13 , which will not be repeated here.

[0319] S24: Based on the video description information, a target scene material and a target animation material are obtained from a preset material set.

[0320] For the description of step S24 , please refer to the aforementioned step S14 , which will not be repeated here.

[0321] In one implementation, the material set includes a scene image material set and an animation material set. Step S24 may further include the following steps:

[0322] Obtain target scene material from the scene image material set, and obtain target animation material from the animation material set.

[0323] To facilitate the management of different types of materials, the material collection can be divided into a scene image material collection and an animation material collection, and the scene image materials and image animations are stored in the scene image material collection and the animation material collection respectively. Afterwards, the target scene material and target animation material can be obtained from the scene image material collection and the animation material collection respectively.

[0324] S25: Combine the target scene material and the target animation material to obtain a target video.

[0325] For the description of step S25 , please refer to the aforementioned step S15 , which will not be repeated here.

[0326] In one implementation, step S25 may be followed by the following steps:

[0327] S26: Setting the target video as the lock screen video of the electronic device.

[0328] For the description of step S26 , please refer to the aforementioned step S16 , which will not be repeated here.

[0329] The following combination Figure 17 , which introduces the lock screen video setting process.

[0330] Figure 17 This is a schematic diagram of a lock screen video setting provided in an embodiment of the present application.

[0331] like Figure 17 As shown in (a), first, click the setting button on the application interface 10 of the electronic device to jump to the setting interface 20.

[0332] like Figure 17 As shown in (b), the setting interface 10 of the electronic device is opened, and the desktop personalization option is clicked in the setting interface 10 to jump to the desktop personalization interface 30.

[0333] like Figure 17 As shown in (c), in the desktop personalized interface 30, a variety of optional lock screen types are displayed, such as a digital human lock screen, a static landscape lock screen, a simple lock screen, etc. In the embodiment of the present application, a digital human lock screen can be selected. After selecting the lock screen type, the lock screen application is started to generate the target video. In the process of generating the target video, it is first determined whether the electronic device has a digital image resource of the digital human, that is, the default digital image. If not, the digital image resource is downloaded from the cloud. Thereafter, the digital image can be edited and used. At the same time, the electronic device intelligently perceives the current scene information such as the user status and festivals to obtain the video description information. According to the video description information, the scene image material and the image animation are randomly selected from the matching set as the target scene material and the target animation material. Finally, the target scene material and the target animation material are combined to obtain the target video.

[0334] like Figure 17 As shown in (d) in FIG, after obtaining the target video, the interface 40 is jumped to the video display interface 40. The user can click the preview button in the display interface 40 to jump to the preview interface 50.

[0335] like Figure 17 As shown in (e) in FIG, the preview interface 50 can display the lock screen effect corresponding to the target video. If the user is satisfied with the display effect, the user can click the Apply button in the preview interface 50 to set the video as the lock screen video.

[0336] like Figure 17 As shown in (f), after setting the lock screen video, the lock screen interface 60 is displayed in the lock screen state.

[0337] Other embodiments of the present application provide a video generating device.

[0338] Figure 18 This is a schematic diagram of a video generation device provided in an embodiment of the present application.

[0339] like Figure 18 As shown, the video generation device may include: a display screen 1001, a memory 1002, a processor 1003, and a communication module 1004. The aforementioned components may be connected via one or more communication buses 1005. Display screen 1001 may include a display panel 10011 and a touch sensor 10012. Display panel 10011 is used to display images. Touch sensor 10012 may transmit detected touch operations to an application processor to determine the type of touch event and provide visual output related to the touch operation via display panel 10011. Processor 1003 may include one or more processing units, such as an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a video codec, a digital signal processor, a baseband processor, and / or a neural network processor. The different processing units may be independent devices or integrated into one or more processors. Memory 1002 is coupled to processor 1003 and is used to store various software programs and / or computer instructions. Memory 1002 may include volatile memory and / or non-volatile memory. When the processor executes the computer instructions, the video generation device may perform the various functions or steps performed by the above method embodiments.

[0340] When the software program and / or multiple groups of instructions in the memory 1002 are executed by the processor 1003, the video generation device implements the following method steps: obtaining video description information, the video description information is used to describe the scene of the target video and / or the action of the digital image in the target video; based on the video description information, obtaining target scene materials and target animation materials from a preset material set, wherein the target scene materials include scene image materials, the scene image materials include foreground image materials and / or background image materials, and the target animation materials include image animation of the digital image based on motion data; combining the target scene materials and target animation materials to obtain the target video.

[0341] The present application also provides an electronic device, comprising: a processor, a memory and a touch screen; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the video generation method in any implementation manner in the above embodiments.

[0342] An embodiment of the present application also provides a chip system, which includes at least one processor and at least one interface circuit. The processor and the interface circuit can be interconnected via lines. For example, the interface circuit can be used to receive signals from other devices (such as the memory of an electronic device). For another example, the interface circuit can be used to send signals to other devices. Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can perform the various steps in the above embodiments. Of course, the chip system can also include other discrete devices, which is not specifically limited in the embodiment of the present application.

[0343] The embodiment of the present application further provides a computer-readable storage medium, which includes computer instructions. When the computer instructions are executed in the above-mentioned electronic device (such as Figure 3 When the method is executed on the electronic device 100 shown in the figure, the electronic device is enabled to perform each function or step in the above method embodiment.

[0344] The embodiment of the present application further provides a computer program product, which, when executed on a computer, enables the computer to execute the functions or steps executed by the mobile phone in the above method embodiment.

[0345] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0346] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0347] The units described as separate components may or may not be physically separate, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0348] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0349] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0350] The above content is only a specific embodiment of this application, but the scope of protection of this application is not limited to this. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A video generation method, characterized in that: The method comprises: Acquiring video description information, wherein the video description information is used to describe a scene of a target video and / or an action of a digital image in the target video; Based on the video description information, target scene materials and target animation materials are obtained from a preset material set, wherein the target scene materials include scene image materials, the scene image materials include foreground image materials and / or background image materials, and the target animation materials include image animation of the digital image moving based on the motion data; The target scene material and the target animation material are combined to obtain the target video.

2. The method according to claim 1, characterized in that Before acquiring the target scene material and the target animation material from the preset material set, the method further includes: The material set is constructed, wherein the materials in the material set include at least one of the image animation and the scene image material.

3. The method according to claim 2, characterized in that The constructing of the material set includes: Acquire at least one of the materials and a material description tag corresponding to the material, wherein the material description tag is used to identify a scene or action corresponding to the material; The material set is constructed based on the correspondence between the materials and the material description tags.

4. The method according to claim 3, characterized in that The acquiring of at least one of the materials and a material description tag corresponding to the material includes: generating at least one of the scene image materials according to a preset scene tag, and using the scene tag as the material description tag corresponding to the scene image material, wherein the scene tag is used to identify the scene corresponding to the scene image material; At least one of the image animations is generated according to a preset animation tag, and the animation tag is used as the material description tag corresponding to the image animation, wherein the animation tag is used to identify an action or an action occurrence scene in the image animation.

5. The method according to claim 4, characterized in that Generating at least one scene image material according to a preset scene tag includes: Inputting at least one scene description text corresponding to the scene label into a semantic model to obtain a scene description feature corresponding to the scene description text, wherein the scene description text is used to describe the scene identified by the scene label; The scene description feature is input into an image generation model to obtain the scene image material corresponding to the scene description feature.

6. The method according to claim 5, characterized in that Before inputting at least one scene description text corresponding to the scene label into the semantic model, the method further includes: For each scene category, set at least one scene label; For each of the scene tags, at least one scene description text is determined.

7. The method according to claim 4, characterized in that After generating at least one scene image material according to the preset scene tag, the method further includes: If the scene image material does not match a preset screening rule, the scene image material is discarded, wherein the screening rule includes: the scene image material matches the scene label, and / or the naturalness of the drawing of the scene image material meets a preset condition.

8. The method according to claim 4, characterized in that Generating at least one of the image animations according to the preset animation tag includes: Designing an action sequence according to at least one animation description text corresponding to the animation tag, wherein the animation description text is used to describe the action identified by the animation tag or the scene in which the action occurs; capturing the motion data generated by the user during the motion sequence; The motion data is redirected to the digital image, and the redirected motion data is adjusted to obtain the image animation.

9. The method according to claim 8, characterized in that Before designing an action sequence based on at least one animation description text corresponding to the animation tag, the method further includes: Setting at least one animation tag corresponding to an action category, wherein the animation tag corresponding to the action category is used to identify the action in the image animation; and / or, Setting at least one animation tag corresponding to a category of an action occurrence scene, wherein the animation tag corresponding to the category of the action occurrence scene is used to identify the action occurrence scene corresponding to the image animation; For each of the animation tags, at least one animation description text is determined.

10. The method according to claim 9, characterized in that The determining of at least one animation description text corresponding to the animation tag includes: Determining a character corresponding to the animation tag, wherein the character is a human character or an animal character; In the case that there are multiple characters, interactive actions between the multiple characters are determined, and the animation description text corresponding to the animation tag is determined according to the interactive actions.

11. The method according to claim 3, characterized in that After constructing the material set, the method further includes: Adding the material information of each material to the material set, wherein the material information includes at least one of the material description tag, material description text, index, default material identification bit, extension information, and material version number corresponding to the material; The material description text corresponding to the scene image material is a scene description text, and the material description text corresponding to the image animation is an animation description text.

12. The method according to claim 3, characterized in that After constructing the material set based on the correspondence between the material and the material description tag, the method further includes: constructing at least one material subset based on the scene image material and the image animation in the material set; A mapping relationship table is constructed based on the corresponding relationship between each of the material subsets and the material description tags.

13. The method according to claim 12, characterized in that The constructing a mapping relationship table based on the correspondence between each of the material subsets and the material description tags includes: For each of the material subsets, determining a correspondence between the material subset and the material description tag according to a correspondence between each of the materials in the material subset and the material description tag; The mapping relationship table is constructed according to the corresponding relationship between the material subsets and the material description tags.

14. The method according to claim 1, wherein The step of obtaining target scene material and target animation material from a preset material set includes: Determining a material description tag corresponding to the video description information; Based on a pre-built mapping relationship table, the target scene material and the target animation material corresponding to the material description tag are obtained from the material set.

15. The method according to claim 14, characterized in that The obtaining, based on the mapping relationship table, the target scene material and the target animation material corresponding to the material description tag from the material set includes: Based on the mapping relationship table, determining the material subset corresponding to the material description tag as the target material subset; In the target material subset, the target scene material and the target animation material are determined.

16. The method according to claim 15, characterized in that Determining the target scene material and the target animation material in the target material subset includes: Obtaining a first target index and a second target index, wherein the first target index is used to distinguish different scene image materials corresponding to the same scene tag, and the second target index is used to distinguish different image animations corresponding to the same animation tag; In the material subset, the scene image material corresponding to the first target index is determined to be the target scene material, and the image animation corresponding to the second target index is determined to be the target animation material.

17. The method according to claim 1, wherein The material set includes a scene image material set and an animation material set; and obtaining target scene material and target animation material from the preset material set includes: The target scene material is obtained from the scene image material set, and the target animation material is obtained from the animation material set.

18. The method according to claim 1, wherein Combining the target scene material and the target animation material to obtain the target video includes: Superimposing the target scene material and the target animation material in sequence according to a preset order to obtain the target video; According to the illumination change information and / or the camera movement trajectory information, each frame image in the target video is rendered respectively to obtain the rendered target video.

19. The method according to claim 1, wherein After obtaining the target video, the method further includes: The target video is set as the lock screen video of the electronic device.

20. The method according to claim 1, wherein Before acquiring the target scene material and the target animation material from the preset material set, the method further includes: Responding to a video generation instruction, acquiring the digital image; In response to the editing instruction, the digital image is edited.

21. The method according to claim 1, wherein The obtaining of video description information includes: The video description information is determined based on current scene information, wherein the current scene information includes at least one of current user state information, current date information, and user input information.

22. An electronic device, characterized in that: It comprises a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code includes computer instructions, and when the processor executes the computer instructions, the electronic device executes the video generation method according to any one of claims 1 to 21.

Citation Information

Patent Citations

  • Control method and device of mobile terminal and mobile terminal

    CN107239211A

  • Animation generation method, device and system and storage medium

    CN111968207A

  • Data processing method and related product

    CN116468827A

  • Text-based image generation method and device, electronic equipment and storage medium

    CN117493599A

  • Page display method and electronic equipment

    CN118295616A