A video generation method and electronic device
By generating dynamic lock screen videos based on digital avatars, the problem of lock screen interfaces being unable to achieve dynamic effects and individual element changes has been solved, improving user experience and the playability of personalized settings.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HONOR DEVICE CO LTD
- Filing Date
- 2024-09-06
- Publication Date
- 2026-05-12
AI Technical Summary
Digital avatar-based lock screen interfaces cannot achieve dynamic effects or modify individual elements within the lock screen, resulting in low playability and failing to meet users' personalized needs.
By acquiring video description information, and using preset materials, foreground image materials, background image materials, and digital avatar animations are selected and combined to generate the target video, achieving dynamic effects for the digital avatar and allowing changes to individual elements.
It implements dynamic effects and personalized settings for the lock screen interface, enhancing its playability and fun.
Smart Images

Figure CN120455755B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, and in particular to a video generation method and an electronic device. Background Technology
[0002] With technological advancements and the increasing prevalence of electronic devices, the importance of electronic device security is constantly rising. Screen lock is a crucial function that enhances the security of electronic devices, preventing unauthorized access. Specifically, when an electronic device is locked, the screen displays a lock screen interface, including a lock screen image and basic information such as the date and time. Users can verify their access by entering a password, pattern, fingerprint, or facial recognition. Only after successful verification can the user access the content or functions of the electronic device; otherwise, the screen remains locked.
[0003] The lock screen interface can be a natural landscape lock screen, an abstract art lock screen, a digital avatar lock screen, etc. The digital avatar lock screen interface contains a cartoon or anime-style digital avatar, which is usually brightly colored and cute, providing a pleasant visual experience.
[0004] However, lock screens based on digital avatars typically display basic information such as date and time on a static image containing the avatar. The static image results in fixed content, making dynamic effects impossible and leading to poor display quality. Furthermore, when editing elements in a lock screen based on cartoon avatars, only the entire static image can be replaced; individual elements cannot be modified, hindering personalized lock screen customization. Summary of the Invention
[0005] This application provides a video generation method and an electronic device to solve the problems that lock screen interfaces based on digital images cannot achieve dynamic effects and cannot modify individual elements in the lock screen interface.
[0006] To achieve the above objectives, in a first aspect, embodiments of this application provide a video generation method, the method comprising:
[0007] Obtain video description information, which describes the scene of the target video and / or the actions of the digital character in the target video; based on the video description information, obtain target scene material and target animation material from a preset material set, wherein the first target scene material includes scene image material, the scene image material includes foreground image material and / or background image material, and the target animation material includes image animation of the digital character moving based on motion data; combine the target scene material and the target animation material to obtain the target video.
[0008] The video generation method disclosed in this application first obtains video description information to describe the lock screen video. Then, based on the video description information, it selects foreground image material, background image material, and an animated image containing a digital avatar from a pre-set material set containing multiple materials, and combines them to generate the target video. Since the video generated by this application uses an animated image of a moving digital avatar, it can achieve dynamic effects based on the digital avatar. Furthermore, this application implements decoupled componentization of the various materials in the video, using multiple materials to combine to obtain the target video. Therefore, when it is necessary to change an element in the lock screen interface, only the material corresponding to that element needs to be changed. Customizing any material facilitates personalized settings of the lock screen interface, enhancing playability and fun.
[0009] In one implementation, before acquiring the target scene and target animation materials from a pre-set material set, the method further includes: constructing a material set, wherein the materials in the material set include at least one character animation and scene image material. Using this implementation, multiple different materials can be generated in advance to construct the material set. When using the combination of materials to obtain a video, it is only necessary to select the pre-generated materials from the material set.
[0010] In one implementation, constructing a media set includes: obtaining at least one media clip and its corresponding description tag, wherein the description tag identifies the scene or action corresponding to the media clip; and constructing a media set based on the correspondence between the media clip and its description tag. Using this implementation, the pre-constructed media set includes not only the media clips but also the correspondence between the media clips and their description tags. Subsequently, based on the description tags, the corresponding media clip can be found in the media set, resulting in a higher degree of matching between the video generated based on that media clip and the video description information.
[0011] In one implementation, obtaining at least one source material and its corresponding material description tag includes: generating at least one scene image material based on preset scene tags, and using the scene tags as the material description tags corresponding to the scene image material, wherein the scene tags are used to identify the scene corresponding to the scene image material; generating at least one animated character based on preset animation tags, and using the animation tags as the material description tags corresponding to the animated character, wherein the animation tags are used to identify the action or the scene in which the action occurs in the animated character. Using this implementation, scene image materials are generated using scene tags, and animated character pieces are generated using animation tags. Since the source material is generated based on tags, the correspondence between tags and source material can be guaranteed. Therefore, based on scene tags, scene image materials that meet the requirements can be found; based on animation tags, animated character pieces that meet the requirements for action or the scene in which the action occurs can be found, ensuring that the scene and action of the combined target video both meet the requirements of the video description information.
[0012] In one implementation, at least one scene image asset is generated based on preset scene tags. This includes: inputting at least one scene description text corresponding to the scene tag into a semantic model to obtain scene description features corresponding to the scene description text, wherein the scene description text describes the scene identified by the scene tag; and inputting the scene description features into an image generation model to obtain scene image assets corresponding to the scene description features. This implementation utilizes a semantic model and an image generation model to generate foreground image assets and / or background image assets. Through the above generative artificial intelligence model, a large number of foreground image assets and / or background image assets can be quickly obtained, which is beneficial to improving the production efficiency of asset resources, while also reducing the need for designers and lowering costs.
[0013] In one implementation, before inputting at least one scene description text corresponding to a scene tag into the semantic model, the method further includes: setting at least one scene tag for each scene category; and determining at least one scene description text for each scene tag. Using this implementation, multiple scene tags can be set for each scene category. Since scene tags are used to identify scenes, each scene category can correspond to multiple different scenes. The resulting scene image materials can cover multiple different scenes, providing users with a wider range of choices. Furthermore, determining multiple scene description texts for each scene tag allows for a more comprehensive and detailed description of the scene corresponding to the scene tag, improving the matching degree between the generated foreground and / or background image materials and the scene tag. Moreover, setting multiple scene description texts also facilitates the generation of a larger number of foreground and / or background image materials.
[0014] In one implementation, after generating at least one scene image material based on preset scene tags, the method further includes: discarding scene image materials if they do not match preset filtering rules, wherein the filtering rules include: the scene image material matches the scene tag, and / or the naturalness of the scene image material's drawing meets preset conditions. Using this implementation, by discarding substandard materials after generating scene image materials, the quality of the generated target video can be guaranteed.
[0015] In one implementation, at least one avatar animation is generated based on preset animation tags. This includes: designing an action sequence based on at least one animation description text corresponding to the animation tag, wherein the animation description text describes the action or action scene identified by the animation tag; capturing motion data generated by the user moving according to the action sequence; redirecting the motion data to the digital avatar; and adjusting the redirected motion data to obtain the avatar animation. Using this implementation, motion data generated by the user's movements is captured using motion capture technology and redirected to the digital avatar to obtain the avatar animation. In the avatar animation generated in this way, the movements of the digital avatar are more natural and realistic, which helps improve the viewing experience for video viewers. Furthermore, generating avatar animations through motion capture technology eliminates the need for frame-by-frame drawing and modeling, thus improving the generation efficiency of avatar animations.
[0016] In one implementation, before designing the action sequence based on at least one animation description text corresponding to an animation tag, the method further includes: setting at least one animation tag corresponding to an action category, wherein the animation tag corresponding to the action category is used to identify the action in the character animation; and / or, setting at least one animation tag corresponding to an action occurrence scene category, wherein the animation tag corresponding to the action occurrence scene category is used to identify the action occurrence scene corresponding to the character animation; and determining at least one animation description text for each animation tag. Using this implementation, action categories and animation scene categories are designed for the character animation, and multiple animation tags are set for each action category and animation scene category. Since the animation tag of the action occurrence scene category is used to identify the action occurrence scene, the action occurrence scene category can correspond to multiple different action occurrence scenes. The resulting character animation can cover multiple different action occurrence scenes, providing users with a wider range of choices. Since the animation tag of the action category is used to identify the action of the digital character, the action category can correspond to multiple different actions. The resulting character animation can cover multiple different actions, providing users with a wider range of choices. Furthermore, this implementation defines multiple animation description texts for each animation tag, enabling a more comprehensive and detailed description of the scene or action corresponding to the animation tag, thereby improving the matching degree between the generated animated character and the animation tag. Additionally, setting multiple animation description texts also facilitates the generation of a larger number of animated characters.
[0017] In one implementation, determining at least one animation description text corresponding to an animation tag includes: determining the character corresponding to the animation tag, wherein the character is a human character or an animal character; if there are multiple characters, determining the interactive actions between the multiple characters, and determining the animation description text corresponding to the animation tag based on the interactive actions. Using this implementation, the action description text can describe the interactive actions between multiple characters. Through the design of interactive actions, it is possible to generate animated images of multiple characters interacting, ultimately resulting in a target video of multiple digital characters interacting. Furthermore, the characters are not limited to human characters; they can also be animal characters, increasing fun and playability.
[0018] In one implementation, after constructing the resource set, the method further includes adding resource information for each resource to the resource set. This resource information includes at least one of the following: a resource description tag, resource description text, index, default resource identifier, extended information, and resource version number. Specifically, the resource description text for scene image resources is the scene description text, and the resource description text for character animations is the animation description text. This implementation adds detailed resource information to each resource, facilitating resource management. Furthermore, the index can improve retrieval efficiency, and the resource version number enables version control and tracking. Setting up resource information allows for better utilization of resource resources.
[0019] In one implementation, after constructing a material set based on the correspondence between the materials and their description tags, the method further includes: constructing at least one material subset based on the scene image materials and character animations in the material set; and constructing a mapping table based on the correspondence between each material subset and its description tags. This implementation constructs multiple material subsets and determines the mapping relationship between each subset and its description tags. Therefore, based on the mapping table, the corresponding material subset can be found through its description tags, and the target video can then be generated using the materials in the subset.
[0020] In one implementation, a mapping table is constructed based on the correspondence between each subset of materials and its description tags. This includes: for each subset of materials, determining the correspondence between the subset and its description tags based on the correspondence between each material and its description tags; and constructing the mapping table based on the correspondence between the subset and its description tags. This implementation transforms the correspondence between materials and their description tags into a correspondence between subsets of materials and their description tags. Therefore, when searching for target scene materials and target animation materials, it is only necessary to determine the correspondence between the subset and the description tags, rather than determining the correspondence between each individual material and its description tag. Thus, this implementation simplifies the material search process, allowing for quick retrieval and application of the required materials.
[0021] In one implementation, target scene materials and target animation materials are obtained from a pre-defined material set, including: determining the material description tags corresponding to the video description information; and obtaining the target scene materials and target animation materials corresponding to the material description tags from the material set based on a pre-built mapping table. This implementation method, by using a table lookup approach, retrieves target scene materials and target animation materials from the material set, which is simple, fast, and helps improve the efficiency of target video generation.
[0022] In one implementation, based on a mapping table, target scene materials and target animation materials corresponding to material description tags are obtained from the material set. This includes: determining, based on the mapping table, a subset of materials corresponding to material description tags as a subset of target materials; and within the subset of target materials, determining the target scene materials and target animation materials. Using this implementation, determining the target scene materials and target animation materials within the subset of target materials corresponding to material description tags ensures that the target scene materials and target animation materials are the materials corresponding to the material description tags.
[0023] In one implementation, determining target scene assets and target animation assets within a subset of target assets includes: obtaining a first target index and a second target index, wherein the first target index is used to distinguish different scene image assets corresponding to the same scene tag, and the second target index is used to distinguish different character animations corresponding to the same animation tag; within the asset subset, the scene image asset corresponding to the first target index is determined as the target scene asset, and the character animation corresponding to the second target index is determined as the target animation asset. This implementation, by using the first and second target indices to search for target scene assets and target animation assets within the target asset subset, improves the efficiency of determining target scene assets and target animation assets.
[0024] In one implementation, the resource set includes a scene image resource set and an animation resource set. Retrieving target scene resources and target animation resources from the preset resource set includes: retrieving target scene resources from the scene image resource set and retrieving target animation resources from the animation resource set. Using this implementation, scene image resources and animation resources are stored separately in the scene image resource set and the animation resource set, respectively. This facilitates systematic management of the resources, allows for quick location of target scene resources and target animation resources, and also facilitates archiving and backup operations.
[0025] In one implementation, target scene footage and target animation footage are combined to obtain a target video. This includes: sequentially overlaying the target scene footage and target animation footage in a preset order to obtain the target video; and rendering each frame of the target video based on lighting change information and / or camera movement trajectory information to obtain the rendered target video. Using this implementation, sequentially overlaying the target scene footage and target animation footage helps to obtain a video with depth, making the target video more three-dimensional and realistic. After obtaining the stereoscopic video, rendering each frame of the video further improves video quality and enhances the user's visual experience.
[0026] In one implementation, after obtaining the target video, the method further includes setting the target video as the lock screen video of the electronic device. This implementation provides users with lock screen videos containing dynamic digital avatars, improving the user experience.
[0027] In one implementation, before acquiring target scene and target animation materials from a preset material set, the method further includes: acquiring a digital avatar in response to a video generation command; and editing the digital avatar in response to an editing command. This implementation first acquires a default digital avatar, and then edits it based on the user's editing commands, which increases user engagement and enhances the playability of the target video. Furthermore, the user-customized appearance of the digital avatar better aligns with user aesthetics.
[0028] In one implementation, obtaining video description information includes: determining video description information based on current scene information, wherein the current scene information includes at least one of current user state information, current date information, and user input information. By using this implementation, determining video description information based on the current scene results in a target video that better matches the current scene, providing the user with an immersive experience.
[0029] Secondly, this application also provides an electronic device, including a memory and a processor; the memory and the processor are coupled; wherein the memory is used to store computer program code, the computer program code including computer instructions, and when the processor executes the computer instructions, it causes the electronic device to perform the video generation method as described in the first aspect and any implementation thereof.
[0030] Thirdly, this application also provides a chip system, which includes a processor; the processor is coupled to a memory, the memory being used to store computer program code, the computer program code including computer instructions, and when the processor executes the computer instructions, the video generation method in the first aspect and any of the implementations described above is executed.
[0031] Fourthly, this application also provides a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the video generation method as described in the first aspect and any of its implementations above.
[0032] Fifthly, this application also provides a computer program product, which includes: a computer program or instructions that, when executed on a computer, cause the computer to perform the video generation method as described in the first aspect and any of its implementations above.
[0033] Understandably, the beneficial effects that the technical solutions provided in the second to fifth aspects described above can be achieved by referring to the beneficial effects of the first aspect and any of its optional implementation methods, which will not be repeated here. Attached Figure Description
[0034] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0035] Figure 1 It is a lock screen display effect based on digital images;
[0036] Figure 2 This is a diagram illustrating changes to elements in a lock screen interface;
[0037] Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;
[0038] Figure 4 This is a schematic diagram of the software structure of the electronic device provided in the embodiments of this application;
[0039] Figure 5 This is the first flowchart of the video generation method provided in the embodiments of this application;
[0040] Figure 6 This is the second flowchart of the video generation method provided in the embodiments of this application;
[0041] Figure 7 This is a schematic diagram of the process for generating scene image materials provided in an embodiment of this application;
[0042] Figure 8 This is a schematic diagram of the process for generating an animated image provided in an embodiment of this application;
[0043] Figure 9 This is the third flowchart of the video generation method provided in the embodiments of this application;
[0044] Figure 10 This is the fourth flowchart of the video generation method provided in the embodiments of this application;
[0045] Figure 11 This is the fifth flowchart of the video generation method provided in the embodiments of this application;
[0046] Figure 12 This is the sixth flowchart of the video generation method provided in the embodiments of this application;
[0047] Figure 13 This is a schematic diagram showing the overlay of target scene materials and target animation materials provided in the embodiments of this application;
[0048] Figure 14 This is a schematic diagram illustrating how a target video is obtained by overlaying target scene material and target animation material according to an embodiment of this application;
[0049] Figure 15 This is the seventh flowchart of the video generation method provided in the embodiments of this application;
[0050] Figure 16 This is the eighth flowchart of the video generation method provided in the embodiments of this application;
[0051] Figure 17 This is a schematic diagram of a lock screen video setting provided in an embodiment of this application;
[0052] Figure 18 This is a schematic diagram of a video generation device provided in an embodiment of this application. Detailed Implementation
[0053] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are all within the protection scope of this application.
[0054] In the following description, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0055] Furthermore, in this application, directional terms such as "upper," "lower," "inner," and "outer" are defined relative to the indicated placement of the components in the accompanying drawings. It should be understood that these directional terms are relative concepts, used for relative description and clarification, and can change accordingly depending on the placement of the components in the accompanying drawings.
[0056] The following explanations of the technical terms mentioned in the embodiments of this application are provided to facilitate understanding by those skilled in the art.
[0057] A digital avatar is a virtual character created using digital technology. Digital avatars can be two-dimensional (such as cartoon characters) or three-dimensional (such as 3D animated characters), and can be applied to environments such as animation, games, virtual reality (VR), and augmented reality (AR). In this embodiment, a digital avatar can be used in lock screen videos, so that when the electronic device is in lock screen mode, it can display a dynamic lock screen effect containing a digital avatar.
[0058] The application scenarios of the embodiments of this application will be described below with reference to the accompanying drawings.
[0059] With technological advancements and the increasing prevalence of electronic devices, the importance of electronic device security is constantly rising. Screen lock is a crucial function that enhances the security of electronic devices, preventing unauthorized access. Specifically, when an electronic device is locked, the screen displays a lock screen interface, including a lock screen image and basic information such as the date and time. Users can verify their permissions by entering a password, pattern, fingerprint, or facial recognition. Only after successful verification can the user unlock the electronic device and access its content or functions; otherwise, the screen remains locked.
[0060] The lock screen interface can be a natural landscape lock screen, an abstract art lock screen, a digital avatar lock screen, etc. The digital avatar lock screen interface contains a cartoon or anime-style digital avatar, which is usually brightly colored and cute, providing a pleasant visual experience.
[0061] Figure 1 It is a lock screen display effect based on digital images.
[0062] like Figure 1 As shown in (a) above, in the locked state, the electronic device displays a lock screen interface 10 containing digital avatars. Figure 1 As shown in (b), when the electronic device receives a user's touch, swipe, or other operation, it can display an unlock interface 20, which shows a numeric keypad for entering the lock screen password. Figure 1As shown in (c), after the user enters the correct lock screen password via the numeric keypad, the verification is successful, and the electronic device unlocks. It is understood that this embodiment only uses password unlocking as an example; in practical applications, patterns or other methods can also be used for unlocking, and the lock screen display effect will change accordingly.
[0063] However, this digitally based lock screen interface has the following problems.
[0064] Currently, lock screen interfaces based on digital avatars typically display basic information such as date and time on a static image containing a digital avatar. Due to the limitations of static images, only fixed lock screen content can be displayed on the screen, making dynamic effects impossible. Therefore, when the screen is locked, users can only see a static screen, resulting in a poor viewing experience.
[0065] Furthermore, since digital avatar-based lock screens consist of a single static image and basic information such as date and time, editing a specific element requires replacing the entire static image. This method of replacing the entire image cannot modify individual elements within the lock screen, hindering personalized customization.
[0066] Figure 2 This is a diagram illustrating changes to elements in a lock screen interface.
[0067] Figure 2 (a) in the diagram is a schematic of the lock screen interface before the elements were changed; Figure 2 (b) in the diagram is a schematic of the lock screen interface after the elements have been changed.
[0068] like Figure 2 As shown in (a), the lock screen image is a static image containing a digital avatar and a willow tree, with the avatar standing under the willow tree. In the lock screen state, the electronic device displays a lock screen interface that includes the digital avatar and the willow tree element. If it is necessary to change the willow tree in the lock screen interface to a pine tree, the current lock screen image must be replaced with an image containing both a cartoon digital avatar and a pine tree. Figure 2 As shown in (b), replacing the lock screen image with a static image containing a cartoon avatar and a pine tree, with the static image showing the avatar standing under the pine tree, will result in a different lock screen interface displayed on the electronic device's screen in the lock screen state. This lock screen interface will include the pine tree element in addition to the avatar.
[0069] In conclusion, current digital avatar-based lock screen interfaces lack dynamic effects and the ability to freely modify elements within the screen. Therefore, they offer limited playability, fail to meet users' personalized needs, negatively impact user experience, and lack appeal.
[0070] To address the aforementioned issues, this application provides a video generation method for generating lock screen videos based on digital avatars.
[0071] The video generation method provided in this application first obtains video description information describing the lock screen video. Then, based on the video description information, it selects foreground image material, background image material, and an animated image containing a digital avatar from a pre-set material set containing multiple materials, and combines them to generate the target video. Since the video generated by this application uses an animated image of a moving digital avatar, it can achieve dynamic effects based on the digital avatar. Furthermore, since the video generated by this application is a combination of multiple materials, when it is necessary to change an element in the lock screen interface, only the material corresponding to that element needs to be changed. For example, for a lock screen interface containing a digital avatar and a willow tree, if it is necessary to change the willow tree in the lock screen interface to a pine tree, only the material corresponding to the willow tree needs to be replaced with the material corresponding to the pine tree. Therefore, this application embodiment is more conducive to the personalized settings of the lock screen interface, enhancing playability and fun.
[0072] The video generation method provided in this application can be applied to electronic devices. In some embodiments, the electronic device may be a mobile phone, tablet computer, handheld computer, personal computer (PC), ultra-mobile personal computer (UMPC), netbook, as well as cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, vehicle-mounted device, and other mobile terminals. This application does not impose any special restrictions on the specific type of electronic device.
[0073] Figure 3 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.
[0074] like Figure 3As shown, the electronic device 100 may include a processor 110, a memory 120, an antenna 1, an antenna 2, a mobile communication module 130, a wireless communication module 140, a sensor module 150, a display screen 160, etc. The sensor module 150 may include a pressure sensor 150A, a touch sensor 150B, a fingerprint sensor 150C, an ambient light sensor 150D, a temperature sensor 150E, a gyroscope sensor 150F, a proximity sensor 150G, etc. In this embodiment, video generation commands from the user can be received through the pressure sensor 150A and the touch sensor 150B.
[0075] It is understood that the structures illustrated in the embodiments of the present invention do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0076] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.
[0077] The memory 120 can be used to store computer executable program code, including instructions. The memory 120 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of the foldable electronic device 100 (such as audio data, phonebook, etc.). Furthermore, the memory 120 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the foldable electronic device 100 by running instructions stored in the memory 120 and / or instructions stored in memory located within the processor.
[0078] In this embodiment, the code implementing the video generation method of this embodiment can be stored in non-volatile memory. When generating video, the electronic device 100 can load the executable code stored in the non-volatile memory into random access memory.
[0079] The wireless communication function of electronic device 100 can be implemented through antenna 1, antenna 2, mobile communication module 130, wireless communication module 14, modem, and baseband processor.
[0080] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be reused to improve antenna utilization.
[0081] The mobile communication module 130 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 130 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 130 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 130 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 130 and at least some modules of the processor 110 may be housed in the same device.
[0082] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device or displays an image or video through the display screen 160. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 130 or other functional modules.
[0083] Video codecs are used to compress or decompress digital video. Terminal 100 may support one or more video codecs. Thus, terminal 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG 2, MPEG 3, MPEG 4, etc. In this embodiment, the video codec can be used to generate video and display the generated video on the display screen 160 of electronic device 100.
[0084] The wireless communication module 140 can provide solutions for wireless communication applications in electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 140 can be one or more devices integrating at least one communication processing module. The wireless communication module 140 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 140 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.
[0085] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 130, and antenna 2 is coupled to wireless communication module 140, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.
[0086] Electronic device 100 implements display functions through a GPU, a display screen 160, and an application processor. The GPU is a microprocessor for image processing, connecting the display screen 160 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.
[0087] In this embodiment, the electronic device 100 implements the video generation method provided in this embodiment, which mainly relies on the video codec and the image computing and processing capabilities provided by the GPU.
[0088] Display screen 160 is used to display images, videos, etc. Display screen 160 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 100 may include one or N displays 160, where N is a positive integer greater than 1.
[0089] In this embodiment, the ability of the electronic device 100 to display lock screen video depends on the display functions provided by the GPU, display screen 160, and application processor.
[0090] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments of this application, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.
[0091] Generally, the implementation of the functions of electronic device 100 requires not only hardware support but also software cooperation. The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment takes the layered architecture Android system as an example to illustrate the software structure of electronic device 100.
[0092] Figure 4This is a schematic diagram of the software structure of the electronic device provided in the embodiments of this application.
[0093] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android Runtime and system libraries, and the kernel layer.
[0094] The application layer can include a series of application packages.
[0095] like Figure 4 As shown, the application package may include applications such as camera, calendar, map, WLAN, music, SMS, gallery, call, navigation, video, lock screen, etc.
[0096] Lock screen applications are applications that provide a lock screen interface. When an electronic device is locked by a lock screen application, the screen of the electronic device displays the lock screen interface. Users must unlock the screen by entering a password, pattern, fingerprint recognition, or facial recognition in order to access the content and functions on the electronic device.
[0097] The lock screen application can provide a lock screen resource download path, allowing users to download corresponding lock screen images or videos and set the electronic device's lock screen interface using these images or videos. Furthermore, the lock screen application can also provide the function of generating lock screen images or videos. In this embodiment, the lock screen application can generate a video matching the video description information using materials from a resource set, based on the video description information, as the lock screen video for the electronic device.
[0098] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.
[0099] like Figure 4 As shown, the application framework layer may include a window manager, content provider, phone manager, resource manager, notification manager, view system, lock screen management service, etc.
[0100] Content providers are used to store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, etc.
[0101] The lock screen management service provides download paths for lock screen resources and also manages and sets lock screen resources. It can include functions such as lock screen image library, lock screen video library, personalized recommendations, download and settings, community and user uploads.
[0102] The Android Runtime consists of core libraries and a virtual machine. The Android Runtime is responsible for the scheduling and management of the Android system.
[0103] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0104] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0105] System libraries can include multiple functional modules. For example: surface manager, 3D graphics processing library (e.g., OpenGL ES (OpenGL for Embedded Systems)), 2D graphics engine (e.g., SGL (Simple Graphics Library)), media libraries, etc.
[0106] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.
[0107] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0108] A 2D graphics engine is a graphics engine for 2D drawing.
[0109] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.
[0110] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.
[0111] The following example, using a lock screen video setting scenario, illustrates the workflow of the software and hardware of electronic device 100.
[0112] First, the lock screen application calls the lock screen management service in the application framework layer to launch the lock screen application. Then, the lock screen application calls the 2D graphics engine and / or 3D graphics processing library in the system library to generate a target video based on a digital avatar. After the lock screen application sets the target video as the electronic device's lock screen video, it can call the media library in the system library to play the lock screen video while the device is locked.
[0113] Figure 5 This is a flowchart of one embodiment of the video generation method provided in this application.
[0114] like Figure 5 As shown, this embodiment includes the following steps S11-S15.
[0115] S11: Construct a resource set.
[0116] The materials in the material collection include at least one character animation and scene image material.
[0117] Electronic devices can construct a media set using multiple different types of materials. It is understood that the target video is generated by combining animated characters and scene image materials. Therefore, the materials in the media set include animated characters, as well as at least one of foreground image materials and background image materials, to ensure that there are sufficient materials to choose from. It should be noted that the media set can be constructed by an electronic device, a server, or other devices; this application does not limit this.
[0118] The embodiments of this application pre-construct a material set including multiple materials, so that when generating the target video, the pre-generated materials can be selected from the material set, thereby improving the efficiency of video generation.
[0119] S12: Construct a mapping relationship table based on the materials in the material set.
[0120] Electronic devices can construct a mapping table to locate the material corresponding to the video description information in subsequent steps. For example, this can be done by constructing key-value pairs, where each material is the value, and the material's description or ID can be the corresponding key. For instance, for a scene image material containing a water glass, the key can be set to "glass". Then, the scene image material can be found using "glass". It should also be noted that the mapping table can be constructed by the electronic device, a server, or other devices; this application does not limit this.
[0121] This application embodiment constructs a mapping table for materials, which allows materials in the material set to be found directly, improving the efficiency of material retrieval. Furthermore, the structure of the mapping table is typically flexible, making it easy to add or modify mapping relationships when the material set is updated or expanded.
[0122] S13: Obtain video description information.
[0123] Electronic devices can acquire video description information to describe a target video, facilitating the generation of a target video corresponding to the video description information in subsequent steps. This video description information can be used to describe the scene of the target video and / or the actions of the digital characters within it. The video description information can be information input by the user through text or voice, or it can be determined by the electronic device based on the user's current state, the current environment, etc. It can be understood that video description information can be used to describe the scene of the target video to be generated and the actions of the digital characters within it. For example, video description information such as "birthday" and "Mid-Autumn Festival" describes the scene of the target video, while video description information such as "listening to music" describes the actions of the digital characters in the target video. It can also be understood that video description information can simultaneously describe both the scene and the actions of the digital characters; for example, "eating mooncakes on Mid-Autumn Festival" describes both the scene of the target video as the Mid-Autumn Festival and the action of the digital characters in the target video as eating mooncakes.
[0124] S14: Based on the video description information and mapping relationship table, obtain the target scene material and target animation material from the preset material set.
[0125] The electronic device acquires target scene footage and target animation footage that match the content described in the video description information. For example, a footage set containing multiple footage items can be pre-set, and then the target scene footage and target animation footage can be selected from the footage set by querying a mapping table.
[0126] The target scene materials include scene image materials, which include foreground image materials and / or background image materials. Foreground and background image materials are used to form different layers in the target video. Foreground image materials are the layers in the video that are closer to the viewer, usually located in the foreground of the video, and can be elements such as objects and plants; background image materials are the layers in the video that are farther away from the viewer, usually located in the background of the video, and can be primary colors such as walls and sky.
[0127] The target animation assets include digital avatars that move based on motion data. In the animation, the digital avatar performs specified actions according to the motion data. The digital avatars in the animation are pre-set avatars, which can be cartoon or anime style. The actions corresponding to the motion data in the animation match the video description information. For example, if the video description information is "birthday," then the motion data could be data corresponding to actions such as blowing out candles or eating cake.
[0128] Furthermore, it's understandable that in the target video, the digital avatar needs to perform specified actions; therefore, the animation of the avatar is dynamic. Elements in the scene image assets can be dynamic (such as twinkling stars, drifting clouds, etc.) or static (such as a stationary table, etc.). Therefore, scene image assets can be dynamic video or static images.
[0129] Optionally, in one implementation, the target scene assets and target animation assets may also include other types of assets such as audio assets. For example, for a birthday scene, the target scene assets may include background image assets displaying the background of the birthday party, foreground image assets displaying elements such as the birthday cake, and may also include audio assets as background music for playing the happy birthday song.
[0130] This application's embodiments determine target scene materials and target animation materials based on video description information, ensuring that the generated target video matches the content described in the video description information. Furthermore, using a mapping table to determine target scene materials and target animation materials simplifies the material search logic, thus improving query efficiency.
[0131] S15: Combine the target scene materials and target animation materials to obtain the target video.
[0132] The electronic device combines the target scene materials and target animation materials obtained in the aforementioned steps to obtain the target video. For example, the target scene materials and target animation materials can be superimposed in a certain order. For instance, background image materials, animated characters, and foreground image materials can be superimposed in order of distance from the viewer, from farthest to closest, to obtain the target video.
[0133] This application's embodiments generate a target video by combining target scene materials and target animation materials, thus achieving decoupling between different materials in the target video. When a user needs to modify the content in the target video, they only need to modify the material corresponding to that content.
[0134] In one implementation, step S15 may be followed by the following steps:
[0135] S16: Set the target video as the lock screen video for the electronic device.
[0136] Electronic devices set a target video as their lock screen video, enabling the device to play the target video while locked, thus providing users with a personalized visual experience.
[0137] Optionally, in one implementation, after step S15, the target video can be saved locally in the lock screen resource library of the electronic device, so that it can be retrieved from the lock screen resource library at any time as the lock screen video. It can also be uploaded to the lock screen resource community of the lock screen management service, allowing users to communicate and share their created videos with others, increasing social interaction. If the generated target video is not satisfactory, the user can exit directly without saving the target video.
[0138] The following is combined Figures 6-8 The technical solutions involved in step S1 will be further described.
[0139] Figure 6 This is a flowchart of another embodiment of the video generation method provided in this application.
[0140] like Figure 6 As shown, in this embodiment, the aforementioned step S11 may include the following steps S111-S113:
[0141] S111: Obtain at least one material and its corresponding material description tag.
[0142] The material description tag is used to identify the scene or action corresponding to the material.
[0143] The electronic device identifies the corresponding material description tags, which can then be used in subsequent steps to construct a material set. Afterward, the corresponding material can be found within the material set using its description tags.
[0144] Multiple different assets can be pre-set for the same scene or action, and asset description tags can identify the scene or action corresponding to the asset. Therefore, different assets can correspond to the same scene or action, that is, to the same asset description tag. For example, an asset containing a foreground image of a full moon can be labeled "Mid-Autumn Festival"; an asset containing a foreground image of a mooncake can also be labeled "Mid-Autumn Festival". For example, an animation of a digital character waving its right hand can be labeled "greeting"; an animation of a digital character waving its left hand can also be labeled "greeting".
[0145] This application embodiment obtains the materials and their corresponding material description tags so that the required materials can be found based on the material description tags in subsequent steps.
[0146] In one implementation, step S111 may include the following steps:
[0147] Step 1: Generate at least one scene image asset based on the preset scene tags. In addition, the scene tags can be used as descriptive tags for the corresponding scene image assets.
[0148] Scene tags are used to identify the scene corresponding to the scene image material.
[0149] Taking foreground image material as an example, foreground image material is first generated based on pre-set scene tags. Since scene tags can identify the scene corresponding to the foreground image material, the scene image material generated based on the scene tags contains elements from the specified scene. For example, for a birthday scene, the scene tag could be "birthday". The foreground image material generated based on the birthday scene tag could contain elements that typically appear when celebrating a birthday, such as a birthday cake.
[0150] It is understandable that since the foreground image material is generated based on preset scene tags, the elements displayed in the foreground image material match the scene tags. Based on this, the scene tags can be used as the material description tags corresponding to the foreground image material. For example, in this step, the foreground image material generated based on the birthday scene tag displays elements such as a birthday cake that match the birthday scene. Therefore, the birthday scene tag can be used as the material description tag corresponding to this foreground image material to establish a correspondence between the birthday scene tag and the foreground image material. Afterwards, the foreground image material can be found in the material library using the birthday scene tag.
[0151] Furthermore, to increase the amount of material in the resource library and provide a wider range of choices for the same target video description information, multiple foreground image materials can be generated for the same scene tag. When generating the target video, users can choose foreground image materials according to their preferences, which is conducive to the personalized design of the target video; they can also randomly select from foreground image materials that meet the requirements or select according to certain rules, bringing freshness to users and avoiding visual fatigue caused by looking at the same target video.
[0152] This application generates multiple scene image materials so that a target scene material can be selected from the generated scene image materials in subsequent steps. Furthermore, since the scene image materials are generated using scene tags, only the scene tags need to be limited to obtain the desired scene image materials.
[0153] The generation method and scene label determination method for background image materials are similar to those for foreground image materials, and will not be repeated here.
[0154] The following is combined Figure 7The process of generating scene image materials will be further introduced.
[0155] Figure 7 This is a schematic diagram of the process for generating scene image materials provided in the embodiments of this application.
[0156] like Figure 7 As shown, firstly, multiple scene tags are defined for each scene category, with each scene tag corresponding to a scene. This method achieves scene definition. Then, several scene description texts are defined for each scene. In some application scenarios, the scene description texts are input into an artificial intelligence model that includes a semantic model and an image generation model, resulting in scene image materials output by the AI model. In other application scenarios, scene image materials are created manually based on the scene description texts. Finally, a foreground image material library is constructed using foreground image materials, and a background image material library is constructed using background image materials. The resulting foreground and background image material libraries store several foreground and background image materials respectively, which can be used as part of a material set.
[0157] The following is about Figure 7 To further describe the process, in one implementation, the aforementioned first step may include the following steps:
[0158] For each scene category, set at least one scene label. For each scene label, determine at least one scene description text. Input the at least one scene description text corresponding to the scene label into a semantic model to obtain the scene description features corresponding to the scene description text. Input the scene description features into an image generation model to obtain the scene image materials corresponding to the scene description features.
[0159] In some embodiments, multiple scene categories can be predefined, and scene labels can be set for each scene category. The scene label is a brief scene description.
[0160] For example, the following scenario categories can be defined: personal-specific scenario category, activity operation scenario category, personal status scenario category, natural scenario category, and daily general scenario category.
[0161] For specific personal scenarios, the following scenario tags can be set: birthday scenario tags, anniversary scenario tags, etc. For event / operation scenario categories, the following scenario tags can be set: holiday scenario tags, event scenario tags, spring scenario tags, summer scenario tags, autumn scenario tags, winter scenario tags, etc. For personal status scenario categories, the following scenario tags can be set: life service scenario tags, music listening scenario tags, exercise scenario tags, work scenario tags, rest scenario tags, etc. For nature scenario categories, the following scenario tags can be set: weather scenario tags, sunrise / sunset scenario tags, etc.
[0162] It should be noted that the embodiments of this application only take the aforementioned scene categories and scene tags as examples. In some implementations, other scene categories and scene tags can also be defined, and this application does not limit this.
[0163] This application's embodiments design a variety of different scene categories and scene tags, which can cover a variety of scenarios and meet the personalized needs of different customers.
[0164] In some embodiments, when determining the scene description text for each scene tag, a relatively detailed scene description can be designed for the scene identified by the scene tag and used as the scene description text.
[0165] For example, for the birthday scene tag, the following scene description text can be set: birthday cake, birthday candles, balloons, gifts, etc. For the spring scene tag, the following scene description text can be set: flowers, grass, trees, picnic basket, outdoor barbecue grill, kites, etc. For the music listening tag, the following scene description text can be set: headphones, record player, cassette tape, music poster, musical instrument model, sheet music, etc.
[0166] It should be noted that the embodiments of this application only take the aforementioned scene tags and scene description text as examples. In some implementations, other scene description texts can also be defined.
[0167] It should be noted that the foregoing embodiments only provide one method for determining scene tags and scene description text. In some implementations, other methods can also be used to determine scene tags and scene description text. For example, models such as GPT (Generative Pre-trained Transformer) can be used to generate scene tags and scene description text; alternatively, historical lock screen interfaces from multiple users can be collected, statistically analyzed, and clustered to obtain scene tags and scene description text. It is understood that the above is merely an exemplary description of the method for determining scene tags and scene description text; in practical applications, other methods can also be used, and this application does not limit them.
[0168] This application's embodiments design multiple different scene description texts for each scene tag, enhancing the diversity of scene image materials. In some embodiments, taking foreground image materials as an example, when obtaining scene description features corresponding to the scene description text using a semantic model, the scene description text can be input into semantic models capable of recognizing the meaning expressed by the text, such as BERT (Bidirectional Encoder Representations from Transformers) or GPT. The semantic model then extracts the corresponding scene description features from the scene description text. These scene description features can be used to generate corresponding images or videos as foreground image materials.
[0169] It is understandable that, for a given scene label, the content described by different scene description texts can be mutually exclusive or appear simultaneously. For example, for the aforementioned birthday scene label, the content described by scene description texts such as birthday cake, birthday candles, balloons, and gifts can appear simultaneously. Therefore, when generating foreground image material, scene description texts such as birthday cake, birthday candles, balloons, and gifts can be simultaneously input into the semantic model. In this way, the semantic description features simultaneously contain information about birthday cake, birthday candles, balloons, and gifts. Therefore, the final foreground image material can simultaneously contain elements such as birthday cake, birthday candles, balloons, and gifts.
[0170] In addition, scene descriptions such as birthday cake, birthday candles, balloons, and gifts can be input into the semantic model separately. This results in multiple foreground image assets, each containing one of the following: birthday cake, birthday candles, balloons, or a gift. It can be understood that by generating foreground image assets separately for each element, any number of suitable foreground image assets can be used when generating the target video, resulting in a wider variety of combinations to choose from. For example, foreground image assets containing a birthday cake and foreground image assets containing birthday candles can be used to obtain a target video containing both a birthday cake and birthday candles; alternatively, only foreground image assets containing a birthday cake can be used to obtain a target video containing only a birthday cake.
[0171] The method for extracting scene description features from background image materials is similar to that for foreground image materials, and will not be repeated here.
[0172] This application embodiment utilizes a semantic model to extract scene description features corresponding to scene description text, which can accurately identify the information expressed by the scene description text, thus helping to obtain scene image materials that are more consistent with the scene description information.
[0173] In some embodiments, taking foreground image material as an example, when obtaining scene image material corresponding to scene description features using an image generation model, the scene description features are input into image generation models such as GAN (Generative Adversarial Network) and cGAN (Conditional Generative Adversarial Network), and the corresponding foreground image material is generated based on the scene description features using the image generation model.
[0174] It should be noted that the embodiments in this application only use the aforementioned image generation model to generate foreground image materials as an example. In some implementations, other models can also be used to generate foreground image materials. For example, a pre-trained artificial intelligence model (such as the Stable Diffusion model) can be used to generate corresponding images or videos based on scene description text, which can then be used as foreground image materials.
[0175] In particular, for some special and detailed scenes, foreground image materials can also be generated manually.
[0176] The method for generating background image materials is similar to that for foreground image materials, and will not be repeated here.
[0177] This application embodiment utilizes an image generation model to quickly generate a large number of scene image materials, improving material production efficiency and reducing material production costs.
[0178] In some implementations, after obtaining the scene image materials, the following steps may also be included:
[0179] If the scene image material does not match the preset filtering rules, the scene image material is removed. The filtering rules include: the scene image material matches the scene tag, and / or the naturalness of the scene image material meets the preset conditions.
[0180] Considering the possibility of inconsistencies between the scene image assets generated by the model and the model input, the scene image assets and the model input, i.e., the scene labels, are compared to determine if they match. If they match, the scene image asset is considered to display the scene identified by the scene label, and therefore the scene image asset can be retained. If they do not match, the content displayed by the scene image asset is considered to be inconsistent with the scene identified by the scene label, and therefore the scene image asset is discarded.
[0181] Furthermore, considering the potential for inconsistent quality in the scene image materials generated by the model, lower-quality scene image materials can be discarded. For example, scene image materials with unnatural drawing patterns can be discarded, while those with natural drawing patterns can be retained to obtain a more realistic target video.
[0182] It should be noted that the embodiments in this application only provide an exemplary description of the method for filtering scene image materials based on the naturalness of the drawing. In the actual filtering process, other filtering conditions can also be set. For example, the realism, style, and detail richness of the scene image materials can be used as filtering conditions, and scene image materials that do not meet the filtering conditions can be removed accordingly.
[0183] This application embodiment improves the quality of scene image materials in the material set by screening and removing those with poor performance, thereby obtaining a better target video.
[0184] Step 2: Generate at least one character animation based on the preset animation tags, and use the animation tags as the material description tags corresponding to the character animation.
[0185] Among them, the animation tag is used to identify the action or the scene in which the action occurs in the character animation.
[0186] Electronic devices generate animated visuals and determine their corresponding material description tags. Optionally, the animated visuals are first generated based on pre-set animation tags. Since animation tags can identify actions or action scenes in the animated visuals, the animated visuals generated based on the animation tags contain the specified actions or action scenes. For example, a greeting animation tag can identify a greeting action in the animated visuals. The animated visuals generated based on the greeting animation tag may include greeting actions, such as waving. For example, a work animation tag can identify a work scene in the animated visuals. The animated visuals generated based on the work animation tag may include actions that occur in the work scene, such as operating a computer or making a phone call.
[0187] It's understandable that since the character animations are generated based on preset animation tags, the actions or scenes depicted in the animation match the animation tags. Therefore, the animation tags can be used as the material description tags corresponding to the character animation. For example, in this step, the character animation generated based on the "greeting" animation tag shows a waving gesture that matches a greeting. Therefore, the "greeting" animation tag can be used as the material description tag corresponding to this character animation to establish a correspondence between the "greeting" animation tag and the character animation. Afterward, the character animation can be found in the material library using the "greeting" animation tag. Similarly, in this step, the character animation generated based on the "work" animation tag shows actions that match a work scene. Therefore, the "work" animation tag can be used as the material description tag corresponding to this character animation to establish a correspondence between the "work" animation tag and the character animation. Afterward, the animation can be found in the material library using the "work" animation tag.
[0188] Furthermore, to increase the amount of material in the resource library and provide a wider range of choices for the same target video description information, multiple character animations can be generated for the same animation tag. For example, for the greeting animation tag, multiple character animations can be generated, showing actions such as waving, nodding, and shaking hands.
[0189] The following is combined Figure 8 The process of generating animated characters will be further explained.
[0190] Figure 8 This is a schematic diagram of the process for generating an animated image provided in an embodiment of this application.
[0191] like Figure 8 As shown, firstly, multiple animation tags are defined for action categories and action occurrence scene categories. Each animation tag corresponds to an action or action occurrence scene, thus achieving action definition and action occurrence scene definition. Then, several animation description texts are defined for each action or action occurrence scene. If the animation description text includes a pet character, then the interaction actions between the human character and the pet character need to be designed when designing the action sequence; otherwise, only the human character's action sequence needs to be designed separately. After the action sequence design is completed, motion data generated based on the movement of the action sequence is captured, i.e., motion capture data. The motion capture data is redirected to the digital avatar to obtain the avatar's animation based on the motion data. Finally, the avatar animations are used to build an avatar animation library. This avatar animation library stores several avatar animations and can be used as part of a resource set.
[0192] The following is about Figure 8 To further describe the process, in one implementation, the aforementioned second step may include the following steps:
[0193] Set at least one animation tag corresponding to an action category, and / or, set at least one animation tag corresponding to an action scene category. For each animation tag, determine at least one animation description text. Based on the at least one animation description text corresponding to the animation tag, design an action sequence, where the animation description text describes the action or action scene identified by the animation tag. Capture motion data generated by the user moving according to the action sequence. Redirect the motion data to a digital avatar, and adjust the redirected motion data to obtain the avatar animation.
[0194] In some embodiments, multiple animation categories can be predefined, and an animation tag can be set for each animation category. The animation tag is a brief description of the animation. Unlike the multiple scene categories defined for scene image materials in the previous embodiments, this embodiment defines the following two animation categories for character animation: action category and action occurrence scene category.
[0195] For example, for action categories, the following animation tags can be defined to identify actions in character animations: greeting animation tag, heart-making animation tag, fist bump animation tag, hands on hips animation tag, stretching animation tag, etc. For action scene categories, the following animation tags can be defined to identify the scene in which the corresponding action in the character animation occurs: working animation tag, resting animation tag, listening to music animation tag, running animation tag, walking animation tag, cycling animation tag, etc.
[0196] Specifically, it can be understood that the scenarios described by the scene tags in the personal status category typically include the actions of the digital avatar. Meanwhile, the animation tags in the action occurrence scene category typically describe scenarios where the digital avatar's actions occur. Therefore, when defining tags, the scene tags for the personal status category and the animation tags for the action occurrence scene category can be set to be highly consistent. For example, the work scene tags for the personal status category and the work animation tags for the action occurrence scene category correspond to each other.
[0197] It should be noted that the embodiments of this application only take the aforementioned animation categories and animation tags as examples. In some implementations, other animation categories and animation tags can also be defined.
[0198] This application's embodiments design a variety of different animation categories and animation tags, which can cover a variety of motion scenarios and meet the personalized needs of different customers.
[0199] In some embodiments, when determining the animation description text for each animation tag, a relatively detailed description can be designed for the action or the scene in which the action occurs, and this description can be used as the animation description text.
[0200] For example, for the greeting animation tag, the animation description text can be set as follows: waving left hand, waving right hand, nodding, shaking hands, etc. For the work animation tag, the animation description text can be set as follows: typing, making a phone call, reading documents, moving goods, etc.
[0201] It should be noted that the embodiments of this application only take the aforementioned animation tags and animation description text as examples. In some implementations, other animation description texts can also be defined.
[0202] This application's embodiments design multiple different animation description texts for each animation tag, which can improve the diversity of character animations.
[0203] In one implementation, the aforementioned "determining at least one animation description text for each animation tag" may include the following steps:
[0204] Identify the characters corresponding to the animation tags. If there are multiple characters, determine the interaction actions between them. Based on the interaction actions, determine the animation description text corresponding to the animation tags.
[0205] When designing animated descriptive text for animated labels on electronic devices, interactive actions can be designed for multiple characters to increase the fun of the animated characters. These multiple characters can include human characters or animal characters.
[0206] For example, a running animation tag can correspond to a human character and a pet character. When designing the animation description text, in addition to describing the running actions of the human character and the pet character separately, interactive actions between the human character and the pet character can also be designed. For example, if the human character makes a command gesture, the pet character will perform actions such as jumping or stopping according to the command gesture.
[0207] Optionally, for the same animation tag, it can be associated with different characters to generate several character-based animations. For example, for the running animation tag, it can be associated with a single human character to generate several character-based running animations; simultaneously, it can be associated with a human character and a pet character to generate several character-based and pet-based running animations together. Using both types of animations as the animations corresponding to the running animation tag can increase the diversity of the generated target videos.
[0208] Optionally, for cases with multiple characters, separate actions can be designed for each character, without designing interactive actions between them. This can also generate character animations containing multiple characters, and the target video generated using these character animations can contain multiple characters, with each character's actions being independent of the others.
[0209] Optionally, individual character animations can be generated for each character separately. When generating the target video, several character animations are selected and combined to obtain the target video. The resulting target video also contains multiple characters, with each character's movements being independent of the others.
[0210] This application's embodiments enhance the richness of content in the target video through the design of interactive actions by multiple characters. Furthermore, the addition of animal characters is more appealing to pet lovers, increasing its attractiveness to animal enthusiasts.
[0211] In some embodiments, when designing an action sequence based on animation description text, the actions described in the animation description text can first be broken down into multiple steps. Alternatively, based on the scenario in which the actions described in the animation description text occur, the actions corresponding to that scenario can be determined, and then the actions can be broken down into multiple steps.
[0212] As you can understand, animation description text is a more detailed description of the animation than the animation tag, describing the specific action or the scene in which the action occurs. For example, if the animation tag is "greeting," the corresponding animation description text could be "waving," or it could be "shaking hands," "nodding," "hugging," etc.
[0213] An action sequence consists of a series of ordered steps that can realize the content described in the animation description text. For example, for the animation description text of waving, the following action sequence can be designed: [body facing the viewer of the animated image; raising the right hand; swinging the wrist up and down; lowering the arm]. By executing each step in the action sequence in sequence, the action of waving can be completed. The animated image of waving based on this action sequence can include the waving animation.
[0214] Optionally, for a single animation tag, the content described by different animation description texts can be mutually exclusive or appear simultaneously. For example, for the aforementioned greeting animation tag, the content described by animation description texts such as waving and hugging can appear simultaneously. Therefore, the following action sequence can be designed by combining the waving and hugging animation description texts: [Body facing the viewer of the animated image; raising right hand; wrist gently swinging up and down; walking towards the viewer; opening arms; releasing the hug; returning to a standing posture]. Executing each step in the action sequence sequentially completes the greeting action. The greeting animation obtained based on this action sequence includes both waving and hugging animations.
[0215] In some embodiments, when capturing motion data generated during a user's movement according to an action sequence, the user can first perform the action sequence designed in the aforementioned steps to execute the action corresponding to the animation tag. In some implementations, the user moving according to the action sequence can be a professional motion capture actor. Because motion capture actors can accurately execute complex movements and expressions, the accuracy of the motion data can be guaranteed, resulting in more expressive animated characters. In other implementations, the user moving according to the action sequence can also be a user of the target video, thus obtaining more personalized animated characters to meet the user's individual needs.
[0216] Motion capture technology can be used to capture motion data during a user's movement according to a sequence of actions. Motion capture technology is a technique that can track and record the motion trajectory of a human body or other objects in real three-dimensional space. In the embodiments of this application, motion capture technology can be used to capture the user's motion trajectory and convert it into motion data.
[0217] For example, optical motion capture technology can be used to capture motion data. Specifically, reflective or actively luminous markers can be added to specific locations on the user's body, such as joints, head, and hands. During the user's movement, a camera captures the reflected light from the markers to obtain the motion trajectory of specific locations such as joints, head, and hands. The motion trajectory can then be converted into motion data in the form of three-dimensional coordinates.
[0218] It should be noted that the embodiments in this application only use optical motion capture technology as an example. In some implementations, other motion capture technologies can also be used. For example, inertial motion capture technology can be used, where the user wears a gyroscope, and the user's motion data is calculated based on the rotation information of the gyroscope during the user's movement. For example, visual motion capture technology can also be used, which records the user's movements through a camera and uses algorithms such as deep learning to identify the user's joint information to obtain motion data. In practical applications, the appropriate motion capture technology can be selected based on the application scenario. For example, for professional motion capture actors, optical motion capture technology and inertial motion capture technology can be used to capture their motion data; for users of electronic devices, visual motion capture technology can be selected to capture their motion data.
[0219] In some embodiments, when redirecting motion data to a digital avatar, the motion data corresponding to each specific location can first be bound to the corresponding location in the digital avatar. For example, the motion data may include the motion trajectory data of each joint; the motion trajectory data of the wrist joint can be bound to the wrist joint of the digital avatar. Then, the motion trajectory of the wrist joint can be mapped onto the wrist joint of the digital avatar, thus achieving motion data redirection. At this point, the aforementioned motion data is applied to the digital avatar, and the digital avatar can perform corresponding actions based on the motion data to obtain animation effects.
[0220] Optionally, considering that unnatural movements in the motion sequence and acquisition biases may result in less realistic animation effects in the generated characters, the redirected motion data can be fine-tuned to achieve more realistic animation effects.
[0221] Optionally, considering that motion data may not be entirely suitable for a cartoonish style, the redirected motion data can be fine-tuned. For example, by adjusting the motion data, the amplitude and rhythm of the movements can be adjusted, so that the digital character performs corresponding actions based on the fine-tuned motion data, which is more in line with the cartoon style.
[0222] This application embodiment captures user motion data and redirects it to a digital avatar, resulting in avatar animations that accurately reflect actual movements, thus improving the naturalness and realism of the animations. Furthermore, compared to directly designing avatar animations, this application embodiment also improves the production efficiency of avatar animations.
[0223] S112: Construct a set of materials based on the correspondence between the materials and their description tags.
[0224] A material set is constructed based on the correspondence between the material and its description tag. Subsequently, the corresponding material can be quickly found in the material set based on the description tag.
[0225] In some implementations, after constructing the resource set, the following step S113 is also included:
[0226] S113: Add the material information for each material to the material set.
[0227] When storing media in a media set, add corresponding media information to each media. Media information includes at least one of the following: media description tag, media description text, index, default media identifier, extended information, and media version number.
[0228] In some implementations, the material information is shown in Table 1 below:
[0229] Table 1 Material Information Table
[0230] Fields type meaning Classification string Scene tags scene string Animation Tags description string Material description text index int index defaultflag int Default material identifier extend json-string Extended Information version string Material version number
[0231] In Table 1, the material description tags corresponding to scene image materials and character animations are scene tags and animation tags, respectively. For an explanation of scene tags, please refer to the first step of S111 above; they can be personal-specific, holiday-specific, personal status, nature, general daily use, etc. For an explanation of animation tags, please refer to the second step of S111 above; they can be birthday, holiday, listening to music, sports, work, rest, sunrise, sunset, weather-sunny, weather-rainy, weather-snowy, weather-fog, weather-haze, weather-windy, general animation tags, etc., and will not be elaborated further here. The material description text is used to briefly describe the material. The material description text corresponding to scene image materials is the scene description text, and the material description text corresponding to character animations is the animation description text. The index is the unique identifier of the material, such as the file code corresponding to the material, and can be used to distinguish different materials corresponding to the same tag. The default material identifier is used to identify whether the material can be used as the default material. When generating the target video, a default material can be randomly selected based on the default material identifier. Optionally, the default material is the material corresponding to the general daily scene tag. Extended information includes the material's effective time and expiration time, which can be determined based on the time zone. Within the timeframes defined by the effective and expiration times, the material is in an active state. During other timeframes, the material is in an expiration state. By setting the effective and expiration states, materials can be configured to be available only during specific periods and unavailable at other times. For example, some materials may only be available during the day, while others may only be available at night. The material version number is an identifier added when the material is modified or updated. It can be used to record different historical versions of the same material, avoiding confusion between different versions.
[0232] This application embodiment records the material information of each material and adds it to the material set, which facilitates the management of materials and is especially suitable for situations with a large number of materials.
[0233] The following is combined Figure 9 The technical solutions involved in step S12 will be further described.
[0234] Figure 9 This is a flowchart of another embodiment of the video generation method provided in this application.
[0235] like Figure 9 As shown, in this embodiment, the aforementioned step S12 may include the following steps S121-S122:
[0236] S121: Based on scene image materials and character animations in the material set, construct at least one subset of materials.
[0237] The resource set includes multiple scene image assets and character animations, while the resource subset is obtained by selecting certain scene image assets and character animations from the resource set. For example, in the resource set {F,B,A}, F = {f1,f2,…,f…} n} represents a set of multiple foreground image materials; B = {b1, b2, ..., b} m} represents a collection of multiple background image materials; A = {a1, a2, ..., a} k} is a collection of multiple animated characters. Then, a subset {F} can be constructed using several assets from the asset set. i B i A i}. Among them, F i B is a subset of F. i A is a subset of B. i This is a subset of A. Therefore, the material subset {F} is a subset of A. i B i A i} is a subset of the material set {F,B,A}.
[0238] Optionally, the materials in the same subset can be related materials. For example, the visual animation corresponding to the running action, the scene image material corresponding to the treadmill, and the scene image material corresponding to the playground can be placed in the same subset.
[0239] S122: Construct a mapping table based on the correspondence between each material subset and the material description tag.
[0240] It can be understood that a subset of materials includes several materials, and there is a correspondence between the material description tags and the material subsets. Based on this, the correspondence between material subsets and material description tags can be obtained to construct a mapping table. Subsequently, the material description tags can be queried in the mapping table to obtain the corresponding material subsets.
[0241] For example, as mentioned above, a subset of materials that simultaneously includes visual animations corresponding to running actions, scene image materials corresponding to treadmills, and scene image materials corresponding to playgrounds can correspond to the material description tag "running".
[0242] The mapping table constructed in this embodiment contains the correspondence between material subsets and material description tags. Therefore, when querying material description tags, all materials in the material subset can be obtained directly, which helps to improve search efficiency.
[0243] In one implementation, step S122 may include the following steps:
[0244] For each subset of materials, determine the correspondence between the subset and its description tags based on the existing correspondence between materials and their description tags. Then, construct a mapping table based on this correspondence.
[0245] Considering that a subset of materials consists of multiple materials, the correspondence between the subset and its corresponding description tags can be determined based on the existing correspondence between the materials and their description tags, thus constructing a mapping table. For example, the mapping table can be in the form of key-value pairs, where the key and value are the subset of materials and its corresponding description tag, respectively.
[0246] Optionally, within a subset of materials, the description tag for each material corresponds to the subset itself. For example, for the subset {F}... i B i A i}, F i The material in the text corresponds to the material description tag l fi B i The material in the text corresponds to the material description tag l bi A i The material in the text corresponds to the material description tag l ai Then the subset of materials {F i B i A i} and corresponding tag l fi l bi and l ai In other words, there is a one-to-many relationship between the subset of materials and the material description tags.
[0247] Optionally, within a subset of materials, all the material description tags corresponding to the materials form a composite tag, and the subset of materials corresponds to this composite tag. For example, for the material subset {F... i B i A i}, F i The material in the text corresponds to the material description tag l fi B i The material in the text corresponds to the material description tag l bi A i The material in the text corresponds to the material description tag l ai Then the subset of materials {F i B i A i} Corresponding to the compound tag {l fi ,l bi ,l ai}
[0248] Optionally, after determining the material description tags corresponding to the material subset, it can also be determined that these tags are the material description tags corresponding to each material within the material subset. For example, the material subset {F} i B i A i} and corresponding tag l fi l bi and l ai Then F can be determined. i The material in the text also corresponds to the tag l fi l bi and l ai B i The material in the text also corresponds to the tag l fi l bi and l ai A i The material in the text also corresponds to the tag l fi l bi and l ai That is, a single asset corresponds to both a scene tag and an animation tag. For example, a subset of assets {F} u B i A i} Corresponding to the compound tag {l fi ,l bi ,l ai}, then F can be determined i B i and A i The materials in the text all correspond to the compound tag {l fi ,l bi ,l ai}
[0249] The material subset in this application embodiment includes both scene image materials and character animations. Therefore, based on a material description tag corresponding to the material subset, both different types of materials, such as scene image materials and character animations, can be found simultaneously.
[0250] The following is combined Figure 10 The technical solutions involved in step S13 will be further described.
[0251] Figure 10 This is a flowchart of another embodiment of the video generation method provided in this application.
[0252] like Figure 10 As shown, in this embodiment, the aforementioned step S13 may include the following step S131:
[0253] Step S131: Determine video description information based on the current scene information.
[0254] Electronic devices determine video description information based on current scene information to obtain a target video that matches the current scene. Current scene information may include current user status information, current date information, user input information, etc.
[0255] Current user status information can be used to identify a user's activity status and can be obtained through various channels. For example, the current user status information can be determined using built-in sensors in an electronic device; for instance, if a phone's built-in sensors detect that the user is running, then the current user status can be determined as running. Alternatively, the status of applications installed on an electronic device can be monitored, and the current user status information can be determined based on the application status; for instance, if a music application is detected playing music, then the current user status can be determined as listening to music. Finally, the current user status information can be determined based on information sent by other electronic devices; for example, if a cycling information message is received from a fitness watch, then the current user status can be determined as cycling.
[0256] Current date information can be used to determine seasons, holidays, anniversaries, etc., and can be obtained from calendar applications installed on electronic devices. For example, if the current date is January 1st, then the current date can be determined to be New Year's Day.
[0257] User input information can be entered by the user through text or voice, and is entered by the user according to their actual needs.
[0258] It is understood that video description information can be determined by a single piece of current scene information, or by a combination of multiple pieces of current scene information. For example, if the current date is New Year's Day, the video description information can be determined to be New Year's Day, and a target video related to New Year's Day can be generated. For example, if the current user's state is running and the current date is winter, the video description information can be determined to be {winter, running}, and a target video of running in winter can be generated.
[0259] It is understood that the embodiments of this application only use the aforementioned current scene information as an example. In some implementations, other current scene information may also be used. For example, the current scene information may include current time information, used to determine day / night and daily scheduled activities, etc. For example, if the current time is determined to be 12:00 noon, then the current time can be determined to be lunchtime, and "eating lunch" can be used as the current scene information to generate a target video related to "eating lunch" in subsequent steps. For example, if the weather application detects that the current weather is light rain, then "light rain" can be used as the current scene information to generate a target video related to "light rain" in subsequent steps.
[0260] This application provides various methods for determining video description information, which can yield target videos that better match the current actual scenario or user preferences.
[0261] The following is combined Figure 11 The technical solutions involved in step S14 will be further described.
[0262] Figure 11 This is a flowchart of another embodiment of the video generation method provided in this application.
[0263] like Figure 11 As shown, in this embodiment, the aforementioned step S14 may include the following steps S141-S142:
[0264] S141: Determine the material description tags corresponding to the video description information.
[0265] The electronic device determines the material description tag corresponding to the video description information, and uses the material description tag to determine the corresponding target scene material and target animation material. Optionally, semantic features of the video description information can be extracted, and then a material description tag matching the semantic features can be found among a number of predefined material description tags. For example, for the video description information "birthday", the material description tag "birthday" that matches the semantics of "birthday" can be found among a number of material description tags. At this time, the material description tag corresponding to the video description information "birthday" is "birthday". Optionally, in the technical solutions involved in steps S13 and S131, the video description information is directly generated using words in the material description tag, so in the technical solution involved in step S141, the video description information can be directly used as the corresponding material description tag. For example, after detecting that the current date is the user's birthday, the video description information "birthday" can be directly generated using the word "birthday" in the material description tag. The video description information generated in this way can be directly used as its corresponding material description tag.
[0266] It is understood that video description information can correspond to one material description tag or multiple material description tags. For example, for the aforementioned video description information "running", it can be determined that it corresponds to the "running" animation tag. For example, for the aforementioned video description information {winter, running}, it can be determined that it corresponds to a "winter" scene tag and a "running" animation tag. That is, the video description information {winter, running} corresponds to the composite tag "winter running".
[0267] This application embodiment obtains material description tags based on complex video description information that lacks a standardized format. Since these material description tags are recorded in a mapping table, this application embodiment establishes a connection between the video description information and the mapping table through these tags. This allows for subsequent steps to query the mapping table and obtain materials that match the video description information.
[0268] S142: Based on the pre-built mapping table, obtain the target scene material and target animation material corresponding to the material description tag corresponding to the video description information from the material set.
[0269] The electronic device determines the target scene material and the target animation material based on the material description tags corresponding to the video description information. Since each material has a direct correspondence with a material description tag, the material corresponding to the material description tag corresponding to the video description information can be found in the mapping relationship table by looking up a table, which will then be used as the target scene material and the target animation material. The mapping relationship table is pre-built and records the correspondence between materials and material description tags. For an explanation and construction method, please refer to the technical solution involved in step S12 above, which will not be repeated here.
[0270] This embodiment of the application determines the target scene material and the target animation material by querying a mapping relationship table. Since the aforementioned steps have established a connection between the video description information and the mapping relationship table, this embodiment of the application can ensure that the material queried based on the mapping relationship table conforms to the content described in the video description information.
[0271] In some implementations, step S142 may include the following steps:
[0272] Based on the mapping table, the subset of materials corresponding to the material description tags that correspond to the video description information is determined as the target material subset. Within the target material subset, target scene materials and target animation materials are identified.
[0273] The material subset includes several materials from the material set, and in the aforementioned embodiments, a mapping table has been constructed based on the correspondence between each material subset and the material description tag. Therefore, the material description tag corresponding to the target video information can be found in the mapping table to obtain its corresponding material subset, which is then used as the target material subset.
[0274] Optionally, if the target video corresponds to one material description tag, or multiple material description tags, and these multiple tags correspond to the same material subset, then that material subset can be used as the target material subset, from which the target scene material and target animation material can be determined. For example, if the target video corresponds to one material description tag "running," and a certain material subset {F1, B1, A1} corresponds to both the scene tag "summer" and the animation tag "running," then this material subset {F1, B1, A1} can be determined as the target material subset. Therefore, the target scene material can be determined from F1 and B1, and the target animation material can be determined from A1.
[0275] Optionally, if the target video corresponds to multiple material description tags, and these tags correspond to different material subsets, then the corresponding materials can be selected from each material subset based on the material description tags corresponding to the target video. For example, the target video corresponds to two material description tags, "winter" and "running," but none of the existing material subsets simultaneously correspond to both tags. Only the material subset {F1,B1,A1} corresponds to both the scene tag "summer" and the animation tag "running," and the material subset {F2,B2,A2} corresponds to both the scene tag "winter" and the animation tag "cycling." In this case, it can be determined that both material subsets {F1,B1,A1} and {F2,B2,A2} are target material subsets, and the target animation material is determined in A1, while the target scene material is determined in F2 and B2.
[0276] It is understandable that, since the target material subset is pre-established, compared to directly searching for the target material, this embodiment first searches for the target material subset and then determines the target material within the target material subset, which can effectively improve the search efficiency.
[0277] In one implementation, "determining the target scene assets and target animation assets within the target asset subset" may include the following steps:
[0278] Obtain the first target index and the second target index. In the subset of materials, determine the scene image material corresponding to the first target index as the target scene material, and the character animation corresponding to the second target index as the target animation material.
[0279] Considering that the target asset set includes multiple assets, and that an index can distinguish between different assets within the target asset set, the target scene asset and the target animation asset can be uniquely specified within the target asset set based on the target index.
[0280] The target index includes a first target index and a second target index. The first target index is used to distinguish different scene image materials corresponding to the same scene tag, and the second target index is used to distinguish different character animations corresponding to the same animation tag. For example, the target material set includes scene image materials f1, f2, b1, b2 and character animations a1, a2, with corresponding indices f01, f02, b01, b02, a01, and a02, respectively. Then, with the first target index being f01 and b01 and the second target index being a02, the target scene material f1, b1 and the target animation material a2 can be uniquely identified.
[0281] Optionally, the target index can be entered by the user in the form of text, voice, or a selection box. For example, materials from the target material set can be displayed on the screen. Users can click on any scene image material to use that scene image material as the target scene material, and click on any character animation to use that character animation as the target animation material.
[0282] Optionally, the target index can be randomly generated. When the target material set includes multiple scene image materials and multiple animated characters, a first target index and a second target index can be randomly generated, and then the target scene material and target animation material can be determined using the randomly generated target index.
[0283] This application embodiment utilizes a target index to uniquely and quickly identify target scene materials and target animation materials, thereby improving video generation efficiency and accuracy.
[0284] It is understood that the foregoing embodiments only illustrate the method of determining target scene materials and target animation materials using target indexes. In some implementations, other methods may also be used to determine target scene materials and target animation materials from a target material set. For example, target materials can be determined based on default identifier bits in the material information. For instance, by querying the default identifier bit information of each material in the target material set, materials whose default identifier bits meet the conditions are identified as default materials, and the default materials can be determined as either target scene materials or target animation materials.
[0285] The following is combined Figures 12 to 14 The technical solutions involved in step S15 will be further described.
[0286] Figure 12 This is a flowchart of another embodiment of the video generation method provided in this application.
[0287] like Figure 12 As shown, in this embodiment, the aforementioned step S15 may include the following steps S151-S152:
[0288] S151: Superimpose the target scene material and the target animation material in a preset order to obtain the target video.
[0289] Electronic devices can sequentially overlay target scene materials and target animation materials in a certain order, while maintaining a certain interval between adjacent materials, in order to obtain a target video with a sense of depth.
[0290] It is understood that the target scene materials can include any number of foreground image materials. The following will combine... Figure 13 This indicates two scenarios: the target scene material includes multiple foreground image materials and the target scene material does not include foreground image materials.
[0291] Figure 13This is a schematic diagram showing the overlay of target scene materials and target animation materials provided in the embodiments of this application.
[0292] like Figure 13 As shown in (a) above, for example, the target scene material includes two foreground image materials: a water glass image material and a grass image material. The background image material, the animated character, the water glass image material, and the grass image material can be superimposed sequentially according to their distance from the viewer, from farthest to closest, to obtain the target video. For example... Figure 13 As shown in (b) in the target video, part of the water glass is obscured by grass.
[0293] like Figure 13 As shown in (c) above, for example, the target scene material includes two foreground image materials: a water glass image material and a grass image material. The background image material, the animated image, the grass image material, and the water glass image material can be overlaid in order of distance from the viewer, from farthest to closest, to obtain the target video. For example... Figure 13 As shown in (d) in the target video, part of the grass is obscured by the water glass.
[0294] like Figure 13 As shown in (e) above, for example, the target scene material does not include foreground image material, only background image material. The background image material and the animated image can be overlaid in order of distance from the viewer, from farthest to closest, to obtain the target video. Figure 13 As shown in (f), the target video at this point contains only digital images and grass and clouds in the background.
[0295] This application embodiment can overlay multiple foreground image materials in different orders, thereby obtaining different videos using the same foreground image materials. Furthermore, this application embodiment can also omit foreground image materials, thus obtaining a video without foreground elements. This design increases the diversity of the target video.
[0296] It is understood that the target scene material can include any number of background image materials, and the superposition method is similar to the superposition example of any number of foreground image materials, which will not be repeated here.
[0297] S152: Render each frame of the target video according to the lighting change information and / or camera movement trajectory information to obtain the rendered target video.
[0298] It is understandable that in three-dimensional space, the position, intensity, and direction of a light source affect the distribution of light and shadow on an object's surface, thus creating a visual sense of depth. Therefore, a target video can be rendered based on lighting changes, enhancing the sense of depth through contrast and lighting effects. For example, a light source can be set for the target video. Points at different locations in the target video have different relative positions to the light source, thus receiving different amounts of light, resulting in varying degrees of brightness. Based on this, brightness variations can be applied to different locations in the target video to simulate the shape and depth of each element in the target video in three-dimensional space.
[0299] It is understandable that in three-dimensional space, visual layers can be created and the sense of depth in an image can be enhanced through camera movement. Therefore, a target video can be rendered based on camera movement information, and the sense of depth can be enhanced through the movement of the lens. For example, a camera can be set up for the target video. The relative positions of points in different locations in the target video to the camera are different. Rendering the target video based on perspective and overlapping occlusion effects at each location can improve the spatial sense of the target video.
[0300] Optionally, the lighting change information and camera movement information can be information pre-stored in the renderer or custom information input by the user.
[0301] This application embodiment takes into account the influence of lighting and camera movement on the video, and renders the target video to increase the three-dimensionality of objects in the target video and improve the realism of the target video.
[0302] Figure 14 This is a schematic diagram illustrating how a target video is obtained by overlaying target scene material and target animation material according to an embodiment of this application.
[0303] like Figure 14 As shown, background image materials, digital human animation, and foreground image materials are sequentially overlaid, and each frame is rendered according to lighting changes and camera movement to obtain the final target video. For the material overlay method and image rendering method, please refer to the technical solutions involved in steps S151-S152 above. The target video includes grass and cloud elements in the background, an animation of the digital human waving, and a water glass element in the foreground. In addition, the image animation of a digital pet can also be overlaid, which is the same as the digital human's image animation and will not be described again here.
[0304] The following is combined Figure 15 and Figure 16 We will now introduce another implementation method.
[0305] Figure 15 This is a flowchart of another embodiment of the video generation method provided in this application.
[0306] like Figure 15As shown, the system first searches for video resources locally on the electronic device. These resources can include digital avatars, scene images, and animations. If no resources are available locally, a prompt message is generated to encourage the user to download them. If resources are available locally, the user edits the digital avatar to their preferred appearance. Then, based on the edited avatar, the target video is generated. If video generation fails, a corresponding prompt message is generated to alert the user. Otherwise, the generated target video is saved to a specified location. This specified location can be either on the electronic device or on a cloud server; no specific limitation is made here.
[0307] The following is combined Figure 16 ,right Figure 15 The process will be further explained below.
[0308] Figure 16 This is a flowchart of another embodiment of the video generation method provided in this application.
[0309] like Figure 16 As shown, this embodiment includes the following steps S21-S26.
[0310] S21: In response to the video generation command, acquire the digital image.
[0311] Electronic devices respond to video generation instructions by acquiring a default digital image, which can then be used to personalize the image.
[0312] Optionally, the video generation command can be user-inputted, allowing the user to actively generate a target video. The video generation command can also be timed, such as generating a target video daily as the lock screen video, ensuring users see a new lock screen video every day and increasing user engagement. Furthermore, the video generation command can be generated based on changes in the current scene information. For example, if the current scene information changes from "sunny" to "rainy," a video generation command can be generated in response to this change, producing a target video corresponding to the new current scene information.
[0313] Optionally, a default digital avatar can be pre-stored locally on the electronic device. Therefore, before generating the digital avatar, it is determined whether the digital avatar is stored locally on the electronic device. If not, the digital avatar is downloaded to the electronic device from the cloud server.
[0314] Optionally, in addition to determining whether the digital image is stored locally on the electronic device, it can also be determined whether a set of source materials is stored locally on the electronic device. If not, the source materials are downloaded from the cloud server to the electronic device. Alternatively, this step can be skipped, and instead, in subsequent steps of acquiring target scene materials and target animation materials, it can be directly determined whether the electronic device has materials matching the video description information. If not, the materials are downloaded from the cloud server to the electronic device.
[0315] S22: Responds to editing commands to edit digital images.
[0316] Electronic devices modify the appearance of digital avatars based on user editing instructions to achieve customization. For example, based on the default avatar in the aforementioned technical solution S21, the hairstyle, hair color, skin tone, etc. of the digital avatar can be modified to obtain a digital avatar that meets the user's personalized needs, thus increasing the user's degree of freedom.
[0317] S23: Obtain video description information.
[0318] For an explanation of step S23, please refer to step S13 above, which will not be repeated here.
[0319] S24: Based on the video description information, acquire the target scene material and target animation material from the preset material set.
[0320] For an explanation of step S24, please refer to step S14 above, which will not be repeated here.
[0321] In one implementation, the resource set includes a scene image resource set and an animation resource set, and step S24 may further include the following steps:
[0322] Obtain the target scene material from the scene image material set, and obtain the target animation material from the animation material set.
[0323] To facilitate the management of different types of materials, the material set can be divided into a scene image material set and an animation material set, with the scene image materials and character animations stored in the scene image material set and animation material set respectively. Subsequently, the target scene material and target animation material can be retrieved from the scene image material set and animation material set respectively.
[0324] S25: Combine the target scene material and the target animation material to obtain the target video.
[0325] For an explanation of step S25, please refer to step S15 above, which will not be repeated here.
[0326] In one implementation, step S25 may be followed by the following steps:
[0327] S26: Set the target video as the lock screen video for the electronic device.
[0328] For an explanation of step S26, please refer to step S16 above, which will not be repeated here.
[0329] The following is combined Figure 17 This section introduces the process of setting up lock screen video.
[0330] Figure 17 This is a schematic diagram of a lock screen video setting provided in an embodiment of this application.
[0331] like Figure 17 As shown in (a), first, in the application interface 10 of the electronic device, click the settings button to jump to the settings interface 20.
[0332] like Figure 17 As shown in (b), open the settings interface 10 of the electronic device, and click the desktop personalization option in the settings interface 10 to jump to the desktop personalization interface 30.
[0333] like Figure 17 As shown in (c), the desktop personalization interface 30 displays several selectable lock screen types, such as digital human lock screen, static landscape lock screen, and minimalist lock screen. In this embodiment, a digital human lock screen can be selected. After selecting the lock screen type, the lock screen application starts to generate the target video. During the generation of the target video, it first determines whether the electronic device has a digital human avatar resource locally, i.e., the default digital avatar. If not, the digital avatar resource is downloaded from the cloud. The digital avatar can then be edited and used. Simultaneously, the electronic device intelligently senses the user's status and current scene information such as holidays to obtain video description information. Based on the video description information, scene image materials and avatar animations are randomly selected from the matching set as target scene materials and target animation materials. Finally, the target scene materials and target animation materials are combined to obtain the target video.
[0334] like Figure 17 As shown in (d), after obtaining the target video, the user is redirected to the video display interface 40. The user can click the preview button on the display interface 40 to jump to the preview interface 50.
[0335] like Figure 17 As shown in (e), the preview interface 50 can display the lock screen effect corresponding to the target video. If the user is satisfied with the display effect, they can click the Apply button in the preview interface 50 to set the video as the lock screen video.
[0336] like Figure 17 As shown in (f), after setting the lock screen video, the lock screen interface 60 is displayed in the lock screen state.
[0337] Other embodiments of this application provide a video generation apparatus.
[0338] Figure 18 This is a schematic diagram of a video generation device provided in an embodiment of this application.
[0339] like Figure 18 As shown, the video generation device may include: a display screen 1001, a memory 1002, a processor 1003, and a communication module 1004. These devices can be connected via one or more communication buses 1005. The display screen 1001 may include a display panel 10011 and a touch sensor 10012. The display panel 10011 is used to display images, and the touch sensor 10012 can transmit detected touch operations to the application processor to determine the touch event type, providing visual output related to the touch operation through the display panel 10011. The processor 1003 may include one or more processing units, such as an application processor, a modem processor, a graphics processor, an image signal processor, a controller, a video codec, a digital signal processor, a baseband processor, and / or a neural network processor. Different processing units may be independent devices or integrated into one or more processors. The memory 1002 is coupled to the processor 1003 and is used to store various software programs and / or computer instructions. The memory 1002 may include volatile memory and / or non-volatile memory. When the processor executes computer instructions, the video generation device can perform the various functions or steps performed in the above method embodiments.
[0340] When the software program and / or multiple sets of instructions in the memory 1002 are executed by the processor 1003, the video generation device performs the following method steps: acquiring video description information, which describes the scene of the target video and / or the actions of the digital image in the target video; based on the video description information, acquiring target scene material and target animation material from a preset material set, wherein the target scene material includes scene image material, which includes foreground image material and / or background image material, and the target animation material includes an animation of the digital image moving based on motion data; and combining the target scene material and the target animation material to obtain the target video.
[0341] This application also provides an electronic device, including: a processor, a memory, and a touch screen; the memory stores program instructions, which, when executed by the processor, cause the electronic device to perform the video generation method in any of the above embodiments.
[0342] This application also provides a chip system including at least one processor and at least one interface circuit. The processor and the interface circuit are interconnected via lines. For example, the interface circuit can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit can be used to send signals to other devices. Exemplarily, the interface circuit can read instructions stored in the memory and send the instructions to the processor. When the instructions are executed by the processor, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete devices, and this application does not specifically limit this.
[0343] This application embodiment also provides a computer-readable storage medium, which includes computer instructions, and when the computer instructions are used in the above-mentioned electronic device (such as...). Figure 3 When the electronic device 100 shown is run, it causes the electronic device to perform the various functions or steps in the above method embodiments.
[0344] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments.
[0345] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0346] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0347] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0348] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0349] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0350] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video generation method, characterized in that, Applied to electronic devices, the method includes: Display a lock screen video and display a digital avatar and a first scene in the lock screen video, wherein the digital avatar is displayed with dynamic effects, and the first scene includes at least one first element. When the first scene changes, at least one target element in the first element changes to a second element, and the first element and the second element are different. The digital avatar is generated based on target animation material and displays the dynamic effect based on changes in the target animation material. The first scene is generated based on target scene material and changes based on changes in the target scene material. The target animation material includes an animation of the digital avatar moving based on motion data. The target scene material includes scene image material, which includes foreground image material and / or background image material. The display of the lock screen video specifically includes: obtaining video description information; obtaining the target animation material and the target scene material from a preset material set based on the video description information; combining the target scene material and the target animation material to obtain a target video; and setting the target video as the lock screen video after obtaining the target video. The video description information is used to describe the scene of the lock screen video and / or the actions of the digital avatar in the lock screen video. The video description information is determined based on current scene information, which includes at least one of the current user status, current date information, and user input information.
2. The method according to claim 1, characterized in that, Before acquiring the target scene material and the target animation material from the preset material set, the method further includes: Construct the material set, wherein the materials in the material set include at least one of the character animations and the scene image materials.
3. The method according to claim 2, characterized in that, The construction of the material set includes: Obtain at least one of the aforementioned materials and a corresponding material description tag, wherein the material description tag is used to identify the scene or action corresponding to the material; Based on the correspondence between the materials and the material description tags, the material set is constructed.
4. The method according to claim 3, characterized in that, The step of obtaining at least one of the materials and the corresponding material description tags includes: Based on preset scene tags, at least one scene image material is generated, and the scene tags are used as the material description tags corresponding to the scene image material, wherein the scene tags are used to identify the scene corresponding to the scene image material; Based on preset animation tags, at least one of the character animations is generated, and the animation tags are used as the material description tags corresponding to the character animations, wherein the animation tags are used to identify the actions or action scenes in the character animations.
5. The method according to claim 4, characterized in that, The step of generating at least one scene image material based on preset scene tags includes: Input at least one scene description text corresponding to the scene label into the semantic model to obtain the scene description features corresponding to the scene description text, wherein the scene description text is used to describe the scene identified by the scene label; The scene description features are input into the image generation model to obtain the scene image material corresponding to the scene description features.
6. The method according to claim 5, characterized in that, Before inputting at least one scene description text corresponding to the scene label into the semantic model, the method further includes: For each scene category, set at least one scene label; For each of the scene tags, at least one scene description text is determined.
7. The method according to claim 4, characterized in that, After generating at least one scene image material based on preset scene tags, the method further includes: If the scene image material does not match the preset filtering rules, the scene image material is removed. The filtering rules include: the scene image material matches the scene tag, and / or the naturalness of the scene image material meets the preset conditions.
8. The method according to claim 4, characterized in that, The step of generating at least one of the character animations based on preset animation tags includes: Based on at least one animation description text corresponding to the animation tag, an action sequence is designed, wherein the animation description text is used to describe the action identified by the animation tag or the scene in which the action occurs; Capture the motion data generated during the user's movement according to the said action sequence; The motion data is redirected to the digital image, and the redirected motion data is adjusted to obtain the image animation.
9. The method according to claim 8, characterized in that, Before designing the action sequence based on at least one animation description text corresponding to the animation tag, the method further includes: At least one animation tag corresponding to the action category is set, wherein the animation tag corresponding to the action category is used to identify the action in the character animation; and / or, At least one animation tag is set that corresponds to the category of the action occurrence scene, wherein the animation tag corresponding to the action occurrence scene category is used to identify the action occurrence scene corresponding to the image animation; For each of the animation tags, at least one animation description text is determined.
10. The method according to claim 9, characterized in that, Determining at least one animation description text corresponding to the animation tag includes: Identify the character corresponding to the animation tag, wherein the character is a human character or an animal character; When there are multiple characters, determine the interaction actions between the multiple characters, and determine the animation description text corresponding to the animation tag based on the interaction actions.
11. The method according to claim 3, characterized in that, After constructing the material set, the method further includes: Add the material information of each material to the material set, wherein the material information includes at least one of the following: material description tag, material description text, index, default material identifier, extended information, and material version number; Wherein, the material description text corresponding to the scene image material is the scene description text, and the material description text corresponding to the image animation is the animation description text.
12. The method according to claim 3, characterized in that, After constructing the material set based on the correspondence between the material and the material description tags, the method further includes: Based on the scene image materials and character animations in the material set, at least one material subset is constructed; A mapping table is constructed based on the correspondence between each of the material subsets and the material description tags.
13. The method according to claim 12, characterized in that, The step of constructing a mapping relationship table based on the correspondence between each of the material subsets and the material description tags includes: For each of the material subsets, the correspondence between the material subset and the material description tag is determined based on the correspondence between each material in the material subset and the material description tag; Based on the correspondence between the material subset and the material description tags, construct the mapping table.
14. The method according to claim 1, characterized in that, The step of acquiring target scene materials and target animation materials from a preset material set includes: Determine the material description tags corresponding to the video description information; Based on a pre-built mapping table, the target scene material and the target animation material corresponding to the material description tag are obtained from the material set.
15. The method according to claim 14, characterized in that, The step of obtaining the target scene material and the target animation material corresponding to the material description tag from the material set based on the mapping relationship table includes: Based on the mapping table, the subset of materials corresponding to the material description tags is determined as the target material subset; Within the target material subset, the target scene material and the target animation material are determined.
16. The method according to claim 15, characterized in that, The step of determining the target scene material and the target animation material in the target material subset includes: Obtain a first target index and a second target index, wherein the first target index is used to distinguish different scene image materials corresponding to the same scene tag, and the second target index is used to distinguish different character animations corresponding to the same animation tag; In the subset of materials, the scene image material corresponding to the first target index is determined as the target scene material, and the image animation corresponding to the second target index is determined as the target animation material.
17. The method according to claim 1, characterized in that, The material set includes a set of scene image materials and a set of animation materials; the step of acquiring the target scene materials and target animation materials from the preset material set includes: The target scene material is obtained from the scene image material set, and the target animation material is obtained from the animation material set.
18. The method according to claim 1, characterized in that, The combination of the target scene material and the target animation material to obtain the target video includes: The target scene material and the target animation material are superimposed in a preset order to obtain the target video; Based on the lighting change information and / or camera movement trajectory information, each frame of the target video is rendered to obtain the rendered target video.
19. The method according to claim 1, characterized in that, Before acquiring the target scene material and target animation material from the preset material set, the method further includes: In response to a video generation command, the digital image is acquired; In response to the edit command, edit the digital image.
20. An electronic device, characterized in that, The device includes a memory and a processor; the memory and the processor are coupled; the memory is used to store computer program code, the computer program code including computer instructions, which, when executed by the processor, cause the electronic device to perform the video generation method as described in any one of claims 1-19.