Dynamic image generation method and device, electronic equipment and medium
By adding different effects to multiple target areas of a panoramic image to generate dynamic images, the problem of monotonous visual effects in dynamic photos is solved, achieving visual diversity and user creative flexibility, and enhancing the immersive and personalized experience of dynamic images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-13
AI Technical Summary
In existing technologies, the visual effects of dynamic photos lack variation and depth, making it difficult to meet users' complex needs for personalized and creative images.
By receiving user input of panoramic images, different effects are added to multiple target areas to generate dynamic images. The artificial intelligence model applies effects to each area, including seasonal, time and artistic style transformation effects, and the dynamic effect is achieved by controlling the layer transparency through the timeline.
It enhances the visual diversity of moving images and the flexibility of user creation, providing an immersive and personalized visual experience.
Smart Images

Figure CN121661209A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, specifically to dynamic image generation methods, apparatus, electronic devices, and media. Background Technology
[0002] A live photo is a technology that records approximately three seconds of video and sound before and after the moment it was captured, allowing a still photo to be presented as a short video. With the development of photography and artificial intelligence technologies, adding dynamic visual effects to still photos to create live photos has become a common feature.
[0003] In existing technologies, dynamic photos created by adding dynamic visual effects to static photos usually lack variation and depth, making it difficult to meet users' complex needs for personalized and creative images. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, electronic device, and medium for generating dynamic images, thereby enhancing the visual diversity of dynamic images.
[0005] In a first aspect, embodiments of this application provide a dynamic image generation method, the method comprising: receiving a first input from a user to a panoramic image, the panoramic image including multiple target regions; responding to the first input, adding different special effects to different target regions among the multiple target regions to obtain multiple special effect regions; and generating a dynamic image based on the panoramic image and the multiple special effect regions.
[0006] Secondly, embodiments of this application provide a dynamic image generation apparatus, which includes: a receiving unit for receiving a first input from a user to a panoramic image, the panoramic image including multiple target regions; an adding unit for adding different special effects to different target regions in response to the first input, thereby obtaining multiple special effect regions; and a generating unit for generating a dynamic image based on the panoramic image and the multiple special effect regions.
[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0008] Fourthly, embodiments of this application provide a readable storage medium on which a computer program is stored, and when executed by a processor, the computer program implements the steps of the method described in the first aspect above.
[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method described in the first aspect.
[0010] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0011] In this embodiment, after receiving the user's first input of a panoramic image, different effects are first added to different target areas in the panoramic image to obtain multiple effect areas. Then, based on the panoramic image and the multiple effect areas, a dynamic image is generated. By adding different effects to multiple different target areas in the panoramic image, multiple effects are integrated and dynamically displayed in the dynamic image, thereby enhancing the visual diversity of the dynamic image. Attached Figure Description
[0012] Figure 1 This is a flowchart of the dynamic image generation method provided in the embodiments of this application; Figure 2 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 3 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 4 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 5 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 6 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 7 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 8 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 9 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 10 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 11 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 12This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 13 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 14 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 15 This is a schematic diagram illustrating an application scenario of the dynamic image generation method provided in the embodiments of this application; Figure 16 This is a schematic diagram of the structure of the dynamic image generation device provided in the embodiments of this application; Figure 17 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application; Figure 18 This is a schematic diagram of the hardware structure of an electronic device suitable for implementing the embodiments of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0015] The dynamic image generation method and apparatus provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0016] Please refer to Figure 1 This document illustrates one of the flowcharts for a dynamic image generation method provided in an embodiment of this application. The dynamic image generation method provided in this application can be applied to electronic devices. In practice, the aforementioned electronic devices can be smartphones, tablets, laptops, wearable devices, etc.
[0017] The flow of the dynamic image generation method provided in this application embodiment includes the following steps: Step 101: Receive the user's first input on the panoramic image, which includes multiple target areas.
[0018] In this embodiment, a panoramic image refers to a static image with a wide field of view, which can be generated by a camera's panoramic shooting mode or a software stitching algorithm. Typically, panoramic images have a significantly larger aspect ratio than regular photos. The panoramic image here can be a historically captured panoramic image selected by the user from their photo album, or a panoramic image currently captured by the user using panoramic shooting mode.
[0019] The first input can be used to initiate the dynamic image generation process. The first input can be touch input, voice command, a specific gesture input by the user, or other feasible input methods; the specific method can be determined according to actual usage needs, and this application embodiment does not impose limitations. The specific gesture in this application embodiment can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area change gesture, a double-press gesture, or a double-tap gesture. The click input in this application embodiment can be a single-click input, a double-tap input, or any number of clicks, and can also be a long-press input or a short-press input. As an example, the first input can be a specific gesture performed by the user on a panoramic image, or an operation input on a specific functional control; no specific limitations are imposed here.
[0020] The target area can be a local region of the panoramic image. Target areas may be adjacent or non-adjacent, and they do not overlap. In practice, the aforementioned multiple target areas can be multiple image segments acquired in chronological order and spatial location during panoramic shooting, or multiple local areas arbitrarily selected by the user on the panoramic image through the interactive interface during post-editing of the panoramic image; no specific limitations are imposed here.
[0021] Step 102: In response to the first input, different effects are added to different target areas in multiple target areas to obtain multiple effect areas.
[0022] In this embodiment, special effects refer to the effects presented after visual transformation of an image using an artificial intelligence (AI) model, with the aim of changing the visual style of the original image. The AI model mentioned above may include, but is not limited to, large models, deep learning-based style transfer networks, generative adversarial networks, and diffusion models.
[0023] Specifically, different effect parameters can be automatically set, randomly set, or manually set by the user for different target areas. Effect parameters can include the name, keywords, and description of the effect. For each target area, a local or cloud-based AI model can be invoked, using the target area and its corresponding effect parameters as input. The model performs forward inference calculations and outputs the effect area corresponding to the target area. Alternatively, a panoramic image and the effect parameters corresponding to the target area can be used as input to the model. The model performs forward inference calculations and outputs a panoramic image with effects, which will include the effect area corresponding to the target area.
[0024] As an example, different seasonal change effects can be added to different target areas, such as Springtime, Summer, Autumn, Winter, and Snow Village effects; different time change effects can also be added to different target areas, such as Dawn, Dusk, and Night effects; different artistic style effects can also be added to different target areas, such as Oil Painting and Ink Painting effects; and different scene content change effects can also be added to different target areas, such as Urban, Snow Village, and Seaside effects.
[0025] Step 103: Generate a dynamic image based on the panoramic image and multiple special effects areas.
[0026] In this embodiment, a moving image refers to a multimedia file containing multiple frames. Common formats include, but are not limited to, GIF (Graphics Interchange Format), MP4 (Moving Picture Experts Group-4 Part 14), and other video file formats, or moving photo formats specific to mobile operating systems. Essentially, it creates a visually dynamic effect through the rapid switching of consecutive frames.
[0027] Specifically, the original panoramic image can be used as the base layer, and all effect areas can be loaded as independent, precisely aligned layers. The timeline controls the graphics rendering pipeline to set the transparency properties of these layers, allowing different layers to change from opaque to transparent or vice versa at different times, achieving an alternating display effect between the target areas and the effect areas. Finally, all rendered frame sequences are encoded together with audio to synthesize the final animated image. The cover of this animated image can use a panoramic image containing all the aforementioned effect areas.
[0028] Optionally, audio data can be acquired simultaneously during the panoramic image acquisition process. When generating a dynamic image, the audio data can be merged to obtain a dynamic image with sound. Furthermore, during or after generating the dynamic image, the user can update the audio data, or automatically generate or retrieve matching audio data based on elements in the image.
[0029] The method provided in the above embodiments of this application, after receiving the user's first input of a panoramic image, first adds different special effects to different target areas in the panoramic image to obtain multiple special effect areas, and then generates a dynamic image based on the panoramic image and the multiple special effect areas. By adding different special effects to multiple different target areas in the panoramic image, multiple special effects are integrated and dynamically displayed in the dynamic image, thereby improving the visual diversity of the dynamic image.
[0030] In some optional embodiments, before step 101 is executed, the following step may also be performed: during the process of the user capturing panoramic images in segments, the multiple image segments captured by the user in segments are regarded as multiple target regions. At this time, the display order of each special effect region in the dynamic image is the shooting order of the corresponding target region.
[0031] Panoramic images can be captured in panoramic shooting mode. Panoramic shooting mode is a specific operating mode in camera applications. In this mode, multiple frames of images with overlapping fields of view are captured continuously or intermittently by the image sensor, and the user is guided to smoothly move the device along a specific trajectory to acquire image data covering a wide field of view.
[0032] During panoramic shooting, the entire scene is not processed all at once. Instead, the continuous shooting process is divided into several discrete acquisition stages based on temporal sequence and spatial location. The set of images captured and temporarily cached in each stage constitutes an image segment. Specifically, after panoramic shooting mode is initiated, the camera module begins to continuously capture image frames. Simultaneously, the user interface provides real-time feedback, guiding the user to move or rotate the electronic device at a constant speed. In practice, the user interface may include an indicator, such as a circle or progress bar guiding the user's shooting. Each time the electronic device moves or rotates to a new predetermined position, the captured image frame is recorded and marked as an image segment. This process repeats until the user finishes shooting.
[0033] As an example, see Figure 2After the camera app is launched and switched to "Panorama" mode, the user interface updates to a style specific to panorama shooting, typically including guide arrows, horizontal reference lines, and a shutter button. With the "AI Visual Effects" function enabled, when the user taps the shutter control, a circle 1 and a positioning point 2 appear on the screen. When the user moves the phone so that circle 1 aligns with positioning point 2 for the first time, image segment 1 is obtained. The user continues to move the phone, and when circle 1 aligns with the next positioning point 2, image segment 2 is obtained. This process continues until image segments 3 and 4 are obtained. This completes the segmented shooting of four image segments.
[0034] After shooting, computer vision and image processing algorithms can be used to fuse the multiple panoramic image segments to generate a seamless and complete panoramic image. The multiple image segments captured by the user in the panoramic image represent multiple target regions. The display order of each special effects region in the dynamic image corresponds to the shooting order of the corresponding target region, i.e., the shooting order of the image segments.
[0035] By deeply binding the panoramic shooting process with the dynamic image generation logic, the display order of dynamic effects is naturally synchronized with the user's shooting trajectory, providing users with an intuitive, automated, and immersive image shooting workflow.
[0036] In some optional embodiments, step 102 may further include: Step S21: Obtain the metadata of the panoramic image. The metadata includes at least one of the following: shooting date, shooting time, and shooting location. The metadata can be used to describe the image features and shooting conditions of the panoramic image, and includes background information beyond the image content itself.
[0037] Step S22: Based on metadata, match different special effects for different target areas. Specifically, based on the shooting date, information such as season (spring, summer, autumn, winter, etc.) can be inferred; based on the shooting time, the shooting period (dawn, noon, dusk, night, etc.) can be inferred; based on the shooting location, the environment type (urban, snowy, seaside, grassland, etc.) can be inferred. The inferred season, period, and environment type can be combined to form a set of context labels, such as "spring + dawn + urban". Then, search the special effects library for the preset special effect "Spring Sunshine and Sea Dream" that best matches the context label, and use it as the special effect corresponding to one of the target areas. Afterward, change one or more labels, for example, change "spring" to "summer", "autumn", and "winter" in sequence to obtain different context labels, which are used as special effects corresponding to other target areas.
[0038] Step S23: Add matching effects to different target areas to obtain multiple effect areas.
[0039] After completing the special effects matching, the image data of each target region and the special effects parameters matched for it can be used as model input to perform forward inference calculations to obtain the special effects region corresponding to each target region.
[0040] By utilizing the metadata carried by the image itself, semantic contextual information such as season, time of day, and environment type can be obtained. By selecting special effects based on this information, appropriate special effects that conform to the logic of the physical world and the user's psychological expectations can be automatically assigned to the target area, thereby improving the efficiency and accuracy of special effects selection.
[0041] In some optional embodiments, step 102 may further include: Step S31: Receive the user's fourth input.
[0042] The fourth input can be used to select the corresponding special effect for each target area among multiple target areas. The fourth input can be touch input, voice command, a specific gesture entered by the user, or other feasible input, which can be determined according to actual usage needs, and is not limited in the embodiments of this application.
[0043] Step S32: In response to the fourth input above, determine the special effects corresponding to each target region in the plurality of target regions.
[0044] In practice, a first control corresponding to different target areas can be displayed. This first control can provide effect options; users can select the corresponding effect for each target area by interacting with the first control, thus completing the fourth input. The first control can take the form of a button, drop-down menu, icon list, slider, or thumbnail preview panel, etc.
[0045] As an example, see Figure 3The effects editing interface shown depicts a panoramic image comprising four target areas. Each target area displays a "Select AI Effect" control, which can be used as the primary control. The user wishes to set different seasonal effects for each of the four target areas of the panoramic photo. First, clicking the "Select AI Effect" control 301 below area 1 brings up a selection list containing various available effects such as "Springtime," "Summer," "Autumn," "Winter," and "Snowy Village." Selecting the "Springtime" effect from the list establishes a correspondence between area 1 and the "Springtime" effect. Next, clicking the "Select AI Effect" control below area 2 and selecting the "Summer" effect establishes a correspondence between area 2 and the "Summer" effect. Then, clicking the "Select AI Effect" control below area 3 and selecting the "Autumn" effect establishes a correspondence between area 3 and the "Autumn" effect. Finally, clicking the "Select AI Effect" control below area 4 and selecting the "Winter" effect establishes a correspondence between area 4 and the "Winter" effect. Once the effects for all target areas have been determined using the above method, the user clicks the "OK" control to complete the effect selection process, which is the fourth input step. At this point, based on the correspondence between each target area and its corresponding effect, the effect for each of the multiple target areas can be determined.
[0046] Step S33: Add corresponding special effects to multiple target areas to obtain multiple special effect areas.
[0047] Continuing with the example above, after the user selects the effects for the four target areas (Springtime, Summer, Autumn, Winter) and clicks "OK," a pop-up window prompting the system to process the AI effects appears. See [link / reference]. Figure 4 As shown in reference numeral 401, the image was then input into an AI model for processing. The AI model applied the "Springtime" effect to region 1, the "Summer" effect to region 2, the "Autumn" effect to region 3, and the "Winter" effect to region 4. After processing, four special effects regions were obtained.
[0048] Furthermore, after processing is complete, see [link to documentation]. Figure 5 This can display a "Generate Live Photo" control 501 and a "Save to Album" control 502. After the user clicks the "Generate Live Photo" control 501, step 103 can be executed, and a pop-up window indicating that the live photo is being generated will be displayed. (See [link]). Figure 6 As shown in reference numeral 601, after the user clicks the "Save to Album" control 502, the panoramic image with effects added to each target area can be directly saved to the album as a still image. Continuing the example above, this still image can include scenery from all four seasons.
[0049] By providing effects selection for each target area, the system ensures that users can control the effects of each target area and supports arbitrary combinations of effects, thereby enabling users to create an infinite variety of unique and customized visual works, greatly improving the flexibility of image generation and user engagement.
[0050] In some optional embodiments, in step 103 above, the dynamic image can be generated through the following process: First, obtain the initial animation file. The initial animation file includes particle motion effect layers, transition motion effect layers, multiple user layers, and parameter information. The parameter information includes the opacity of each user layer at different time intervals.
[0051] Specifically, the open-source animation framework PAG, based on OpenGL ES rendering technology, can be used in advance to generate .pag format files containing particle motion effect layers, transition motion effect layers, and multiple user layers. The order of each layer is as follows: Figure 7 As shown. The particle animation layer and transition animation layer in the file are used to place pre-designed particle animations and transition animations. In practice, transition animation files in .mp4 format can be created in Adobe After Effects. After creating the transition animation file, a new AE project can be created in Adobe After Effects to produce a .pag format file. Create a new layer, using multiple leaf .png format files as an example, to form a sequence of frames to create the effect of falling leaves, and place it in this layer as the particle animation layer. Then create a new layer and place the transition animation file in this layer as the transition animation layer. Finally, create multiple layers as user layers, specify the relationship between each layer and the transparency changes, and then export it as a .pag file, which can be built into the electronic device. The electronic device can generate dynamic images based on this .pag file.
[0052] Next, the panoramic image and the different effect areas from the multiple effect areas are input to different user layers within the multiple user layers to obtain the target animation file. For example, the panoramic image can be input to user layer 1 of the .pag file, and the images containing the effect areas can be input to other different user layers within the .pag file. For instance, if there are 5 user layers and 4 effect areas, the image containing effect area 1 is input to user layer 2, the image containing effect area 2 is input to user layer 3, the image containing effect area 3 is input to user layer 4, and the image containing effect area 4 is input to user layer 5.
[0053] Finally, the target animation file is rendered to generate dynamic images. Specifically, each frame of the .pag file within its rendering duration, after adding effects, is encoded to control changes in layer transparency, particle animations, and transition effects. Audio encoding is also integrated into the export process, resulting in a final .mp4 video file with sound, which users can manually replace with audio later.
[0054] Continuing with the example above, see [link to example]. Figure 7 The panoramic image includes four target regions, denoted as Region 1_a, Region 2_a, Region 3_a, and Region 4_a. After adding corresponding effects to each target region, the resulting four effect regions are denoted as Region 1_b, Region 2_b, Region 3_b, and Region 4_b. The video file content is as follows: The first stage displays user layer 1, which is the original panoramic image, including region 1_a, region 2_a, region 3_a, and region 4_a.
[0055] In the second stage, user layer 1 becomes transparent and disappears, while user layer 2 becomes opaque, causing area 1_a to become transparent and disappear, and area 1_b to become opaque.
[0056] In the third stage, user layer 2 becomes transparent and disappears, and user layer 3 becomes opaque, causing region 2_a to become transparent and disappear, and region 2_b to become opaque. At the same time, region 1_b becomes transparent and disappears, and region 1_a becomes opaque.
[0057] In the fourth stage, user layer 3 becomes transparent and disappears, and user layer 4 becomes opaque, causing area 3_a to become transparent and disappear, and area 3_b to become opaque. At the same time, area 2_b becomes transparent and disappears, and area 2_a becomes opaque.
[0058] In the fifth stage, user layer 4 becomes transparent and disappears, and user layer 5 becomes opaque, causing region 4_a to become transparent and disappear, and region 4_b to become opaque. At the same time, region 3_b becomes transparent and disappears, and region 3_a becomes opaque.
[0059] After the above stages, the final effect is that four special effect areas appear and disappear sequentially from left to right. This creates a video showing the dynamic changes of four special effects from left to right: Springtime Splendor, Summer Brilliance, Autumn Gold, and Winter Serenity. Finally, this video is combined with a panoramic photo featuring the special effects to obtain a dynamic image. The panoramic photo with the special effects can be used as the cover of this dynamic image. Users can preview the dynamic image by long-pressing or other actions.
[0060] By alternating the display of multiple original image areas and multiple special effects areas in a dynamic image, the image content in different areas and the corresponding special effects can be presented in a time-sharing and orderly manner, enriching the visual expression of the dynamic image.
[0061] In some optional embodiments, step 103 may further include: Step S41: Identify visual elements in each of the multiple target regions.
[0062] Visual elements refer to specific objects, scenes, or landscape components identified from images using computer vision techniques, such as object detection based on convolutional neural networks and semantic segmentation models. Examples include mountains, water, trees, clouds, buildings, people, and animals. These elements define the main content of an image.
[0063] Step S42: Based on the visual elements and the special effects added to each target area, generate the audio segment corresponding to each target area.
[0064] Visual elements determine the content of the audio, that is, the type of sound it contains. For example, if "mountain" and "stream" are identified, the audio may contain the sound of flowing water and possibly wind, birdsong, etc. The added effects determine the style and emotional tone of the audio. For example, the "Springtime" effect corresponds to crisp, lively, and vibrant sound effects, while the "Wintertime" effect corresponds to quiet, ethereal, and bleak sound effects.
[0065] In practice, a control can be displayed to trigger the generation of audio clips. After clicking the control, the user can search a tagged audio library for the most matching pre-stored audio file based on visual elements and special effects keywords. Alternatively, visual elements and special effects can be used as text prompts, inputting them into an AI audio generation model to generate a new, context-matched audio clip. The AI audio generation model can convert object categories, colors, and emotions into audio element parameters such as pitch, rhythm, and emotional timbre. For example, if a special effects area includes the "Springtime" effect, and the area also features the sun, mountains, and an orange and pink sunrise, the audio can include sounds of birdsong and flowing water, reflecting spring. By recognizing the overall atmosphere of early morning, sunlight, and natural tranquility, the AI audio generation model establishes an emotional mapping from visual to auditory perception. The generated bird songs are the clear calls of cuckoos and partridges, and the stream sounds are the clear, resonant sounds of flowing water, embodying the freshness and vitality of a spring morning.
[0066] Furthermore, after matching the audio to each target region, the duration of each audio segment can be calculated: First, obtain the rendering duration of the .pag file, then divide this rendering duration by the number of target regions to obtain the duration of each audio segment. For example, if the rendering duration is 6 seconds and there are 4 target regions, then the duration of each audio segment is 1.5 seconds. If the local audio is less than 1.5 seconds, it is slowed down to 1.5 seconds; if it is more than 1.5 seconds, the first 1.5 seconds are extracted to achieve audio-visual synchronization. Then, fade-in / fade-out effects are added to each 1.5-second audio segment. Finally, the audio segments are spliced together, so that the synthesized audio can be synchronized with the rendered screen and transition smoothly.
[0067] Step S43: Generate a dynamic image based on the panoramic image, multiple special effects areas, and the audio clips corresponding to each target area.
[0068] Specifically, a silent video frame sequence is first generated using a graphics rendering engine, while an empty audio timeline is created simultaneously. Then, according to the display order and duration of each target area, audio clips are arranged sequentially on the audio timeline, with fade-in and fade-out audio processing applied at clip transitions to ensure smooth transitions. Finally, using a video encoder and an audio encoder, the rendered video frame sequence is multiplexed with the processed complete audio track, encapsulating it into a single, synchronized audio-visual dynamic image.
[0069] Continuing the example above, with four effect areas—Springtime, Summer, Autumn, and Winter—four 1.5-second audio clips can be generated for each area. When generating the animated image, these four audio clips can be sequentially stitched together into a single 6-second audio clip with fade-in / fade-out effects. Then, the rendered 6-second video, showcasing the changing effects of the four seasons, is synchronized with this 6-second audio clip. Finally, when the user plays this animated image, they will hear birdsong and flowing water simultaneously with the "Springtime" effect, and hear cicadas and wind sounds when they see the "Summer" effect, achieving precise audio-visual synchronization.
[0070] By dynamically generating and synchronously matching contextualized audio with the visual content and effects for each target area, the immersiveness, emotional expressiveness, and artistic integrity of the dynamic images are greatly enhanced.
[0071] In some optional embodiments, after step 103 is performed, the following operations may also be performed: Step S51: Receive the user's fifth input.
[0072] The fifth input can be used to trigger the replacement of an audio segment. The fifth input can be touch input, voice command, a specific gesture input by the user, or other feasible input. The specific input can be determined according to actual usage needs, and this application embodiment does not limit it.
[0073] Step S52, in response to the fifth input, update the audio segment corresponding to at least one of the multiple effect regions.
[0074] In practice, a second control corresponding to different effect areas can be displayed. This second control is a visual interactive element used to trigger the audio clip replacement function; it can be displayed after the dynamic image is generated or in its editing preview interface. For each of the multiple effect areas mentioned above, after the user clicks the second control corresponding to that effect area, the audio clip corresponding to that effect area is updated. The new audio clip can be regenerated using an AI audio generation model, or it can use audio synchronously acquired during the original panoramic image acquisition process, or an option can be provided for the user to manually select it; no specific limitations are specified here.
[0075] As an example, the second control can be found here. Figure 8 The "Replace Audio" control, indicated by reference numeral 801, displays an audio selection interface when clicked by the user. (See also...) Figure 9 As shown in the image. In the audio selection interface, users can choose the original sound recorded when capturing each target area of the panoramic image, other local audio, or audio regenerated using an AI audio generation model. They can also click the "Cut" control to select audio segments of the required length. For example, if the entire rendering time is 6 seconds, and the audio segment length corresponding to each effect area is 1.5 seconds, selecting audio shorter than 1.5 seconds will be slowed down to 1.5 seconds; if longer than 1.5 seconds, users can click the "Cut" control to manually extract 1.5 seconds of content, achieving audio-visual synchronization. Afterwards, fade-in and fade-out effects can be applied. Finally, the dynamic image is re-encoded and synthesized, replacing the original audio with the new audio.
[0076] By providing independent audio clip replacement functionality for each special effects area, it enables refined and personalized post-editing of audio for dynamic images, improving the flexibility of dynamic image generation and editing.
[0077] In some optional embodiments, after performing step 103, the special effects in the panoramic image or dynamic image can be edited, such as deleted or replaced. Controls for editing the special effects can be displayed in interfaces such as the panoramic photo preview interface, the special effects editing interface, or the dynamic image preview interface. Users can edit the special effects by operating these controls.
[0078] As an example, see Figure 10In the album interface, users can open a panoramic photo that has undergone AI visual effects processing. Clicking the "Clear Effects" control 1001 will clear all effects from the panoramic photo, displaying the original image. Users can then save the original image to their album using buttons or gestures. When the user clicks the "Overall Effect Replacement" control 1002, the following will be displayed: Figure 3 The special effects editing interface is shown. Therefore, by providing special effects editing functions, users can flexibly adjust the generated dynamic images, thereby further meeting their needs.
[0079] In some optional embodiments, prior to step 101, the following steps may also be performed to obtain multiple target regions in the panoramic image: Step S61: Receive the user's second input on the panoramic image in the album interface.
[0080] Step S62, in response to the second input, determines multiple target regions.
[0081] The second input can be used to select multiple target regions in the panoramic image. The second input can be touch input, voice command, a specific gesture input by the user, or other feasible input methods; the specific method can be determined according to actual usage needs, and this application embodiment does not limit this. In practice, a region division tool, such as a rectangular or free-form selection tool, can be provided to receive multiple closed regions outlined by the user in the panoramic image through touch gestures, and then multiple target regions can be obtained based on the coordinate information of these regions.
[0082] As an example, see Figure 11 Users can open a panoramic photo without AI visual effects processing in the photo album application, and after clicking the "AI visual effects processing" control 1101, they can select the target area of the panoramic photo to complete the second input.
[0083] As yet another example, see Figure 10 In the album interface, users can open a panoramic photo that has already undergone AI visual effects processing. Clicking the "Clear Effects" control 1001 will clear all effects from the panoramic photo, displaying the original image browsing interface. (See also...) Figure 12 As shown. The original image browsing interface may include a "Partial Selection AI Visual Effect Processing" control 1201. After clicking on this control 1201, the user can enter the target area editing interface for the panoramic photo. See [link to target area editing interface] for details. Figure 13 As shown. Users can select at least one target area in the panoramic photo within the target area editing interface, for example... Figure 13 The second input is completed by referring to "Region 1", "Region 2", and "Region 3".
[0084] Furthermore, users can specify AI effects for each target area in the target area editing interface. After clicking the "OK" control, step 102 can be executed. After the background processing is complete, a panoramic photo with the corresponding effects applied to each target area will be obtained, and "Save to Album" and "Generate Live Photo" controls will be displayed below the panoramic photo. If the user clicks the "Save to Album" control, the panoramic image with effects can be saved directly. If the user clicks the "Generate Live Photo" control, step 103 can be executed to generate a dynamic image.
[0085] Furthermore, the following steps can be performed to set the sorting order of the target areas: Step S63: Receive third input from the user regarding the aforementioned multiple target areas.
[0086] Step S64: In response to the third input, determine the arrangement order of the above multiple target regions, and the display order of each special effect region in the dynamic image is the arrangement order of the corresponding target regions.
[0087] The third input can be used to set the arrangement order of the multiple target areas. The third input can be touch input, voice command, a specific gesture input by the user, or other feasible input. The specific input can be determined according to actual usage needs, and this application embodiment does not limit it.
[0088] As an example, after selecting the target area, a sortable control corresponding to each target area can be displayed. Users can long-press and drag the sortable controls to reset the order of the target areas, thus enabling third-party input. See also Figure 14 The interface displays three sortable controls: "Area 1", "Area 2", and "Area 3". Users can choose to display effects in the order of Area 2 before Area 1, and Area 1 before Area 3. To do this, long-press the "Area 2" control and drag it to the first position; then long-press the "Area 1" control and drag it to the second position; the "Area 3" control will automatically be placed at the end.
[0089] Subsequently, when different effects were added to these three areas and animated images were generated, the display order of the effect areas in the animated images was: area 2 with effects, area 1 with effects, and area 3 with effects. Specifically, continuing with the example of a .pag file built into an electronic device, see [link to relevant documentation]. Figure 15 First, delete "User Layer 5" from the .pag file. Then, control the transparency of User Layers 1-4 according to the display order of User-specified Regions 1, 2, and 3 to create the effect of each effect region appearing and disappearing in sequence. Finally, export a .mp4 format video file. The video content is as follows: The first stage involves displaying user layer 1.
[0090] In the second stage, user layer 1 becomes transparent and disappears, while user layer 2 becomes opaque and is displayed.
[0091] In the third stage, user layer 2 becomes transparent and disappears, while user layer 3 becomes opaque and is displayed.
[0092] In the fourth stage, user layer 3 becomes transparent and disappears, while user layer 4 becomes opaque and is displayed.
[0093] The final result is the sequential appearance and disappearance of areas 2, 1, and 3 with special effects. Finally, the video file and the panoramic image with special effects are combined into a single animated image, which serves as the cover of the animated image. Users can long-press on this screen to preview the animated image.
[0094] It should be noted that steps S63 and S64 above can be performed before step 103.
[0095] Through the above process, users can define the special effects area and its display order independently in the later editing stage, realizing a high degree of flexibility and personalization in dynamic image creation, thus enabling users to create dynamic images freely.
[0096] It should be noted that the dynamic image generation method provided in this application embodiment can be executed by a dynamic image generation device. This application embodiment uses a dynamic image generation device executing the dynamic image generation method as an example to illustrate the dynamic image generation device provided in this application embodiment.
[0097] like Figure 16 As shown, the dynamic image generation device 1600 of this embodiment includes: a receiving unit 1601, used to receive a first input from a user to a panoramic image, the panoramic image including multiple target regions; an adding unit 1602, used to add different special effects to different target regions in response to the first input, to obtain multiple special effect regions; and a generating unit 1603, used to generate a dynamic image based on the panoramic image and the multiple special effect regions.
[0098] In some optional implementations of this embodiment, the generation unit 1603 is further configured to: obtain an initial animation file, the initial animation file including a particle motion effect layer, a transition motion effect layer, multiple user layers and parameter information, the parameter information including the transparency of each user layer in the multiple user layers at different time periods; input the panoramic image and different special effect areas in the multiple special effect areas to different user layers in the multiple user layers to obtain a target animation file; render the target animation file to generate the dynamic image.
[0099] In some optional implementations of this embodiment, the device further includes a setting unit, configured to: during the process of the user capturing the panoramic image in segments, use multiple image segments captured by the user in segments as the multiple target regions, and the display order of each special effect region in the dynamic image is the shooting order of the corresponding target regions; or, receive a second input from the user on the panoramic image in the album interface; in response to the second input, determine the multiple target regions; receive a third input from the user on the multiple target regions; in response to the third input, determine the arrangement order of the multiple target regions, and the display order of each special effect region in the dynamic image is the arrangement order of the corresponding target regions.
[0100] In some optional implementations of this embodiment, the adding unit 1602 is further configured to: acquire metadata of the panoramic image, the metadata including at least one of the following: shooting date, shooting time, shooting location; match different special effects for the different target areas based on the metadata; add the matched special effects to the different target areas to obtain multiple special effect areas.
[0101] In some optional implementations of this embodiment, the adding unit 1602 is further configured to: receive a fourth input from the user; in response to the fourth input, determine the special effects corresponding to each target region in the plurality of target regions; add the corresponding special effects to the plurality of target regions to obtain the plurality of special effect regions.
[0102] In some optional implementations of this embodiment, the generation unit 1603 is further configured to: identify visual elements in each of the plurality of target regions; generate an audio segment corresponding to each target region based on the visual elements and the special effects added to each target region; and generate the dynamic image based on the panoramic image, the plurality of special effect regions, and the audio segment corresponding to each target region.
[0103] In some optional implementations of this embodiment, the device further includes an update unit, configured to: receive a fifth input from the user; and update the audio segment corresponding to at least one of the multiple effect regions in response to the fifth input.
[0104] The apparatus provided in the above embodiments of this application, after receiving a user's first input of a panoramic image, first adds different special effects to different target areas in the panoramic image to obtain multiple special effect areas, and then generates a dynamic image based on the panoramic image and the multiple special effect areas. By adding different special effects to multiple different target areas in the panoramic image, multiple special effects are integrated and dynamically displayed in the dynamic image, thereby enhancing the visual diversity of the dynamic image.
[0105] The dynamic image generation device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0106] The motion image generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.
[0107] The dynamic image generation device provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0108] Optionally, such as Figure 17 As shown, this application embodiment also provides an electronic device 1700, including a processor 1701 and a memory 1702. The memory 1702 stores a program or instructions that can run on the processor 1701. When the program or instructions are executed by the processor 1701, they implement the various steps of the above-described dynamic image generation method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0109] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.
[0110] Figure 18 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application. The electronic device 1800 includes, but is not limited to, components such as: radio frequency unit 1801, network module 1802, audio output unit 1803, input unit 1804, sensor 1805, display unit 1806, user input unit 1807, interface unit 1808, memory 1809, and processor 1810.
[0111] Those skilled in the art will understand that the electronic device 1800 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1810 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 18 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here. The processor 1810 is configured to receive a first input from a user to a panoramic image via a user input unit 1807, the panoramic image including multiple target regions; in response to the first input, add different effects to different target regions within the multiple target regions to obtain multiple effect regions; and generate a dynamic image based on the panoramic image and the multiple effect regions.
[0112] By adding different effects to multiple target areas in a panoramic image, various effects can be integrated and dynamically displayed in a dynamic image, thereby enhancing the visual diversity of the dynamic image.
[0113] In some optional implementations of this embodiment, the processor 1810 is further configured to acquire an initial animation file, the initial animation file including a particle motion effect layer, a transition motion effect layer, multiple user layers and parameter information, the parameter information including the transparency of each user layer in the multiple user layers at different time periods; input the panoramic image and different effect areas in the multiple effect areas to different user layers in the multiple user layers to obtain a target animation file; render the target animation file to generate the dynamic image.
[0114] In some optional implementations of this embodiment, the processor 1810 is further configured to, during the process of the user capturing the panoramic image in segments, use multiple image segments captured by the user in segments as the multiple target regions, and the display order of each special effect region in the dynamic image is the shooting order of the corresponding target regions; or, receive a second input from the user on the panoramic image in the album interface through the user input unit 1807; determine the multiple target regions in response to the second input; receive a third input from the user on the multiple target regions through the user input unit 1807; determine the arrangement order of the multiple target regions in response to the third input, and the display order of each special effect region in the dynamic image is the arrangement order of the corresponding target regions.
[0115] In some optional implementations of this embodiment, the processor 1810 is further configured to acquire metadata of the panoramic image, the metadata including at least one of the following: shooting date, shooting time, and shooting location; based on the metadata, match different special effects for the different target areas; add the matched special effects to the different target areas to obtain multiple special effect areas.
[0116] In some optional implementations of this embodiment, the processor 1810 is further configured to receive a fourth input from a user via the user input unit 1807; in response to the fourth input, determine the special effects corresponding to each target region in the plurality of target regions; and add corresponding special effects to the plurality of target regions to obtain the plurality of special effect regions.
[0117] In some optional implementations of this embodiment, the processor 1810 is further configured to identify visual elements in each of the plurality of target regions; generate an audio segment corresponding to each target region based on the visual elements and the special effects added to each target region; and generate the dynamic image based on the panoramic image, the plurality of special effect regions, and the audio segment corresponding to each target region.
[0118] In some optional implementations of this embodiment, the processor 1810 is further configured to receive a fifth input from a user via the user input unit 1807; and in response to the fifth input, update the audio segment corresponding to at least one of the multiple effect regions.
[0119] It should be understood that, in this embodiment, the input unit 1804 may include a graphics processing unit (GPU) 18041 and a microphone 18042. The GPU 18041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1806 may include a display panel 18061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1807 includes at least one of a touch panel 18071 and other input devices 18072. The touch panel 18071 is also called a touch screen. The touch panel 18071 may include a touch detection device and a touch controller. Other input devices 18072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0120] The memory 1809 can be used to store software programs and various data. The memory 1809 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1809 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1809 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0121] Processor 1810 may include one or more processing units; optionally, processor 1810 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1810.
[0122] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described dynamic image generation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0123] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0124] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described dynamic image generation method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0125] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0126] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described dynamic image generation method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0127] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0129] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method for generating dynamic images, characterized in that, The method includes: Receive the user’s first input on a panoramic image, which includes multiple target regions; In response to the first input, different special effects are added to different target regions in the plurality of target regions to obtain multiple special effect regions; Based on the panoramic image and the multiple special effects areas, a dynamic image is generated.
2. The method according to claim 1, characterized in that, The generation of dynamic images based on the panoramic image and the multiple special effects regions includes: Obtain an initial animation file, which includes a particle motion effect layer, a transition motion effect layer, multiple user layers, and parameter information. The parameter information includes the transparency of each user layer in the multiple user layers at different time periods. The panoramic image and different special effects regions in the multiple special effects regions are input into different user layers in the multiple user layers to obtain the target animation file; The target animation file is rendered to generate the dynamic image.
3. The method according to claim 1, characterized in that, Prior to receiving the user's first input on the panoramic image, the method further includes: During the process of the user capturing the panoramic image in segments, the multiple image segments captured by the user in segments are used as the multiple target areas, and the display order of each special effect area in the dynamic image is the shooting order of the corresponding target area. or, Receive a second input from the user regarding the panoramic image displayed in the album interface; In response to the second input, the plurality of target regions are determined; Receive third input from the user regarding the multiple target regions; In response to the third input, the arrangement order of the plurality of target regions is determined, and the display order of each special effect region in the dynamic image is the arrangement order of the corresponding target regions.
4. The method according to claim 1, characterized in that, The process involves adding different special effects to different target regions within the multiple target regions to obtain multiple effect regions, including: Obtain the metadata of the panoramic image, which includes at least one of the following: shooting date, shooting time, and shooting location; Based on the metadata, different special effects are matched for the different target areas; Add matching effects to the different target areas to obtain multiple effect areas.
5. The method according to claim 1, characterized in that, The process involves adding different special effects to different target regions within the multiple target regions to obtain multiple effect regions, including: Receive the user's fourth input; In response to the fourth input, the special effects corresponding to each target region in the plurality of target regions are determined; Add corresponding special effects to the multiple target areas to obtain the multiple special effect areas.
6. The method according to claim 1, characterized in that, The generation of dynamic images based on the panoramic image and the multiple special effects regions includes: Identify visual elements in each of the plurality of target regions; Based on the visual elements and the special effects added to each target area, an audio segment corresponding to each target area is generated; The dynamic image is generated based on the panoramic image, the multiple special effects areas, and the audio segment corresponding to each target area.
7. The method according to claim 6, characterized in that, After generating the dynamic image, the method further includes: Receive the user's fifth input; In response to the fifth input, the audio segment corresponding to at least one of the multiple effect regions is updated.
8. A dynamic image generation device, characterized in that, The device includes: A receiving unit is used to receive a user's first input to a panoramic image, the panoramic image including multiple target areas; An adding unit is configured to, in response to the first input, add different effects to different target regions within the plurality of target regions, thereby obtaining a plurality of effect regions; The generation unit is used to generate a dynamic image based on the panoramic image and the multiple special effects areas.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the dynamic image generation method as described in any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the dynamic image generation method as described in any one of claims 1-7.
Citation Information
Cited By
Image processing method and apparatus, electronic device, medium, and program product
US20250315993A1