Dynamic photo generation method and device and electronic equipment

By generating dynamic photos with audio duration exceeding video duration, the problems of large memory usage and slow transmission of video files are solved, enabling efficient recording and sharing of on-site audio and video information.

CN120416615APending Publication Date: 2025-08-01VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510536744.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing technologies produce large video files when filming concerts, karaoke sessions, and other similar events. These files consume a lot of electronic device memory and take a long time to share and transmit, causing inconvenience to users.

Method used

Generate and store dynamic photos that combine cover images, videos, and audio. The audio duration is longer than the video duration, and the dynamic photo file is smaller than the video file. Record on-site video and audio information by taking dynamic photos.

Benefits of technology

It reduces the memory usage of electronic devices, shortens sharing and transmission time, and allows dynamic photos to carry more audio information and record more on-site information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416615A_ABST
    Figure CN120416615A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic photo generation method and device and electronic equipment, and belongs to the technical field of photography. The method comprises the following steps: receiving first input; outputting a first dynamic photo in response to the first input; the first dynamic photo comprises a first cover image, a first video and a first audio which are stored in an associated manner; playing the first video and the first audio under the condition that a second input for the first cover image is received; wherein the audio duration of the first audio is greater than the video duration of the first video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of camera technology, and particularly relates to a method, apparatus, and electronic device for generating dynamic photos. Background Art

[0002] With the rapid development of electronic device photography technology, people increasingly like to record their lives by taking photos or videos. Currently, for shooting scenarios such as concerts, KTVs, talk shows, etc., if a user wants to record the on-site images and sound information, they can only record a video with a relatively long duration. However, a video with a long duration will occupy a large amount of memory resources on the electronic device, and due to the large size of the video file, when sharing it on Moments or sharing it with friends, the time for uploading or transmitting the video file is relatively long, which brings inconvenience to users. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a method, apparatus, and electronic device for generating dynamic photos, which can capture dynamic photos that are associated with and store a cover image, a video, and an audio, and the audio duration of the audio is longer than the video duration of the video. The dynamic photo can carry more audio information, and the file size of the dynamic photo is much smaller than that of the video file. It can record the images and more audio information of the shooting scene. When sharing the dynamic photo, the time for uploading or transmitting the dynamic photo is much shorter than that of the video file.

[0004] In a first aspect, the embodiments of this application provide a method for generating dynamic photos, which includes:

[0005] Receiving a first input;

[0006] In response to the first input, displaying a first dynamic photo; the first dynamic photo includes an associated and stored first cover image, a first video, and a first audio;

[0007] When a playback control input for the first cover image is received, playing the first video and the first audio;

[0008] Wherein, the audio duration of the first audio is longer than the video duration of the first video.

[0009] In a second aspect, the embodiments of this application provide a device for generating dynamic photos, which includes:

[0010] A receiving module, configured to receive a first input;

[0011] A display module, configured to display a first dynamic photo in response to the first input; the first dynamic photo includes an associated and stored first cover image, a first video, and a first audio;

[0012] A playback module, configured to play the first video and the first audio when receiving a playback control input for the first cover image; wherein, an audio duration of the first audio is greater than a video duration of the first video.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.

[0014] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.

[0015] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is configured to run a program or instruction to implement the method described in the first aspect.

[0016] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the method described in the first aspect.

[0017] In the embodiment of the present application, by responding to the user's first input, a first live photo including an associated stored first cover image, a first video, and a first audio can be displayed. In this way, when the user wants to record the on-site pictures and sound information of shooting scenes such as concerts, KTVs, talk shows, etc., it can be achieved by taking a live photo, without taking a video, reducing the occupation of memory resources in the electronic device. At the same time, in the live photo generated by the embodiment of the present application, the audio duration of the first audio is greater than the video duration of the first video. In this way, the generated live photo can carry more audio information, and the file size of the live photo is much smaller than that of the video file, capable of recording the images at the shooting scene and more audio information at the shooting scene. When sharing the live photo, the upload or transmission time of the live photo is much shorter than that of the video file. Description of the Drawings

[0018] Figure 1 is a schematic flowchart of a method for generating a live photo provided by some embodiments of the present application;

[0019] Figure 2 is a schematic diagram of a photo-taking preview interface when the live photo-taking function is not enabled provided by some embodiments of the present application;

[0020] Figure 3It is a schematic diagram of a photo preview interface when the dynamic photo-taking function is enabled provided by some embodiments of the present application;

[0021] Figure 4 It is a schematic diagram of the principle for generating a first dynamic photo based on an electronic device provided by some embodiments of the present application;

[0022] Figure 5 It is a schematic diagram of the interface of a file management application provided by some embodiments of the present application;

[0023] Figure 6 It is a schematic diagram of an audio recording duration interface provided by some embodiments of the present application;

[0024] Figure 7 It is a schematic diagram of the main interface of an album application provided by some embodiments of the present application;

[0025] Figure 8 It is a schematic diagram of the display of a first dynamic photo provided by some embodiments of the present application;

[0026] Figure 9 It is a schematic diagram of an editing interface for a first dynamic photo provided by some embodiments of the present application;

[0027] Figure 10 It is a schematic diagram of an audio editing interface for a first dynamic photo provided by some embodiments of the present application;

[0028] Figure 11 It is a schematic diagram of an audio editing interface for a first dynamic photo provided by some embodiments of the present application;

[0029] Figure 12 It is a schematic diagram of the interface of a file management application provided by some embodiments of the present application;

[0030] Figure 13 It is a schematic diagram of the main interface of an album application provided by some embodiments of the present application;

[0031] Figure 14 It is a schematic diagram of the display of a second dynamic photo provided by some embodiments of the present application;

[0032] Figure 15 It is a schematic diagram of an editing interface for a second dynamic photo provided by some embodiments of the present application;

[0033] Figure 16 It is a schematic diagram of an audio editing interface for a second dynamic photo provided by some embodiments of the present application;

[0034] Figure 17 It is a schematic diagram of an audio editing interface for a second dynamic photo provided by some embodiments of the present application;

[0035] Figure 18 It is a schematic structural diagram of a dynamic photo generation device provided by some embodiments of the present application;

[0036] Figure 19 It is a schematic structural diagram of an electronic device provided by some embodiments of the present application;

[0037] Figure 20 It is a schematic hardware structure diagram of an electronic device provided by some embodiments of the present application. Specific Embodiments

[0038] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application belong to the scope of protection of the present application.

[0039] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or N. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means an "or" relationship between the associated objects before and after.

[0040] The terms used in the embodiments part of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.

[0041] Next, the terms related to the embodiments of the present invention will be explained.

[0042] The identifiers in the present application are used to indicate information such as text, symbols, images, etc., and can use controls or other containers as carriers for displaying information, including but not limited to text identifiers, symbol identifiers, and image identifiers.

[0043] Control: It refers to a graphical element that can be directly operated or perceived by the user during the interaction process, used to receive user input, trigger functions, or display real-time information. It is a bridge for the user to interact with the device or application functions, and realizes instruction transmission and status feedback through a visual form.

[0044] Dynamic Photo: A dynamic photo generally includes a static cover image and a short video. A dynamic photo can be understood as a data set composed of a static cover image and a short video. The static cover image and the short video can be regarded as a whole to understand the data composition of the dynamic photo. A two-way pointer is established between the static cover image and the short video through a lightweight data format or an index file of Extensible Markup Language (XML) to associate the data structures of the static cover image and the short video. Users can play the short video by operating on the static image, such as the common long-press operation. For example, a general dynamic photo includes a static image in the Joint Photographic Experts Group (JEPG) format and a video in the Moving Picture Experts Group 4 (MP4) format, which are stored associatively. A dynamic photo can also have other names, such as Live Photo, dynamic picture, live photo, etc. When users browse photos in the photo library or album, the first thing they see is this static cover image, which is equivalent to the home page or entrance of the short video, similar to the cover of a book or magazine. This static cover image provides a quick visual overview, allowing users to understand the theme or scene of the content without playing the short video. Therefore, the "cover image" not only technically identifies the static component of the entire dynamic photo but also provides a visual or interactive entrance in the user experience. If a user wants to delete a dynamic photo, both the static cover image and the short video that make up the dynamic photo will be deleted. Correspondingly, if a user wants to share a dynamic photo, the static cover image and the short video that make up the dynamic photo will be sent to the sharing destination together.

[0045] Interface: Refers to the graphical interaction layer that users see through the screen of an electronic device. That is, the "user interface (UI)", which is the media interface for interaction and information exchange between an application or an operating system and users, and it realizes the conversion between the internal form of information and the form that users can accept. The user interface is the source code written in specific computer languages such as Java and XML. The interface source code is parsed and rendered on the electronic device and finally presented as content that users can recognize. The common manifestation form of the user interface is the graphical user interface (GUI), which refers to the user interface related to computer operations displayed in a graphical way. It can be visible interface elements such as text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and web widgets (Widgets) displayed on the display screen of the electronic device.

[0046] Application: It refers to a computer program developed and run on an operating system to complete a certain or certain specific tasks. The application runs in user mode, can interact with users, and has a visual user interface.

[0047] Photo-taking preview interface: It is the interface of the real-time viewfinder screen that the user sees through the device screen before taking a photo. It is a visual interaction area that is rendered in real time after the image data captured by the camera sensor is processed. It can be a GUI, on which visual interface elements such as buttons, navigation bars, and Widgets can be displayed.

[0048] The technical solution of the embodiment of the present application can be applied to the scenario of generating dynamic photos with different video durations and audio durations. For example, the user is playing with friends Xiaohong and Xiaoming in a karaoke. At this time, the user wants to take dynamic photos of Xiaohong singing song A and Xiaoming singing song B respectively. However, in the two dynamic photos that the user wants to take, the user wants the video duration of Xiaohong and Xiaoming to be shorter, but the audio duration of Xiaohong and Xiaoming singing to be longer.

[0049] The following will combine the accompanying drawings to explain in detail the dynamic photo generation method provided by the embodiment of the present application through specific embodiments and their application scenarios.

[0050] Figure 1 It is a schematic flowchart of a dynamic photo generation method provided by an embodiment of the present application. The execution subject of this dynamic photo generation method can be an electronic device, which can be but is not limited to a personal computer (PC), a smart phone, a tablet computer, or a personal digital assistant (PDA), etc.

[0051] As Figure 1 shown, the dynamic photo generation method provided by the embodiment of the present application may include step 110 to step 130.

[0052] Step 110: Receive a first input.

[0053] Among them, the first input may be the user's input to the camera preview interface. The above first input is used to output a first live photo, and the first input may be a first operation. Exemplarily, the above first input includes but is not limited to: the user's touch input to the shooting control in the camera preview interface through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above first input may be: the user's touch input to the shooting control in the camera preview interface. For example, the above first input may be: the user's click input to the shooting control in the camera preview interface.

[0054] In some embodiments of the present application, the camera preview interface may include a toolbar, a preview window, a camera mode option, an album quick control, a shooting control, and a camera flip control, etc. As Figure 2 shown, the camera preview interface includes a toolbar 21, a preview window 22, a camera mode option, an album quick control 24, a shooting control 25, and a camera flip control 26.

[0055] The toolbar 21 can be used to display one or more function controls. As Figure 2 shown, the one or more function controls may include a flash control 211, a live photo control 212, and a settings control 213, etc. It can be understood that the toolbar may also include more or fewer function controls.

[0056] The live photo function indicated by the live photo control 212 can capture video and audio within a period of time before and after the image is taken. For example, the live photo function can record the taken image and the video of 1 second before and after the moment when the shooting control is pressed, that is, a total of 2 seconds of video, and the audio of 1 second before and 7 seconds after the moment when the shooting control is pressed, to obtain a live photo containing 2 seconds of video and 8 seconds of audio.

[0057] The preview window 22 can be used to display the preview image obtained after real-time rendering of the image data, and the preview window can also be called a viewfinder.

[0058] The camera mode option can be a preset function module that can be selected by the user, used to adapt to different shooting scenarios or operation requirements, and realize specific shooting effects or creative intentions by adjusting the combination of hardware parameters and software algorithms. As Figure 2As shown, the camera mode options may include a portrait mode option 231, a video recording mode option 232, a photo shooting mode option 233, etc.

[0059] Step 120: In response to a first input, display a first live photo.

[0060] Among them, the first live photo may include an associated stored first cover image, a first video, and a first audio. The first cover image may be an image used to indicate the first live photo. Specifically, the first cover image may be a thumbnail of the first live photo. This first cover image is generally a static picture and is the picture in the first live photo. The first video may be a video of a certain duration intercepted from the video continuously recorded by the camera in the background after the first input. The first audio may be an audio of a certain duration intercepted from the audio recorded by the microphone in the background after the first input.

[0061] In some embodiments of the present application, when the first input is a control input to a photo shooting control, before step 110, the method further includes:

[0062] When the live photo function is turned on, continuously cache the image data collected by the camera, control the microphone to perform audio recording, and continuously cache the digital audio data collected by the microphone;

[0063] Before displaying the first live photo, the method involved above may further include:

[0064] Perform image processing on the image data cached at the reference moment to generate the first cover image;

[0065] Perform video encoding on all the image data cached within the first time period to generate the first video;

[0066] Perform audio encoding on all the digital audio data cached within the second time period to generate the first audio;

[0067] Associatively store the first cover image, the first audio, and the first video.

[0068] Among them, the image data may be the image data collected by the camera when the live photo function is turned on.

[0069] The reference moment may be the input moment of the first input. For example, if the user clicks the shooting control at 10:00 am on March 25, 2024, the reference moment is the system time of the electronic device when the user clicks the shooting control, which is 10:00 am on March 25, 2024.

[0070] The first time period consists of a first duration before the reference time and a second duration after the reference time. The first duration can be a period of time before the reference time. For example, the first duration can be 1 second before the reference time. The second duration can be a period of time after the reference time. For example, the second duration can be 1 second after the reference time.

[0071] The second time period consists of a first duration before the reference time and a third duration after the reference time. The third duration can be a period of time after the reference time, and the third duration can be greater than the second duration. For example, when the second duration is 1 second after the reference time, the third duration can be 7 seconds after the reference time.

[0072] In some embodiments of the present application, when the user needs to take a photo, the camera application can be opened first, and then as Figure 2 shown, the photo-taking preview interface is displayed. At this time, the dynamic photo function is not yet enabled. If the user wants to enable the dynamic photo, the user clicks on Figure 2 the "Dynamic Photo" control 212 in Figure 3 then the dynamic photo function is enabled as Figure 2 shown. After the dynamic photo function is enabled, the "Dynamic Photo" control 212 is updated from the Figure 3 shown form to the

[0073] shown form, and the camera continuously records video in the background, and the microphone continuously records audio in the background. When the dynamic photo function is turned off, the recording of video and audio in the background is stopped, and the normal photo-taking preview display state is restored. Figure 4 Figure 4 Refer to

[0074]

[0075] Figure 4 is a schematic diagram of the principle of generating dynamic photos based on an electronic device. When the dynamic photo function is enabled, the camera 1 continuously sends the RAW format image data collected to the Image Signal Processor (ISP) 2 for image processing such as demosaicing, white balance calibration, noise reduction, sharpening, and dynamic range optimization, and then obtains the RGB format image data. By converting the RGB format image data in the color space, the YUV format image data can be obtained, and then the YUV format image data is sent to the image data cache space 3 for caching. The display module 4 of the electronic device can obtain the cached YUV format image data from the image data cache space 3 in real time, and perform real-time rendering and display on the YUV format image data, so that the preview image can be displayed in real time in the photo-taking preview interface.

[0074] It should be noted that the above image data cache space 3 can be the memory of the electronic device or the video coding buffer of the electronic device.

[0075] Then, after the electronic device receives the first input from the user to the photographing control 5, it can generate a photographing instruction. Based on this photographing instruction, the image encoder 6 can obtain the image data cached at the moment of the first input from the image data cache space 3, and perform discrete cosine transform (DCT) processing, quantization, and encoding on the obtained image data, and then the first cover image can be obtained. The format of the first cover image can be the JPEG format. At the same time, after receiving the first input from the user to the photographing control 5, the video encoder 7 can also obtain all the image data cached within the first time period from the image data cache space 3, and perform video encoding on all the image data cached within the first time period, and then the first video can be generated.

[0076] When the dynamic photo function is enabled, the microphone 8 continuously records audio in the background. Specifically, the microphone 8 can convert the sound waves collected through the diaphragm vibration into analog electrical signals, and then generate digital audio data through analog-to-digital conversion. The format of the digital audio data can be, for example, the PCM format. Then, after receiving the first input from the user to the photographing control 5, the microphone 8 can send the continuously collected digital audio data to the audio cache space 9 for caching. For example, the microphone can send the continuously collected digital audio data to the electronic device memory or the system audio driver buffer for caching. At the same time, all the digital audio data cached within the second time period is sent to the audio encoder 10 for audio encoding. Specifically, it can perform lossy compression on all the digital audio data cached within the second time period, and remove the frequency bands that are insensitive to the human ear, and then the first audio can be generated. The format of the first audio can be the WAV format.

[0077] Continuing to refer to the above example, taking the dynamic photo of Xiaohong singing song A that the user wants to take as an example, which includes the audio of 8 seconds when Xiaohong sings song A and the video of 2 seconds within the 8 - second audio when Xiaohong sings song A, that is, the first duration is 1 second, the second duration is 1 second, and the third duration is 7 seconds. As Figure 2 shown, if the user clicks the "Dynamic Photo" control 212 at 9:59 am on March 25, 2024, then at this time, the camera in the electronic device starts to continuously collect image data, and sends the collected image data to the cache space for caching. At the same time, the microphone starts to continuously record digital audio data, and sends the recorded digital audio data to the cache space for caching.

[0078] If the user clicks the "Shoot" control 25 at 10:00 am on March 25, 2024, the image data cached at 10:00 am on March 25, 2024 can be sent to an image encoder for image encoding processing, generating a JPEG-format image and saving it to the album application. This JPEG-format image is the first cover image. Also, the image data cached from 9:59:59 am to 10:00:01 am on March 25, 2024 is sent to a video encoder, and DCT processing, quantization, and video encoding are performed on the image data cached from 9:59:59 am to 10:00:01 am on March 25, 2024. Then, the video-encoded image data is encapsulated to obtain a video of Xiaohong singing song A from 9:59:59 am to 10:00:01 am on March 25, 2024. The video of Xiaohong singing song A from 9:59:59 am to 10:00:01 am on March 25, 2024 is the first video.

[0079] After the user clicks the "Shoot" control 25 at 10:00 am on March 25, 2024, simultaneously, the digital audio data cached from 9:59:59 am to 10:00:07 am on March 25, 2024 can be sent to an audio encoder for lossy compression of the digital audio data cached from 9:59:59 am to 10:00:07 am on March 25, 2024, obtaining the audio data of Xiaohong singing song A from 9:59:59 am to 10:00:07 am on March 25, 2024. The audio data of Xiaohong singing song A from 9:59:59 am to 10:00:07 am on March 25, 2024 is the first audio.

[0080] In some embodiments of the present application, after generating the first cover image, the first video, and the first audio, the first cover image, the first video, and the first audio can be respectively associated and stored in the file management application of the electronic device. As Figure 5 shown, the first video is stored in the video folder 41 of the file management application, the first audio is stored in the music folder 42 of the file management application, and the first cover image is stored in the image folder 43 of the file management application.

[0081] In an embodiment of the present application, when the dynamic photo function is enabled, image data captured by the camera can be continuously cached, the microphone is controlled to perform audio recording, and the digital audio data captured by the microphone is continuously cached. Thus, after receiving the first input from the user, image processing can be performed on the image data cached at the moment when the first input is executed to generate a first cover image, all the image data cached within a first time period is video-encoded to generate a first video, and all the digital audio data cached within a second time period is audio-encoded to generate a first audio. Then, the first cover image, the first video, and the first audio are associated and stored. In this way, the audio duration and the video duration in the obtained dynamic photo may not be of equal length, and the audio duration of the first audio is greater than the video duration of the first video. Therefore, compared with the dynamic photo obtained by the traditional method, the dynamic photo of the embodiment of the present application can carry more audio information, and the file size of the dynamic photo is much smaller than that of the video file. It not only records the image of the shooting scene but also records more audio information of the shooting scene, and is also convenient for post-editing of the first audio. In addition, the dynamic photo generated by the embodiment of the present application also occupies less memory resources than the long video captured, reducing the consumption of memory resources in the electronic device. When sharing the dynamic photo, the upload or transmission time of the dynamic photo is much shorter than that of the video file.

[0082] It should be noted that after the user clicks the shooting control, the duration of the video intercepted from the recorded video and the duration of the audio intercepted from the audio recorded by the microphone can be determined based on the setting operation of the user before executing the first input. The specific determination methods for the duration of the video intercepted from the recorded video and the duration of the audio intercepted from the audio recorded by the microphone are as follows:

[0083] Receive the sixth input from the user;

[0084] In response to the sixth input, display an audio duration setting interface;

[0085] Receive the control input of the user for any one of at least one audio recording duration option in the audio duration setting interface;

[0086] In response to the control input, determine a first duration, a second duration, and a third duration according to the duration corresponding to the audio recording duration option selected by the control input.

[0087] Among them, the sixth input may be the user's input to the setting control in the photo-taking preview interface. The above-mentioned sixth input is used to display the audio duration setting interface, and the sixth input may be a sixth operation. Exemplarily, the above-mentioned sixth input includes, but is not limited to: the user's touch input to the setting control in the photo-taking preview interface through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned sixth input may be: the user's touch input to the setting control in the photo-taking preview interface. For example, the above-mentioned sixth input may be: the user's click input to the setting control in the photo-taking preview interface.

[0088] The audio duration setting interface may be an interface for setting the duration of the audio in the live photo. At least one audio recording duration option may be included in the audio duration setting interface.

[0089] The control input may be the user's input to any one of the at least one audio recording duration option in the audio duration setting interface. The above-mentioned control input is used to determine the first duration, the second duration, and the third duration. Exemplarily, the above-mentioned control input includes, but is not limited to: the user's touch input to any one of the at least one audio recording duration option in the audio duration setting interface through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned control input may be: the user's touch input to any one of the at least one audio recording duration option in the audio duration setting interface. For example, the above-mentioned control input may be: the user's click input to any one of the at least one audio recording duration option in the audio duration setting interface.

[0090] In some embodiments of the present application, in response to a sixth input from the user, an audio duration setting interface may be displayed, and then in response to a control input from the user for any one of at least one audio recording duration option in the audio duration setting interface, the first duration, the second duration, and the third duration may be determined according to the duration corresponding to the audio recording duration option selected by the control input.

[0091] Continuing to refer to Figure 2 , after the user clicks on the "Settings" control 213 in the photo preview interface, an audio duration setting interface may be displayed as Figure 6 shown. In this audio duration setting interface, there is at least one audio recording duration option. As Figure 6 shown, 3 audio recording duration options are displayed, namely the "2s audio recording duration" option 51, the "5s audio recording duration" option 52, and the "8s audio recording duration" option 53.

[0092] It should be noted that for each audio recording duration option, the duration of the video and the duration of the audio in the corresponding live photo may be the same or different, and can be specifically set according to the user's needs. For example, in the above Figure 6 , the "2s audio recording duration" option 51 may correspond to both the video and the audio in the live photo having a duration of 2s. The "5s audio recording duration" option 52 may correspond to both the video and the audio in the live photo having a duration of 5s. The "8s audio recording duration" option 53 may correspond to the video in the live photo having a duration of 2s and the audio having a duration of 8s.

[0093] Continuing to refer to Figure 6 , taking the "8s audio recording duration" option 53 corresponding to the video in the live photo having a duration of 2s and the audio having a duration of 8s as an example, if the user selects the "8s audio recording duration" option 53, the first duration may be determined to be 1 second, the second duration may be determined to be 1 second, and the third duration may be determined to be 8 seconds.

[0094] In the embodiments of the present application, according to the user's needs, by responding to a seventeenth input from the user for any one of at least one audio recording duration option in the audio duration setting interface, the first duration, the second duration, and the third duration may be determined. In this way, according to the user's needs, videos and audios with the durations required by the user can be selected, further improving the flexibility of generating live photos.

[0095] It should be noted that at least one audio recording duration option in the audio duration setting interface is in a mutually exclusive relationship, that is, the user can only select one audio recording duration option in the audio duration setting interface. After the user selects an audio recording duration option, before performing the first input, the selected audio recording duration option can be changed according to the user's needs.

[0096] Before step 110, the method described above may further include:

[0097] Receiving a duration setting input for the audio duration setting interface;

[0098] In response to the duration setting input, updating the audio recording duration of the first live photo.

[0099] Wherein, the duration setting input may be a selection input for at least one audio recording duration option in the audio duration setting interface.

[0100] In some embodiments of the present application, after the user has selected a certain audio recording duration option in the audio duration setting interface, if the user re - selects another audio recording duration option, that is, the electronic device receives the user's duration setting input for the audio duration setting interface, the audio recording duration of the first live photo can be updated.

[0101] Continuing to refer to Figure 6 , the user can select the "2s audio recording duration" option 51 by clicking the switch control 511. When the user selects the "2s audio recording duration" option 51, the electronic device stores the video from one second before to one second after the moment when the user presses the photo - taking control, and the audio from one second before to one second after the moment when the user presses the photo - taking control. The user can click the switch control 521 to select the "5s audio recording duration" option 52. When the user selects the "5s audio recording duration" option 52, the electronic device stores the video from one second before to four seconds after the moment when the user presses the photo - taking control, and the audio from one second before to four seconds after the moment when the user presses the photo - taking control. The user can select the "8s audio recording duration" option 53 by clicking the switch control 531. When the user selects the "8s audio recording duration" option 53, the electronic device stores the video from one second before to one second after the moment when the user presses the photo - taking control, and the audio from one second before to seven seconds after the moment when the user presses the photo - taking control. If the user initially selects the "8s audio recording duration" option 53, at this time the audio recording duration of the first live photo is 8 seconds. At this time, if the user wants to change the audio recording duration of the first live photo from 8 seconds to 5 seconds, the user can click the switch control 521 behind the "5s audio recording duration" option 52, and the audio recording duration of the first live photo can be changed from 8 seconds to 5 seconds.

[0102] In the embodiments of the present application, according to the user's needs, by responding to the user's duration setting input for the audio duration setting interface, the audio recording duration of the first live photo can be updated, thus improving the flexibility of the audio recording duration of the first live photo.

[0103] Step 130: When a playback control input for the first cover image is received, play the first video and the first audio.

[0104] Among them, the playback control input may be a control input of the user for the first cover image, and the above second input is used to play the first video and the first audio. Exemplarily, the above playback control input includes but is not limited to: a touch input of the user for the first cover image through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, and may also be a long press input or a short press input. For example, the above playback control input may be: a touch input of the user for the first cover image. For example, the above playback control input may be: a click input of the user for the first cover image.

[0105] In some embodiments of the present application, after the first cover image is generated, the first cover image may be displayed on the main interface of the album application of the electronic device. The captured live photo can be viewed through the album application program. Then, when the electronic device responds to the playback control input of the user for the first cover image, the first video and the first audio associated with the first cover image can be directly played. For example, continuing to refer to the above example, on the main interface of the album application, a thumbnail of the live photo of Xiaohong singing song A taken by the user is displayed. As Figure 7 shown, after the thumbnail 62 displayed in the main interface 61 of the album application program, the electronic device can respond to the operation of the user on the thumbnail of the live photo and can display as Figure 8 shown in the album interface 71. The live photo can be displayed on the album interface 71. The user can trigger the electronic device to display the video in the live photo and play the audio in the live photo by performing a specific operation. As Figure 8 shown, when the electronic device responds to a long press operation of the user on the live photo on the album interface 71, it can trigger the playback of the video in the live photo and the playback of the audio in the live photo.

[0106] It should be noted that other forms of controls for triggering the playback of the video and audio in the live photo may also be included in the album interface 71.

[0107] It should be noted that after the user clicks on the first cover image 62 displayed in the main interface 61 of the album application program, it is displayed as Figure 8When presenting the first dynamic photo as shown, it can be statically displayed, such as showing a thumbnail of the first dynamic photo, or it can be dynamically displayed, such as playing a video of the first dynamic photo. However, when playing the video of the first dynamic photo, the audio of the first dynamic photo will not be played. Only when the user clicks Figure 8 any position of the first dynamic photo shown in

[0108] In some embodiments of the present application, in order to enhance the visual experience of playing the first video and the first audio, the playing of the first video and the first audio may specifically include:

[0109] Synchronously playing the first video and the first audio.

[0110] In some embodiments of the present application, when playing the first video and the first audio, the first video and the first audio need to be synchronously played. Specifically, the first video and the first audio start playing at the same time point, and the playing mode of the first video is a loop playing mode, and the total playing duration of the first video is the same as the playing duration of the first audio. That is to say, the playing start time stamp of the first audio is the same as the playing start time stamp of the first video when the first video is played for the first time. In this way, the problem of out-of-sync audio and video during the playing of the first video and the first audio can be avoided.

[0111] Continuing to refer to Figure 7 , after generating the first cover image, the first cover image 62 can be displayed on the main interface 61 of the album application. Then, when the user clicks on the first cover image 62, the album interface 71 can be displayed as shown in Figure 8 . On this album interface 71, the first dynamic photo can be displayed. The first video in the first dynamic photo is a video of Xiaohong singing song A a cappella in the KTV, and the first audio in the first dynamic photo is the audio of Xiaohong singing song A a cappella in the KTV. As shown in Figure 8 , what can be displayed on the album interface is the cover image of the first dynamic photo, which is an image of Xiaohong singing song A a cappella in the KTV. Then, when the user clicks on any position of the first dynamic photo, the 8-second audio of Xiaohong singing song A and the 2-second video of Xiaohong singing song A can be played simultaneously. And when playing the 8-second audio of Xiaohong singing song A and the 2-second video of Xiaohong singing song A, the 2-second video of Xiaohong singing song A is looped 4 times, that is, the 2-second video of Xiaohong singing song A is played a total of 4 times for a total playing duration of 8 seconds.

[0112] In an embodiment of the present application, the first video and the first audio are played synchronously. In this way, during the playing process of the first video and the first audio, the temporal coordination and consistency can be maintained, thereby avoiding the problem of audio-video asynchrony during the playing process of the first video and the first audio, and improving the visual experience of the played first video and first audio.

[0113] In some embodiments of the present application, in order to improve the playing flexibility of the first dynamic photo, after step 120, the method described above may further include:

[0114] Display the first dynamic photo on the album interface;

[0115] Receive a control input for the audio playback control;

[0116] In response to the control input, play the first audio.

[0117] Wherein, the audio playback control may be a control for indicating the playback of the audio of the first dynamic photo, such as Figure 8 in control 72.

[0118] The control input may be an input by the user to the audio playback control, and the above control input is used to play the first audio. Exemplarily, the above control input includes but is not limited to: a touch input by the user on the audio playback control using a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, and may also be a long press input or a short press input. For example, the above control input may be: a touch input by the user on the audio playback control. For example, the above control input may be: a click input by the user on the audio playback control.

[0119] In some embodiments of the present application, the first dynamic photo displayed on the album interface may include an audio playback control. In response to a control input by the user to the audio playback control, only the audio of the first dynamic photo is played.

[0120] Continue to refer to Figure 8 , if the user only wants to listen to the 8-second audio of the song A sung by Xiaohong at this time, the user can click control 72, and only the 8-second audio of the song A sung by Xiaohong is played.

[0121] It should be noted that in Figure 8When the first dynamic photo is displayed dynamically, if the user only wants to listen to the 8-second audio of Xiaohong singing song A at this time, the user clicks After the control 72, the 2-second video when Xiaohong sings song A can be stopped, that is, in Figure 8 The first dynamic photo is displayed statically, and only the 8-second audio of Xiaohong singing song A is played.

[0122] It should be noted that when the first dynamic photo is displayed on the album interface, an identifier for characterizing that the first dynamic photo is a dynamic photo rather than a static photo can be displayed in the first dynamic photo. For example, as Figure 8 shown, an identifier 73 is displayed in the first dynamic photo, and the identifier 73 is used to indicate that the first dynamic photo is a dynamic photo rather than a static photo.

[0123] In the embodiment of the present application, when the user only wants to play the first audio, by responding to the user's control input on the audio playback control in the first dynamic photo, only the first audio segment can be played. In this way, according to the user's needs, only the audio of the first dynamic photo can be played, improving the playback flexibility of the first dynamic photo.

[0124] In some embodiments of the application, in order to meet the differentiated requirements of the user for the playback volume and playback effect of the first audio, after step 120, the above-mentioned method may further include:

[0125] Receiving a second input;

[0126] Responding to the second input, displaying an audio editing interface of the first dynamic photo;

[0127] Receiving a control input for the volume adjustment control;

[0128] Responding to the control input, updating the playback volume of the first audio.

[0129] Among them, the second input may be the user's input to the audio editing control. The above-mentioned second input is used to display the audio editing interface of the first live photo, and the second input may be a second operation. Exemplarily, the above-mentioned second input includes but is not limited to: the user's touch input to the audio editing control through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and the embodiments of the present invention do not make limitations. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned second input may be: the user's touch input to the audio editing control. For example, the above-mentioned second input may be: the user's click input to the audio editing control.

[0130] The audio editing interface may be an interface for editing the first audio. The audio editing interface may include a volume adjustment control for the first audio. The volume adjustment control may be a control for adjusting the volume of the first audio.

[0131] The control input may be the user's input to the volume adjustment control. The above-mentioned control input is used to update the playback volume of the first audio. Exemplarily, the above-mentioned control input includes but is not limited to: the user's touch input to the volume adjustment control through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and the embodiments of the present invention do not make limitations. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned control input may be: the user's touch input to the volume adjustment control. For example, the above-mentioned control input may be: the user's click input to the volume adjustment control.

[0132] In some embodiments of the present application, the first live photo displayed on the album interface may further include an editing control. By responding to the user's selection input to the editing control, an editing interface can be displayed. The audio editing control is included in the editing interface. By responding to the second input of the user to the audio editing control in the editing interface, the audio editing interface of the first live photo can be displayed. The volume adjustment control may be included in the audio editing interface. By responding to the control input of the user to the volume adjustment control, the playback volume of the first audio can be updated.

[0133] Continue to refer to Figure 8 , if the user wants to adjust the playback volume of the 8 - second audio of Song A sung by Xiaohong, the user can click on Figure 8 the "Edit" control 74 in, and an edit interface as shown in Figure 9 can be displayed. In this edit interface, there may be an "Audio Edit" control 81. When the user clicks on this "Audio Edit" control 81, an audio edit interface of the first dynamic photo as shown in Figure 10 can be displayed. In this audio edit interface, there may be a volume adjustment control, and the volume adjustment control includes Figure 10 the "Volume Up" control 911 and the "Volume Down" control 912 in. When the user clicks on the "Volume Up" control 911, the playback volume of the 8 - second audio of Song A sung by Xiaohong can be increased. If the user clicks on the "Volume Down" control 912, the playback volume of the 8 - second audio of Song A sung by Xiaohong can be decreased.

[0134] In an embodiment of the present application, by responding to the second input of the user, an audio edit interface of the first dynamic photo can be displayed. In this audio edit interface, there may be a volume adjustment control. In this way, according to the user's needs, in response to the user's control input to the volume adjustment control, the playback volume of the first audio can be updated. In this way, according to the user's needs, the playback volume of the first audio can be flexibly adjusted, meeting the differentiated needs of the user for the playback effect of the playback volume of the first audio.

[0135] In some embodiments of the present application, when the first audio includes audio of multiple timbres, in response to the fifth input of the user to the volume adjustment control, it may be to adjust the playback volume of the audio of each timbre in the first audio simultaneously.

[0136] Continuing to refer to the above example, the 8 - second audio of Song A sung by Xiaohong includes audio of two timbres. One timbre of the audio is the audio of Xiaohong singing Song A by herself, and the other timbre of the audio is the accompaniment audio of Song A. After the user clicks on the "Volume Up" control 911, it may be to increase the playback volume of the audio of Xiaohong singing Song A by herself and the playback volume of the accompaniment audio of Song A simultaneously.

[0137] In some embodiments of the present application, when the first audio includes audio of multiple timbres, it is also possible to adjust the playback volume of one or several of the audio of multiple timbres. Specifically, at least two audio tracks can be displayed in the audio edit interface, with one audio track corresponding to the audio of one timbre in the first audio. After the audio edit interface of the first dynamic photo is displayed and before the control input to the volume adjustment control is received, the method involved above may further include:

[0138] Receive user input for selecting the first audio track among at least two audio tracks;

[0139] Updating the playback volume of the first audio may specifically include:

[0140] Updating the playback volume of the audio corresponding to the timbre of the first audio track.

[0141] Among them, the first audio track can be any one of at least two audio tracks. For example, the first audio track can be Figure 10 audio track 92 in.

[0142] In some embodiments of the present application, in the case where the first audio includes audio of multiple timbres, if the user wants to adjust the playback volume of the audio of a certain timbre in the first audio, the user input for selecting the audio track corresponding to the audio of that timbre can be received, and then in response to the fifth input of the user to the volume adjustment control, the playback volume of the audio of that timbre can be updated.

[0143] Continuing to refer to the above example, as Figure 10 shown, there are two audio tracks in the audio editing interface, namely audio track 92 and audio track 93. Among them, audio track 92 corresponds to the audio of Xiaohong singing song A by herself, and audio track 93 corresponds to the accompaniment audio of song A. If the user wants to turn down the playback volume of the accompaniment audio of song A, the user can first click on audio track 93, and then click on the "Volume Key -" control 912, and then the playback volume of the accompaniment audio of song A can be turned down, while the playback volume of the audio of Xiaohong singing song A by herself remains unchanged.

[0144] In the embodiments of the present application, in the case where the first audio includes audio of multiple timbres, the playback volume of any timbre audio in the first audio can be adjusted according to the user's needs, thus improving the flexibility of adjusting the playback volume of the first audio.

[0145] In some embodiments of the present application, for any audio track, the playback volume value of the audio corresponding to the timbre of that audio track can be displayed at a preset position of the audio track. As Figure 10 shown, the playback volume value of the audio corresponding to the timbre of audio track 92 is 15 decibels, and the playback volume value of the audio corresponding to the timbre of audio track 93 is 7 decibels. After updating the playback volume of the audio corresponding to the timbre of the first audio track, the methods involved above may further include:

[0146] Displaying the updated playback volume value of the audio corresponding to the timbre of the first audio track at the preset position of the first audio track.

[0147] Among them, the preset position can be a certain position of the first audio track set in advance. For example, it can be the rightmost position of the display area corresponding to the first audio track. For example, as Figure 10As shown, the playback volume value "15" decibels of the audio of the timbre corresponding to track 92 is displayed at the rightmost position of the display area corresponding to track 92. The playback volume value "7" decibels of the audio of the timbre corresponding to track 93 is displayed at the rightmost position of the display area corresponding to track 93.

[0148] In some embodiments of the present application, after the playback volume of the audio of the timbre corresponding to the first track is updated, the updated playback volume value may be displayed at a preset position of the first track.

[0149] Continuing to refer to the above example, after the user clicks the "Volume Key -" control 912 and reduces the playback volume of the accompaniment audio of song A, for example, reduces the playback volume of the accompaniment audio of song A from 7 decibels to 5 decibels, then it can be as Figure 11 shown, the updated playback volume value "5" decibels of the audio of the timbre corresponding to track 93 is displayed at the rightmost position of the display area corresponding to track 93.

[0150] It should be noted that for any track, the playback volume value of the audio of the timbre corresponding to the track may be the average value of the playback volume of the audio of the timbre corresponding to the track.

[0151] In the embodiments of the present application, by displaying the updated playback volume value of the audio of the timbre corresponding to the first track at a preset position of the first track, the user can intuitively view the playback volume values of the audio of the timbre corresponding to each track, so that the user can accurately adjust the playback volume of the audio of the timbre corresponding to each track according to the displayed playback volume value, improving the adjustment accuracy of the playback volume of the audio corresponding to each track.

[0152] In some embodiments of the present application, in order to improve the playback effect of dynamic photos, the audio editing interface may further include a music import control, such as Figure 10 the "Import Music" control 95 in. After the audio editing interface for displaying the first dynamic photo is displayed, the methods involved above may further include:

[0153] Receiving a control input to the music import control;

[0154] In response to the control input, displaying at least one audio identifier;

[0155] Receiving a selection input from the user for at least one audio identifier;

[0156] In response to the selection input, displaying an audio segment truncation control for the third audio;

[0157] Receiving a setting input to the audio segment truncation control;

[0158] In response to a setting input, an audio segment of a fourth duration set by the setting input is intercepted from the third audio, and the audio segment of the fourth duration is stored in an associated manner with the first cover image;

[0159] When a playback control input for the first cover image is received, the first video, the first audio, and the audio segment of the fourth duration are played synchronously.

[0160] Among them, the control input can be an input of the user to the music import control, and the above control input is used to display at least one audio identifier. Exemplarily, the above control input includes, but is not limited to: a touch input of the user to the music import control by a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements, and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input. For example, the above control input can be: a touch input of the user to the music import control. For example, the above control input can be: a click input of the user to the music import control.

[0161] The audio identifier can be an identifier used to indicate an audio, and one audio identifier is used to indicate one audio. For example, Figure 12 , Figure 12 is the file management application interface of the electronic device, and a plurality of audio identifiers are displayed in the file management application interface. Each audio identifier corresponds to an audio. For example, audio identifier 111 corresponds to audio C, audio identifier 112 corresponds to audio D, and audio identifier 113 corresponds to audio E.

[0162] The third audio can be the audio indicated by the audio identifier selected when a selection input for at least one audio identifier is made. For example, when the audio identifier selected by the user during the selection input is Figure 12 the audio identifier 111 in, the third audio is audio C. The above audio segment intercept control can be a control used to intercept the audio segment of the third audio, such as Figure 12 the control 114 in.

[0163] The setting input can be the input of the user to the audio segment intercepting control. The above setting input is used to intercept an audio segment with a fourth duration set by the setting input from the third audio. Exemplarily, the above setting input includes but is not limited to: the touch input of the user to the audio segment intercepting control through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input. For example, the above setting input can be: the touch input of the user to the audio segment intercepting control. For example, the above setting input can be: the swipe input of the user to the audio segment intercepting control.

[0164] The fourth duration is the same as the audio duration of the audio included in the first live photo. For example, if the duration of the audio included in the first live photo is 8 seconds, then the fourth duration is 8 seconds.

[0165] The playback control input can be the input of the user to the first cover image. The above playback control input is used to synchronously play the first video, the first audio, and the audio segment with the fourth duration. Exemplarily, the above playback control input includes but is not limited to: the touch input of the user to the first cover image through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input. For example, the above playback control input can be: the touch input of the user to the first cover image. For example, the above playback control input can be: the click input of the user to the first cover image.

[0166] In some embodiments of the present application, if a user wants to add additional music to a first dynamic photo, at least one audio identifier may be displayed in response to a control input of the user to a music import control. Then, in response to a selection input of the user to at least one audio identifier, an audio segment truncation control for a third audio indicated by the selected audio identifier of the selection input may be displayed. In the case of receiving a setting input to the audio segment truncation control, an audio segment with a fourth duration set by the setting input may be truncated from the third audio, and then the audio segment with the fourth duration may be associated and stored with the first cover image. Thus, in the case of receiving a playback control input to the first cover image, the first video, the first audio, and the audio segment with the fourth duration may be played synchronously. Specifically, the playback start timestamp of the audio segment with the fourth duration may be the same as the playback start timestamp of the first audio, that is, the audio segment with the fourth duration and the first audio start playing simultaneously.

[0167] Continuing to refer to the above example, if the user wants to add other music to the dynamic photo of the song A sung by Xiaohong obtained, the user can click Figure 10 the "Import Music" control 95 in Figure 12 to display the file management application interface of the electronic device shown in Figure 12 . In this file management application interface, there is an audio identifier 111 corresponding to audio C, an audio identifier 112 corresponding to audio D, and an audio identifier 113 corresponding to audio E. If the user wants to add audio C to the dynamic photo of Xiaohong singing song A, the user can click the audio identifier 111, and then, as shown in Figure 12 , an audio segment truncation control 114 for audio C may be displayed. On the audio segment truncation control 114 for audio C, there is a start truncation slider 1141 and an end truncation slider 1142. The start truncation slider 1141 indicates the start truncation position when truncating an audio segment with a fourth duration from audio C, and the end truncation slider 1142 indicates the end truncation position when truncating an audio segment with a fourth duration from audio C. The user can hold down the start truncation slider 1141 and position it at the start truncation position where the user wants to truncate an audio segment with a fourth duration from audio C, and then hold down the end truncation slider 1142 and position it at the end truncation position where the user wants to truncate an audio segment with a fourth duration from audio C. In this way, the audio segment between the start truncation slider 1141 and the end truncation slider 1142 in audio C can be truncated. Specifically, the duration of the audio segment truncated by the start truncation slider 1141 and the end truncation slider 1142 is 8 seconds. Then, click the "Confirm Selection" control 115. In this way, the 8-second audio segment can be associated and stored with the first cover image. Furthermore, when the user clicks Figure 7When the first cover image 62 displayed on the main interface 61 of the photo album application is shown, an 8-second audio of Xiaohong singing song A, a 2-second video of Xiaohong singing song A, and an 8-second audio clip intercepted from audio C can be played simultaneously, and the 2-second video of Xiaohong singing song A is looped 4 times.

[0168] In an embodiment of the present application, the user can intercept the audio clip required by the user from other audio according to the need, and the duration of the audio clip is the same as the audio duration of the audio in the first dynamic photo, and add the audio clip to the first dynamic photo to add other audio except the audio of the first dynamic photo to the first dynamic photo. In this way, according to the user's needs, audio clips with different emotional expressions can be added to the first dynamic photo, and then the playback effect with different emotional expressions can be added to the first dynamic photo. For example, adding a lively audio clip to the first dynamic photo can make the playback effect of the first dynamic photo more relaxed and pleasant, and adding a heavy audio clip to the first dynamic photo can make the playback effect of the first dynamic photo more profound and urgent, improving the playback effect of the dynamic photo.

[0169] In some embodiments of the present application, all the dynamic photos stored in the photo album program may include a second dynamic photo, and the second dynamic photo may include an associated second cover image, a second video, and a second audio. Here, the second cover image may be a thumbnail of the second dynamic photo, the second video is the video of the second dynamic photo, and the second audio is the audio of the second dynamic photo.

[0170] After step 120, the method involved above may further include:

[0171] Receiving a third input from the user for the first cover image and the second cover image;

[0172] In response to the third input, displaying a third dynamic photo;

[0173] Receiving a playback control input from the user for the third cover image;

[0174] In response to the playback control input, playing the first video, the first audio, the second video, and the second audio in the input order of the third input.

[0175] Among them, the third input may be the user's input for the first cover image and the second cover image. The above-mentioned third input is used to control the electronic device to synthesize at least two live photos, and the third input may be a third operation. Exemplarily, the above-mentioned third input includes but is not limited to: the touch input of the first cover image and the second cover image by the user through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned third input may be: the touch input of the first cover image and the second cover image by the user. For example, the above-mentioned third input may be: the click input of the first cover image and the second cover image by the user.

[0176] It should be noted that the synthesis of live photos may be to synthesize at least two live photos. Specifically, it may be to determine the cover image of the synthesized live photo from the cover images or videos included in at least two live photos respectively, and then associate and store the cover image, as well as the videos and audios included in at least two live photos respectively. For example, for the live photo of Xiaohong singing song A and the live photo of Xiaoming singing song B in the above example, if the live photo of Xiaohong singing song A and the live photo of Xiaoming singing song B are synthesized into a new live photo, then first determine a new cover image from the live photo of Xiaohong singing song A and the live photo of Xiaoming singing song B, and then associate and store the cover image, as well as the 2s video and 8s audio when Xiaohong sings song A, and the 2s video and 8s audio when Xiaoming sings song B.

[0177] The third live photo may include an associated and stored third cover image, a first video, a first audio, a second video, and a second audio. Here, the third cover image may be the thumbnail of the third live photo.

[0178] The playback control input may be the user's input to the third cover image, and the above-mentioned ninth input is used to play the first video, the first audio, the second video, and the second audio. Exemplarily, the above-mentioned playback control input includes, but is not limited to: the user's touch input to the third cover image through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned playback control input may be: the user's touch input to the third cover image. For example, the above-mentioned playback control input may be: the user's click input to the third cover image.

[0179] In some embodiments of the present application, when the second cover image displayed on the main interface of the album application is a thumbnail of the second live photo, the live photos corresponding to the multiple cover images may be merged to obtain a new live photo. Specifically, in response to the user's third input to the first cover image and the second cover image, the third live photo may be displayed. The third live photo may include the associated stored third cover image, the first video, the first audio, the second video, and the second audio. Then, in response to the user's playback control input to the third cover image, the first video, the first audio, the second video, and the second audio may be played in the input order of the third input.

[0180] Specifically, when playing the first video, the first audio, the second video, and the second audio in the input order of the third input, the first video and the first audio may be played simultaneously, and the playback mode of the first video is a loop playback mode, that is, the playback start timestamp of the first audio is the same as the playback start timestamp of the first video's first playback. After playing the first video and the first audio, the second video and the second audio may be played simultaneously, that is, the playback start timestamp of the second audio is later than the playback end timestamp of the first audio. When playing the second video and the second audio, the playback mode of the second video is a loop playback mode, that is, the playback start timestamp of the second audio is the same as the playback start timestamp of the second video's first playback.

[0181] Continuing to refer to the above example, the first live photo is a live photo of Xiaohong singing song A, and this first live photo includes Figure 7The first cover image 62 shown, 8-second audio of Xiaohong singing song A, 2-second video of Xiaohong singing song A, and the second dynamic photo is a dynamic photo of Xiaoming singing song B a cappella. Xiaoming is inviting Xiaohong to sing together while singing song B a cappella. The second dynamic photo includes as Figure 7 shown the second cover image 63, 8-second audio of Xiaoming singing song B, and 2-second video of Xiaoming singing song B. For example, the second video in the second dynamic photo is a 2-second video of Xiaoming singing song B a cappella in a KTV, and the second audio in the second dynamic photo is an 8-second audio of Xiaoming singing song B a cappella in a KTV. As Figure 7 shown, the second cover image 63 of the second dynamic photo is also included in the main interface 61 of the photo album application. The second cover image 63 is a thumbnail of the second dynamic photo. For example, the second cover image 63 can be an image of Xiaoming inviting Xiaohong to sing together while singing song A a cappella in a KTV. If the user wants to play the 8-second audio of Xiaohong singing song A, the 2-second video of Xiaohong singing song A, the 8-second audio of Xiaoming singing song B, and the 2-second video of Xiaoming singing song B together, the user can click on the first cover image 62 and the second cover image 63, and then click the "Merge" control 65. In this way, the 8-second audio of Xiaohong singing song A, the 2-second video of Xiaohong singing song A, the 8-second audio of Xiaoming singing song B, and the 2-second video of Xiaoming singing song B can be associated and stored in the file management application, and as Figure 13 shown, the third cover image 121 is displayed in the main interface 61 of the photo album application. At this time, if the user clicks on the third cover image 121, the 8-second audio of Xiaohong singing song A and the 2-second video of Xiaohong singing song A can be played simultaneously first. When playing the 8-second audio of Xiaohong singing song A and the 2-second video of Xiaohong singing song A, the 2-second video of Xiaohong singing song A loops 4 times, that is, the 2-second video of Xiaohong singing song A is played for a total of 8 seconds. Then, the 8-second audio of Xiaoming singing song B and the 2-second video of Xiaoming singing song B are played simultaneously. When playing the 8-second audio of Xiaoming singing song B and the 2-second video of Xiaoming singing song B, the 2-second video of Xiaoming singing song B loops 4 times, that is, the 2-second video of Xiaoming singing song B is played for a total of 8 seconds.

[0182] It should be noted that in the above example, after the user clicks on the first cover image 62 and the second cover image 63 and associates and stores the 8-second audio of Xiaohong singing song A, the 2-second video of Xiaohong singing song A, the 8-second audio of Xiaoming singing song B, and the 2-second video of Xiaoming singing song B in the file management application, as Figure 13 shown, the first cover image 62 and the second cover image 63 in the main interface 61 of the photo album application will disappear, and only the third cover image 121 will be displayed.

[0183] It should be noted that after the user performs a third input on the first cover image 62 and the second cover image 63, corresponding selected marks can be added to the first cover image 62 and the second cover image 63 to indicate that the user has selected the first cover image 62 and the second cover image 63. As Figure 7 shown, after the user clicks on the first cover image 62 and the second cover image 63, the "○" marks 60 on the first cover image 62 and the second cover image 63 are both filled to indicate that the user has clicked on the first cover image 62 and the second cover image 63, so that the user can intuitively view the selected cover images and avoid misoperations.

[0184] In an embodiment of the present application, by responding to the user's third input on the first cover image and the second cover image, a third dynamic photo can be displayed, and then by responding to the user's playback control input on the third cover image, the first video, the first audio, the second video, and the second audio can be played. In this way, according to the user's needs, the videos and audios of the dynamic photos corresponding to multiple cover images can be played together without the user performing a playback control input on each cover image, improving the playback convenience of the videos and audios of the dynamic photos corresponding to multiple cover images. At the same time, the dynamic photos corresponding to multiple cover images can be combined into one dynamic photo according to the user's needs, further improving the flexibility of generating dynamic photos.

[0185] In some embodiments of the present application, in order to improve the diversity of generating the third cover image, before displaying the third dynamic photo, the method described above may further include:

[0186] Determine the third cover image according to the input order of the third input; or, in response to the selection input on the first cover image and the second cover image, use the cover image selected by the selection input as the third cover image; or, select any video frame from the first video and the second video as the third cover image;

[0187] Associate and store the third cover image, the first video, the first audio, the second video, and the second audio.

[0188] Among them, the selection input can be the user's input for the first cover image and the second cover image, and the above selection input is used to use the cover image selected by the selection input as the third cover image. Exemplarily, the above selection input includes but is not limited to: the user's touch input on the first cover image and the second cover image through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input. For example, the above selection input can be: the user's touch input on the first cover image and the second cover image. For example, the above selection input can be: the user's double click input on the first cover image and the second cover image.

[0189] In some embodiments of the present application, when determining the third cover image, the third cover image can be determined according to the input order of the third input performed on the first cover image and the second cover image. For example, it can be that the cover image that first performs the third input among the first cover image and the second cover image is used as the third cover image, or it can be that the cover image that later performs the third input among the first cover image and the second cover image is used as the third cover image. Specifically, it can be set according to the user's needs and is not limited in the embodiments of the present application.

[0190] Continue to refer to Figure 7 , taking the cover image that first performs the third input among the first cover image and the second cover image as the third cover image as an example. When the user clicks on the first cover image 62 and the second cover image 63, if the user first clicks on the second cover image 63 and then clicks on the first cover image 62, then as Figure 13 shown, the third cover image 121 is the second cover image 63.

[0191] In some embodiments of the present application, when determining the third cover image, it can also be that according to the user's needs, the cover image specified by the user among the first cover image and the second cover image is used as the third cover image, that is, by responding to the user's selection input for the first cover image and the second cover image, the cover image selected by the selection input is used as the third cover image.

[0192] Continue to refer to the above Figure 7, taking double - click input as an example of the selection input, if the user wants to use the second cover image 63 as the third cover image, after the user clicks on the first cover image 62 and the second cover image 63, the user can double - click on the second cover image 63 again, that is, select the second cover image 63 as the third cover image. Thus, after the user clicks on the "merge" control 65, it can be as Figure 13 shown, using the second cover image 63 as the third cover image 121.

[0193] In some embodiments of the present application, when determining the third cover image, any video frame can also be selected from the first video and the second video as the third cover image. For example, the user can select any video frame from a 2 - second video of Xiaohong singing song A as the third cover image.

[0194] In the embodiments of the present application, by determining the third cover image according to the order of performing the third input on the first cover image and the second cover image, or by responding to the user's selection input on the first cover image or the second cover image, or by selecting any video frame from the first video and the second video as the third cover image, in this way, the third cover image can be determined in different ways, improving the diversity of the generation of the third cover image.

[0195] In some embodiments of the present application, in order to meet the user's differentiated requirements for the playback effects of the playback volume of the first audio and the second audio, after displaying the third dynamic photo, the above - mentioned method may further include:

[0196] Receiving a fourth input;

[0197] In response to the fourth input, displaying an audio editing interface of the third dynamic photo;

[0198] When receiving a control input for the first volume adjustment control, updating the playback volume of the first audio;

[0199] When receiving a control input for the second volume adjustment control, updating the playback volume of the second audio.

[0200] Among them, the fourth input may be an input by the user to the third cover image, and the above-mentioned fourth input is used to display an audio editing interface for the third live photo. The fourth input may be a fourth operation. Exemplarily, the above-mentioned fourth input includes but is not limited to: a touch input by the user on the third cover image through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application may be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application may be a single click input, a double click input, or a click input of any number of times, etc., and may also be a long press input or a short press input. For example, the above-mentioned fourth input may be: a touch input by the user on the third cover image. For example, the above-mentioned fourth input may be: a click input by the user on the third cover image.

[0201] The audio editing interface may be an interface for editing the audio of the third live photo. The audio editing interface may include a first editing area and a second editing area. One editing area corresponds to an area for editing one of the first audio and the third audio. The first editing area may include a first volume adjustment control for the first audio, and the first volume adjustment control may be used to adjust the playback volume of the first audio. The second editing area may include a second volume adjustment control for the second audio, and the second volume adjustment control may be used to adjust the playback volume of the second audio.

[0202] In some embodiments of the present application, after the user clicks on the third cover image displayed on the main interface of the photo album application, an album interface may be displayed. The album interface may display the third live photo, and the third live photo may include editing controls. By responding to the user's selection input to the editing controls, an editing interface may be displayed. The editing interface includes audio editing controls. In response to the fourth input by the user to the audio editing controls in the editing interface, an audio editing interface for the third live photo may be displayed. The audio editing interface may include a first editing area and a second editing area. The first editing area may include a first volume adjustment control for adjusting the playback volume of the first audio, and the second editing area may include a second volume adjustment control for adjusting the playback volume of the second audio. If the user wants to adjust the playback volume of the first audio, a control input may be executed on the first volume adjustment control. In this way, the electronic device responds to the control input and may update the playback volume of the first audio. If the user wants to adjust the playback volume of the second audio, a control input may be executed on the second volume adjustment control. In this way, the electronic device responds to the control input and may update the playback volume of the second audio.

[0203] Continuing to refer to the above example, asFigure 13 As shown, when the user clicks on the third cover image 121, Figure 14 as shown, an album interface can be displayed. In this album interface, the third dynamic photo corresponding to the third cover image 121 can be displayed. An "Edit" control 122 is displayed in the third dynamic photo. When the user clicks on the "Edit" control 122, an edit interface as shown Figure 15 can be displayed. In this edit interface, an "Audio Edit" control 131 can be included. When the user clicks on the "Audio Edit" control 131, an audio edit interface of the third dynamic photo as shown Figure 16 can be displayed. In this audio edit interface, an editing area 142 for editing the playback volume of the first audio and an editing area 143 for editing the playback volume of the second audio can be included. That is, the editing area 142 is an area for editing the playback volume of the 8-second audio of Xiaohong singing song A, and the editing area 143 is an area for editing the playback volume of the 8-second audio of Xiaoming singing song B. A first volume adjustment control can be included in the editing area 142, and a second volume adjustment control 1402 can be included in the editing area 143. Among them, the first volume adjustment control includes a "Volume Key +" control 14011 and a "Volume Key -" control 14012, and the second volume adjustment control includes a "Volume Key +" control 14021 and a "Volume Key -" control 14022.

[0204] If the user wants to turn down the playback volume of the 8-second audio of Xiaohong singing song A, the user can click on the "Volume Key -" control 14012, and the playback volume of the 8-second audio of Xiaohong singing song A can be turned down.

[0205] It should be noted that when adjusting at least one of the playback volumes of the first audio and the second audio, in addition to adjusting by clicking on at least one of the first volume adjustment control and the second volume adjustment control, adjustment can also be made by clicking on the physical volume keys of the electronic device. Specifically, which method to select for adjusting at least one of the playback volumes of the first audio and the second audio can be freely chosen according to the user's needs and is not limited in the embodiments of the present application.

[0206] In an embodiment of the present application, by responding to a fourth input of a user, an audio editing interface for a third live photo can be displayed. The audio editing interface may include a first editing area and a second editing area. The first editing area may include a first volume adjustment control for a first audio, and the second editing area may include a second volume adjustment control for a second audio. Then, when a control input to the first volume adjustment control is received, the playback volume of the first audio can be updated. When a control input to the second volume adjustment control is received, the playback volume of the second audio can be updated. In this way, at least one of the playback volumes of the first audio and the second audio can be flexibly adjusted according to user needs, and the differentiated requirements of the user for the playback effects of the playback volumes of the first audio and the second audio can be met.

[0207] In some embodiments of the present application, when adjusting at least one of the playback volumes of the first audio and the second audio, it may also be to adjust the playback volume of the audio corresponding to a certain track in the first audio or a certain track in the second audio, that is, for each editing area, at least one track of the audio corresponding to the editing area can be displayed in the editing area, and one track corresponds to the audio of a certain timbre in the audio corresponding to the editing area. When a control input to the first volume adjustment control is received, updating the playback volume of the first audio may specifically include:

[0208] When a control input to the second track and the first volume adjustment control is received, update the playback volume of the audio corresponding to the timbre of the second track.

[0209] Among them, the second track may be any track among the at least one track displayed in the first editing area. For example, it may be Figure 16 track 1421 in

[0210] In some embodiments of the present application, when the first audio corresponding to the first editing area includes audio of multiple timbres, if the user wants to adjust the playback volume of the audio of a certain timbre in the first audio corresponding to the first editing area, after the electronic device receives the selection input of the user for the track corresponding to the audio of the timbre and the first volume adjustment control, the playback volume of the audio corresponding to the track can be updated.

[0211] Continuing to refer to the above example, such as Figure 16As shown in the figure, the editing area 142 includes two audio tracks, namely audio track 1421 and audio track 1422. Among them, audio track 1421 corresponds to the audio of Xiaohong singing song A by herself, and audio track 1422 corresponds to the accompaniment audio of song A. If the user wants to reduce the playback volume of the accompaniment audio of song A, the user can first click on audio track 1421 in the editing area 142, and then click on the "Volume Down" control 14012, and then the playback volume of the accompaniment audio of song A can be reduced, while the playback volume of the audio of Xiaohong singing song A by herself remains unchanged.

[0212] It should be noted that in Figure 16 , in the editing area 142, "Audio 1" is the audio of the timbre corresponding to audio track 1421, and "Audio 2" is the audio of the timbre corresponding to audio track 1422. That is to say, "Audio 1" in the editing area 142 is the audio of Xiaohong singing song A by herself, and "Audio 2" in the editing area 142 is the accompaniment audio of song A. Similarly, in the editing area 143, "Audio 1" is the audio of the timbre corresponding to audio track 1431, and "Audio 2" is the audio of the timbre corresponding to audio track 1432. That is to say, "Audio 1" in the editing area 143 is the audio of Xiaoming singing song B by herself, and "Audio 2" in the editing area 143 is the accompaniment audio of song B.

[0213] In the embodiments of the present application, in the case where the first audio corresponding to the first editing area includes audio of multiple timbres, the playback volume of the audio of any timbre in the first audio can be adjusted according to the user's needs, thus improving the flexibility of adjusting the playback volume of the first audio.

[0214] In some embodiments of the present application, for any audio track, the playback volume value of the audio of the timbre corresponding to the audio track can be displayed at a preset position of the audio track. For example, Figure 16 as shown, the playback volume value of the audio of the timbre corresponding to audio track 1421 is 15 decibels, the playback volume value of the audio of the timbre corresponding to audio track 1422 is 7 decibels, the playback volume value of the audio of the timbre corresponding to audio track 1431 is 20 decibels, and the playback volume value of the audio of the timbre corresponding to audio track 1432 is 14 decibels. Correspondingly, after adjusting the playback volume of the audio of the timbre corresponding to a certain audio track, the adjusted playback volume value of the audio of the timbre corresponding to the audio track can be displayed at the preset position of the audio track.

[0215] In some embodiments of the present application, in order to improve the playback effect of the third dynamic photo, the audio editing interface of the third dynamic photo may further include a music import control, such as Figure 16 the "Import Music" control 145 in

[0216] Receive the control input of the user for the music import control;

[0217] In response to the control input, display at least one audio identifier;

[0218] Receive the selection input of the user for at least one audio identifier;

[0219] In response to the selection input, display the audio segment truncation control for the fourth audio;

[0220] Receive the setting input for the audio segment truncation control;

[0221] In response to the setting input, truncate the audio segment of the fifth duration set by the setting input from the fourth audio, and associate and store the audio segment of the fifth duration with the third cover image;

[0222] When receiving the playback control input for the third cover image, play the first video, the first audio, the second video, the second audio, and the audio segment of the fifth duration.

[0223] Among them, the control input can be the input of the user for the music import control, and the above control input is used to display at least one audio identifier. Exemplarily, the above control input includes but is not limited to: the touch input of the user for the music import control through a touch device such as a finger or a stylus, or the voice command input by the user, or the specific gesture input by the user, or other feasible inputs, which can be determined according to the actual usage requirements, and the embodiments of the present invention do not make limitations. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, etc., and can also be a long press input or a short press input. For example, the above control input can be: the touch input of the user for the music import control. For example, the above control input can be: the click input of the user for the music import control.

[0224] The audio identifier can be an identifier used to indicate an audio, and one audio identifier can indicate one audio. For example Figure 12 , Figure 12 is the file management application interface of the electronic device. In this file management application interface, multiple audio identifiers are displayed, and each audio identifier corresponds to an audio. For example, audio identifier 111 corresponds to audio C, audio identifier 112 corresponds to audio D, and audio identifier 113 corresponds to audio E.

[0225] The fourth audio can be the audio indicated by the audio identifier selected when making the selection input for at least one audio identifier. For example, when the user makes the selection input, the selected audio identifier is Figure 12If the audio identifier 111 in it, then the fourth audio is Audio C. The above-mentioned audio segment intercepting control can be a control for intercepting audio segments of the fourth audio, such as Figure 12 the control 114 in it.

[0226] The setting input can be the user's input to the audio segment intercepting control. The above-mentioned setting input is used to intercept an audio segment of the fifth duration set by the setting input from the second audio. Exemplarily, the above-mentioned setting input includes but is not limited to: the user's touch input to the audio segment intercepting control through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs, which can be determined according to actual usage requirements and are not limited in the embodiments of the present invention. The specific gesture in the embodiments of the present application can be any one of a click gesture, a swipe gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double press gesture, and a double click gesture; the click input in the embodiments of the present application can be a single click input, a double click input, or a click input of any number of times, and can also be a long press input or a short press input. For example, the above-mentioned setting input can be: the user's touch input to the audio segment intercepting control. For example, the above-mentioned setting input can be: the user's swipe input to the audio segment intercepting control.

[0227] The fifth duration is equal to the sum of the audio duration of the first audio and the audio duration of the second audio. For example, if the audio included in the third live photo is 8 seconds of the song A sung by Xiaohong and 8 seconds of the song B sung by Xiaoming, then the fifth duration is 16 seconds.

[0228] In some embodiments of the present application, if the user wants to add additional music to the third live photo, at least one audio identifier can be displayed in response to the user's control input to the music import control, and then in response to the user's selection input of at least one audio identifier, an audio segment intercepting control of the fourth audio indicated by the selected audio identifier of the selection input can be displayed. In the case of receiving the setting input to the audio segment intercepting control, an audio segment of the fifth duration set by the setting input can be intercepted from the fourth audio, and the audio segment of the fifth duration can be associated and stored with the third cover image. Thus, in the case of receiving the playback control input to the third cover image, the first video, the first audio, the second video, the second audio, and the audio segment of the fifth duration can be played.

[0229] Specifically, when playing the first video, the first audio, the second video, the second audio, and the audio clip of the fifth duration, the first video, the first audio, and the audio clip of the fifth duration can start playing simultaneously, and the playing mode of the first video is the loop playing mode, that is, the playing start timestamp of the first audio is the same as the playing start timestamp of the first video's first play. After playing the first video and the first audio, the second video and the second audio can start playing simultaneously, that is, the playing start timestamp of the second audio is later than the playing end timestamp of the first audio. When playing the second video and the second audio, the playing mode of the second video is the loop playing mode, that is, the playing start timestamp of the second audio is the same as the playing start timestamp of the second video's first play.

[0230] It should be noted that during the process of playing the first video, the first audio, the second video, and the second audio, the audio clip of the fifth duration is also playing simultaneously all the time, and the playing duration of the audio clip of the fifth duration is the sum of the audio durations of the first audio and the second audio.

[0231] Continuing to refer to the above example, if the user wants to add other music to the third dynamic photo formed by combining the dynamic photos of Xiaohong singing song A and Xiaoming singing song B, the user can click on Figure 16 the "Import Music" control 145 in Figure 12 to display the file management application interface of the electronic device shown in Figure 12As shown, an audio segment capture control 114 of audio C is displayed. A start capture slider 1141 and an end capture slider 1142 are provided on the audio segment capture control 114 of audio C. The start capture slider 1141 indicates the start capture position when the fifth duration audio segment of audio C is captured, and the end capture slider 1142 indicates the end capture position when the fifth duration audio segment of audio C is captured. The user can press and hold the start capture slider 1141 to position it at the start position where the user wants to capture the fifth duration audio segment of audio C. The user then presses the end capture slider 1142 and positions it at the end capture position where the user wants to capture the fifth duration of the audio segment of Audio C. This allows the audio segment between the start capture slider 1141 and the end capture slider 1142 of Audio C to be captured. Specifically, the duration of the audio segment captured by the start capture slider 1141 and the end capture slider 1142 is 16 seconds. The user then clicks the "Confirm Selection" control 115, which associates the 16-second audio segment with the third cover image and stores it. Figure 13 When the third cover image 121 is displayed, the 8-second audio of Xiaohong singing song A, the 2-second video of Xiaohong singing song A and the 16-second audio clip intercepted from audio C can be played simultaneously. When the 8-second audio of Xiaohong singing song A and the 2-second video of Xiaohong singing song A are played, the 2-second video of Xiaohong singing song A is looped 4 times, that is, the 2-second video of Xiaohong singing song A is played for a total of 8 seconds. Then, the 8-second audio of Xiaoming singing song B and the 2-second video of Xiaoming singing song B are played simultaneously. When the 8-second audio of Xiaohong singing song A and the 2-second video of Xiaoming singing song B are played, the 2-second video of Xiaoming singing song B is looped 4 times, that is, the 2-second video of Xiaoming singing song B is played for a total of 8 seconds. During the process of playing the 8-second audio of Xiaohong singing song A and the 2-second video of Xiaohong singing song A, as well as the 8-second audio of Xiaoming singing song B and the 2-second video of Xiaoming singing song B, the 16-second audio clip intercepted from audio C is always playing.

[0232] In an embodiment of the present application, the user can, as needed, intercept an audio clip with the same length as the audio playback length of the third dynamic photo from other audios and add it to the third dynamic photo. In this way, according to user needs, the required audio clips with different emotional expressions can be added to the third dynamic photo, thereby increasing the playback effects of different emotional expressions for the third dynamic photo. For example, adding a cheerful audio clip to the third dynamic photo can make the playback effect of the third dynamic photo more relaxed and enjoyable, and adding a heavy audio clip to the third dynamic photo can make the playback effect of the third dynamic photo deeper and more urgent, thereby improving the playback effect of the third dynamic photo.

[0233] In some embodiments of the present application, in addition to adding other music to the entire third dynamic photo, other music can also be added to a certain segment that makes up the third dynamic photo. For example, the user can add other music during the process of Xiaoming singing song B in the third dynamic photo formed by combining the dynamic photo of Xiaohong singing song A and the dynamic photo of Xiaoming singing song B.

[0234] That is, the audio editing interface of the third dynamic photo may further include a music import control, such as Figure 16 the "Import Music" control 145 in

[0235] Receiving a fifth input to the first editing area and the music import control;

[0236] In response to the fifth input, displaying at least one audio identifier;

[0237] Receiving a selection input from the user for at least one audio identifier;

[0238] In response to the selection input, displaying an audio segment truncation control for the fifth audio;

[0239] Receiving a setting input to the audio segment truncation control;

[0240] In response to the setting input, truncating an audio segment of the sixth duration set by the setting input from the fifth audio, and associatively storing the audio segment of the sixth duration with the third cover image;

[0241] When a playback control input for the third cover image is received, playing the first video, the first audio, the second video, the second audio, and the audio segment of the sixth duration.

[0242] Among them, the fifth input can be the user's input to the first editing area and the music import control, and the above-mentioned fifth input is used to display at least one audio identifier, and the fifth input can be the fifth operation. Exemplarily, the above-mentioned fifth input includes but is not limited to: the user's touch input to the first editing area and the music import control through a touch device such as a finger or a stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible input, which can be determined according to actual use needs and is not limited in the embodiment of the present invention. The specific gesture in the embodiment of the present application can be any one of a single-click gesture, a sliding gesture, a drag gesture, a pressure recognition gesture, a long press gesture, an area change gesture, a double-press gesture, and a double-click gesture; the click input in the embodiment of the present application can be a single-click input, a double-click input, or any number of click inputs, etc., and can also be a long press input or a short press input. For example, the above-mentioned fifth input can be: the user's touch input to the first editing area and the music import control, for example, the above-mentioned fifth input can be: the user's click input to the first editing area and the music import control.

[0243] The fifth audio may be the audio indicated by the audio identifier selected by the selection input of at least one audio identifier, for example, the audio identifier selected by the user when selecting the input is Figure 12 The audio identifier 111 in the audio codec is the fifth audio codec. The audio segment interception control may be a control for intercepting the audio segment of the fifth audio codec, such as Figure 12 Control 114 in.

[0244] The sixth duration mentioned above is the same as the duration of the first audio. For example, if the first editing area is Figure 16 In the editing area 143, the audio corresponding to the editing area 143 is 8 seconds of audio of Xiao Ming singing song B, so the sixth duration is 8 seconds.

[0245] In some embodiments of the present application, when a user wants to add additional music during the playback of the first dynamic photo content when playing the third dynamic photo, at least one audio identifier can be displayed in response to the user's control input of the first editing area and the music import control, and then in response to the user's selection input of at least one audio identifier, an audio segment capture control of the fifth audio indicated by the audio identifier selected by the selection input can be displayed. When a setting input of the audio segment capture control is received, an audio segment of the sixth length set by the setting input can be captured from the fifth audio, and the audio segment of the sixth length can be associated with the third cover image and stored. In this way, when a playback control input of the third cover image is received, the first video, the first audio, the second video, the second audio and the audio segment of the sixth length can be played.

[0246] Specifically, when playing the first video, the first audio, the second video, the second audio, and the sixth-long audio clip, the first video, the first audio, and the sixth-long audio clip may start playing at the same time, and the playing duration of the sixth-long audio clip is the same as the playing duration of the first audio, and the playing mode of the first video is the loop playback mode, that is, the playing start timestamp of the first audio is the same as the playing start timestamp of the first video. After the first video, the first audio, and the sixth-long audio clip are played, the second video and the second audio may start playing simultaneously, that is, the playing start timestamp of the second audio is later than the playing end timestamp of the first audio. When playing the second video and the second audio, the playing mode of the second video is the loop playback mode, that is, the playing start timestamp of the second audio is the same as the playing start timestamp of the first video.

[0247] Continuing with the above example, if the user wants to merge the dynamic photo of Xiaohong singing song A and the dynamic photo of Xiaoming singing song B into a third dynamic photo, and add other music to the process of Xiaoming singing song B, the user can first click Figure 16 Then click the "Import Music" control 145 to display the audio area 143 in the Figure 12 The file management application interface of the electronic device shown in the figure has an audio identifier 111 corresponding to audio C, an audio identifier 112 corresponding to audio D, and an audio identifier 113 corresponding to audio E. If the user wants to add audio C to the dynamic photo of Xiaohong singing song A, the user can click on the audio identifier 111, and then the audio identifier 113 can be added. Figure 12 As shown, an audio segment interception control 114 of audio C is displayed. A start interception slider 1141 and an end interception slider 1142 are provided on the audio segment interception control 114 of audio C. The start interception slider 1141 indicates the start interception position when intercepting the sixth-length audio segment of audio C, and the end interception slider 1142 indicates the end interception position when intercepting the sixth-length audio segment of audio C. The user can press and hold the start interception slider 1141 to position it at the position where the user wants to intercept the sixth-length audio segment of audio C. The user then presses the end interception slider 1142 and positions it at the end interception position where the user wants to intercept the sixth duration of the audio segment of audio C. In this way, the audio segment between the start interception slider 1141 and the end interception slider 1142 in audio C can be intercepted. Specifically, the audio segment intercepted by the start interception slider 1141 and the end interception slider 1142 is an 8-second audio segment. In this way, the 8-second audio segment can be associated with the third cover image and stored, and then when the user clicks Figure 13When the third cover image 121 shown is played, the 8-second audio of Xiaohong singing song A, the 2-second video of Xiaohong singing song A, and the 8-second audio clip intercepted from audio C can be played simultaneously first. When playing the 8-second audio of Xiaohong singing song A and the 2-second video of Xiaohong singing song A, the 2-second video of Xiaohong singing song A is looped 4 times, that is, the 2-second video of Xiaohong singing song A is played for a total of 8 seconds. After the 8-second audio of Xiaohong singing song A, the 2-second video of Xiaohong singing song A, and the 8-second audio clip intercepted from audio C are played, then the 8-second audio of Xiaoming singing song B and the 2-second video of Xiaoming singing song B are played simultaneously. When playing the 8-second audio of Xiaoming singing song B and the 2-second video of Xiaoming singing song B, the 2-second video of Xiaoming singing song B is looped 4 times, that is, the 2-second video of Xiaoming singing song B is played for a total of 8 seconds.

[0248] In the embodiments of the present application, the user can intercept an audio clip with the same duration as the audio playback duration of the first audio or the second audio in the third dynamic photo from other audio according to the need, and add it to the third dynamic photo. In this way, according to the user's needs, audio clips with different emotional expressions can be added to the third dynamic photo, thereby adding playback effects with different emotional expressions to the third dynamic photo. For example, adding a lively audio clip to the third dynamic photo can make the playback effect of the third dynamic photo more relaxed and pleasant, and adding a heavy audio clip to the third dynamic photo can make the playback effect of the third dynamic photo more profound and have a sense of urgency, improving the playback effect of the third dynamic photo.

[0249] In some embodiments of the present application, in order to improve the convenience of adjusting the playback volume of the fourth audio, after intercepting and setting the audio clip with the fifth duration set from the fourth audio and associatively storing the audio clip with the fifth duration with the third cover image, the method involved above may further include:

[0250] Display a third editing area in the audio editing interface;

[0251] When receiving a selection input for the third volume adjustment control, update the playback volume of the audio clip with the fifth duration.

[0252] Among them, the third editing area can be an additional editing area outside the first editing area and the second editing area. The third editing area includes a third volume adjustment control for the audio clip with the fifth duration.

[0253] In some embodiments of the present application, after associating and storing the audio segment of the fifth duration with the third cover image, a third editing area may be further displayed in the audio editing interface. The third editing area may include a third volume adjustment control for the audio segment of the fifth duration. When the user wants to adjust the playback volume of the audio segment of the fifth duration, the playback volume of the audio segment of the fifth duration may be updated when a selection input for the third volume adjustment control is received.

[0254] Continuing to refer to the above example, as Figure 17 shown, the third editing area 151 may be displayed in the audio editing interface. The third editing area includes a third volume adjustment control, and the third volume adjustment control includes a "Volume -" control 152 and a "Volume +" control 153. If the user clicks the "Volume -" control 152 at this time, the playback volume of the 16 - second audio segment intercepted from Audio C can be reduced.

[0255] In the embodiments of the present application, by displaying the third editing area in the audio editing interface, and the third editing area includes a third volume adjustment control for the audio segment of the fifth duration. In this way, when a selection input for the third volume adjustment control is received, the playback volume of the third volume adjustment control of the audio segment of the fifth duration can be updated, without the user having to call out the audio segment of the fifth duration from the file management application and then adjust the playback volume of the audio segment of the fifth duration, which improves the convenience of adjusting the playback volume of the audio segment of the fifth duration.

[0256] For the dynamic photo generation method provided by the embodiments of the present application, the execution subject may be a dynamic photo generation device. In the embodiments of the present application, taking the dynamic photo generation device executing the dynamic photo generation method as an example, the dynamic photo generation device provided by the embodiments of the present application is described.

[0257] Figure 18 is a schematic structural diagram of a dynamic photo generation device shown according to an exemplary embodiment. As Figure 18 shown, the dynamic photo generation device 1600 may include:

[0258] A receiving module 1610, configured to receive a first input;

[0259] A display module 1620, configured to display a first dynamic photo in response to the first input; the first dynamic photo includes an associated and stored first cover image, a first video, and a first audio;

[0260] A playback module 1630, configured to play the first video and the first audio when a playback control input for the first cover image is received; wherein, the audio duration of the first audio is greater than the video duration of the first video.

[0261] In an embodiment of the present application, by responding to a first input from a user, a first live photo including an associated stored first cover image, a first video, and a first audio can be displayed. In this way, when the user wants to record the live scene and sound information of shooting scenarios such as concerts, KTVs, talk shows, etc., it can be achieved by taking a live photo, without the need to shoot a long video, reducing the occupation of memory resources in the electronic device. At the same time, in the live photo generated by the embodiment of the present application, the audio duration of the first audio is longer than the video duration of the first video. In this way, compared with the live photos in the prior art, more audio information can be carried, and it is also convenient for the later editing of the first audio.

[0262] In some embodiments of the present application, the playback module 1630 is specifically configured to:

[0263] Play the first video and the first audio synchronously;

[0264] Wherein, the playback mode of the first video is loop playback, and the total playback duration of the first video is the same as the playback duration of the first audio; the playback start timestamp of the first audio is the same as the playback start timestamp of the first video when the first video is played for the first time.

[0265] In some embodiments of the present application, the first input is a control input to a photographing control; the device further includes:

[0266] A caching module, configured to continuously cache the image data collected by the camera, control the microphone to perform audio recording, and continuously cache the digital audio data collected by the microphone before receiving the first input when the live photo function is enabled;

[0267] A generating module, configured to perform image processing on the image data cached at a reference moment to generate a first cover image before displaying the first live photo; perform video encoding on all the image data cached within a first time period to generate a first video; wherein, the first time period consists of a first duration before the reference moment and a second duration after the reference moment; perform audio encoding on all the digital audio data cached within a second time period to generate a first audio; wherein, the second time period consists of the first duration before the reference moment and a third duration after the reference moment, the third duration is greater than the second duration, and the reference moment is the input moment of the first input;

[0268] A storage module, configured to store the first cover image, the first audio, and the first video in an associated manner.

[0269] In some embodiments of the present application, the display module 1620 is further configured to display the first live photo on the album interface; an audio playback control is included on the live photo;

[0270] The receiving module 1610 is further configured to receive a control input to the audio playback control;

[0271] The playing module 1630 is further configured to play the first audio in response to the control input.

[0272] In some embodiments of the present application, the receiving module 1610 is further configured to receive a duration setting input for the audio duration setting interface;

[0273] The device also includes: an updating module, configured to update the audio recording duration of the first dynamic photo in response to the duration setting input.

[0274] In some embodiments of the present application, the receiving module 1610 is further configured to receive a second input from the user;

[0275] The display module 1620 is further configured to display an audio editing interface for the first dynamic photo in response to the second input, the audio editing interface including a volume adjustment control for the first audio;

[0276] The receiving module 1610 is further configured to receive a control input to the volume adjustment control;

[0277] The device further includes an updating module configured to update the playback volume of the first audio in response to the control input.

[0278] In some embodiments of the present application, the audio editing interface of the first dynamic photo further includes: a music import control; the receiving module 1610 is further configured to: after displaying the audio editing interface of the first dynamic photo, receive a control input to the music import control;

[0279] The display module 1620 is further configured to display at least one audio identifier in response to the control input; wherein one audio identifier indicates one audio;

[0280] The receiving module 1610 is further configured to receive a user's selection input of the at least one audio identifier;

[0281] The display module 1620 is further configured to display an audio segment capture control for a third audio in response to the selection input, where the third audio is the audio indicated by the audio identifier selected by the selection input;

[0282] The receiving module 1610 is further configured to receive a setting input for the audio segment interception control;

[0283] The device further includes: a cutting module, configured to, in response to the setting input, cut an audio segment of a fourth duration set by the setting input from the third audio, and store the audio segment of the fourth duration in association with the first cover image, wherein the fourth duration is the same as the audio duration of the first audio;

[0284] The playback module 1630 is also used to synchronously play the first video, the first audio and the audio segment of the fourth length upon receiving a playback control input for the first cover image; wherein the playback start timestamp of the audio segment of the fourth length is the same as the playback start timestamp of the first audio.

[0285] In some embodiments of the present application, all dynamic photos stored in the album program include a second dynamic photo, and the second dynamic photo includes a second cover image, a second video, and a second audio stored in association;

[0286] The receiving module 1610 is further configured to receive a third input from a user regarding the first cover image and the second cover image;

[0287] The display module 1620 is further configured to display a third dynamic photo in response to the third input, the third dynamic photo including the third cover image, the first video, the first audio, the second video, and the second audio stored in association;

[0288] The receiving module 1610 is further configured to receive a user input for controlling the playback of the third cover image;

[0289] The playing module 1630 is further configured to, in response to the play control input, play the first video, the first audio, the second video, and the second audio in the order of the third input;

[0290] Among them, the playback start timestamp of the first audio is the same as the playback start timestamp of the first video; the playback start timestamp of the second audio is the same as the playback start timestamp of the second video; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio.

[0291] In some embodiments of the present application, the apparatus further comprises:

[0292] a determination module configured to, before displaying the third dynamic photo, determine a third cover image based on the input order of the third input; or, in response to a selection input of the first cover image or the second cover image, use the cover image selected by the selection input as the third cover image; or, select any one video frame from the first video and the second video as the third cover image;

[0293] A storage module is used to store the third cover image, the first video, the first audio, the second video and the second audio in an associated manner.

[0294] In some embodiments of the present application, the receiving module 1610 is further configured to: after displaying the third dynamic photo, receive a fourth input from the user;

[0295] The display module 1620 is further configured to display, in response to the fourth input, an audio editing interface for the third dynamic photo, the audio editing interface comprising a first editing area and a second editing area, the first editing area comprising a first volume adjustment control for the first audio, and the second editing area comprising a second volume adjustment control for the second audio;

[0296] The device further includes: an updating module configured to update the playback volume of the first audio when a control input to the first volume adjustment control is received; and to update the playback volume of the second audio when a control input to the second volume adjustment control is received.

[0297] In some embodiments of the present application, the audio editing interface of the third dynamic photo further includes a music import control;

[0298] The receiving module 1610 is further configured to: after displaying the audio editing interface of the third dynamic photo, receive a user's control input on the music import control;

[0299] The display module 1620 is further configured to display at least one audio identifier in response to the control input; wherein one audio identifier indicates one audio;

[0300] The receiving module 1610 is further configured to receive a user's selection input of the at least one audio identifier;

[0301] The display module 1620 is further configured to display an audio segment capture control for a fourth audio in response to the selection input, where the fourth audio is the audio indicated by the audio identifier selected by the selection input;

[0302] The receiving module 1610 is further configured to receive a setting input for the audio segment interception control;

[0303] The device further includes: a cutting module, configured to, in response to the setting input, cut an audio segment of a fifth duration set by the setting input from the fourth audio, and store the audio segment of the fifth duration in association with the third cover image, wherein the fifth duration is equal to the sum of the audio duration of the first audio and the audio duration of the second audio;

[0304] The playback module 1630 is further configured to play the first video, the first audio, the second video, the second audio, and the audio clip of the fifth duration when receiving a playback control input for the third cover image;

[0305] Wherein, the playback timestamp of the first audio is the same as the playback start timestamp of the first video when the first video is played for the first time; the playback start timestamp of the second audio is the same as the playback start timestamp of the second video when the second video is played for the first time; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio, and the playback start timestamp of the audio clip of the fifth duration is the same as the playback start timestamp of the first audio.

[0306] In some embodiments of the present application, the audio editing interface of the third live photo further includes a music import control;

[0307] The receiving module 1610 is further configured to receive a fifth input for the first editing area and the music import control after the audio editing interface of the third live photo is displayed;

[0308] The display module 1620 is further configured to display at least one audio identifier in response to the fifth input; wherein, one audio identifier indicates one audio.

[0309] The receiving module 1610 is further configured to receive a selection input from the user for the at least one audio identifier;

[0310] The display module 1620 is further configured to display an audio segment intercept control for the fifth audio in response to the selection input, where the fifth audio is the audio indicated by the selected audio identifier of the selection input;

[0311] The receiving module 1610 is further configured to receive a setting input for the audio segment intercept control;

[0312] The device further includes: an intercepting module, configured to intercept an audio segment of a sixth duration set by the setting input from the fifth audio in response to the setting input, and associate and store the audio segment of the sixth duration with the third cover image, where the sixth duration is the same as the audio duration of the first audio;

[0313] The playback module 1630 is further configured to play the first video, the first audio, the second video, the second audio, and the audio segment of the sixth duration when receiving a playback control input for the third cover image;

[0314] Among them, the playback timestamp of the first audio is the same as the playback start timestamp of the first play of the first video; the playback start timestamp of the second audio is the same as the playback start timestamp of the first play of the second video; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio, and the playback start timestamp of the audio segment with the sixth duration is the same as the playback start timestamp of the first audio.

[0315] The dynamic photo generation device in the embodiments of the present application may be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device may be a terminal or other devices other than terminals. Exemplarily, the electronic device may be a mobile phone, a tablet computer, a laptop computer, a handheld computer, an in-vehicle electronic device, a Mobile Internet Device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and may also be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.

[0316] The dynamic photo generation device in the embodiments of the present application may be a device with an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems. The embodiments of the present application do not make specific limitations.

[0317] The dynamic photo generation device provided in the embodiments of the present application can implement Figure 1 each process implemented by the method embodiments. To avoid repetition, it will not be elaborated here.

[0318] Optionally, as Figure 19 shown, the embodiments of the present application further provide an electronic device 1700, including a processor 1701 and a memory 1702. A program or instruction that can run on the processor 1701 is stored on the memory 1702. When the program or instruction is executed by the processor 1701, it implements each step of the above-mentioned dynamic photo generation method embodiment and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0319] It should be noted that the electronic devices in the embodiments of the present application include the above-mentioned mobile electronic devices and non-mobile electronic devices.

[0320] Figure 20 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present application.

[0321] The electronic device 1800 includes, but is not limited to: a radio frequency unit 1801, a network module 1802, an audio output unit 1803, an input unit 1804, a sensor 1805, a display unit 1806, a user input unit 1807, an interface unit 1808, a memory 1809, and a processor 1810 and other components.

[0322] Those skilled in the art can understand that the electronic device 1800 may further include a power supply (such as a battery) for supplying power to each component. The power supply can be logically connected to the processor 1810 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. Figure 20 The structure of the electronic device shown in does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0323] Among them, the user input unit 1807 is used to receive a first input;

[0324] The display unit 1806 is used to display a first dynamic photo in response to the first input; the first dynamic photo includes a first cover image, a first video, and a first audio stored in association.

[0325] The processor 1810 is used to play the first video and the first audio when receiving a playback control input for the first cover image; wherein, the audio duration of the first audio is greater than the video duration of the first video.

[0326] In this way, by responding to the user's first input, a first dynamic photo including a first cover image, a first video, and a first audio stored in association can be displayed. In this way, when the user wants to record the live pictures and sound information of shooting scenes such as concerts, KTVs, and talk shows, it can be achieved by taking dynamic photos, without taking a very long video, reducing the occupation of memory resources in the electronic device. At the same time, in the dynamic photo generated in the embodiment of the present application, the audio duration of the first audio is greater than the video duration of the first video. In this way, compared with the dynamic photos in the prior art, more audio information can be carried, and it is also convenient for post-editing of the first audio.

[0327] Optionally, the processor 1810 is further configured to synchronously play the first video and the first audio; wherein, the playback mode of the first video is loop playback, and the total playback duration of the first video is the same as the playback duration of the first audio; the playback start timestamp of the first audio is the same as the playback start timestamp of the first playback of the first video.

[0328] In this way, the first video and the first audio are synchronously played, so that during the playback of the first video and the first audio, temporal coordination and consistency can be maintained, thus avoiding the problem of audio-video asynchrony during the playback of the first video and the first audio, and improving the visual experience of the played first video and first audio.

[0329] Optionally, the first input is a control input for a photographing control; the processor 1810 is further configured to continuously cache the image data collected by the camera, control the microphone to perform audio recording, and continuously cache the digital audio data collected by the microphone when the dynamic photo function is enabled; perform image processing on the image data cached at the reference moment to generate a first cover image; perform video encoding on all the image data cached within a first time period to generate a first video; wherein, the first time period consists of a first duration before the reference moment and a second duration after the reference moment; perform audio encoding on all the digital audio data cached within a second time period to generate a first audio; wherein, the second time period consists of the first duration before the reference moment and a third duration after the reference moment, the third duration is greater than the second duration, and the reference moment is the input moment of the first input; and associatively store the first cover image, the first audio, and the first video.

[0330] In this way, when the dynamic photo function is enabled, each piece of image data collected by the camera can be continuously cached, the microphone can be controlled to perform audio recording, and the digital audio data collected by the microphone can be continuously cached. Thus, after receiving the first input from the user, image processing can be performed on the image data cached at the moment when the first input is executed to generate a first cover image, video encoding can be performed on all the image data cached within a first time period to generate a first video, audio encoding can be performed on all the digital audio data cached within a second time period to generate a first audio, and then the first cover image, the first video, and the first audio can be associatively stored. In this way, the audio duration and the video duration in the obtained dynamic photo can be of unequal length, and the audio duration of the first audio is greater than the video duration of the first video. Thus, compared with the dynamic photos obtained by traditional methods, the dynamic photos of the embodiments of the present application can carry more audio information, and at the same time, it is also convenient for the later editing of the first audio. In addition, the dynamic photos generated by the embodiments of the present application also occupy less memory resources than the long videos taken, reducing the consumption of memory resources in the electronic device.

[0331] Optionally, a display unit 1806 is configured to display the first dynamic photo in an album interface; an audio playback control is included on the dynamic photo;

[0332] A user input unit 1807 is further configured to receive a control input for the audio playback control;

[0333] A processor 1810 is further configured to play the first audio in response to the control input.

[0334] In this way, when the user only wants to play the first audio, by responding to the user's control input for the audio playback control in the first dynamic photo, only the first audio segment can be played. Thus, according to the user's needs, only the audio of the first dynamic photo can be played, improving the playback flexibility of the first dynamic photo.

[0335] Optionally, the user input unit 1807 is further configured to receive a duration setting input for an audio duration setting interface;

[0336] The processor 1810 is further configured to update the audio recording duration of the first dynamic photo in response to the duration setting input.

[0337] In this way, according to the user's needs, by responding to the user's duration setting input for the audio duration setting interface, the audio recording duration of the first dynamic photo can be updated, thus improving the flexibility of the audio recording duration of the first dynamic photo.

[0338] Optionally, the user input unit 1807 is further configured to receive a second input;

[0339] The display unit 1806 is further configured to display an audio editing interface of the first dynamic photo in response to the second input, and the audio editing interface includes a volume adjustment control for the first audio;

[0340] The user input unit 1807 is further configured to receive a control input for the volume adjustment control;

[0341] The processor 1810 is further configured to update the playback volume of the first audio in response to the control input.

[0342] In this way, by responding to the user's second input, an audio editing interface of the first dynamic photo can be displayed, and a volume adjustment control can be included in the audio editing interface. Thus, according to the user's needs, in response to the user's control input for the volume adjustment control, the playback volume of the first audio can be updated. Thus, according to the user's needs, the playback volume of the first audio can be flexibly adjusted to meet the different requirements of the user for the playback effect of the playback volume of the first audio.

[0343] Optionally, the audio editing interface of the first live photo further includes: a music import control; the user input unit 1807 is further configured to receive a control input for the music import control;

[0344] The display unit 1806 is further configured to display at least one audio identifier in response to the control input; wherein, one audio identifier indicates one audio;

[0345] The user input unit 1807 is further configured to receive a selection input of the user for the at least one audio identifier;

[0346] The display unit 1806 is further configured to display an audio segment truncation control for a third audio in response to the selection input, where the third audio is the audio indicated by the audio identifier selected by the selection input;

[0347] The user input unit 1807 is further configured to receive a setting input for the audio segment truncation control;

[0348] The processor 1810 is further configured to, in response to the setting input, truncate an audio segment of a fourth duration set by the setting input from the third audio, and associate and store the audio segment of the fourth duration with the first cover image, where the fourth duration is the same as the audio duration of the first audio; in the case of receiving a play control input for the first cover image, synchronously play the first video, the first audio, and the audio segment of the fourth duration; wherein, the play start timestamp of the audio segment of the fourth duration is the same as the play start timestamp of the first audio.

[0349] In this way, the user can intercept an audio segment from other audio according to needs, and the duration of the audio segment is the same as the audio duration of the audio in the first live photo, and add the audio segment to the first live photo, adding other audio except the audio of the first live photo to the first live photo. In this way, according to the user's needs, an audio segment with different emotional expressions can be added to the first live photo, thereby adding a playback effect with different emotional expressions to the first live photo. For example, adding a lively audio segment to the first live photo can make the playback effect of the live photo more relaxed and pleasant, and adding a heavy audio segment to the first live photo can make the playback effect of the live photo more profound and urgent, improving the playback effect of the first live photo.

[0350] Optionally, among all the live photos stored in the album program, there is a second live photo, and the second live photo includes an associated second cover image, a second video, and a second audio; the user input unit 1807 is further configured to receive a third input from the user for the first cover image and the second cover image;

[0351] The processor 1810 is further configured to display a third dynamic photo in response to the third input, where the third dynamic photo includes the associated and stored third cover image, the first video, the first audio, the second video, and the second audio;

[0352] The user input unit 1807 is further configured to receive a playback control input from the user for the third cover image;

[0353] The processor 1810 is further configured to play the first video, the first audio, the second video, and the second audio in the input order of the third input in response to the playback control input; wherein, the playback start timestamp of the first audio is the same as the playback start timestamp of the first video's first playback; the playback start timestamp of the second audio is the same as the playback start timestamp of the second video's first playback; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio.

[0354] In this way, by responding to the user's third input for the first cover image and the second cover image, the third dynamic photo can be displayed, and then by responding to the user's playback control input for the third cover image, the first video, the first audio, the second video, and the second audio can be played. Thus, according to the user's needs, the videos and audios of the dynamic photos corresponding to multiple cover images can be played together without the user having to perform a playback control input for each cover image, improving the playback convenience of the videos and audios of the dynamic photos corresponding to multiple cover images. At the same time, according to the user's needs, the dynamic photos corresponding to multiple cover images can be combined into one dynamic photo, further improving the flexibility of generating dynamic photos.

[0355] Optionally, the processor 1810 is further configured to determine the third cover image according to the input order of the third input; or, in response to a selection input for the first cover image or the second cover image, use the cover image selected by the selection input as the third cover image; or, select any video frame from the first video and the second video as the third cover image; and associate and store the third cover image, the first video, the first audio, the second video, and the second audio.

[0356] In this way, by determining the third cover image according to the order of performing the third input on the first cover image and the second cover image, or by determining the third cover image in response to the user's selection input for the first cover image or the second cover image, or by selecting any video frame from the first video and the second video as the third cover image, the generation diversity of the third cover image can be improved in different ways.

[0357] Optionally, the user input unit 1807 is further configured to receive a fourth input;

[0358] The display unit 1806 is further configured to, in response to the fourth input, display an audio editing interface for the third live photo, where the audio editing interface includes a first editing area and a second editing area, the first editing area includes a first volume adjustment control for the first audio, and the second editing area includes a second volume adjustment control for the second audio;

[0359] The processor 1810 is further configured to, when receiving a control input for the first volume adjustment control, update the playback volume of the first audio; and when receiving a control input for the second volume adjustment control, update the playback volume of the second audio.

[0360] In this way, by responding to the user's fourth input, an audio editing interface for the third live photo can be displayed. The audio editing interface can include a first editing area and a second editing area. The first editing area can include a first volume adjustment control for the first audio, and the second editing area can include a second volume adjustment control for the second audio. Then, when receiving a control input for the first volume adjustment control, the playback volume of the first audio can be updated. When receiving a control input for the second volume adjustment control, the playback volume of the second audio can be updated. In this way, at least one of the playback volumes of the first audio and the second audio can be flexibly adjusted according to the user's needs, and the differentiated requirements of the user for the playback effects of the playback volumes of the first audio and the second audio can be met.

[0361] Optionally, the audio editing interface for the third live photo further includes a music import control; the user input unit 1807 is further configured to receive a control input from the user for the music import control;

[0362] The display unit 1806 is further configured to, in response to the control input, display at least one audio identifier, where one audio identifier indicates one audio;

[0363] The user input unit 1807 is further configured to receive a selection input from the user for the at least one audio identifier;

[0364] The display unit 1806 is further configured to, in response to the selection input, display an audio segment truncation control for a fourth audio, where the fourth audio is the audio indicated by the audio identifier selected by the selection input;

[0365] The user input unit 1807 is further configured to receive a setting input for the audio segment truncation control;

[0366] The processor 1810 is further configured to, in response to the setting input, intercept an audio segment of a fifth duration set by the setting input from the fourth audio, and store the audio segment of the fifth duration in association with the third cover image, where the fifth duration is equal to the sum of the audio duration of the first audio and the audio duration of the second audio; when receiving a playback control input for the third cover image, play the first video, the first audio, the second video, the second audio, and the audio segment of the fifth duration; where the playback timestamp of the first audio is the same as the playback start timestamp of the first play of the first video; the playback start timestamp of the second audio is the same as the playback start timestamp of the first play of the second video; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio, and the playback start timestamp of the audio segment of the fifth duration is the same as the playback start timestamp of the first audio.

[0367] In this way, the user can intercept an audio segment with the same duration as the audio playback duration of the third live photo from other audio according to the need, and add it to the third live photo. In this way, according to the user's needs, an audio segment with different emotional expressions can be added to the third live photo, thereby adding a playback effect with different emotional expressions to the third live photo. For example, adding a lively audio segment to the third live photo can make the playback effect of the live photo more relaxed and pleasant, and adding a heavy audio segment to the third live photo can make the playback effect of the live photo more profound and urgent, improving the playback effect of the third live photo.

[0368] Optionally, the audio editing interface of the third live photo further includes a music import control; the user input unit 1807 is further configured to receive a fifth input for the first editing area and the music import control;

[0369] The display unit 1806 is further configured to, in response to the fifth input, display at least one audio identifier; where one audio identifier indicates one audio;

[0370] The user input unit 1807 is further configured to receive a selection input from the user for the at least one audio identifier;

[0371] The display unit 1806 is further configured to, in response to the selection input, display an audio segment intercept control for a fifth audio, where the fifth audio is the audio indicated by the audio identifier selected by the selection input;

[0372] The user input unit 1807 is further configured to receive a setting input for the audio segment intercept control;

[0373] The processor 1810 is further configured to, in response to the setting input, intercept an audio segment of a sixth duration set by the setting input from the fifth audio, and associate and store the audio segment of the sixth duration with the third cover image, where the sixth duration is the same as the audio duration of the first audio; in the case of receiving a play control input for the third cover image, play the first video, the first audio, the second video, the second audio, and the audio segment of the sixth duration; where the play timestamp of the first audio is the same as the play start timestamp of the first play of the first video; the play start timestamp of the second audio is the same as the play start timestamp of the first play of the second video; the play start timestamp of the second audio is later than the play end timestamp of the first audio, and the play start timestamp of the audio segment of the sixth duration is the same as the play start timestamp of the first audio.

[0374] In this way, the user can intercept an audio segment with the same duration as the audio play duration of the first audio or the audio play duration of the second audio in the third dynamic photo from other audio according to the demand, and add it to the third dynamic photo. In this way, different audio segments with different emotional expressions required by the user can be added to the third dynamic photo, and then different emotional expression play effects can be added to the third dynamic photo. For example, adding a lively audio segment to the third dynamic photo can make the play effect of the dynamic photo more relaxed and pleasant, and adding a heavy audio segment to the third dynamic photo can make the play effect of the dynamic photo more profound and urgent, improving the play effect of the third dynamic photo.

[0375] It should be understood that in the embodiments of the present application, the input unit 1804 may include a Graphics Processing Unit (GPU) 18041 and a microphone 18042. The graphics processor 18041 processes the image data of static pictures or videos obtained by an image capture device (such as a color camera) in the video capture mode or the image capture mode. The display unit 1806 may include a display panel 18061, and the display panel 18061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1807 includes at least one of a touch panel 18071 and other input devices 18072. The touch panel 18071 is also called a touch screen. The touch panel 18071 may include two parts: a touch detection device and a touch controller. The other input devices 18072 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which will not be elaborated here.

[0376] The memory 1809 can be used to store software programs and various data. The memory 1809 mainly includes a first storage area for storing programs or instructions and a second storage area for storing data. Among them, the first storage area can store an operating system, application programs or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1809 can include a volatile memory or a non-volatile memory, or the memory 1809 can include both a volatile memory and a non-volatile memory. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link dynamic random access memory (SLDRAM), and a direct rambus random access memory (DRRAM). The memory 1809 in the embodiments of the present application includes, but is not limited to, these and any other suitable types of memories.

[0377] The processor 1810 can include one or more processing units; optionally, the processor 1810 integrates an application processor and a modem processor. Among them, the application processor mainly processes operations related to the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication signals, such as a baseband processor. It can be understood that the above modem processor may not be integrated into the processor 1810.

[0378] The embodiments of the present application also provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above embodiments of the dynamic photo generation method and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0379] Among them, the processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media such as computer read-only memory ROM, random access memory RAM, magnetic disks, or optical discs.

[0380] Another embodiment of the present application provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement each process of the above-mentioned embodiment of the dynamic photo generation method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0381] It should be understood that the chip mentioned in the embodiments of the present application may also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.

[0382] The embodiments of the present application provide a computer program product. The program product is stored in a storage medium and is executed by at least one processor to implement each process of the above-mentioned embodiment of the dynamic photo generation method, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.

[0383] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including that element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0384] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a computer software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0385] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A method for generating a dynamic photo, characterized in that, The method includes: Receiving a first input; In response to the first input, displaying a first live photo; the first live photo includes a first cover image, a first video, and a first audio that are associated and stored; When a play control input for the first cover image is received, playing the first video and the first audio; Wherein, the audio duration of the first audio is greater than the video duration of the first video.

2. The method according to claim 1, wherein The playing of the first video and the first audio includes: Synchronously playing the first video and the first audio; Wherein, the playing mode of the first video is loop play, and the total playing duration of the first video is the same as the playing duration of the first audio; the playing start timestamp of the first audio is the same as the playing start timestamp of the first video when the first video is played for the first time.

3. The method according to claim 1, wherein The first input is a control input for a camera control; Before receiving the first input, the method further includes: When the live photo function is enabled, continuously caching the image data collected by the camera, controlling the microphone to perform audio recording, and continuously caching the digital audio data collected by the microphone; Before displaying the first live photo, it further includes: Performing image processing on the image data cached at the reference moment to generate a first cover image; Performing video encoding on all the image data cached within a first time period to generate a first video; wherein, the first time period consists of a first duration before the reference moment and a second duration after the reference moment; Performing audio encoding on all the digital audio data cached within a second time period to generate a first audio; wherein, the second time period consists of the first duration before the reference moment and a third duration after the reference moment; Associatively storing the first cover image, the first audio, and the first video; Wherein, the third duration is greater than the second duration, and the reference moment is the input moment of the first input.

4. The method according to claim 1, characterized in that, The method further includes: Displaying the first live photo in the album interface; the live photo includes an audio play control; Receiving a control input for the audio play control; In response to the control input, playing the first audio.

5. The method according to claim 1, characterized in that, The method further includes: Receiving a duration setting input for an audio duration setting interface; In response to the duration setting input, updating the audio recording duration of the first live photo.

6. The method according to claim 1, characterized in that, The method further includes: Receiving a second input; In response to the second input, displaying an audio editing interface of the first live photo, the audio editing interface includes a volume adjustment control for the first audio; Receiving a control input for the volume adjustment control; ​ 7. The method according to claim 6, wherein ​ ​ ​ ​ ​ In response to the selection input, display an audio segment truncation control for a third audio, where the third audio is the audio indicated by the audio identifier selected by the selection input; Receive a setting input for the audio segment truncation control; In response to the setting input, truncate an audio segment of a fourth duration set by the setting input from the third audio, and associatively store the audio segment of the fourth duration with the first cover image, where the fourth duration is the same as the audio duration of the first audio; When a playback control input for the first cover image is received, synchronously play the first video, the first audio, and the audio segment of the fourth duration; where the playback start timestamp of the audio segment of the fourth duration is the same as the playback start timestamp of the first audio.

8. The method according to claim 1, characterized in that, Among all the dynamic photos stored in the album program, there is a second dynamic photo, and the second dynamic photo includes an associated second cover image, a second video, and a second audio; The method further includes: Receive a third input from the user for the first cover image and the second cover image; In response to the third input, display a third dynamic photo, where the third dynamic photo includes an associated third cover image, the first video, the first audio, the second video, and the second audio; Receive a playback control input from the user for the third cover image; In response to the playback control input, play the first video, the first audio, the second video, and the second audio in the input order of the third input; where the playback start timestamp of the first audio is the same as the playback start timestamp of the first play of the first video; the playback start timestamp of the second audio is the same as the playback start timestamp of the first play of the second video; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio.

9. The method according to claim 8, characterized in that, Before displaying the third dynamic photo, it further includes: Determine the third cover image according to the input order of the third input; or, in response to a selection input for the first cover image or the second cover image, use the cover image selected by the selection input as the third cover image; or, select any video frame from the first video and the second video as the third cover image; Associatively store the third cover image, the first video, the first audio, the second video, and the second audio.

10. The method according to claim 9, characterized in that, After displaying the third dynamic photo, the method further includes: Receive a fourth input; In response to the fourth input, display an audio editing interface for the third dynamic photo, where the audio editing interface includes a first editing area and a second editing area, the first editing area includes a first volume adjustment control for the first audio, and the second editing area includes a second volume adjustment control for the second audio; When a control input for the first volume adjustment control is received, update the playback volume of the first audio; When a control input for the second volume adjustment control is received, update the playback volume of the second audio.

11. The method according to claim 10, wherein The audio editing interface of the third dynamic photo also includes a music import control; After displaying the audio editing interface of the third dynamic photo, the method further includes: Receiving a user's control input to the music import control; In response to the control input, displaying at least one audio identifier; wherein one audio identifier indicates one audio; receiving a user selection input of the at least one audio identifier; In response to the selection input, displaying an audio segment interception control for a fourth audio, the fourth audio being the audio indicated by the audio identifier selected by the selection input; receiving a setting input for the audio segment interception control; In response to the setting input, an audio segment of a fifth duration set by the setting input is intercepted from the fourth audio, and the audio segment of the fifth duration is associated with the third cover image and stored, wherein the fifth duration is equal to the sum of the audio duration of the first audio and the audio duration of the second audio; When receiving a play control input for the third cover image, play the first video, the first audio, the second video, the second audio, and the fifth duration audio clip; Among them, the playback timestamp of the first audio is the same as the playback start timestamp of the first video playback; the playback start timestamp of the second audio is the same as the playback start timestamp of the second video playback; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio, and the playback start timestamp of the fifth-length audio segment is the same as the playback start timestamp of the first audio.

12. The method according to claim 10, wherein The audio editing interface of the third dynamic photo also includes a music import control; After displaying the audio editing interface of the third dynamic photo, the method further includes: receiving a fifth input to the first editing area and the music import control; In response to the fifth input, displaying at least one audio identifier; wherein one audio identifier indicates one audio; receiving a user selection input of the at least one audio identifier; In response to the selection input, displaying an audio clip capture control for a fifth audio, wherein the fifth audio is the audio indicated by the audio identifier selected by the selection input; receiving a setting input for the audio segment interception control; In response to the setting input, extracting an audio segment of a sixth duration set by the setting input from the fifth audio, and storing the audio segment of the sixth duration in association with the third cover image, wherein the sixth duration is the same as the audio duration of the first audio; When receiving a play control input for the third cover image, play the first video, the first audio, the second video, the second audio, and the audio clip of the sixth duration; Among them, the playback timestamp of the first audio is the same as the playback start timestamp of the first play of the first video; the playback start timestamp of the second audio is the same as the playback start timestamp of the first play of the second video; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio, and the playback start timestamp of the audio segment with the sixth duration is the same as the playback start timestamp of the first audio.

13. A dynamic photo generation device, characterized in that, The device includes: a receiving module, configured to receive a first input; a display module, configured to display a first live photo in response to the first input; the first live photo includes an associated first cover image, a first video, and a first audio; a playback module, configured to play the first video and the first audio when a playback control input for the first cover image is received; wherein, the audio duration of the first audio is greater than the video duration of the first video.

14. The device according to claim 13, wherein Specifically, the playback module is configured to: synchronously play the first video and the first audio; wherein, the playback mode of the first video is loop playback, and the total playback duration of the first video is the same as the playback duration of the first audio; the playback start timestamp of the first audio is the same as the playback start timestamp of the first play of the first video.

15. The device according to claim 13, characterized in that, The first input is a control input for a photographing control; the device further includes: a caching module, configured to continuously cache image data collected by a camera, control a microphone to perform audio recording, and continuously cache digital audio data collected by the microphone before receiving the first input when the live photo function is enabled; a generating module, configured to perform image processing on the image data cached at a reference moment to generate a first cover image before displaying the first live photo; perform video encoding on all the image data cached within a first time period to generate a first video; wherein, the first time period consists of a first duration before the reference moment and a second duration after the reference moment; perform audio encoding on all the digital audio data cached within a second time period to generate a first audio; wherein, the second time period consists of the first duration before the reference moment and a third duration after the reference moment, the third duration is greater than the second duration, and the reference moment is the input moment of the first input; a storage module, configured to associatively store the first cover image, the first audio, and the first video.

16. The device according to claim 13, characterized in that, The display module is further configured to display the first live photo in an album interface; an audio playback control is included on the live photo; the receiving module is further configured to receive a control input for the audio playback control; the playback module is further configured to play the first audio in response to the control input.

17. The device according to claim 13, wherein, the receiving module is further configured to receive a duration setting input for an audio duration setting interface; the device further includes: an updating module, configured to update the audio recording duration of the first live photo in response to the duration setting input.

18. The device according to claim 13, characterized in that, the receiving module is further configured to receive a second input from a user; The display module is further configured to display an audio editing interface of the first dynamic photo in response to the second input, the audio editing interface including a volume adjustment control for the first audio; The receiving module is further configured to receive a control input to the volume adjustment control; The device further includes an updating module configured to update the playback volume of the first audio in response to the control input.

19. The device according to claim 18, characterized in that, The audio editing interface of the first dynamic photo further includes a music import control; the receiving module is further configured to receive a control input to the music import control after the audio editing interface of the first dynamic photo is displayed; The display module is further configured to display at least one audio identifier in response to the control input; wherein one audio identifier indicates one audio; The receiving module is further configured to receive a user's selection input of the at least one audio identifier; The display module is further configured to display an audio segment capture control for a third audio in response to the selection input, where the third audio is the audio indicated by the audio identifier selected by the selection input; The receiving module is further configured to receive a setting input for the audio segment interception control; The device further includes: a cutting module, configured to, in response to the setting input, cut an audio segment of a fourth duration set by the setting input from the third audio, and store the audio segment of the fourth duration in association with the first cover image, wherein the fourth duration is the same as the audio duration of the first audio; The playback module is also used to synchronously play the first video, the first audio and the audio segment of the fourth length upon receiving a playback control input for the first cover image; wherein the playback start timestamp of the audio segment of the fourth length is the same as the playback start timestamp of the first audio.

20. The device according to claim 13, characterized in that, All dynamic photos stored in the album program include a second dynamic photo, and the second dynamic photo includes a second cover image, a second video, and a second audio stored in association; The receiving module is further configured to receive a third input from a user regarding the first cover image and the second cover image; The display module is further configured to display a third dynamic photo in response to the third input, the third dynamic photo including a third cover image, the first video, the first audio, the second video, and the second audio stored in association; The receiving module is further configured to receive a user's input for controlling the playback of the third cover image; The playback module is further configured to, in response to the playback control input, play the first video, the first audio, the second video, and the second audio in the input order of the third input; Among them, the playback start timestamp of the first audio is the same as the playback start timestamp of the first video; the playback start timestamp of the second audio is the same as the playback start timestamp of the second video; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio.

21. The device according to claim 20, characterized in that, The device further comprises: a determination module configured to, before displaying the third dynamic photo, determine a third cover image based on the input order of the third input; or, in response to a selection input of the first cover image or the second cover image, use the cover image selected by the selection input as the third cover image; or, select any one video frame from the first video and the second video as the third cover image; A storage module is used to store the third cover image, the first video, the first audio, the second video and the second audio in an associated manner.

22. The device according to claim 21, characterized in that, The receiving module is further configured to: after displaying the third dynamic photo, receive a fourth input from the user; The display module is further configured to display, in response to the fourth input, an audio editing interface for the third dynamic photo, the audio editing interface comprising a first editing area and a second editing area, the first editing area comprising a first volume adjustment control for the first audio, and the second editing area comprising a second volume adjustment control for the second audio; The device further includes: an updating module configured to update the playback volume of the first audio when a control input to the first volume adjustment control is received; and to update the playback volume of the second audio when a control input to the second volume adjustment control is received.

23. The device according to claim 22, characterized in that, The audio editing interface of the third dynamic photo also includes a music import control; The receiving module is further configured to: after displaying the audio editing interface of the third dynamic photo, receive a user's control input on the music import control; The display module is further configured to display at least one audio identifier in response to the control input; wherein one audio identifier indicates one audio; The receiving module is further configured to receive a user's selection input of the at least one audio identifier; The display module is further configured to display an audio segment capture control for a fourth audio in response to the selection input, where the fourth audio is the audio indicated by the audio identifier selected by the selection input; The receiving module is further configured to receive a setting input for the audio segment interception control; The device further includes: a cutting module, configured to, in response to the setting input, cut an audio segment of a fifth duration set by the setting input from the fourth audio, and store the audio segment of the fifth duration in association with the third cover image, wherein the fifth duration is equal to the sum of the audio duration of the first audio and the audio duration of the second audio; The playback module is further configured to, upon receiving a playback control input for the third cover image, play the first video, the first audio, the second video, the second audio, and the audio clip of the fifth duration; Among them, the playback timestamp of the first audio is the same as the playback start timestamp of the first play of the first video; the playback start timestamp of the second audio is the same as the playback start timestamp of the first play of the second video; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio, and the playback start timestamp of the audio segment with the fifth duration is the same as the playback start timestamp of the first audio.

24. The device according to claim 22, characterized in that, The audio editing interface of the third live photo further includes a music import control. The receiving module is further configured to, after the audio editing interface of the third live photo is displayed, receive a fifth input to the first editing area and the music import control. The display module is further configured to, in response to the fifth input, display at least one audio identifier; wherein, one audio identifier indicates one audio. The receiving module is further configured to receive a selection input from the user to the at least one audio identifier. The display module is further configured to, in response to the selection input, display an audio segment intercept control for the fifth audio, where the fifth audio is the audio indicated by the audio identifier selected by the selection input. The receiving module is further configured to receive a setting input to the audio segment intercept control. The device further includes: an intercepting module, configured to, in response to the setting input, intercept an audio segment with a sixth duration set by the setting input from the fifth audio, and associate and store the audio segment with the sixth duration with the third cover image, where the sixth duration is the same as the audio duration of the first audio. The playback module is further configured to, when a playback control input to the third cover image is received, play the first video, the first audio, the second video, the second audio, and the audio segment with the sixth duration. Among them, the playback timestamp of the first audio is the same as the playback start timestamp of the first play of the first video; the playback start timestamp of the second audio is the same as the playback start timestamp of the first play of the second video; the playback start timestamp of the second audio is later than the playback end timestamp of the first audio, and the playback start timestamp of the audio segment with the sixth duration is the same as the playback start timestamp of the first audio.

25. An electronic device, characterized in that, It includes a processor and a memory, and the memory stores a program or instruction that can run on the processor. When the program or instruction is executed by the processor, the steps of the live photo generation method according to any one of claims 1 to 12 are implemented.