Live photo shooting method, apparatus, device, storage medium and program product

By encoding the image groups of video frames and using a sliding window mechanism during the dynamic photo shooting process, the problem of excessive memory usage in dynamic photo shooting is solved, and efficient dynamic photo generation and system performance optimization are achieved.

WO2025200510A1PCT designated stage Publication Date: 2025-10-02BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/134336
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2024-11-25
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

During the dynamic photo shooting process, the electronic device cannot predict when the user will press the shooting control, resulting in a large number of video frame images being cached in the dynamic photo shooting mode, occupying a large amount of memory.

Method used

By performing image group encoding on the video frame images within the target time window, a video frame sequence of the image group is generated. When a trigger operation is received, it is moved to the video file to synthesize a dynamic photo. The sliding window mechanism is combined to delete outdated frame sequences and optimize memory usage.

Benefits of technology

It reduces the memory usage in dynamic shooting mode, improves system performance and shooting efficiency, ensures that the first frame of the shooting file is consistent with the first frame of the final product, and avoids time-consuming search and secondary encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024134336_02102025_PF_FP_ABST
    Figure CN2024134336_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a live photo shooting method, an apparatus, a device, a storage medium and a program product. The method comprises: in response to a received trigger operation for a shooting control, shooting a still photo; acquiring at least one image group video frame sequence, the at least one image group video frame sequence being obtained by coding in units of image groups video image frames collected in a target time window, and the target time window comprising the moment when the trigger operation is received; and, on the basis of the still photo and the at least one image group video frame sequence, compositing a live photo.
Need to check novelty before this filing date? Find Prior Art

Description

Dynamic photo shooting method, device, equipment, storage medium and program product

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese invention patent application with application number 202410382805.0, entitled “A similarity determination method, device, equipment, medium, product” and filing date March 29, 2024, and the entire application is incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the field of terminal technology, and in particular to a method, device, electronic device, computer-readable storage medium, and program product for shooting dynamic photos. Background Art

[0004] As people's living standards continue to improve, static photos no longer meet their photography needs. Therefore, a new shooting mode has emerged: Live Photo. Live Photos are a composite of the static photo taken when the user presses the capture button and a dynamic video of a preset length (e.g., 1.5 seconds) recorded before (or before or after) the static photo. When the user long-presses on the Live Photo, the photo automatically plays a dynamic effect, recreating the dynamic moment of capture. Summary of the Invention

[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a method, device, electronic device, computer-readable storage medium and program product for taking dynamic photos.

[0006] According to a first aspect of an embodiment of the present disclosure, a method for shooting dynamic photos is provided, which includes: shooting a static photo in response to a trigger operation received on a shooting control; obtaining at least one image group video frame sequence, wherein the at least one image group video frame sequence is obtained by encoding video frame images captured within a target time window in units of image groups, and the target time window includes the moment when the trigger operation is received; and synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence.

[0007] In some embodiments of the present disclosure, before receiving the trigger operation, the at least one image group video frame sequence is stored in a target cache; obtaining the at least one image group video frame sequence includes: moving the at least one image group video frame sequence from the target cache to a video file; synthesizing the dynamic photo based on the static photo and the at least one image group video frame sequence includes: synthesizing the dynamic photo based on the static photo and the at least one image group video frame sequence stored in the video file.

[0008] In some embodiments of the present disclosure, the at least one image group video frame sequence is an image group video frame sequence in an image group queue; before synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence, the method further includes: deleting the first image group video frame sequence in the image group queue when the time interval between the current frame encoded image saved in the image group queue and the first frame of the second image group video frame sequence in the image group queue is greater than the target duration; obtaining the at least one image group video frame sequence includes: obtaining all image group video frame sequences in the image group queue to obtain the at least one image group video frame sequence; wherein the duration of the target time window is the duration corresponding to the image group queue, the duration corresponding to the image group queue is greater than the target duration, and is less than or equal to the sum of the target duration and the first duration, and the first duration is the duration corresponding to the first image group video frame sequence in the image group queue.

[0009] In some embodiments of the present disclosure, before synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence, the method further includes: acquiring a current frame image, the current frame image being any frame image saved in the image group queue after encoding; in the case where the current frame image is an I frame, encoding the current frame image to obtain a current frame encoded image, and saving the current frame encoded image as the first frame of the next image group video frame sequence in the image group queue to the image group queue; in the case where the current frame image is not an I frame and the number of frames of the last image group video frame sequence in the image group queue is greater than or equal to In the case of the target frame number, the current frame image is intra-encoded to obtain the current frame coded image, and the current frame coded image is used as the first frame of the next picture group video frame sequence of the picture group queue and saved in the picture group queue; in the case that the current frame image is not an I frame and the number of frames of the last picture group video frame sequence in the picture group queue is less than the target frame number, the current frame image is encoded to obtain the current frame coded image, and the current frame coded image is used as a frame of the last picture group video frame sequence and saved in the picture group queue; wherein the duration corresponding to the target frame number is less than the target duration.

[0010] In some embodiments of the present disclosure, synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence includes: using the static photo as a cover and synthesizing it with the at least one image group video frame sequence to form the dynamic photo.

[0011] Some embodiments of the present disclosure synthesize a dynamic photo based on the static photo and the at least one image group video frame sequence, including: displaying multiple candidate images, the multiple candidate images including the static photo, and the candidate images other than the static photo among the multiple candidate images are pictures in the at least one image group video frame sequence; in response to a selection operation of a target candidate image among the multiple candidate images, using the target candidate image as a cover and synthesizing it with the at least one image group video frame sequence to form the dynamic photo.

[0012] According to a second aspect of an embodiment of the present disclosure, a dynamic photo shooting device is provided, which includes: a shooting module for shooting a static photo in response to a trigger operation received on a shooting control; an acquisition module for acquiring at least one image group video frame sequence, wherein the at least one image group video frame sequence is obtained by encoding video frame images captured within a target time window in units of image groups, and the target time window includes the moment when the trigger operation is received; and a synthesis module for synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence.

[0013] In some embodiments of the present disclosure, before receiving the trigger operation, the at least one image group video frame sequence is stored in a target cache; the acquisition module is specifically used to move the at least one image group video frame sequence from the target cache to a video file; and the synthesis module is specifically used to synthesize the dynamic photo based on the static photo and the at least one image group video frame sequence stored in the video file.

[0014] In some embodiments of the present disclosure, the at least one image group video frame sequence is an image group video frame sequence in an image group queue; the device also includes: a deletion module for deleting the first image group video frame sequence in the image group queue before synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence, if the time interval between the current frame encoded image saved in the image group queue and the first frame of the second image group video frame sequence in the image group queue is greater than the target duration; the acquisition module is specifically used to acquire all image group video frame sequences in the image group queue to obtain the at least one image group video frame sequence; wherein the duration of the target time window is the duration corresponding to the image group queue, the duration corresponding to the image group queue is greater than the target duration, and less than or equal to the sum of the target duration and the first duration, and the first duration is the duration corresponding to the first image group video frame sequence in the image group queue.

[0015] In some embodiments of the present disclosure, the device further includes: an acquisition module for acquiring a current frame image before synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence, the current frame image being any frame image saved in the image group queue after encoding; an encoding and saving module for encoding the current frame image to obtain a current frame encoded image when the current frame image is an I frame, and saving the current frame encoded image as the first frame of the next image group video frame sequence in the image group queue to the image group queue; when the current frame image is not an I frame and the last image group video frame sequence in the image group queue When the number of frames in the column is greater than or equal to the target number of frames, the current frame image is intra-coded to obtain the current frame coded image, and the current frame coded image is used as the first frame of the next picture group video frame sequence of the picture group queue and saved in the picture group queue; when the current frame image is not an I frame and the number of frames of the last picture group video frame sequence in the picture group queue is less than the target number of frames, the current frame image is encoded to obtain the current frame coded image, and the current frame coded image is used as a frame of the last picture group video frame sequence and saved in the picture group queue; wherein the duration corresponding to the target number of frames is less than the target duration.

[0016] In some embodiments of the present disclosure, the synthesis module is specifically configured to use the static photo as a cover and synthesize it with the at least one image group video frame sequence into the dynamic photo.

[0017] In some embodiments of the present disclosure, the synthesis module is specifically used to display multiple candidate images, where the multiple candidate images include the static photo, and the candidate images other than the static photo among the multiple candidate images are pictures in the video frame sequence of the at least one image group; in response to a selection operation of a target candidate image among the multiple candidate images, the target candidate image is used as a cover and synthesized with the at least one image group video frame sequence to form the dynamic photo.

[0018] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the dynamic photo shooting method as described in the first aspect is implemented.

[0019] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the dynamic photo shooting method as described in the first aspect is implemented.

[0020] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided, wherein the computer program product includes a computer program. When the computer program product runs on a processor, the processor executes the computer program to realize the dynamic photo shooting method as described in the first aspect. According to a sixth aspect of the embodiments of the present disclosure, a chip is provided, which includes a processor and a communication interface, the communication interface is coupled to the processor, and the processor is used to run program instructions to realize the dynamic photo shooting method as described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0022] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0023] FIG1 is a flow chart of a method for taking dynamic photos according to an embodiment of the present disclosure;

[0024] FIG2 is a schematic diagram of an image group queue according to an embodiment of the present disclosure;

[0025] FIG3 is a second schematic diagram of an image group queue provided by an embodiment of the present disclosure;

[0026] FIG4 is a second flow chart of a method for taking dynamic photos according to an embodiment of the present disclosure;

[0027] FIG5 is a structural block diagram of a dynamic photo shooting device provided by an embodiment of the present disclosure;

[0028] FIG6 is a structural block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0029] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0030] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0031] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects related to each other are in an "or" relationship.

[0032] The electronic devices in the embodiments of the present disclosure may be mobile electronic devices or non-mobile electronic devices. Mobile electronic devices may include mobile phones, tablet computers, laptop computers, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs); non-mobile electronic devices may include personal computers (PCs), televisions (TVs), ATMs, or self-service kiosks; and the embodiments of the present disclosure do not specifically limit these.

[0033] As mentioned above, dynamic photos are increasingly used. However, because dynamic photos include images before the photo is taken, and the electronic device cannot predict when the user will press the capture control to take the photo, when the electronic device enters dynamic photo mode, the electronic device needs to continuously cache the real-time images captured by the camera into a video file. If the difference between the time when the user presses the capture control and the time when the electronic device enters dynamic photo mode is large, the electronic device will cache a large number of video frames in the video file before taking a still photo (or before and after taking a still photo), which takes up a large amount of memory.

[0034] The execution subject of the dynamic photo shooting method provided by the embodiment of the present disclosure can be the above-mentioned electronic device (including mobile electronic devices and non-mobile electronic devices), or it can be a functional module and / or functional entity in the electronic device that can implement the dynamic photo shooting method. The specific one can be determined according to actual usage requirements and is not limited by the embodiment of the present disclosure.

[0035] The following describes in detail the dynamic photo shooting method provided by the embodiment of the present disclosure through specific embodiments and their application scenarios in conjunction with the accompanying drawings.

[0036] As shown in FIG1 , an embodiment of the present disclosure provides a method for shooting dynamic photos, which may include the following steps 101 to 103 .

[0037] 101. In response to a received triggering operation on a shooting control, take a still photo.

[0038] It can be understood that before executing the dynamic photo shooting method provided by the embodiment of the present disclosure, the camera of the electronic device has been started and the dynamic shooting mode has been turned on.

[0039] The received triggering operation on the shooting control is an operation that triggers taking a picture. The shooting control may be a control such as a shutter key for triggering a shooting operation.

[0040] Among them, the triggering operation of the shooting control can include any of the following methods: the user can trigger the shooting control by clicking the shooting control component displayed on the touch screen of the electronic device; the user can trigger the shooting control through non-contact methods such as voice control or hovering gestures; the user can trigger the shooting control by pressing a physical button or a combination of physical buttons on the electronic device; the user can trigger the shooting control through a control device connected to the electronic device (for example: the electronic device is connected to a selfie stick through a USB interface, and the selfie stick has physical buttons that can control the electronic device to take pictures; for example, the electronic device is connected to a tripod shooting stick through a USB interface, and the tripod shooting stick is connected to a wireless remote control, and the electronic device can be controlled to take pictures through the wireless remote control).

[0041] 102. Obtain at least one GOP video frame sequence.

[0042] The at least one image group video frame sequence is obtained by encoding video frame images captured within a target time window in units of image groups, and the target time window includes the moment when the trigger operation is received.

[0043] In some embodiments of the present disclosure, the moment of receiving the trigger operation may be any moment within the target time window, which may be determined based on actual conditions.

[0044] Exemplarily, the end moment of the target time window may be the moment when the trigger operation is received, the start moment of the target time window may be the moment when the trigger operation is received, and the middle moment of the target time window includes the moment when the trigger operation is received.

[0045] In some embodiments of the present disclosure, when the end moment of the target time window is the moment when the trigger operation is received, the electronic device captures video frame images in real time before receiving the trigger operation, and encodes the captured video frame images.

[0046] In some embodiments of the present disclosure, when the starting moment of the target time window is the moment when the trigger operation is received, the electronic device captures video frame images in real time after receiving the trigger operation, and encodes the captured video frame images.

[0047] In some embodiments of the present disclosure, when the middle moment of the target time window includes the moment when the trigger operation is received, the electronic device captures video frame images in real time before receiving the trigger operation, and encodes the captured video frame images. After receiving the trigger operation, the electronic device still captures video frame images in real time for a period of time, and encodes the captured video frame images.

[0048] In some embodiments of the present disclosure, the start time, end time, and duration of the target time window can be determined according to actual conditions and are not limited here.

[0049] In some embodiments of the present disclosure, the target time window can be a time window with a fixed length, or a time window whose length changes with the changes of the actual captured video frame image. The specific length can be determined according to actual conditions and is not limited here.

[0050] In some embodiments of the present disclosure, the electronic device can encode the collected video frame images in real time, that is, collect one frame of video frame image and encode the one frame of video frame image; the electronic device can also collect a preset number of video frame images and then encode the preset number of video frame images; the specific determination is based on actual conditions and is not limited here.

[0051] A GOP is also called a Group of Pictures (GOP). A GOP is a set of consecutive pictures. The first frame of a GOP must be an I-frame, so that the GOP can be decoded independently without reference to other pictures.

[0052] The first frame of each GOP video frame sequence in the at least one GOP video frame sequence is an I frame.

[0053] In the embodiment of the present disclosure, since the first frame of at least one image group video frame sequence is an I frame, when a dynamic photo is synthesized based on a static photo and at least one image group video frame sequence, the first frame of the dynamic photo can be the first frame of at least one image group video frame sequence, thereby ensuring that the first frame of the shooting file (cached sequence frame, i.e., at least one image group video frame sequence) is consistent with the first frame of the final product (the first frame of the actually generated video frame, i.e., the dynamic photo), thereby avoiding the time consumption of searching (seeking) and the time consumption of secondary encoding caused by the inconsistency between the first frame of the shooting file and the final product.

[0054] It can be understood that if the first frame of the target video frame sequence used to synthesize the dynamic photo is not an I frame, the video frames between the first frame and the first I frame of the target video frame sequence cannot be decoded independently. Therefore, if the first I frame of the target video frame sequence is used as the first frame of the dynamic photo, it is necessary to find the first I frame in the target video frame sequence, and then delete the video frame sequence before the first I frame in the target video frame sequence, and generate the dynamic photo based on the deleted video frame sequence and the static photo. This will increase the extra time for finding the first I frame; if the first frame of the target video frame sequence is used as the first frame of the dynamic photo, it is necessary to find the first I frame before the target video frame sequence, and then at least decode the video frames between the first frame and the first I frame of the target video frame sequence based on the first I frame before the target video frame sequence, and then re-encode the decoded video frame sequence. This not only increases the extra time for finding the first I frame before the target video frame sequence, but also increases the extra time for secondary encoding.

[0055] 103. Synthesize a dynamic photo based on the static photo and the at least one image group video frame sequence.

[0056] Among them, the dynamic photo includes a cover picture and a target video clip. The cover picture can be a static photo or any frame of video image in the target video clip. The specific cover picture can be determined according to actual conditions and is not limited here.

[0057] In the embodiment of the present disclosure, by encoding and compressing the video frame images within the target time window collected before taking a static photo (or before and after taking a static photo) in units of image groups, at least one image group video frame sequence is obtained, and then a dynamic photo is synthesized based on the static photo and the at least one image group video frame sequence. On the one hand, the memory usage of the cached video frame images in the dynamic photo mode can be reduced; on the other hand, it can ensure that the first frame of the shooting file (cached sequence frame, that is, at least one image group video frame sequence) is consistent with the first frame of the final product (actually generated video frame, that is, the dynamic photo), which can avoid the time-consuming search and secondary encoding caused by the inconsistency between the shooting file and the first frame of the final product.

[0058] In some embodiments of the present disclosure, before receiving the trigger operation, the at least one image group video frame sequence is stored in the target cache; the above step 102 can be specifically implemented by the following step 102a, and the above step 103 can be specifically implemented by the following step 103a.

[0059] 102a. Move the at least one GOP video frame sequence from the target buffer to a video file.

[0060] It can be understood that at least one GOP video frame sequence is copied from the target cache to the video file, and at least one GOP video frame sequence in the target cache is deleted.

[0061] 103a. Synthesize the dynamic photo based on the static photo and the at least one image group video frame sequence stored in the video file.

[0062] In the disclosed embodiment, in dynamic photography mode, at least one image group video frame sequence after encoding the captured video frame images is stored in a target cache rather than in a video file. Upon receiving a trigger operation on the capture control, the at least one image group video frame sequence is moved from the target cache to the video file. Compared to storing the at least one image group video frame sequence in a file, this can reduce file input and output (IO) operations, i.e., file read and write operations, thereby improving system performance. Furthermore, by encoding the captured video frame images and storing them in the target cache rather than in a scratch disk, and only saving the resulting video file (dynamic photo) in the scratch disk, the draft disk usage can be reduced.

[0063] In some embodiments of the present disclosure, the at least one GOP video frame sequence is a GOP video frame sequence in a GOP queue.

[0064] In some embodiments of the present disclosure, before the above step 103, the dynamic photo shooting method provided by the embodiment of the present disclosure may further include the following step 104; the above step 102 may be specifically implemented through the following step 102b.

[0065] 104. When the time interval between the current frame coded image saved in the GOP queue and the first frame of the second GOP video frame sequence in the GOP queue is greater than the target duration, delete the first GOP video frame sequence in the GOP queue.

[0066] It can be understood that in the dynamic photography mode, the captured video frames are continuously encoded and cached in a queue (image group queue) in the memory, and outdated image group video frame sequences are continuously discarded in the form of a sliding window (target duration) (when the time interval between the current frame encoded image saved in the image group queue and the first frame of the second image group video frame sequence in the image group queue is greater than the target duration, the first image group video frame sequence is an outdated image group video frame sequence), and the encoded video frame sequence within the most recent period is kept in the memory (when the time interval between the current frame encoded image saved in the image group queue and the first frame of the second image group video frame sequence in the image group queue is less than or equal to the target duration, all image group video frame sequences saved in the image group queue are not outdated image group video frame sequences and can be used to synthesize dynamic photos). In this way, memory usage can be further reduced.

[0067] The target duration can be determined based on actual conditions and is not limited here.

[0068] 102b. Acquire all GIP video frame sequences in the GIP queue to obtain the at least one GIP video frame sequence.

[0069] Among them, the duration of the target time window is the duration corresponding to the image group queue, the duration corresponding to the image group queue is greater than the target duration, and less than or equal to the sum of the target duration and the first duration, and the first duration is the duration corresponding to the first image group video frame sequence in the image group queue.

[0070] The duration corresponding to the GOP queue is the sum of the durations of all GOP video frame sequences stored in the GOP queue. The duration corresponding to the first GOP video frame sequence in the GOP queue is the sum of the durations of all video frames included in the first GOP video frame sequence in the GOP queue. If each GOP video frame sequence in the GOP queue includes a different number of video frames, the first duration changes in accordance with the duration corresponding to the first GOP video frame sequence in the GOP queue.

[0071] It can be understood that since the number of video frame sequences stored in the image group queue follows the number of image group video frame sequences in the image group queue, and the number of video frames included in each image group video frame sequence is constantly changing and is not a fixed value, the target time window is not a fixed value, but changes with the corresponding duration of the image group queue.

[0072] For example, as shown in Figure 2, if the time interval between the current coded image frame in the GOP queue and the first frame of GOP 2 in the GOP queue is greater than the target duration (S), GOP 1 in the GOP queue is deleted. If at least one GOP video frame sequence is obtained at this time, the at least one GOP video frame sequence is GOP 2 to GOP n. As shown in Figure 3, if the time interval between the current coded image frame in the GOP queue and the first frame of GOP 2 is less than the target duration (S), the GOP video frame sequence in the GOP queue remains unchanged. If at least one GOP video frame sequence is obtained at this time, the at least one GOP video frame sequence is GOP 1 to GOP n. Combining Figures 2 and 3, it can be seen that the final captured product (dynamic photo) is a video with a cover image and a duration of T seconds, where T is greater than S and less than or equal to the sum of S and the duration corresponding to GOP 1.

[0073] In the embodiment of the present disclosure, since there are no outdated image group video frame sequences in the image group queue, what is always stored are the encoded video frame sequences within the most recent period. Therefore, when it is necessary to generate dynamic photos, all image group video frame sequences in the image group queue can be directly obtained, and there is no need to select the video frame sequence required for synthesizing dynamic photos from the image group queue, thereby simplifying the process of generating dynamic photos, reducing time consumption, and improving the efficiency of generating dynamic photos.

[0074] In some embodiments of the present disclosure, if the end moment of the target time window is the moment when the trigger operation is received (i.e., the moment when the dynamic photo is triggered), and the static photo is the cover of the dynamic photo, then in response to the received trigger operation, all image group video frame sequences in the image group queue can be obtained as at least one image group video frame sequence, and synthesized with the static photo to form a dynamic photo. Compared with delaying the shooting time after receiving the trigger operation, collecting the video frame images for encoding and storing them in the image group queue until the required at least one image group video frame sequence is obtained, and then synthesizing the dynamic photo, the process of shooting dynamic photos can be simplified, time consumption can be reduced, and shooting efficiency can be improved.

[0075] In some embodiments of the present disclosure, before the above-mentioned step 103, the dynamic photo shooting method provided by the embodiments of the present disclosure may further include the following steps 105 to 108.

[0076] 105. Capture the current frame image.

[0077] The current frame image is any frame image saved in the image group queue after encoding, and is any video frame image captured at a preset frame rate corresponding to the dynamic photography mode.

[0078] 106. When the current frame image is an I frame, encode the current frame image to obtain a current frame encoded image, and save the current frame encoded image as the first frame of the next picture group video frame sequence in the picture group queue to the picture group queue.

[0079] It can be understood that the current frame image is an I frame, which means that the current frame image and the last image group video frame sequence of the image group queue are no longer a set of continuous pictures. Therefore, it is necessary to perform intra-frame encoding on the current frame image to obtain the current frame encoded image, and use the current frame encoded image as the first frame of the next image group video frame sequence of the image group queue and save it to the image group queue.

[0080] 107. When the current frame image is a non-I frame and the number of frames of the last image group video frame sequence in the image group queue is greater than or equal to the target number of frames, intra-frame encoding is performed on the current frame image to obtain a current frame encoded image, and the current frame encoded image is used as the first frame of the next image group video frame sequence in the image group queue and saved in the image group queue.

[0081] 108. When the current frame image is a non-I frame and the number of frames of the last image group video frame sequence in the image group queue is less than the target number of frames, the current frame image is encoded to obtain a current frame encoded image, and the current frame encoded image is saved as a frame of the last image group video frame sequence in the image group queue.

[0082] The duration corresponding to the target number of frames is less than the target duration.

[0083] It can be understood that in order to ensure that the number of video frames included in each image group video frame sequence in the image group queue does not differ too much, and to ensure that the number of frames of at least one image group video frame sequence of the final synthesized dynamic photo is maintained within a certain frame number range, in an embodiment of the present disclosure, when the current frame image is not an I frame and the number of frames of the last image group video frame sequence in the image group queue is greater than or equal to the target number of frames, the current frame image is forced to be an I frame, or in other words, the current frame image is requested to be an I frame, and the current frame image is intra-encoded to obtain the current frame coded image, and the current frame coded image is used as the first frame of the next image group video frame sequence in the image group queue and saved to the image group queue; when the current frame image is not an I frame and the number of frames of the last image group video frame sequence in the image group queue is less than the target number of frames, the current frame image is encoded to obtain the current frame coded image, and the current frame coded image is used as a frame of the last image group video frame sequence and saved to the image group queue.

[0084] It can be understood that after entering the dynamic photography mode, the electronic device repeatedly executes the above steps 105 to 108 and step 104 until at least one image group video frame sequence is obtained, and then stops executing the above steps 105 to 108 and step 104.

[0085] For example, assuming that the end of the target time window is the moment when the trigger operation for the shooting control is received, and the cover of the dynamic photo is a static photo, as shown in FIG4 , the process of shooting a dynamic photo is as follows: before entering the dynamic shooting mode and before receiving the trigger operation, the camera captures the current frame image at a preset acquisition frame rate, encodes the current frame image to obtain the current frame encoded image, caches the current frame encoded image in a queue (image group queue) in the memory (target cache), and continuously discards outdated image group video frame sequences in the form of a sliding window (target duration), maintaining the encoded video frame sequences within the most recent period in the memory. When the trigger operation is received, the acquisition of video frame images stops, a static image is captured, and at least one encoded image group video frame sequence cached in the memory is rewritten to the video file (MP4 file). The captured static image is used as the cover, and at least one image group video frame sequence is linked to the end of the static image to form a dynamic photo.

[0086] In some embodiments of the present disclosure, when the current frame image is an I frame, the current frame image is encoded to obtain a current frame encoded image, and the current frame encoded image is saved in the image group queue as the first frame of the next image group video frame sequence of the image group queue; when the current frame image is not an I frame, the current frame image is encoded to obtain a current frame encoded image, and the current frame encoded image is saved in the image group queue as a frame of the last image group video frame sequence; the specific details can be determined according to actual conditions and are not limited here.

[0087] In some embodiments of the present disclosure, the above 103 can be specifically implemented through the following step 103b.

[0088] 103b: Use the static photo as a cover and combine it with the at least one image group video frame sequence to form the dynamic photo.

[0089] In the disclosed embodiment, using a static photo as the cover can, on the one hand, simplify the process of taking dynamic photos and improve shooting efficiency. On the other hand, since the static photo is taken in response to the received trigger operation, the shooting effect of the static photo will be better and the user will be more satisfied with the photo, which can improve the user's satisfaction with the dynamic photo.

[0090] In some embodiments of the present disclosure, the above 103 can be specifically implemented through the following steps 103c and 103d.

[0091] 103c. Display multiple candidate images.

[0092] The multiple candidate images include the static photo, and the candidate images other than the static photo in the multiple candidate images are pictures in the video frame sequence of the at least one image group.

[0093] In some embodiments of the present disclosure, before encoding the current frame, a determination is made as to whether the current frame is a candidate image. For example, a captured video frame image may be used as a candidate image at intervals of a preset duration, or the determination as to whether the captured video frame image is a candidate image may be made based on the image content of the captured video frame image. The specific determination may be based on actual circumstances and is not limited herein.

[0094] 103d. In response to a selection operation on a target candidate image from the plurality of candidate images, the target candidate image is used as a cover and combined with the at least one image group video frame sequence to form the dynamic photo.

[0095] In the disclosed embodiment, by displaying multiple candidate images, the user can select a more satisfactory candidate image as the cover, and the human-computer interaction experience can be improved.

[0096] Figure 5 is a structural block diagram of a dynamic photo shooting device shown in an embodiment of the present disclosure. As shown in Figure 5, it includes: a shooting module 501, used to shoot a static photo in response to a received trigger operation on the shooting control; an acquisition module 502, used to acquire at least one image group video frame sequence, and the at least one image group video frame sequence is obtained by encoding the video frame images collected within a target time window in units of image groups, and the target time window includes the moment when the trigger operation is received; a synthesis module 503, used to synthesize a dynamic photo based on the static photo and the at least one image group video frame sequence.

[0097] In some embodiments of the present disclosure, before receiving the trigger operation, the at least one image group video frame sequence is stored in the target cache; the acquisition module 502 is specifically used to move the at least one image group video frame sequence from the target cache to the video file; the synthesis module is specifically used to synthesize the dynamic photo based on the static photo and the at least one image group video frame sequence stored in the video file.

[0098] In some embodiments of the present disclosure, the at least one image group video frame sequence is an image group video frame sequence in an image group queue; the device also includes: a deletion module for deleting the first image group video frame sequence in the image group queue before synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence, if the time interval between the current frame encoded image saved in the image group queue and the first frame of the second image group video frame sequence in the image group queue is greater than the target duration; the acquisition module 502 is specifically used to acquire all image group video frame sequences in the image group queue to obtain the at least one image group video frame sequence; wherein the duration of the target time window is the duration corresponding to the image group queue, the duration corresponding to the image group queue is greater than the target duration, and is less than or equal to the sum of the target duration and the first duration, and the first duration is the duration corresponding to the first image group video frame sequence in the image group queue.

[0099] In some embodiments of the present disclosure, the device further includes: an acquisition module for acquiring a current frame image before synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence, the current frame image being any frame image saved in the image group queue after encoding; an encoding and saving module for encoding the current frame image to obtain a current frame encoded image when the current frame image is an I frame, and saving the current frame encoded image as the first frame of the next image group video frame sequence in the image group queue to the image group queue; when the current frame image is not an I frame and the last image group video frame sequence in the image group queue When the number of frames in the column is greater than or equal to the target number of frames, the current frame image is intra-coded to obtain the current frame coded image, and the current frame coded image is used as the first frame of the next picture group video frame sequence of the picture group queue and saved in the picture group queue; when the current frame image is not an I frame and the number of frames of the last picture group video frame sequence in the picture group queue is less than the target number of frames, the current frame image is encoded to obtain the current frame coded image, and the current frame coded image is used as a frame of the last picture group video frame sequence and saved in the picture group queue; wherein the duration corresponding to the target number of frames is less than the target duration.

[0100] In some embodiments of the present disclosure, the synthesis module 503 is specifically configured to use the static photo as the cover and synthesize it with the at least one image group video frame sequence into the dynamic photo.

[0101] In some embodiments of the present disclosure, the synthesis module 503 is specifically used to display multiple candidate images, where the multiple candidate images include the static photo, and the candidate images other than the static photo among the multiple candidate images are pictures in the video frame sequence of the at least one image group; in response to a selection operation of a target candidate image among the multiple candidate images, the target candidate image is used as a cover and synthesized with the at least one image group video frame sequence to form the dynamic photo.

[0102] In the embodiments of the present disclosure, each module can implement the dynamic photo shooting method provided by the above method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described here.

[0103] FIG6 is a schematic structural diagram of an electronic device provided in an embodiment of the present disclosure, which is used to exemplify an electronic device that implements any dynamic photo shooting method in an embodiment of the present disclosure and should not be understood as a specific limitation on the embodiment of the present disclosure.

[0104] As shown in Figure 6, the electronic device 600 may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 are also stored in the RAM 603. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0105] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or by wire to exchange data. Although the electronic device 600 is shown as having various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may alternatively be implemented or have.

[0106] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processor 601, the functions defined in any dynamic photo shooting method provided by the embodiment of the present disclosure can be executed.

[0107] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. Computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0108] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0109] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0110] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device: takes a static photo in response to a received trigger operation on the shooting control; obtains at least one image group video frame sequence, and the at least one image group video frame sequence is obtained by encoding the video frame images captured within a target time window in units of image groups, and the target time window includes the moment when the trigger operation is received; synthesizes a dynamic photo based on the static photo and the at least one image group video frame sequence.

[0111] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the computer, partially on the computer, as a separate software package, partially on the computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0113] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0114] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0115] In the context of the present disclosure, a computer-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a computer-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0116] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0117] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0118] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for shooting dynamic photos, wherein the method comprises: In response to receiving a trigger operation on a capture control, capturing a still photo; Acquire at least one GOP video frame sequence, where the at least one GOP video frame sequence is obtained by encoding video frame images captured within a target time window in units of GOPs, where the target time window includes the moment when the trigger operation is received; A dynamic photo is synthesized based on the static photo and the at least one image group video frame sequence.

2. The method of claim 1 , wherein before receiving the triggering operation, the at least one GOP video frame sequence is stored in a target buffer; The acquiring of at least one GOP video frame sequence comprises: moving the at least one GOP video frame sequence from the target buffer to a video file; The synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence includes: The dynamic photo is synthesized based on the static photo and the at least one image group video frame sequence stored in the video file.

3. The method according to claim 1 , wherein the at least one GOP video frame sequence is a GOP video frame sequence in a GOP queue; and before synthesizing the dynamic photo based on the still photo and the at least one GOP video frame sequence, the method further comprises: deleting the first GOP video frame sequence in the GOP queue when a time interval between a current frame coded image saved in the GOP queue and a first frame of a second GOP video frame sequence in the GOP queue is greater than a target duration; The acquiring of at least one GOP video frame sequence comprises: Acquire all GOP video frame sequences in the GOP queue to obtain the at least one GOP video frame sequence; Among them, the duration of the target time window is the duration corresponding to the image group queue, the duration corresponding to the image group queue is greater than the target duration, and less than or equal to the sum of the target duration and the first duration, and the first duration is the duration corresponding to the first image group video frame sequence in the image group queue.

4. The method according to claim 3, wherein before synthesizing the dynamic photo based on the static photo and the at least one image group video frame sequence, the method further comprises: Capturing a current frame image, wherein the current frame image is any frame image saved in the image group queue after being encoded; When the current frame image is an I frame, encoding the current frame image to obtain a current frame encoded image, and using the current frame encoded image as the first frame of the next picture group video frame sequence in the picture group queue, and saving the current frame encoded image to the picture group queue; If the current frame image is not an I-frame and the number of frames of the last GOP video frame sequence in the GOP queue is greater than or equal to the target number of frames, intra-coding the current frame image to obtain a current frame coded image, and using the current frame coded image as the first frame of the next GOP video frame sequence in the GOP queue, and saving the image to the GOP queue; If the current frame image is not an I-frame and the number of frames of the last GOP video frame sequence in the GOP queue is less than the target number of frames, encoding the current frame image to obtain a current frame coded image, and saving the current frame coded image as a frame of the last GOP video frame sequence to the GOP queue; The duration corresponding to the target number of frames is smaller than the target duration.

5. The method according to any one of claims 1 to 4, wherein synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence comprises: The static photo is used as a cover and is synthesized with the at least one image group video frame sequence to form the dynamic photo.

6. The method according to any one of claims 1 to 4, wherein synthesizing a dynamic photo based on the static photo and the at least one image group video frame sequence comprises: displaying a plurality of candidate images, wherein the plurality of candidate images include the static photo, and the candidate images other than the static photo among the plurality of candidate images are pictures in the video frame sequence of the at least one image group; In response to a selection operation on a target candidate image from the plurality of candidate images, the target candidate image is used as a cover and synthesized with the at least one image group video frame sequence into the dynamic photo.

7. A dynamic photo shooting device, wherein the device comprises: a shooting module, configured to take a still photo in response to a received triggering operation on a shooting control; an acquisition module, configured to acquire at least one GIP video frame sequence, wherein the at least one GIP video frame sequence is obtained by encoding video frame images captured within a target time window in units of GIP, wherein the target time window includes the moment when the trigger operation is received; A synthesis module is used to synthesize a dynamic photo based on the static photo and the at least one image group video frame sequence.

8. An electronic device, wherein the device comprises: Memory and processor, the memory is used to store computer programs; The processor is used to execute the dynamic photo shooting method described in any one of claims 1 to 6 when calling the computer program.

9. A computer-readable storage medium, wherein a computer program is stored on the storage medium, and when the computer program is executed by a processor, the dynamic photo shooting method according to any one of claims 1 to 6 is implemented.

10. A computer program product, wherein a computer program is stored on the computer program product, and when the computer program is executed by a processor, the dynamic photo shooting method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Video encoding method and video decoding method

    CN110049336A

  • Method, device, and mobile platform for generating dynamic image and storage medium

    CN111034187A

  • Video processing method and device, electronic equipment and computer readable storage medium

    CN111464761A

  • Image processing method and device, storage medium and electronic equipment

    CN112565822A

  • Video processing method and device, equipment and storage medium

    CN115914498A

Cited By

  • Shooting method and device

    CN121619506A