Dynamic image generation method and device, computer equipment and program product
By capturing video frames and generating dynamic images based on user instructions after video capture is completed, the device lag problem caused by high processor usage is solved, and dynamic image generation with higher fluency and quality is achieved.
Patent Information
- Application Number
- CN202510990934.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-03
AI Technical Summary
In the process of generating dynamic images in the existing technology, the processor usage of the computer device is too high, resulting in problems such as device freezing and frame drops in captured images or videos.
After the video acquisition is completed, the video frames of the corresponding time period are intercepted based on the instructions input by the user to generate dynamic images. The peak-shifting processing method is used to avoid real-time generation and improve the smoothness of the device.
By generating dynamic images through staggered processing, the processor occupancy rate is reduced, the smoothness of computer equipment and the quality of dynamic images are improved, and the user experience is enhanced.
Smart Images

Figure CN120751209A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, apparatus, computer equipment, and program product for generating dynamic images. Background Art
[0002] Advances in hardware and image processing technologies have laid the foundation for the realization of dynamic images. Dynamic images combine static images with motion capture technology, providing users with a richer range of video recording options. Currently, computers must process dynamic images in real time after users capture them, which consumes a lot of processor space. This process can affect other computer functions, leading to issues such as lag, image or video frame drops, and other issues. Summary of the Invention
[0003] This application mainly provides a method, device, computer equipment and program product for generating dynamic images, which can avoid a sharp increase in processor occupancy and thereby improve the fluency of the computer equipment.
[0004] The technical solution of this application is achieved as follows:
[0005] In a first aspect, an embodiment of the present application provides a method for generating a dynamic image, the method comprising:
[0006] During the process of capturing the first video, obtaining at least one first instruction input by a user;
[0007] After the first video is captured, obtaining at least one second video in the first video based on at least one first instruction;
[0008] A target dynamic image related to at least one second video is generated.
[0009] In a second aspect, an embodiment of the present application provides a device for generating a dynamic image, the device comprising:
[0010] An acquisition module, configured to acquire at least one first instruction input by a user during the process of capturing the first video;
[0011] The acquisition module is further configured to acquire at least one second video in the first video based on at least one first instruction after the first video acquisition is completed;
[0012] The processing module is used to generate a target dynamic image related to at least one second video.
[0013] In a third aspect, an embodiment of the present application provides a computer device, including:
[0014] memory for storing computer programs;
[0015] a processor, connected to the memory, and configured to call and execute a computer program from the memory to implement the method of the first aspect;
[0016] A transceiver is used to send and receive information between devices.
[0017] In a fourth aspect, an embodiment of the present application provides a computer program product, comprising computer program instructions, which, when executed by a processor, implement the method of the first aspect.
[0018] The present application provides a method, apparatus, computer device, and program product for generating dynamic images. The method obtains at least one first instruction input by a user during the process of capturing a first video, and after the first video capture is completed, intercepts a second video corresponding to each first instruction in the first video, and finally generates a target dynamic image related to the at least one second video. In this way, through this staggered processing method, the target dynamic image is generated asynchronously after the first video capture is completed, avoiding the problem of a sharp increase in processor occupancy due to real-time dynamic image generation, improving the smoothness of the computer device during the process of generating the target dynamic image, and enhancing the quality of the dynamic image and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 A schematic flow chart of the steps of a method for generating a dynamic image provided in an embodiment of the present application;
[0020] Figure 2 A schematic diagram of a dynamic image acquisition interface provided in an embodiment of the present application;
[0021] Figure 3 A schematic diagram of a dynamic image display interface provided in an embodiment of the present application Figure 1 ;
[0022] Figure 4 A schematic diagram of a dynamic image display interface provided in an embodiment of the present application Figure 2 ;
[0023] Figure 5 A schematic diagram of a dynamic image display interface provided in an embodiment of the present application Figure 3 ;
[0024] Figure 6 A schematic diagram of the structure of a dynamic image generation device provided in an embodiment of the present application;
[0025] Figure 7 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0026] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below with reference to the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0028] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0029] It should also be pointed out that the terms "first\second\third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0030] In addition, references to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of such phrases in various places in the specification does not necessarily refer to the same embodiment, nor does it necessarily refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0031] The following is an introduction to the relevant technologies of this application.
[0032] With the diversification of multimedia content creation methods, users' demands for image media have also become diversified. Static images, dynamic images and videos are suitable for different scenarios due to their respective characteristics, and at the same time have given rise to the core demand for cross-media conversion.
[0033] For static images, users seek to capture decisive moments through stable composition and controlled lighting, fulfilling the demands of artistic creation, cultural symbolism, and efficient social media dissemination. However, static images lack dynamic detail and emotional atmosphere, making it difficult to reproduce the continuity of movement or ambient sound effects.
[0034] Although videos can record long-term events (such as travel records and sports events) and achieve personalized narratives through editing and special effects, the threshold for video production is relatively high and the storage cost is high. Users need to perform post-production editing to obtain highlight moments, and cannot directly extract highlight clips as video materials.
[0035] Motion graphics serve as a bridge between static images and videos, allowing users to capture high-definition dynamic clips of approximately three seconds and share them on social media platforms, while maintaining a compact file size and dynamic effects. However, the current process of generating motion graphics relies on real-time processing on computer devices, which consumes a high processor utilization rate. This process can affect other computer functions, causing issues such as computer lag and frame drops in captured images or videos.
[0036] Based on this, the embodiments of the present application provide a method, apparatus, computer device, and medium for generating dynamic images. These methods obtain at least one first instruction input by a user during the process of capturing a first video, and after the first video capture is complete, intercept a second video corresponding to each first instruction in the first video, and finally generate a target dynamic image associated with the at least one second video. Thus, through this staggered processing method, the target dynamic image is generated asynchronously after the first video capture is complete, avoiding the problem of a sharp increase in processor occupancy caused by real-time dynamic image generation. This improves the smoothness of the computer device during the generation of the target dynamic image, and enhances the quality of the dynamic image and user experience.
[0037] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0038] In one embodiment of the present application, Figure 1 The following is a flow chart of the steps of a method for generating a dynamic image provided in an embodiment of the present application. Figure 1 As shown, the method may include:
[0039] S101: During the process of capturing a first video, obtaining at least one first instruction input by a user.
[0040] Each first instruction is received at a corresponding click moment.
[0041] In an embodiment of the present application, during the recording or playback of a first video, a user may input a first instruction to a computer device through a touch operation (e.g., clicking a capture button on a display screen), a key operation, voice control, or automatic system triggering, etc. Each first instruction corresponds to a time point, which is the click moment when the user performs the touch operation.
[0042] like Figure 2As shown, the method is applied to a computer device for illustration. When a user uses the computer device to capture a first video, the computer device is in video recording mode. A prompt icon a1 on its display interface indicates that the current video is being captured. The user can click the video recording icon a2 in the middle to stop capturing the first video and obtain the first video. During the capture of the first video, when the user finds a highlight moment (such as a scoring moment or a smiling face), they can enter a first instruction by clicking the dynamic image capture icon a3. The moment when the user clicks the dynamic image capture image is the click moment of the first instruction.
[0043] In addition, when the user is capturing the first video, there may be more than one highlight moment. The user can click the dynamic image shooting icon a3 at each highlight moment to input the first instruction, and input the first instruction to the computer device at different times.
[0044] In an embodiment of the present application, the first video may include a captured video stream and a corresponding audio stream.
[0045] In some embodiments, the first instruction can be generated based on a neural network model, and a node where the first instruction can be input can be recommended to the user. For example, during a live broadcast, the user captures the first video by screen recording, and the moment when the product is displayed or the moment when the purchase link pops up can be recommended to the user as the click time for inputting the first instruction; alternatively, the click time of the first instruction can be marked by the user who captured the first video and displayed when the first video is shared with other users. For example, when a teacher records a teaching video, he or she marks the key paragraphs, and when students watch the first video, the teacher's marked recommended click time for the first instruction can be displayed.
[0046] S102: After the first video is captured, based on at least one first instruction, obtain at least one second video in the first video.
[0047] In this embodiment of the present application, each first instruction input by the user instructs the computer device to generate a corresponding second video based on image frames within a time period adjacent to the first instruction. After the first video is captured, the computer device caches the first video as a raw data stream and asynchronously intercepts the first video. Based on the peak processing mechanism, the computer device obtains the second video corresponding to each first instruction, and the first instruction and the second video have a one-to-one correspondence.
[0048] The first video includes a sequence of multiple first images, and the second video includes multiple first images in a time period adjacent to a click moment of a first instruction in the first video. For example, the click moment of the first instruction includes but is not limited to a start moment, a middle moment, or an end moment of the adjacent time period.
[0049] In an embodiment of the present application, the length of the adjacent time period can be pre-set or set by the user, and is specifically determined according to application requirements and system performance configuration, including a period of time before the click moment and / or a period of time after the click moment, and the lengths of the time periods before and after the click moment are not necessarily equal. For example, the length of the adjacent time period can be about 3 seconds, including about 1.5 seconds before the click moment and about 1.5 seconds after the click moment, and is specifically determined according to the actual application situation. For example, if the user enters the first instruction less than 1.5 seconds after starting to capture the first video, the user can be prompted that there is insufficient time, or the length after the click moment can be extended to meet the length of the adjacent time period.
[0050] In some embodiments, based on the Group of Pictures (GOP) structure in the first video, the capture interval of the second video is extended to the key frame boundary, and the key frame (I frame) closest to the length of the time period before the click moment is located as the first frame of the corresponding second video, and the I frame closest to the length of the time period after the click moment is located as the last frame of the second video. The image sequence between the first frame and the last frame is a plurality of first images in the adjacent time period corresponding to the click moment of the first instruction in the first video. For example, if the time length before and after the click moment is set to 1.5 seconds, there are two adjacent I frames 1.4 seconds before the click moment and 1.7 seconds before the click moment, respectively. Then, the I frame 1.4 seconds before the click moment is used as the first image frame in the second video, that is, the index is recorded starting from this image frame. Alternatively, if the time period corresponding to the I frame closest to 1.5 seconds before the click moment is shorter, for example, about 1 second before the click moment, then the capture range of the I frame after the click moment can be adaptively extended so that the time period between the first frame and the last frame in the second video is close to the length of the adjacent time period.
[0051] It should be noted that the captured first video is composed of a series of continuous first images, each of which represents a static image of the video at a certain moment. When these video frames are played at a certain frame rate, the human eye can see a dynamic effect.
[0052] In addition, encoding continuous video frames of the first video can generate key frames (I frames) and differential frames (P / B frames), wherein the key frames contain complete image information and can be independently encoded and decoded; the differential frames need to refer to other frames when encoding and decoding.
[0053] In addition, the preview frame can be extracted or generated from any video frame in the first video, retaining the main visual features of the corresponding video frame for fast preview or thumbnail display. The preview frame can be a low-quality version of the corresponding video frame.
[0054] In the embodiment of the present application, the multiple first images may be continuous video frames within an adjacent time period corresponding to the first instruction.
[0055] In an embodiment of the present application, the number of acquired second videos is less than or equal to the number of first instructions input by the user. After the first video is acquired, for any first instruction input during the acquisition of the first video, the computer device intercepts a plurality of first image sequences of the adjacent time period of the click moment corresponding to the first instruction as the second video corresponding to the first instruction. After the user inputs the first instruction, the computer device can determine whether the first instruction meets the conditions for acquiring the second video based on the time of the moment when the first instruction is input in the first video, and eliminate the first instructions that do not meet the conditions for acquiring the second video. For example, if the time interval between the input first instruction and the previous first instruction or the next first instruction is too small (e.g., less than 1 second), or the time interval between the moment when the first instruction is input and the start or end time of the first video is too small (e.g., less than 1 second), it can be determined that the first instruction does not meet the conditions for acquiring the second video.
[0056] It should be noted that in the process of capturing multiple first images in the first video, audio data synchronized with the multiple first images can also be obtained. After acquisition, the multiple first images are stored in YUV format in the flash memory or memory, and the audio data is stored in the flash memory or memory in pulse code modulation (PCM) format.
[0057] S103: Generate a target dynamic image related to at least one second video.
[0058] In the embodiment of the present application, the target dynamic image can be understood as the related video, image and related metadata of the dynamic image displayed to the user. The target dynamic image should include at least one second video.
[0059] In an embodiment of the present application, the target dynamic image can be generated based on one or more of the at least one second video, and the target dynamic image is related to the one or more second videos that generate the target dynamic image.
[0060] For example, when the number of the second video is one, the computer device may process the second video into one target dynamic image.
[0061] Alternatively, when there are multiple second videos, the computer device can merge some or all of the multiple second videos into one video, and further process them to obtain the corresponding target dynamic image, or can process the multiple second videos separately to obtain multiple target dynamic images.
[0062] An embodiment of the present application provides a method for generating a dynamic image, wherein a computer device obtains at least one first instruction input by a user during the process of capturing a first video, and after the first video capture is completed, intercepts a second video corresponding to each first instruction in the first video, and finally generates a target dynamic image related to the at least one second video. In this way, through this staggered processing method, the target dynamic image is generated asynchronously after the first video capture is completed, avoiding the problem of a sharp increase in processor occupancy caused by real-time dynamic image generation, improving the smoothness of the computer device during the process of generating the target dynamic image, and enhancing the quality of the dynamic image and user experience.
[0063] In some embodiments, for step S102, based on at least one first instruction, obtaining at least one second video in the first video may include:
[0064] S201: In response to each first instruction, determine a plurality of first image indexes in a time period adjacent to a click moment of the first instruction in a first video.
[0065] As mentioned above, the first video includes multiple first images, and the first image index corresponds one-to-one to the first image, which is used to identify the position of each first image in the first video. When the computer device captures the first video, it assigns a unique first image index to each first image based on the presentation time stamp (PTS) to facilitate the positioning and extraction of the first image.
[0066] In an embodiment of the present application, during the process of capturing the first video, for each first instruction, after the user inputs the first instruction, the computer device synchronously records the indexes corresponding to multiple first images in the time period adjacent to the click moment of the first instruction in the first video (that is, the time period before and after the click moment corresponding to the first instruction) into an independent file to accelerate the processing of subsequent steps. The independent file can be a file in JavaScript Object Notation (JSON) format.
[0067] After the first video is captured, the computer device obtains a plurality of first image indexes corresponding to each first instruction by reading the independent file.
[0068] S202: Determine, based on multiple first image indexes in the first video, multiple first images in the second video corresponding to the first instruction.
[0069] Among them, the multiple first images may include continuous first images within a time period adjacent to the click moment of the first instruction in the first video; or, in some embodiments, the multiple first images may include key frames (I frames) within a time period adjacent to the click moment corresponding to the first video, thereby reducing the number of saved first image indexes.
[0070] In an embodiment of the present application, after the first video is captured, the computer device starts a background thread based on multiple first image indexes read from an independent file, and processes the first video and the corresponding audio based on preset components, such as multimedia extractors (MediaExtractor or MediaMuxer), thereby extracting multiple first images marked by these first image indexes from the first video.
[0071] It should be noted that the multiple first images corresponding to any first instruction obtained by the computer device are usually continuous frames in the first video, which are used to form a second video corresponding to the first instruction, and are used to show the highlight moments of the user before and after entering the first instruction.
[0072] It should also be noted that the computer device can sort and verify the multiple first image indexes corresponding to each obtained first instruction, so as to ensure the continuity and fluency of the dynamic image.
[0073] In an embodiment of the present application, for any first instruction, after obtaining multiple first images corresponding to the first instruction in the first video, the computer device generates a second video corresponding to the first instruction based on the multiple first images.
[0074] An embodiment of the present application provides a method for generating dynamic images. During the capture process of a first video, a method generates a first image index corresponding to each first instruction input by the user in the first video, so that when the corresponding second video is generated after the capture of the first video is completed, it can directly jump to the target position, avoiding the entire scan of the first video, reducing the time consumption of generating the second video, and improving the generation efficiency of the second video.
[0075] In some embodiments, when the first video is encoded using a variable frame rate, before generating a target dynamic image related to at least one second video in step S103, the method may further include:
[0076] S301: Determine motion information corresponding to at least one key frame in a second video.
[0077] It should be noted that variable frame rate encoding can be understood as dynamically adjusting the frame rate according to the video content. The computer device can use a decision-making mechanism, such as reducing the frame rate according to temperature control or adaptively adjusting the frame rate based on the complexity of the scene, to capture the first video at different frame rates in different time periods. This may result in a different number of images between adjacent key frames in the second video captured from the first video.
[0078] In an embodiment of the present application, the second video may include a sequence of multiple image frames, among which the frames compressed using intra-frame coding are key frames. Key frames can also be called I frames, which can be understood as frames in the second video that independently represent the content of a certain picture, contain complete image information, have no motion compensation dependence, and do not rely on other frames for encoding and decoding.
[0079] In the embodiments of the present application, the motion information corresponding to a key frame can be understood as the changes in the image content of the key frame relative to the previous adjacent key frame (or the next adjacent key frame), such as parameters such as the movement direction, movement speed, and displacement of objects in the image. The motion information can be extracted based on optical flow or block matching algorithms.
[0080] It should be noted that the motion information corresponding to the at least one key frame in the second video may be generated based on the at least one key frame in the plurality of first images after the user inputs the first instruction during the capture of the first video, along with the first image index, and stored in a separate file in JSON format, for the computer device to read and obtain from the separate file after completing the capture of the first video. Alternatively, in some embodiments, the motion information corresponding to the at least one key frame may be generated based on the at least one key frame in the second video corresponding to the first instruction after the capture of the first video is completed.
[0081] It should also be noted that for the first key frame in the second video, the corresponding motion data can be the changes in the picture content between the key frame and its adjacent previous key frame in the first video, or the first key frame in the second video can be regarded as a static scene and a reference frame for the adjacent subsequent key frame.
[0082] S302: Determine the number of images between each key frame and an adjacent key frame based on motion information corresponding to at least one key frame.
[0083] In an embodiment of the present application, a motion intensity grading mechanism is employed. Based on the motion information corresponding to each key frame, the degree of dynamic change of each key frame relative to the previous adjacent key frame is evaluated. Combined with a dynamic interpolation strategy, the number of images between the key frame and the previous adjacent key frame is determined. For example, for scenes with intense motion, a higher image density is required between adjacent key frames, and the number of images between the current adjacent key frames is determined based on the density to ensure smooth dynamic images. For scenes with stillness or slow changes, the image density between adjacent key frames is lower, and the number of images between the current adjacent key frames is determined based on the density to reduce storage space.
[0084] It should be noted that, in some embodiments, the number of images between each key frame and the next adjacent key frame may be determined, and its implementation method refers to the step of determining the number of images between each key frame and the previous adjacent image frame.
[0085] S303: Based on the number of images between each key frame and an adjacent key frame and a first preset algorithm, generate at least one second image between each key frame and an adjacent key frame to obtain an expanded second video.
[0086] In the embodiment of the present application, the first preset algorithm can be understood as an algorithm used to fill in frames between adjacent key frames, generate visually natural and continuous intermediate frames, and make the second video look smoother, such as optical flow method, interpolation algorithm, neural network algorithm based on deep learning, etc., which is not specifically limited here.
[0087] In some embodiments, to address the problem of blurred boundaries in video capture, more parameters and methods can be introduced, such as artificial intelligence (AI) technology. Based on a multi-scale generation strategy, the motion vector data between adjacent key frames can be predicted through the optical flow method to generate smooth intermediate image frames and improve the clarity of the picture.
[0088] In an embodiment of the present application, if the actual number of images between a key frame and a previous adjacent key frame in the second video does not reach the number of images determined in step S302, the computer device may generate one or more second images between the key frame and the previous adjacent key frame as intermediate frames based on a first preset algorithm, so that the actual number of images between the key frame and the previous adjacent key frame reaches the number of images determined in step S302, thereby obtaining an expanded second video. The second images may be differential frames.
[0089] Related technologies that directly generate dynamic images from static images based on neural network algorithms or text instructions not only have low control accuracy but also fail to preserve the details of the original image. In the embodiments of the present application, a second video corresponding to the first instruction is obtained by intercepting a first video recorded by the user, and then adding frames to the second video to generate a dynamic image, which has higher fluency and more realistic movements and details.
[0090] It should be noted that since the expanded second video extends the intermediate frames, the time length of the second video can slightly exceed the time length of the adjacent time period. However, based on a large number of user surveys, users are less sensitive to the timeout of a few tenths of a second caused by frame interpolation. Therefore, frame interpolation has little impact on the user experience, but the displayed dynamic image is smoother.
[0091] In some embodiments, a lightweight motion prediction model can be deployed on a computer device to a neural network processor (NPU) to achieve local processing of key frame compensation and interpolation calculations, reduce dependence on cloud computing, and reduce latency.
[0092] S304: Perform audio and video synchronization processing on the expanded second video to obtain a synchronized second video.
[0093] In an embodiment of the present application, since at least one additional second image is introduced through expansion in the aforementioned steps, in order to ensure that the time axes of the audio and video are consistent so that the sound and picture can be accurately matched during playback, the computer device can dynamically remap the timing based on the PTS, parse the timestamp of the original audio corresponding to the second video that has been frame-filled and expanded, recalculate the playback rhythm of the audio based on the expanded second video, adjust the timestamp of the original audio accordingly, and resample, crop, and perform other operations on the audio when necessary to ensure that the expanded second video is synchronized with the audio, thereby obtaining a synchronized second video.
[0094] The present application provides a method for generating dynamic images. Based on the motion information corresponding to at least one key frame in a second video, the method determines the number of images between each key frame and the adjacent key frame. If the number of images is not reached, at least one key image is generated to fill in the gaps, thereby obtaining an expanded second video. The expanded second video is then synchronized with the audio and video to obtain a synchronized second video. This method not only improves the smoothness of the dynamic image and alleviates the timing discontinuity caused by frame rate changes, but also avoids the problem of audio and video asynchrony caused by expansion, thereby improving the user's viewing experience.
[0095] In some embodiments, for the aforementioned step S103, generating a target dynamic image related to at least one second video may include:
[0096] S401: Perform cover selection processing on at least one second video to determine a cover image.
[0097] In an embodiment of the present application, the cover selection process can be understood as obtaining a cover image based on a preset algorithm from multiple frames of images contained in at least one second video, and using it as a static cover for the dynamic image to display or preview the content of the dynamic image to the user in an album or when there is no need to play dynamic effects.
[0098] It should be noted that, considering factors such as encoding characteristics, clarity requirements, and implementation complexity, different implementation methods can be selected for the cover selection process.
[0099] It should also be noted that when the number of second videos contained in at least one second video is one, cover selection processing can be performed based on the second video to determine the cover image; when the number of second videos contained in at least one second video is multiple, the cover image can be determined based on a randomly selected second video, or a second video specified by the user, or a second video determined by other means.
[0100] S402: Determine a target dynamic image based on the cover image and at least one second video.
[0101] In an embodiment of the present application, the target dynamic image may include a cover image and at least one second video. The cover image may be a frame determined based on the at least one second video and may be a static cover frame displayed as a dynamic image in the album.
[0102] In the embodiments of this application, Figure 3 As shown, the target dynamic image is displayed in the album as a cover image with a dynamic image mark b1 (the "Live" mark in the lower left corner of the image). As a static preview, the user can set it to play at least one second video by long pressing the image when the user needs to view the image, or click on the dynamic image mark b1 to play at least one second video, or play at least one second video when clicking on the cover image in the album to enlarge and view it, to show the dynamic effect.
[0103] In an embodiment of the present application, the target dynamic image (including the cover image and at least one second video) can be encapsulated into a High Efficiency Image File Format (HEIF) format and stored in a computer device, so that it can be compatible and adapted to more platforms and systems, and support different display forms of static preview and dynamic playback on different platforms. Moreover, under the storage format in the related art, reselecting the cover frame requires regenerating the dynamic image, which poses the risk of image quality degradation and screen tearing. In the block storage method in the embodiment of the present application, the user only needs to regenerate the cover image when reselecting the cover frame, which has no effect on the video content of at least one second video, thereby improving the stability of the dynamic image.
[0104] In an embodiment of the present application, the target dynamic image may also include metadata, which may include information such as color gamut and chromaticity indicating suitability for different platforms, and may also include information such as a cover image and the storage location of at least one second video. Based on this universal packaging standard, it can be adapted to more platforms, and be compatible with dynamic playback, static preview, and playback on social platforms, thereby solving the current problem of dynamic image display being restricted by platforms.
[0105] The present application provides a method for generating a dynamic image. This method performs cover selection based on at least one second video to determine a cover image. Furthermore, a target dynamic image is determined based on the cover image and the at least one second video. In this manner, the cover image is determined based on the at least one second video and stored separately from the at least one second video. User reselection of the cover image does not affect the content of the second video, thereby improving the stability and robustness of the target dynamic image.
[0106] In some embodiments, step S401, performing cover selection processing on at least one second video to determine a cover image may include:
[0107] S501: Acquire one or more third images in at least one second video.
[0108] In an embodiment of the present application, the third image can be one or more key frames in any second video that are closest to the click moment of the first instruction corresponding to the second video, containing complete image information. Directly extracting the key frames as cover frames does not require additional decoding operations, is suitable for scenarios with high real-time requirements, and the picture loss is controllable.
[0109] Alternatively, in the embodiment of the present application, the third image can be one or more preview frames in any second video that are closest to or corresponding to the click time of the first instruction corresponding to the second video. The preview frames can be directly intercepted from the second video decoding pipeline, and the implementation is simple, no additional hardware support is required, and the cover is highly consistent with the dynamic image content. Figure 4 As shown, the user can drag or slide the preview frame display component c1 to select a preview frame (e.g. Figure 4 The image frame marked with a circle and displayed enlarged is the preview frame).
[0110] S502: Perform a first process on one or more third images to obtain a cover image.
[0111] The first processing includes at least one of the following: super-resolution processing, sharpening processing, and noise reduction processing.
[0112] It should be noted that some key frames extracted from the second video with a low bit rate may have blurred details due to compression, and the resolution of the preview frame is also low. Therefore, it is necessary to perform super-resolution processing on the third image to improve the resolution of the third image.
[0113] Among them, super-resolution processing is understood as an image enhancement algorithm, which can convert low-resolution images into high-resolution images. For example, super-resolution generative adversarial networks (SRGAN) and MobileNet optimization models can be used to implement detail supplements to obtain a third image after super-resolution, thereby improving the clarity of the third image and avoiding jagged or blurred images, making the third image more suitable for cover display.
[0114] In an embodiment of the present application, optionally, the sharpening intensity and noise reduction parameters of the super-resolved third image can be further adjusted based on the structural similarity (SSIM) index requirement to obtain a cover image stored in a computer device.
[0115] Sharpening is used to enhance the edges and detail contrast of the third image, making the picture clearer. Noise reduction is used to remove noise from the image and improve the overall image quality. Sharpening and noise reduction can improve the clarity and naturalness of the third image.
[0116] It should be noted that when the number of third images is one, the first processing can be performed on the third image based on one or more of the above-mentioned super-resolution processing, sharpening processing, and noise reduction processing to obtain a cover image; or, when the number of third images is multiple, the first processing can also include fusion processing, by fusing multiple third images and performing one or more of the above-mentioned super-resolution processing, sharpening processing, and noise reduction processing on the fused third image to obtain a cover image.
[0117] The present application provides a method for generating a dynamic image by extracting the key frame or preview frame closest to the click moment of a first instruction corresponding to any second video, and then performing one or more of the following processing steps: super-resolution, sharpening, noise reduction, and fusion, to obtain a cover image. This makes the cover image displayed as the cover frame clearer, improves the visual quality of the cover image, and thus enhances the user experience.
[0118] In some embodiments, the method may further include: obtaining a plurality of fourth images in the photographic data stream during the process of capturing the first video.
[0119] In an embodiment of the present application, for a computer device that supports dual image signal processing (ISP) functions, when capturing a first video, an independent photo pipeline can be started simultaneously. When capturing the first video, a fourth image is captured synchronously at a preset interval. After the first video capture is completed, a photo data stream corresponding to the video data stream of the first video can be obtained, and the photo data stream includes multiple fourth images.
[0120] Among them, the fourth image is independently taken by the computer device in the photo mode when capturing the first video, and its image quality is significantly higher than the key frame, differential frame, preview frame or video frame in the second video.
[0121] In some embodiments, step S401, performing cover selection processing on at least one second video to determine a cover image may include:
[0122] S601: Acquire, from the plurality of fourth images, a plurality of fifth images closest to the first click moment.
[0123] The first click moment is one of the click moments of at least one first instruction.
[0124] In an embodiment of the present application, the first click moment can be understood as the click moment corresponding to the first instruction corresponding to the second video specified for determining the cover image in at least one second video. The second video for determining the cover image in at least one second video can be randomly determined based on a computer device, or the second video for determining the cover image can be determined based on user specification. It can be determined according to a specific strategy and can be any one of the click moments of at least one first instruction.
[0125] It should be noted that when capturing multiple fourth images, a corresponding timestamp is generated for each fourth image. When performing cover selection, based on the timestamps corresponding to the fourth image sequence corresponding to the first video, the multiple fourth images closest to the first click moment are selected as the multiple fifth images. The number of fifth images selected can be specified by the user or preset.
[0126] S602: Perform second processing on the plurality of fifth images to obtain a cover image.
[0127] The second processing at least includes fusion processing.
[0128] In an embodiment of the present application, when multiple fifth images are selected in the aforementioned step, a second processing, such as a fusion processing, can be performed on the multiple fifth images to generate a cover image based on the multiple fifth images.
[0129] Among them, the fusion processing can include feature point detection and matching, image alignment, fusion, etc. to generate a cover image. Compared with the fifth image, the cover image can effectively reduce motion blur, enhance detail expression, improve clarity, and be more representative and stable.
[0130] In an embodiment of the present application, the fifth image that has undergone the second processing can be further subjected to the aforementioned first processing, such as one or more of super-resolution processing, sharpening processing, and noise reduction processing, to obtain a cover image.
[0131] It should be noted that which of the above methods is used to obtain the cover image can be determined based on the performance of the computer device. When the performance of the computer device is high, steps S601 to S602 can be used to determine the cover image. When the performance of the computer device is average, steps S501 to S502 can be used to determine the cover image.
[0132] This embodiment of the present application provides a method for generating a dynamic image by obtaining multiple fifth images before and after a first click moment in a fourth image sequence corresponding to a first video and fusing the multiple fifth images to obtain a cover image. This method can produce a clearer and higher-quality cover image, enhancing the user's viewing experience.
[0133] In some embodiments, the cover image obtained as described above can be directly used as a static cover for a dynamic image, or, based on a neural network model, such as a generative adversarial network (GAN) model, dynamic elements such as flowing clouds, swaying leaves, etc. can be added to the static cover image to display a seamless animation effect that is integrated with the original scene of the cover image.
[0134] Alternatively, in some embodiments, combined with augmented reality (AR) technology, a user can scan the cover image through another computer device to trigger superimposed dynamic special effects, such as virtual fireworks, three-dimensional (3D) character actions, etc., to achieve an immersive experience that blends virtual and real.
[0135] In some embodiments, for the aforementioned step S402, determining the target dynamic image based on the cover image and the at least one second video may include:
[0136] S701: Determine timestamps of multiple second videos based on click moments of multiple first instructions.
[0137] In an embodiment of the present application, the timestamp of the second video indicates the time point information of the first instruction input by the user to capture the second video, which can be expressed in a specific time unit, such as seconds, milliseconds, microseconds, etc.
[0138] As mentioned above, during the process of recording the first video, each first instruction is used to instruct the computer device to obtain the second video and the corresponding cover image in the adjacent time period of the click moment of the first instruction in the first video, and generate the corresponding target dynamic image.
[0139] In the case where the user inputs multiple first instructions, the timestamp of each first instruction is used to indicate the position of the target dynamic image corresponding to the first instruction on the time axis of the first video.
[0140] S702: Based on the timestamps corresponding to the multiple second videos, the multiple second videos are merged to obtain a third video.
[0141] In an embodiment of the present application, if a user enters multiple first commands while capturing a first video, multiple corresponding second videos are generated upon completion of the capture. The multiple second videos can be sorted based on their position on the timeline of the first video, as indicated by the timestamp of the click of the first command corresponding to each second video. Through timeline index management, the multiple second videos can be spliced together in the sorted chronological order to generate a coherent third video. The merging process may include frame rate matching, resolution unification, and audio and video synchronization.
[0142] In some embodiments, for the scenario where multiple second videos are merged, batch production technology can be used to asynchronously render high-resolution dynamic collections through distributed clusters to solve the computing power bottleneck of computer equipment.
[0143] It should be noted that the time length of the merged third video is the sum of the time lengths of the multiple second videos, and the content of the third video is the multiple second videos played continuously in chronological order.
[0144] Alternatively, in some embodiments, multiple second videos may be merged according to a preset logic or based on a user operation to obtain a third video.
[0145] In some embodiments, a user historical data recommendation model can also be trained to recommend highlight clips in the first video to the user, and further merge multiple recommended highlight clips to produce a third video. For example, when the content of the first video is a sports event, multiple scoring moments are generated as a third video, thereby reducing the cost of manual screening.
[0146] S703: Determine a target dynamic image using the cover image and the third video.
[0147] In an embodiment of the present application, the target dynamic image may include the aforementioned determined cover image and the third video, wherein the cover image is determined based on any second video or the corresponding fifth image sequence in the third video, serving as a static display cover of the target dynamic image in the album.
[0148] It should be noted that if the user has not set at least one second video to be merged into a third video, then for each second video, a corresponding cover image can be determined based on the above steps, and a target dynamic image can be determined based on each second video and the corresponding cover image. In this way, after the first video is captured, multiple target dynamic images are generated, and the number of target dynamic images corresponds to the number of the first instructions.
[0149] This embodiment of the present application provides a method for generating a dynamic image. When a user enters multiple first instructions while capturing a first video, multiple second videos acquired based on the multiple first instructions are merged to obtain a third video. The cover image and the third video are then used to determine the target dynamic image. This allows for merging multiple second videos within a first video, enabling continuous playback of multiple highlight clips, redefining the duration of the dynamic image, and improving the user experience.
[0150] In some embodiments, the method may further include:
[0151] S801: Determine at least one key frame and at least one differential frame of at least one second video in a target dynamic image.
[0152] Among them, the key frame can be understood as an I frame in the second video or the third video, which is used to ensure the integrity of the picture; the differential frame can be understood as a P frame or B frame in the second video or the third video, which is used to reduce the storage space occupied.
[0153] It should be noted that there is a dependency between key frames and differential frames, meaning that differential frames must rely on the preceding key frames to be correctly decoded. Therefore, in the application embodiment, the GOP structure of a second or third video is first parsed, the locations of all key frames are identified, and multiple key frames are selected as references based on their timestamp distribution. Multiple differential frames following the key frames are then extracted.
[0154] S802: Compress at least one key frame based on a first compression algorithm to obtain compressed image data of the at least one key frame.
[0155] In the embodiment of the present application, the first compression algorithm can be understood as a lossless compression algorithm, which can preserve the details of the key frames as much as possible. The first compression algorithm usually adopts a higher bit rate to ensure that the compressed key frames can still maintain clarity when enlarged or thumbnailed.
[0156] In an embodiment of the present application, at least one key frame is processed based on the first compression algorithm to obtain compressed image data of at least one key frame. Since the data in the key frame is more important, this processing method can minimize the impact of compression on the quality of the target dynamic image.
[0157] S803: Compress at least one differential frame based on a second compression algorithm to obtain compressed image data of the at least one differential frame.
[0158] The compression loss of the first compression algorithm is smaller than the compression loss of the second compression algorithm.
[0159] In the embodiment of the present application, the second compression algorithm may be a lossy compression method with a high compression ratio but allowing a certain loss of image quality, such as the H.264 compression algorithm or the Zstandard compression algorithm.
[0160] It's important to note that since differential frames only contain the difference between them and the keyframes, even if significant data loss occurs during compression, the overall viewing experience won't be significantly affected. By using a second compression algorithm with greater compression loss on differential frames, the storage space required for them can be significantly reduced, resulting in a storage volume reduction of over 40% compared to storing at least one second video in MP4 format.
[0161] In some embodiments, the compressed image data of at least one key frame can be stored in a computer device, and the compressed image data of at least one differential frame can be stored in the cloud. When the user views the target dynamic image, it can be loaded on demand to reduce the storage space occupied by the computer device.
[0162] S804: Encapsulate the compressed image data of at least one key frame of the target dynamic image, the compressed image data of at least one differential frame, and the cover image to obtain storage data of the target dynamic image, and store or transmit the data.
[0163] In an embodiment of the present application, the encapsulation process can be understood as packaging and storing the key frames and differential frames in at least one second video, and the cover image in a preset format, respectively, and storing them in a computer device in the form of storage data of the target dynamic image, or transmitting them to other computer devices.
[0164] In addition, the stored data of the target dynamic image can be unpacked when the user views it and displayed to the user in the form of a static cover image or a dynamic image.
[0165] An embodiment of the present application provides a method for generating a dynamic image. The target dynamic image includes a cover image and at least one second video. Different compression algorithms are used to compress key frames and differential frames in the at least one second video. This method can reduce the storage space of the target dynamic image while ensuring the visual quality of the target dynamic image. Furthermore, by standardizing packaging standards, the method is compatible with different platforms and systems, thereby improving the compatibility and stability of the target dynamic image.
[0166] In some embodiments, for step S101, after obtaining at least one first instruction input by the user, the method further includes:
[0167] In response to each first instruction, a sixth image corresponding to the click moment is generated.
[0168] In an embodiment of the present application, the sixth image may be a preview frame at the click moment corresponding to the first instruction, which is a static image frame at the click moment extracted based on the video stream YUV data of at least one second video.
[0169] While the user is capturing the first video, the first instruction is input through a click or other operation. Before the capture of the first video is completed, the sixth image is displayed in the album as a preview to prevent the user from perceiving a "blank" or "loading" state.
[0170] It should be noted that, for each first instruction, a corresponding sixth image is generated and displayed in the album for user preview.
[0171] In some embodiments, the method further includes: in response to the first operation on the sixth image, executing the step of acquiring at least one second video in the first video based on at least one first instruction.
[0172] In the embodiment of the present application, the first operation may include an interactive behavior performed by the user to view the sixth image, such as clicking, long pressing, sliding, and the like.
[0173] After the first video is captured and before the target dynamic image is generated, if the user performs a first operation, the computer device can respond in a relatively short time, execute the aforementioned first instruction corresponding to the sixth image, intercept the second video corresponding to the first video, further encapsulate the second video, obtain the target dynamic image, and replace the sixth image with the cover frame of the generated target dynamic image to display it to the user, and based on the user's operation, display the dynamic image of the target dynamic image to the user.
[0174] The present embodiment provides a method for generating a dynamic image. After a user inputs a first instruction, a sixth image corresponding to the first instruction is generated and displayed in an album. After the user performs a first operation on the sixth image, the step of obtaining a target dynamic image is performed. This method can reduce user response delays, optimize the user experience, and improve user satisfaction.
[0175] In some embodiments, the method further includes: executing, based on load status information of the computer device, a step of acquiring at least one second video in the first video based on at least one first instruction within a first preset time period.
[0176] In the embodiment of the present application, the load status information of the computer device can be understood as the current system resources of the computer device, such as the usage of the central processing unit (CPU), memory, graphics processing unit (GPU), etc.
[0177] Among them, the first preset duration can be understood as the pre-set maximum time window for generating the target dynamic image. The target dynamic image needs to be generated within the first preset duration. The first preset duration is specifically configured according to the device type, application scenario and user needs.
[0178] Within a first preset time period after the first video is captured, if the user does not perform the first operation, when the load status information of the computer device indicates that the current system resource usage is lower than a preset threshold, the steps of acquiring at least one second video and obtaining the target dynamic image can be performed to obtain the target dynamic image.
[0179] Alternatively, within a first preset time period after the first video capture is completed, the user does not perform the first operation, but the load status information of the computer device indicates that the utilization rate of system resources has been higher than the preset threshold. The computer device can appropriately increase the priority of the task of processing to obtain the target dynamic image, so that the target dynamic image is generated within the first preset time period.
[0180] It should be noted that, if the user performs a first operation after the first video capture is completed, the target dynamic image is generated in response to the first operation without considering the load status information of the computer device.
[0181] Thus, in this embodiment of the present application, the step of obtaining the target dynamic image is executed within a first preset duration based on the load status information of the computer device. Thus, the thread for generating the target dynamic image is started when the load status of the computer device is appropriate, thereby avoiding a sudden increase or overload of system resources. This ensures the quality of the generated target dynamic image while reducing the impact on system operating efficiency.
[0182] In some embodiments, after obtaining the target dynamic image, background music such as cheers, ambient sounds, etc. can be added to the target dynamic image based on the user's settings, according to the multimodal fusion logic, and based on the image content. Alternatively, the user can also add an audio track, such as voice narration, to improve the editability and richness of the target dynamic image.
[0183] In some embodiments, based on the user's settings, the image content and motion information in the target dynamic image can be mapped into tactile signals. For example, when the image content is the jumping moment of an athlete, a short vibration at the moment of jumping can be generated, thereby enhancing the user's multi-sensory interaction experience.
[0184] The following describes in detail the method for generating dynamic images provided in the embodiments of the present application in conjunction with specific application scenarios.
[0185] Currently, users' demands for imaging media are diversified. Still images, dynamic images, and videos, due to their respective characteristics, are suitable for different scenarios, giving rise to the core demand for cross-media conversion. The following discusses the two dimensions of the differentiation of shooting demands and the driving force of conversion demand:
[0186] 1. Differentiation of shooting scenes: precise freezing and artistic expression.
[0187] For static images, the focus is on precise freezing and artistic expression. Users capture decisive moments (such as landscape paintings and architectural landscapes) through stable composition and controlled lighting and shadow, meeting the demands of artistic creation, cultural symbolism, and efficient social media dissemination. However, static images lack dynamic detail and emotional atmosphere, making it difficult to reproduce the continuity of action or ambient sound effects.
[0188] Videos allow users to record long-term events, such as travel vlogs and sporting events, and further personalize their narratives through editing and special effects. However, video production is challenging, requiring careful consideration of camera movement and start and end timing. Unlike photos, which offer the flexibility of switching focal lengths and perspectives, video storage costs are high, and it's difficult to directly extract highlights for quick sharing.
[0189] Motion graphics capture dynamic moments and resonate with emotions. Users can quickly share 1.5-second clips of highlights (such as children's smiles or pets' movements) on social media platforms, balancing small file size (40% smaller than videos) with the realism of "straight-out original images." However, motion graphics cannot capture long-lasting highlights and require capture in shooting mode, which can lead to missed highlights. For already-shot motion graphics, editing functionality is limited, supporting only cover reselection.
[0190] 2. Cross-media conversion needs: technical bottlenecks and user experience gaps.
[0191] For scenarios where dynamic images need to be converted to still images or videos, users need to convert them to static images for printing and document transfer, such as ID photos for resumes, or convert them to videos for playback on devices with different operating systems. Currently, this conversion relies on cloud-based processing, which poses a risk of privacy leakage. Locally reselecting and regenerating the dynamic image and reselecting the cover frame can lead to image quality degradation, which can separate the image from the dynamic image and reduce both image quality and clarity.
[0192] When converting static images into dynamic images, users currently want to inject dynamic elements into static images to generate emojis or creative short films. However, the motion control accuracy of related solutions is low and it is difficult to maintain the details of the original image.
[0193] For scenarios where videos are converted into dynamic images, users need to extract highlight clips from long videos and generate dynamic images for sharing on social platforms, but related technologies currently do not provide a complete solution.
[0194] Therefore, computer equipment needs to have the ability to convert between static images, dynamic images and videos. Although dynamic images are better expression carriers than static images, the image quality, storage format and editability of dynamic images themselves need to continue to evolve and be able to have more expression forms.
[0195] Based on the above-mentioned problems to be solved, an embodiment of the present application provides a method for generating dynamic images, proposes an asynchronous staggered processing architecture and multimodal seamless encapsulation technology, and reconstructs the generation paradigm of dynamic images.
[0196] In the related art, the camera can be enabled in photo mode to synchronously record the video and audio for about 1.5 seconds before and after the shutter is pressed, generating a 3-second dynamic image. However, this solution is a real-time process after the photo is taken, which has a high CPU or GPU occupancy rate and relies on the I-frame alignment at the moment of shooting. Misalignment may cause screen tearing due to the encoding GOP structure. The method for generating dynamic images provided in the embodiment of the present application adopts asynchronous processing. After the first video acquisition is completed, the dynamic image is generated, which reduces the real-time computing pressure, and through the key frame expansion compensation mechanism, such as expanding forward to the nearest I frame, the problem of screen tearing is solved and the picture quality of the dynamic image is improved.
[0197] In the related art, computer devices call dynamic images in the album through the system interface, and support publishing to social platforms and cross-system viewing. However, the current packaging format of dynamic images is not universal and cannot be adapted to computer devices installed with different systems. Due to the sharing bandwidth limitation and bit rate reduction measures, the problem of inconsistency between preview frames and final effects has not yet been solved. The method for generating dynamic images provided in the embodiment of the present application can reduce the time consumed in the later stage of unpacking and improve the picture quality by pre-generating metadata (recording timestamps and key frame indexes), and the packaging format is universal and can be adapted to different systems and platforms.
[0198] In the embodiments of the present application, by asynchronously generating the target dynamic image during the process of capturing the first video, and combining methods such as dynamic keyframe compensation, variable frame rate adaptation, and block storage optimization, the problems of high load, strong keyframe dependence, and low storage efficiency in real-time dynamic image generation in related technologies are solved. The following is a detailed description of the dynamic image generation method provided in the embodiments of the present application:
[0199] 1. Technical architecture.
[0200] The dynamic image generation method of the embodiment of the present application includes the following four core modules, which work together to achieve dynamic capture and efficient generation of highlight moments:
[0201] (1) Timestamp recording and asynchronous processing module.
[0202] In an embodiment of the present application, when a user enters a first command (e.g., clicking "photograph" while capturing a first video), the timestamp corresponding to the first command and the indexes of the first images in the first video are recorded. After the first video is captured, the second video (in YUV format) and audio (in PCM format) from the adjacent time period before and after the click of the first command (e.g., 1.5 seconds before and 1.5 seconds after the click) are cached in memory or flash memory for generating the target dynamic image.
[0203] In this way, a staggered processing mechanism is adopted to asynchronously unpack the first video after the first video recording is completed, thereby avoiding the performance bottleneck caused by real-time encoding.
[0204] (2) Key frame compensation and audio and video synchronization module.
[0205] In an embodiment of the present application, by parsing the GOP structure, the I frame closest to the click moment is located, the intercepted interval is extended forward and backward to the key frame boundary, and the audio and video timing is remapped through timestamps.
[0206] In this way, the problems of screen tearing and audio-visual asynchrony caused by non-I frame capture are solved by optical flow frame interpolation and audio decoding.
[0207] (3) Variable frame rate adaptation module.
[0208] In an embodiment of the present application, the presentation timestamp of each frame is dynamically mapped, and the interpolation strategy is adaptively adjusted according to temperature control frame reduction or scene complexity, such as SSIM motion intensity estimation.
[0209] In this way, through the motion intensity classification mechanism (with reference to motion information), intermediate frames are generated through the interpolation algorithm to alleviate the timing discontinuity caused by the sudden drop in frame rate.
[0210] (4) Block storage and HEIF encapsulation module.
[0211] In an embodiment of the present application, the target dynamic image includes a cover image (static cover frame) and at least one second video, wherein the at least one second video is split into key frames (I frames) and differential frames (P / B frames), and compressed and stored separately using different compression algorithms.
[0212] In this way, the storage of repeated frames is optimized through the Zstandard lossless compression algorithm, and the volume is reduced by more than 40% compared to the MP4 storage solution in related technologies.
[0213] 2. Implementation process.
[0214] In the embodiment of this application, the implementation process is divided into three stages:
[0215] (1) Recording phase. During the capture of the first video, the user enters the first command to trigger the timestamp recording module, which caches the raw data 1.5 seconds before and after the current moment. It also synchronously saves metadata such as keyframe indexes, motion vectors (motion information), and timestamps to a separate file (JSON format) to accelerate post-processing.
[0216] (2) Asynchronous unpacking phase: After the first video is captured, the user starts a background thread, parses the video stream through the MediaExtractor component, and locates the I-frame position corresponding to the click moment.
[0217] In the case where the first video is VFR, the timing is dynamically remapped based on PTS, and missing frames are generated through an interpolation algorithm.
[0218] (3) Generation and storage phase: The second video (3 seconds) corresponding to the click moment of the first instruction in the first video is intercepted, and a cover image is generated and packaged into a HEIF format file to obtain the target dynamic image.
[0219] It should be noted that if the user clicks to take a photo multiple times during the capture of the first video, multiple second videos will be merged according to the timeline index, and a dynamic image containing continuous highlight moments will be generated, which is the third video.
[0220] 3. Solutions to key technical problems.
[0221] (1) Key frame boundary alignment problem.
[0222] In an embodiment of the present application, by generating indexes corresponding to multiple first images in the first video when a first instruction is input during the process of capturing the first video, it is possible to directly jump to the target area where the first image is located during unpacking, avoiding scanning the entire first video. Compared with the full-frame parsing solution in the related art, the generation time is reduced by 50%.
[0223] (2) Variable frame rate adaptation problem.
[0224] In this embodiment of the application, a combination of motion intensity binning and interpolation algorithms is used to dynamically generate intermediate frames and remap timestamps. This can improve the smoothness of dynamic images by 30% in temperature-controlled frame reduction scenarios.
[0225] (3) Storage space optimization issues.
[0226] In the embodiment of this application, HEIF format is used for packaging, differential frames and key frames are stored in blocks, and repeated frames are compressed using the Zstandard algorithm. This storage and packaging method can reduce the size of a target dynamic image of about 3 seconds from 15MB (stored in H.264 format) to 8MB.
[0227] 4. Selection logic for capturing a cover frame (cover image) from at least one second video.
[0228] In the embodiment of the present application, there are three strategies for capturing the cover frame in the video under different platforms and scenarios:
[0229] The first method is to use video keyframes and I-frames to superimpose light super-resolution to improve clarity.
[0230] The second method: When you click Record, an additional photo stream is added in the video mode to obtain the photo frame overlay algorithm processing.
[0231] The third method: directly capture the preview frame when clicking to take a photo.
[0232] It should be noted that the solution selection needs to comprehensively consider encoding characteristics, clarity requirements, and implementation complexity. The following analysis is from three dimensions: technical principles, practical strategies, and recommended solutions:
[0233] (1) Differences in logic and technology for selecting cover frames.
[0234] The first one is the I-frame interception solution. Since the I-frame is a key frame in video encoding (full-frame compression), it contains complete image information, does not rely on motion compensation, and has low decoding complexity. Therefore, direct extraction of the I-frame does not require additional decoding calculations, which is suitable for scenarios with high real-time requirements (such as video preview cover generation), and the image quality loss is controllable, consistent with the original frame quality of the second video. However, if the video GOP is long and the key frame interval is large (such as an I-frame every 30 frames), it may not be able to accurately match the highlight moment expected by the user. The I-frames of some low-bitrate videos have blurred details due to compression and require post-processing enhancement.
[0235] The second is a solution for independent generation of photo streams. By synchronously starting an independent photo pipeline during video recording, at least one sixth image is generated with a higher resolution (such as 48MP). The image quality of the sixth image is significantly higher than the frame in the second video, and it supports multi-frame synthesis and better dynamic range. However, this solution consumes high hardware resources and requires dual ISP parallel processing, which may affect the stability of video recording. In addition, timestamp synchronization errors may cause deviations between the cover frame and the dynamic image content, such as ±50ms jitter.
[0236] The third approach involves capturing video preview frames (low-resolution thumbnails) from the video decoding pipeline as cover images. This approach is simple to implement, requires no additional hardware, and offers full synchronization of the frame rate with the video, ensuring high consistency between the cover image and the dynamic content. However, the preview frame resolution is low (e.g., 720p), resulting in insufficient clarity after zooming (with noticeable jagged edges and noise). Furthermore, the YUV420 color space may result in color shift after conversion to RGB.
[0237] (2) Clarity optimization strategy.
[0238] For I-frame or preview frame extraction strategies, the I-frame closest to the click moment is extracted as the base cover. AI super-resolution (such as SRGAN and ESRGAN) is then used to increase the resolution by 2-4 times. Finally, sharpening and noise reduction parameters are dynamically optimized based on the SSIM metric. This optimization method is suitable for scenarios with high image quality requirements and a certain degree of latency tolerance (such as post-editing).
[0239] For independent photo stream generation, 3-5 consecutive frames (including I / P / B frames) are captured around the click moment, aligned using motion compensation, and then multi-frame stacking for noise reduction and detail fusion. This approach is suitable for reducing motion blur in dynamic scenes (such as sports and dancing).
[0240] It should be noted that in the embodiment of the present application, a metadata dynamic binding method is adopted, and the cover frame is stored separately from the video and associated through the HEIF file header, thereby allowing the user to manually replace the cover frame later without affecting the dynamic content.
[0241] (3)Scheme selection.
[0242] For common video recording scenarios (such as instant sharing on social platforms), I-frame capture + lightweight super-resolution (such as the MobileNet optimization model) is preferred to balance speed and image quality. When reselecting the cover frame in the album, the lightweight algorithm can be called again to optimize the selected frame clarity to the level of the cover frame.
[0243] For high-quality video creation scenarios (such as professional video mode), independent photo stream generation and multi-frame synthesis are preferred, requiring a device with dual ISP support. When taking a photo, the photo data stream is sent down, and the video + photo dual-camera mode is activated, leveraging the advantages of dual ISPs to simultaneously generate the video stream and photo frames. Metadata is associated with the cover and the dynamic clip, allowing the cover frame to have high-definition quality and clarity. When reselecting the cover frame in the album, a lightweight algorithm is used to optimize the selected frame's clarity to the same level as the cover frame.
[0244] For scenarios requiring cross-platform compatibility optimization, cover frames are saved as JPEG / HEIC still images, dynamic videos are encapsulated as MOV / H.265, and then merged and presented via file system symbolic links. On some platforms, MediaExtractor must be called to parse the video keyframe index.
[0245] Table 1
[0246] plan Clarity (SSIM score) Delay (ms) Hardware load I-frame capture 0.82-0.88 10-30 Low Independent photo streaming 0.92-0.95 50-100 high Preview frame + super resolution 0.85-0.90 20-50 middle
[0247] In the embodiment of the present application, Table 1 is a schematic diagram of technical verification and effect comparison of the above-mentioned different strategies for capturing cover frames in videos.
[0248] In summary, for computers with average performance, the preferred solution is I-frame capture + AI enhancement, which ensures real-time performance while improving image quality. If device performance allows, higher-performance computers should use an independent photo stream to generate the optimal cover frame quality. Avoid using low-resolution preview frames directly. For computers with average performance, multi-frame synthesis can be used to compensate for detail loss when necessary.
[0249] The present invention provides a method for generating a dynamic image, which has the following characteristics:
[0250] (1) Asynchronous staggered processing: Separate the process of recording the first video and the generation of the target dynamic image to reduce the real-time load.
[0251] (2) Dynamic key frame compensation: Combine GOP analysis and interpolation algorithm to solve the problem of non-I frame interception.
[0252] (3) Cross-coding format compatibility: Adapting to mixed scenarios of fixed encoding (MediaRecorder) and custom encoding (MediaCodec).
[0253] 5. The method for generating dynamic images provided in this application has the following technical effects:
[0254] The embodiments of the present application significantly improve the user experience by optimizing the video recording and LivePhoto generation processes, combining asynchronous processing, keyframe compensation, and cross-platform adaptation mechanisms. This is specifically reflected in the following aspects:
[0255] (1) Seamless connection between recording and generation, improving operation fluency.
[0256] Real-time, uninterrupted recording: When the user taps "Photo" during the first video recording, only the raw data (YUV+PCM) before and after 1.5 seconds is cached, leaving the recording process unaffected. This reduces CPU / GPU utilization by 40% compared to similar technologies that process video and photos simultaneously.
[0257] Seamless preview experience: Before the dynamic image is generated, a preview screenshot of the moment the user clicks is directly displayed in the album (static frames are quickly extracted based on the YUV data of the video stream), preventing users from perceiving a "blank" or "loading" state.
[0258] For example, when the user clicks to take a photo while shooting a jumping action, the album immediately displays a static preview image of the highest point of the jump, and the background asynchronously generates a dynamic LivePhoto and automatically replaces it, and the user does not perceive the switch.
[0259] (2) Dynamic capture and precise restoration of highlight moments.
[0260] Dynamic range expansion: Through a keyframe compensation mechanism that extends forward to the nearest I-frame and an optical flow interpolation algorithm, dynamic images can be generated without tearing or lag, even if the video capture boundaries do not align with keyframes.
[0261] Timing consistency assurance: Audio timestamps are dynamically remapped based on the expanded video clips to resolve audio and video asynchrony issues caused by frame rate fluctuations (for example, audio delay is ≤ 20ms in temperature-controlled frame reduction scenarios).
[0262] (3) Cross-platform sharing and compatibility optimization.
[0263] HEIF+short video packaging: The cover image and the second video are packaged into the HEIF format (static cover + short video), which is compatible with dynamic playback on different platforms and static preview effects on Android devices, solving the pain point that some platforms only support static images. For example, a dynamic image shot by a user can be long-pressed to display a dynamic effect on a computer device that supports a certain type of system. For a computer device that supports another type of system, the HEIF cover image will be automatically displayed after receiving it, and clicking the play button will play the video.
[0264] Intelligent metadata synchronization: Metadata such as timestamps and keyframe indexes are recorded synchronously during recording to ensure that dynamic effects are not lost when transmitted across devices.
[0265] (4) Storage efficiency and editing flexibility are improved.
[0266] Block compression technology: Using the Zstandard algorithm to compress repeated frames (such as static backgrounds), the size of a 3-second target dynamic image is reduced from 15MB (H.264) to 8MB, saving 46% of storage space.
[0267] Dynamic cover editing: Users can freely replace the cover frame (such as selecting the moment with the most natural expression). After editing, the HEIF file metadata is automatically updated without the need to regenerate the video.
[0268] (5) User scenario coverage and experience extension.
[0269] Multi-segment highlight merging: The segments of dynamic images generated by multiple clicks to take pictures can be automatically spliced through the timeline index to form a continuous dynamic memory (such as merging multiple exciting actions in a travel vlog into a 10-second dynamic image).
[0270] Decoupling of preview and generation: Low-memory computer devices can retain only timestamp metadata and subsequently unpack the original video to generate the target dynamic image on demand, balancing performance and storage flexibility.
[0271] Table 2 is a summary of the technical effects of the method for generating dynamic images provided in the embodiments of the present application.
[0272] Table 2
[0273]
[0274] In summary, the embodiments of the present application achieve "what you see is what you get" highlight capture, "imperceptible" generation transition and "zero threshold" cross-platform sharing, redefining the user experience standard of dynamic images.
[0275] 6. Description of core technical solution.
[0276] In the embodiments of this application, the asynchronous post-processing architecture and the multi-modal seamless switching mechanism are two core innovations that redefine the generation and sharing paradigm of dynamic images, which are specifically reflected in the following dimensions:
[0277] (1) Off-peak post-processing architecture: decoupling of recording and generation.
[0278] Asynchronous task scheduling: When the user clicks to take a photo, only the raw data stream is cached to memory or flash memory. After recording, it is unpacked and processed by a background thread to avoid a sharp increase in CPU / GPU usage caused by real-time encoding.
[0279] Dynamic resource allocation: Intelligently selects processing timing based on device load status (memory, temperature). When memory is sufficient, target dynamic images are generated in real time. When the load is high, only metadata is recorded and processing is delayed.
[0280] (2) Seamless switching mechanism between preview image and dynamic image.
[0281] Instant static frame extraction: Before the target dynamic image is generated, the preview frame at the moment the user clicks is extracted from the cached video data and directly used as the photo cover image, solving the "blank period" problem caused by generation delays.
[0282] Dynamic replacement without perception: After the background generation is completed, the cover is automatically replaced with the complete target dynamic image through metadata update (dynamic tag rewriting of HEIF file), and the user does not need to refresh manually.
[0283] (3) Cross-coding format dynamic compensation technology.
[0284] Keyframe boundary extension: Locates the nearest I-frame by parsing the GOP structure and extends the capture interval forward to avoid screen tearing caused by non-I-frame capture.
[0285] Optical flow interpolation frame supplementation: For variable frame rate scenarios (such as temperature control frame reduction), intermediate frames are generated based on the motion vectors of adjacent P frames to ensure smooth dynamic effects.
[0286] (4) Multi-platform compatible HEIF packaging standard.
[0287] Block storage optimization: The second video is split into key frames (lossless compression) and differential frames (Zstandard algorithm compression), which reduces the volume by more than 40% compared to the compression methods of related technologies.
[0288] Cross-ecosystem metadata synchronization: Embed metadata such as timestamps and keyframe indexes in HEIF files to achieve lossless transmission of dynamic effects between different platforms.
[0289] (5) User scenario-driven interaction design.
[0290] Automatic stitching of multiple highlights: The dynamic image segments (second video) generated by multiple clicks are merged according to the timeline index and support continuous playback.
[0291] Seamless adaptation to social platforms: By packaging it into HEIF+short video links, it solves the problem of confusing formats of dynamic photo sharing on different platforms.
[0292] Compared with the related art, the method for generating dynamic images provided in the embodiment of the present application has the following breakthroughs as shown in Table 3:
[0293] Table 3
[0294]
[0295] In this way, the method for generating dynamic images provided in the embodiments of the present application solves the core problems in the acquisition and generation of dynamic images in related technologies, and by specifying a unified packaging standard for dynamic images, provides a solution for sharing and storing dynamic images across devices and platforms.
[0296] Figure 5 An embodiment of the present application provides a display effect of a target dynamic image in an album. The album displays a static high-definition cover frame generated based on the aforementioned embodiment. After the user selects to enlarge the display, a "live mark" is displayed, indicating that the image is a dynamic image, and a play mark d1 is set in the lower left corner of the dynamic image. The user can click the play mark d1 to play the target dynamic image and display the dynamic effect.
[0297] 7. Expansion of technical solutions.
[0298] The dynamic image generation method provided in the embodiments of the present application is highly scalable. By combining cutting-edge technologies with multimodal interaction logic, it can achieve functional deepening and scenario expansion in the following directions:
[0299] (1) AI-enhanced dynamic generation and intelligent interaction.
[0300] Optical flow frame supplementation optimization: To address the issue of blurred video capture boundaries, a multi-scale generation strategy is introduced. The optical flow method is used to predict the motion vectors of adjacent frames and generate smooth intermediate frames (such as to supplement missing images when temperature control frame reduction is required), thus solving the image distortion problem caused by traditional interpolation algorithms.
[0301] AI generates dynamic backgrounds: Add dynamic elements (such as flowing clouds and swaying leaves) to static cover images, and use the GAN model to generate seamless animation effects that blend with the original scene.
[0302] (2) Enhanced cross-platform and cross-media compatibility.
[0303] Full-ecosystem dynamic format adaptation: Based on universal encapsulation standards, the target dynamic image is expanded into a "HEIF+short video+metadata" combination package, which is compatible with dynamic playback, static preview and social platform playback, solving the platform limitations of related technologies.
[0304] AR augmented reality interaction: Combined with AR technology, users can scan the cover frame with their mobile phones to trigger superimposed dynamic special effects (such as virtual fireworks and 3D character movements), achieving an immersive experience that blends virtual and real.
[0305] Web3 and NFT integration: Write the metadata (timestamp, keyframe index) of the target dynamic image into the blockchain to generate a unique digital certificate, supporting the trading of dynamic photos as NFT collections.
[0306] (3) Dynamic interaction and multimodal fusion.
[0307] Audio emotional enhancement: Based on multimodal fusion logic, background music (such as cheers and ambient sounds) is matched when generating the target dynamic image, and users are supported to customize audio tracks (such as adding voice narration).
[0308] Haptic feedback synchronization: Combined with the device's vibration motor, the motion intensity in dynamic images is mapped to tactile signals (such as a short vibration at the moment of jumping), enhancing the multi-sensory interactive experience.
[0309] AI-assisted editing and recommendation: The recommendation model is trained based on user historical data, automatically extracting highlight clips in the video and generating a collection of dynamic images (such as consecutive scoring moments in sports events), thereby reducing manual screening costs.
[0310] (4) Collaborative optimization of edge computing and cloud.
[0311] Lightweight inference on the device: Deploy a lightweight motion prediction model to the device's NPU to localize keyframe compensation and interpolation calculations, reducing cloud-side reliance (reducing latency from 500ms to 50ms).
[0312] Cloud-based batch rendering: For scenarios where multiple second videos are merged, batch generation technology is used to asynchronously render high-resolution dynamic collections through distributed clusters to solve the computing power bottleneck on mobile devices.
[0313] Block compression transmission: Split HEIF files into key frames (stored on the device) and differential frames (stored in the cloud), and load them on demand when users view them, reducing data consumption.
[0314] (5) Deep customization based on industry scenarios.
[0315] Capturing highlights in e-commerce live streaming: Automatically capture product display moments during live streaming to generate dynamic images, and support one-click sharing to social platforms to increase conversion rates.
[0316] Dynamic courseware in the education field: By combining AR teaching cases with this solution, teachers can mark key sections when recording explanation videos, and students can quickly review experimental steps or formula derivation processes through dynamic images.
[0317] Instant replay of sports events: Expand the motion intensity model, automatically identify athletes' scoring actions, and generate multi-angle dynamic image highlights for rapid dissemination by media and audiences.
[0318] Compared with the related art, the advantages of the expansion solution provided by the embodiment of the present application are shown in Table 4:
[0319] Table 4
[0320]
[0321] In summary, the dynamic image generation method provided in the embodiments of the present application improves the technical advantages of dynamic images in the shooting, generation, interaction and transmission processes, and more effectively adapts to user needs.
[0322] In another embodiment of the present application, Figure 6 This is a schematic diagram of the structure of a dynamic image generation device provided in an embodiment of the present application. Figure 6 As shown, the dynamic image generation device 600 includes:
[0323] An acquisition module 6001 is configured to acquire at least one first instruction input by a user during the process of capturing a first video;
[0324] The acquisition module 6001 is further configured to acquire at least one second video in the first video based on at least one first instruction after the first video is acquired;
[0325] The processing module 6002 is configured to generate a target dynamic image related to at least one second video.
[0326] In some embodiments, the acquisition module 6001 is also used to determine, in response to each first instruction, multiple first image indexes in a time period adjacent to the click moment of the first instruction in the first video; and based on the multiple first image indexes in the first video, determine multiple first images in the second video corresponding to the first instruction.
[0327] In some embodiments, the acquisition module 6001 is also used to determine the motion information corresponding to at least one key frame in the second video when the first video is encoded at a variable frame rate; determine the number of images between each key frame and the adjacent key frame based on the motion information corresponding to the at least one key frame; generate at least one second image between each key frame and the adjacent key frame based on the number of images between each key frame and the adjacent key frame and a first preset algorithm to obtain an extended second video; perform audio and video synchronization processing on the extended second video to obtain a synchronized second video.
[0328] In some embodiments, the processing module 6002 is further configured to perform cover selection processing on at least one second video to determine a cover image; and determine a target dynamic image based on the cover image and the at least one second video.
[0329] In some embodiments, the processing module 6002 is further used to obtain one or more third images in at least one second video; perform first processing on the one or more third images to obtain a cover image.
[0330] In some embodiments, the processing module 6002 is also used to obtain multiple fourth images in the photo data stream during the process of capturing the first video; obtain multiple fifth images closest to the first click moment from the multiple fourth images; the first click moment is one of the click moments of at least one first instruction; perform a second processing on the multiple fifth images to obtain a cover image, and the second processing at least includes fusion processing.
[0331] In some embodiments, the processing module 6002 is also used to determine the timestamps of multiple second videos based on the click moments of multiple first instructions; merge the multiple second videos based on the timestamps corresponding to each of the multiple second videos to obtain a third video; and use the cover image and the third video to determine the target dynamic image.
[0332] In some embodiments, the processing module 6002 is also used to determine at least one key frame and at least one differential frame of at least one second video in the target dynamic image; compress at least one key frame based on a first compression algorithm to obtain compressed image data of at least one key frame; compress at least one differential frame based on a second compression algorithm to obtain compressed image data of at least one differential frame; wherein the compression loss of the first compression algorithm is less than the compression loss of the second compression algorithm; encapsulate the compressed image data of at least one key frame of the target dynamic image, the compressed image data of at least one differential frame and the cover image to obtain the storage data of the target dynamic image, and store or transmit it.
[0333] In some embodiments, the processing module 6002 is also used to generate a sixth image corresponding to the click moment in response to each first instruction; and in response to the first operation on the sixth image, execute the step of obtaining at least one second video in the first video based on at least one first instruction.
[0334] In some embodiments, the processing module 6002 is further used to execute the step of obtaining at least one second video in the first video based on at least one first instruction within a first preset time period based on load status information of the computer device.
[0335] In another embodiment of the present application, Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. Figure 7 As shown, the computer device 700 includes:
[0336] Memory 720, for storing computer programs;
[0337] The processor 710 is connected to the memory 720 and is used to call and run the computer program from the memory to implement the method executed by the terminal device or the method executed by the server;
[0338] The transceiver 730 is used to send and receive information when sending and receiving information with other devices.
[0339] The memory 720 may be a separate device independent of the processor 710 , or may be integrated into the processor 710 .
[0340] Alternatively, as Figure 7 As shown, the computer device 700 may further include a transceiver 730 , and the processor 710 may control the transceiver 730 to communicate with other devices. Specifically, the transceiver 730 may send information or data to other devices, or receive information or data sent by other devices.
[0341] The transceiver 730 may include a transmitter and a receiver. The transceiver 730 may further include an antenna, and the number of antennas may be one or more.
[0342] Optionally, the computer device 700 can implement the corresponding processes in the various methods of the embodiments of the present application, which will not be described here for the sake of brevity.
[0343] An embodiment of the present application also provides a computer-readable storage medium for storing a computer program.
[0344] Optionally, the computer-readable storage medium can be applied to the network device in the embodiments of the present application, and the computer program enables the computer to execute the corresponding processes implemented by the network device in the various methods of the embodiments of the present application. For the sake of brevity, they are not repeated here.
[0345] Optionally, the computer-readable storage medium can be applied to the mobile terminal / terminal device in the embodiments of the present application, and the computer program enables the computer to execute the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0346] An embodiment of the present application also provides a computer program product, including computer program instructions.
[0347] Optionally, the computer program product can be applied to the network device in the embodiments of the present application, and the computer program instructions enable the computer to execute the corresponding processes implemented by the network device in the various methods of the embodiments of the present application. For the sake of brevity, they are not repeated here.
[0348] Optionally, the computer program product can be applied to the mobile terminal / terminal device in the embodiments of the present application, and the computer program instructions enable the computer to execute the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0349] The embodiment of the present application also provides a computer program.
[0350] Optionally, the computer program can be applied to the network device in the embodiments of the present application. When the computer program runs on a computer, the computer executes the corresponding processes implemented by the network device in the various methods of the embodiments of the present application. For the sake of brevity, they are not described here.
[0351] Optionally, the computer program can be applied to the mobile terminal / terminal device in the embodiments of the present application. When the computer program runs on the computer, the computer executes the corresponding processes implemented by the mobile terminal / terminal device in the various methods of the embodiments of the present application. For the sake of brevity, they will not be repeated here.
[0352] It should be understood that "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments. The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other. For the sake of brevity, they will not be repeated here.
[0353] It should also be noted that, in this application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0354] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0355] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0356] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0357] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0358] The above is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A method for generating a dynamic image, characterized in that: The method comprises: During the process of capturing the first video, obtaining at least one first instruction input by a user; After the first video is captured, obtaining at least one second video in the first video based on the at least one first instruction; A target dynamic image related to the at least one second video is generated.
2. The method according to claim 1, characterized in that The acquiring, based on the at least one first instruction, at least one second video in the first video includes: In response to each of the first instructions, determining a plurality of first image indexes within a time period adjacent to a click moment of the first instruction in the first video; Based on the multiple first image indexes in the first video, multiple first images in the second video corresponding to the first instruction are determined.
3. The method according to claim 1, characterized in that In the case where the first video is encoded using a variable frame rate, before generating a target dynamic image related to the at least one second video, the method further includes: Determining motion information corresponding to at least one key frame in the second video; determining the number of images between each key frame and an adjacent key frame based on motion information corresponding to the at least one key frame; Based on the number of images between each key frame and an adjacent key frame and a first preset algorithm, generating at least one second image between each key frame and an adjacent key frame to obtain an expanded second video; Perform audio and video synchronization processing on the expanded second video to obtain a synchronized second video.
4. The method according to claim 1, wherein The generating of a target dynamic image related to the at least one second video includes: performing cover selection processing on at least one of the second videos to determine a cover image; The target dynamic image is determined based on the cover image and at least one of the second videos.
5. The method according to claim 4, characterized in that The performing cover selection processing on at least one of the second videos to determine the cover image includes: Acquire one or more third images in the at least one second video; A first process is performed on one or more of the third images to obtain the cover image.
6. The method according to claim 4, characterized in that The method further comprises: During the process of capturing the first video, obtaining a plurality of fourth images in the photographic data stream; The performing cover selection processing on at least one of the second videos to determine the cover image includes: Acquire, from the plurality of fourth images, a plurality of fifth images closest to a first click moment, wherein the first click moment is one of the click moments of the at least one first instruction; A second processing is performed on the multiple fifth images to obtain the cover image, and the second processing at least includes a fusion processing.
7. The method according to claim 4, characterized in that The determining the target dynamic image based on the cover image and at least one second video includes: Determining timestamps of a plurality of second videos based on click moments of a plurality of first instructions; Merging the plurality of second videos based on respective timestamps corresponding to the plurality of second videos to obtain a third video; The target dynamic image is determined using the cover image and the third video.
8. The method according to claim 4, characterized in that The method further comprises: determining at least one key frame and at least one differential frame of at least one second video in the target dynamic image; compressing the at least one key frame based on a first compression algorithm to obtain compressed image data of the at least one key frame; compressing the at least one differential frame based on a second compression algorithm to obtain compressed image data of the at least one differential frame; wherein a compression loss of the first compression algorithm is less than a compression loss of the second compression algorithm; The compressed image data of the at least one key frame of the target dynamic image, the compressed image data of the at least one differential frame and the cover image are encapsulated to obtain storage data of the target dynamic image, and the storage data is stored or transmitted.
9. The method according to any one of claims 1 to 8, characterized in that After obtaining at least one first instruction input by the user, the method further includes: In response to each of the first instructions, generating a sixth image corresponding to the click moment; The method further comprises: In response to the first operation on the sixth image, the step of acquiring at least one second video in the first video based on the at least one first instruction is performed.
10. The method according to any one of claims 1 to 8, characterized in that The method further comprises: Based on the load status information of the computer device, the step of obtaining at least one second video in the first video based on the at least one first instruction is performed within a first preset time period.
11. A device for generating a dynamic image, characterized in that: The device comprises: An acquisition module, configured to acquire at least one first instruction input by a user during the process of capturing the first video; The acquisition module is further configured to acquire at least one second video in the first video based on the at least one first instruction after the first video is acquired; The processing module is configured to generate a target dynamic image related to the at least one second video.
12. A computer device, characterized in that: include: memory for storing computer programs; a processor, connected to the memory, and configured to call and execute the computer program from the memory to implement the method according to any one of claims 1 to 10; A transceiver is used to send and receive information between devices.
13. A computer program product, characterized in that The method comprises computer program instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 10.