Dynamic photo generation method based on character selection
By using a dynamic photo generation method based on person selection, the problem of unclear people in dynamic photos is solved, achieving the effect of a static background and dynamic people, thus enhancing the user's personalized experience.
Patent Information
- Application Number
- CN202511124224.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-14
AI Technical Summary
The existing dynamic photo function fails to effectively highlight the subject, resulting in cluttered images and unclear visual focus, thus failing to meet users' personalized expression needs.
By selecting the "Dynamic People" mode to take a photo, the system automatically records a video at a specified time before and after the subject, extracts the image frames, performs background removal on the subject, synthesizes the video, and converts it into a dynamic image, ensuring that the background is static while the subject is dynamic, and provides a subject selection and switching function.
This feature allows users to select any person to display dynamic effects when viewing photos, enhancing the personalized expression of photos and improving the user experience.
Smart Images

Figure CN120957008A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically to a method for generating dynamic photos based on person selection. Background Technology
[0002] With societal development, smartphones have become increasingly intelligent. As a crucial function of mobile phones, photography allows people to record their lives through photos and videos. Users' demands for expressive forms in photos are also growing. Traditional photos are static images, lacking the expressive power of movement and time. In recent years, some mobile phone manufacturers have launched "live photos" features, providing users with a new photography experience. Essentially, these features record a few seconds of video footage while taking a photo, saving it as a GIF or short video to showcase the dynamic scenes before and after the moment the photo is taken.
[0003] However, existing animated photos are all in motion, which makes it difficult to highlight the main subject. Furthermore, the dynamic effects make the dynamic process cluttered, the visual focus unclear, and the subject not stand out enough, thus failing to meet the user's needs. Summary of the Invention
[0004] To address the problems in the existing technology, the present invention provides a method for generating dynamic photos based on person selection.
[0005] This invention discloses a method for generating dynamic photos based on person selection, comprising the following steps:
[0006] S1: Select "Dynamic Portrait" mode to take a photo;
[0007] S2: Automatically record image information within a specified time period before and after shooting;
[0008] S3: Extract all image frames from the recorded video;
[0009] S4: Determine if a person exists in the image frame;
[0010] S5: Extract the figures from all image frames to separate the background from the figures.
[0011] S6: Using the first frame as the background, the characters extracted from all the image frames are played in chronological order to create a composite video;
[0012] S7: Convert video into animated images.
[0013] The present invention is further improved by specifying a time of 2 seconds in step S2.
[0014] The present invention is further improved in that if it is determined that there are multiple characters in step S4, then in step S5, the different characters will be cut out separately.
[0015] The present invention is further improved in that if it is determined that there are multiple characters in step S4, the frames of different characters will be combined with the background in step S6 to generate a dynamic video of multiple individual characters.
[0016] The present invention is further improved in that, in step S7, the user can switch between dynamic photos of different people based on the selection of different people.
[0017] The present invention is further improved by directly ending the generation of dynamic photos of people if no task is identified in step S4.
[0018] The present invention is further improved in that, in step S5, the process of cutting out the person in each frame specifically includes the following steps:
[0019] A01: Extract the grayscale values corresponding to all pixels in the image;
[0020] A02: Detect gradient changes in the horizontal and vertical directions in an image;
[0021] A03: Starting from the top left corner of the image, sequentially select specified pixel regions for convolution calculation to calculate the gradient magnitude of that region;
[0022] A04: Set a specified threshold; all pixels whose gradient magnitude exceeds the threshold are identified as edge regions.
[0023] A05: After edge detection, existing object detection models can be used to directly locate people in the image and output the results.
[0024] The edge area of a person;
[0025] A06: The figure can be cut out based on the edge area;
[0026] A07: Output mask image;
[0027] A08: Use a mask to extract the foreground from the original image to obtain a transparent image containing only the person.
[0028] The present invention is further improved in that, in step S6, the video is synthesized by combining the person in each frame with the background frame, and then encoding all the synthesized frames in sequence into a video file.
[0029] The present invention is further improved by reading the video frame by frame in step S3 using the video OpenCV decoding library and saving each frame as image data in RGB format.
[0030] Compared with the prior art, the beneficial effects of the present invention are: by adopting its mechanism, it can effectively solve the problems of excessive dynamic elements and unclear visual focus in the dynamic photo function of the prior art. By using this method, users can select any person when viewing photos, and only that person will present a dynamic effect while the rest of the background remains static, which enhances the personalized expression of the photos and improves the user experience. Attached Figure Description
[0031] To more clearly illustrate the solutions in this invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0032] Figure 1 Here is a flowchart of a method for generating dynamic photos based on person selection; Detailed Implementation
[0033] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings are used to distinguish different objects, not to describe a particular order.
[0034] In this invention, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this invention can be combined with other embodiments.
[0035] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0036] like Figure 1 As shown, this invention provides a dynamic photo generation method based on person selection. This method is applicable to smart devices such as mobile phones, tablets, and watches. The operation method is as follows:
[0037] 1. Select the "Dynamic Portrait" mode to take a photo.
[0038] Users need to manually click the "Dynamic People" function on the camera to use it; otherwise, they can take ordinary dynamic photos.
[0039] 2. Automatically record video information within a specified time period before and after shooting.
[0040] After the user presses the shutter button, the device automatically records and saves video content within each preset time period before and after taking the photo (e.g., 2 seconds before taking the photo to 1 second after taking the photo), forming a continuous video clip.
[0041] 3. Extract all image frames from the recorded video.
[0042] After the user finishes capturing a moving subject, the system decodes the video segment recorded before and after the user presses the shutter button (e.g., the first 2 seconds followed by 1 second) frame by frame. The OpenCV video decoding library can be used to read the video frame by frame, saving each frame as RGB image data. Finally, a series of image frames is obtained.
[0043] 4. Determine if there are people in the image frame;
[0044] If no task is identified, the generation of dynamic photos of people will end immediately.
[0045] 5. Cut out the figures in all image frames to separate the background from the figures.
[0046] all.
[0047] If multiple people are identified, each person will be cut out separately to facilitate the creation of selectable animated photos.
[0048] The specific steps for cutting out the character from each frame are as follows:
[0049] (1) Extract the grayscale values corresponding to all pixels in the image;
[0050] (2) Detect gradient changes in the horizontal and vertical directions in the image;
[0051] (3) Starting from the top left corner of the image, take the specified pixel regions one by one for convolution calculation and calculate the gradient magnitude of the region.
[0052] The pixel region is set to 3x3, containing nine pixels. Starting from the top left corner of the image, two convolution kernels are used to perform calculations on the X and Y axes respectively.
[0053]
[0054] Calculation method, G XThe values in G1, G2, G3, G4, G5, G6, G7, G8 and G9 correspond to grayscale values respectively.
[0055] G X =G1*-1+G2*0+G3*1+G4*-2+G5*0+G6*2+G7*-1+G8*0+G9*1
[0056] G Y With G X The calculation method is the same, so it will not be repeated here.
[0057] G can be derived from the above formula. X and G Y Two sets of data are used, and then the gradient magnitude of the region is calculated using the following formula.
[0058]
[0059] This value indicates whether the region is a marginal area or not. If the value is large, the region may be a marginal area. If the value is small, the region is not a marginal area.
[0060] (4) Set a specified threshold, and all pixels whose gradient magnitude exceeds the threshold are judged as edge regions;
[0061] (5) Once the edge is detected, the existing target detection model can be used to directly locate the person in the image and output the edge region of the person.
[0062] Object detection models include, for example, YELO.
[0063] (6) The figure can be cut out based on the edge area;
[0064] (7) Output the mask image;
[0065] (8) Use the mask to extract the foreground from the original image to obtain a transparent image containing only the person.
[0066] The transparent image is in PNG format and has the same size as the original image, except that the non-human areas are transparent.
[0067] 6. Using the first frame as the background, the characters extracted from all the image frames are played in chronological order to create a composite video.
[0068] like.
[0069] The first frame is the starting image, a static and unchanging scene.
[0070] By default, a certain person is selected, and the cutout results of that person in all frames are extracted to form a dynamic sequence.
[0071] During display, all sequence frames are played continuously. Since all the cut-out images except the first frame contain only one character, when playing the continuous frames, only that character moves, while the rest of the area remains still.
[0072] If multiple characters are detected, frames of different characters will be combined with the background and played together to generate a dynamic video of multiple individual characters.
[0073] The process of compositing video involves combining each frame of the person with a background frame, and then encoding all the composite frames sequentially into a video file.
[0074] The background image and the dynamic sequence of the selected person are combined into a video, which presents a visual effect of "the background is stationary and the person moves" when played.
[0075] All composite frames are encoded sequentially into video files (such as MP4, MOV);
[0076] By blending the dynamic elements of a person with a static background and playing them frame by frame, a photo with a localized dynamic effect can be generated. This allows users to directly view the dynamic effect when previewing images in their gallery, making the photo more vivid and engaging.
[0077] Users can manually select individuals of interest. A "Select Individual" button is provided when playing animated photos of individuals. Clicking this button lists all segmented and marked individuals on the screen, allowing the user to select a specific individual. Upon confirmation, all frames from that individual's sequence will be played automatically.
[0078] 7. Convert the video into animated GIFs.
[0079] Users can switch between different animated photos of different people based on their selection.
[0080] 8. End.
[0081] In summary, the dynamic photo generation method based on person selection provided by this invention can effectively solve the problems of excessive dynamic elements and unclear visual focus in existing dynamic photo functions. By using this method, users can select any person when viewing photos, and only that person will have a dynamic effect while the rest of the background remains static, enhancing the personalized expression of the photos and improving the user experience.
[0082] The specific embodiments described above are preferred embodiments of the present invention and are not intended to limit the specific scope of the present invention. The scope of the present invention includes, but is not limited to, these specific embodiments. All equivalent changes made in accordance with the present invention are within the protection scope of the present invention.
Claims
1. A method for generating dynamic photos based on person selection, characterized in that, Includes the following steps: S1: Select "Dynamic Portrait" mode to take a photo; S2: Automatically record image information within a specified time period before and after shooting; S3: Extract all image frames from the recorded video; S4: Determine if a person exists in the image frame; S5: Extract the figures from all image frames to separate the background from the figures. S6: Using the first frame as the background, the characters extracted from all the image frames are played in chronological order to create a composite video; S7: Convert video into animated images.
2. The dynamic photo generation method based on person selection according to claim 1, characterized in that: In step S2, the specified time is 2 seconds.
3. The dynamic photo generation method based on person selection according to claim 1, characterized in that: If multiple characters are determined to exist in step S4, then in step S5, the different characters will be cut out separately.
4. The dynamic photo generation method based on person selection according to claim 3, characterized in that: If it is determined in step S4 that there are multiple characters, in step S6 the frames of different characters will be combined with the background to generate a dynamic video of multiple individual characters.
5. The dynamic photo generation method based on person selection according to claim 4, characterized in that: In step S7, the user can switch between different animated photos of different people based on the selected person.
6. The method for generating dynamic photos based on person selection according to claim 1, characterized in that: If no task is identified in step S4, the generation of the dynamic photo of the person will end directly.
7. The method for generating dynamic photos based on person selection according to claim 1, characterized in that: Step S5 involves the following steps for cutting out the character from each frame: A01: Extract the grayscale values corresponding to all pixels in the image; A02: Detect gradient changes in the horizontal and vertical directions in an image; A03: Starting from the top left corner of the image, sequentially select specified pixel regions for convolution calculation to calculate the gradient magnitude of that region; A04: Set a specified threshold; all pixels whose gradient magnitude exceeds the threshold are identified as edge regions. A05: After the edge is detected, the existing object detection model can be used to directly locate the person in the image and output the edge region of the person; A06: The figure can be cut out based on the edge area; A07: Output mask image; A08: Use a mask to extract the foreground from the original image to obtain a transparent image containing only the person.
8. The method for generating dynamic photos based on person selection according to claim 1, characterized in that: In step S6, the video synthesis process involves compositing each frame of the person with the background frame, and then encoding all the composite frames sequentially into a video file.
9. The method for generating dynamic photos based on person selection according to any one of claims 1-8, characterized in that: In step S3, the video is read frame by frame using the OpenCV video decoding library, and each frame is saved as image data in RGB format.