Image generation method and device, electronic equipment and storage medium
By calculating the motion vector field of video sequences and generating dynamic region masks, dynamic and static elements are automatically separated, solving the problem of the complexity of creating dynamic low-light images and enabling ordinary users to easily generate and present high-quality images.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-05
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, the creation of dynamic low-light images relies on specialized software and complex frame-by-frame masking editing, which has a high barrier to entry, is complicated to operate, and is difficult for ordinary users to achieve.
By calculating the motion vector field of a video sequence, a dynamic region mask is generated, and dynamic and static elements are automatically separated to generate dynamic low-light images.
It lowers the technical threshold and operational complexity of creating dynamic low-light images, enabling ordinary users to easily and efficiently generate high-quality dynamic low-light images with excellent local motion capture capabilities, presenting a cyclical dynamic effect in local areas.
Smart Images

Figure CN121767487A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of electronic equipment technology, specifically relating to an image generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] Currently, modern electronic devices generally have a variety of dynamic image shooting functions, such as slow-motion video, timelapse photography, and live photos. These shooting functions mainly provide users with rich visual presentation methods by recording the overall dynamics of the scene: slow-motion video magnifies the details of the moment, timelapse photography compresses the time sequence, and live photos capture vivid fragments at the moment of shutter speed.
[0003] However, these shooting functions share a common limitation in visual expression: they showcase the overall movement of the scene, making it difficult to precisely control the viewer's visual focus. In contrast, Cinemagraph achieves a visual breakthrough through its unique "dynamic within stillness" approach. It keeps most of the image static, allowing only subtle elements (such as flickering candlelight or flowing water) to create a looping animation, thus subtly guiding the user's gaze and creating a more immersive and narrative visual experience.
[0004] In related technologies, the creation of dynamic low-light images mainly relies on professional software such as Photoshop and After Effects, requiring users to have professional image processing skills. The production process involves complex frame-by-frame masking editing and loop optimization, making it technically demanding and complex to operate. Summary of the Invention
[0005] The purpose of this application is to provide an image generation method, apparatus, electronic device, and storage medium that can reduce the technical threshold and operational complexity of creating dynamic low-light images.
[0006] In a first aspect, embodiments of this application provide an image generation method, the method comprising: Based on the video sequence, a motion vector field is calculated; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence; Based on the amplitude information of the motion vector field, a dynamic region mask is generated; wherein, the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence; Based on the dynamic region mask, a static background and at least one dynamic element are determined from the video sequence; The dynamic elements and the static background are combined to generate a dynamic low-light image.
[0007] Secondly, embodiments of this application provide an image generation apparatus, the apparatus comprising: A calculation module is used to calculate a motion vector field based on a video sequence; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence; The first generation module is used to generate a dynamic region mask based on the amplitude information of the motion vector field; wherein the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence; A determination module is used to determine a static background and at least one dynamic element from the video sequence based on the dynamic region mask; The compositing module is used to compose the dynamic elements and the static background to generate a dynamic low-light image.
[0008] Thirdly, embodiments of this application provide an electronic device including a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the image generation method as described in the first aspect.
[0009] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the image generation method as described in the first aspect.
[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the image generation method as described in the first aspect.
[0011] In this embodiment, a motion vector field is calculated based on a video sequence; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence; a dynamic region mask is generated based on the amplitude information of the motion vector field; wherein the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence; a static background and at least one dynamic element are determined from the video sequence based on the dynamic region mask; the dynamic element and the static background are synthesized to generate a dynamic low-light image.
[0012] As can be seen, in this embodiment, the introduction of pixel-level motion vector field analysis automates and enhances the precision of the dynamic low-light image generation process. Specifically, by calculating the amplitude information of the motion vector field and automatically generating dynamic region masks, dynamic and static elements in a video sequence can be separated efficiently and accurately. This eliminates the need for traditional manual frame-by-frame editing and masking processes, significantly reducing the technical threshold and operational complexity, allowing ordinary users to conveniently and efficiently generate high-quality dynamic low-light images. Furthermore, the motion vector-based analysis method possesses excellent local motion capture capabilities, effectively identifying subtle dynamic changes such as flowing water and drifting leaves, thereby presenting high-quality cyclical dynamic effects in local areas while ensuring the static stability of the main image subject. Attached Figure Description
[0013] Figure 1 This is one of the flowcharts of an image generation method provided in the embodiments of this application; Figure 2 This is an example diagram of an image generation method provided in an embodiment of this application; Figure 3 This is a second flowchart of an image generation method provided in an embodiment of this application; Figure 4 This is the third flowchart of an image generation method provided in the embodiments of this application; Figure 5 This is a structural block diagram of an image generation apparatus provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the hardware structure of an electronic device that implements an embodiment of this application. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0015] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0016] To facilitate understanding, the relevant concepts involved in the embodiments of this application will be introduced first below.
[0017] Cinemagraph is a dynamic image form that combines local animation with static images. Most of the image is still, with only subtle parts (such as water flow or steam) in motion to create a strong visual impact and immersive experience.
[0018] Slow-motion video is a type of video shot at a frame rate higher than normal playback speed, which presents detailed dynamics through slow playback.
[0019] Timelapse photography is a type of video shot or played at a frame rate lower than normal playback frame rate, compressing time to quickly show changes over a long period of time.
[0020] Field of view (FOV) refers to the range of scenery that a lens can cover.
[0021] A media codec is a software or hardware component used to encode and decode audio or video data.
[0022] The image generation method provided in this application will now be described in detail with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0023] It should be noted that the execution subject of the image generation method provided in this application embodiment can be an electronic device or a server, or it can be executed collaboratively by an electronic device and a server. This application embodiment does not limit this.
[0024] Among them, electronic devices include, but are not limited to: smartphones, tablets, smartwatches, personal digital assistants, in-vehicle devices, augmented reality (AR) devices, virtual reality (VR) devices, and other mobile or fixed terminals.
[0025] A server can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0026] In practical applications, the execution flow of the above image generation method can be flexibly deployed according to specific scenario requirements.
[0027] For example, in scenarios where electronic devices process independently, they can complete all steps from video capture to dynamic low-light image generation independently, making them suitable for situations with high requirements for privacy protection or network conditions.
[0028] In cloud processing scenarios, electronic devices are responsible for acquiring and uploading video sequences. The cloud server executes the core computational steps in the image generation method (such as motion vector field calculation, dynamic region mask generation, background and element determination and compositing), and then transmits the generated dynamic low-light image back to the electronic device. This method can make full use of the powerful computing power of the cloud to process more complex video content and is suitable for occasions with high real-time requirements.
[0029] In edge-cloud collaborative scenarios, electronic devices can perform preprocessing (such as frame downsampling and preliminary screening) and final display, while computationally intensive tasks (such as motion vector field calculation, dynamic region mask generation, background and element determination and compositing) are handled by the cloud, thereby achieving the best balance between computing power and power consumption.
[0030] As can be seen, the embodiments of this application not only significantly reduce the technical threshold for the production of dynamic low-light images and achieve convenient and efficient generation, but also provide a technical foundation for diversified deployments on the edge, cloud, and edge-cloud collaboration due to the flexibility of its method architecture, thus expanding the application scenarios of this technology.
[0031] Figure 1 This is one of the flowcharts of an image generation method provided in the embodiments of this application, such as... Figure 1 As shown, the method may include the following steps: step 101, step 102, step 103 and step 104.
[0032] In step 101, a motion vector field is calculated based on the video sequence; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence.
[0033] In this embodiment of the application, the sources of the video sequence include, but are not limited to: videos captured by electronic devices, video streams captured in real time by surveillance cameras, videos already stored in the album, or videos on the Internet.
[0034] For example, such as Figure 2 As shown, the user first clicks on video 2 in the album of the electronic device 20, and then clicks on the "Dynamic Glow" control, which triggers the electronic device 20 to automatically generate a dynamic glow image based on video 2.
[0035] In this embodiment of the application, the video sequence contains multiple consecutive frames of images, which record the state of a specific scene at different times, providing basic data for subsequent processing.
[0036] In some embodiments, to achieve an optimal balance between computational load and transmission efficiency in scenarios where electronic devices and servers collaborate, the raw video data can be preprocessed by downsampling at the electronic device end to generate a video sequence suitable for transmission and processing. The raw video data is acquired by acquisition devices such as cameras and typically has a high frame rate and large data volume. Directly transmitting it to the cloud would consume a large amount of bandwidth and increase processing latency.
[0037] In this embodiment, the electronic device can perform frame reduction sampling on the original video using a built-in multimedia framework (such as MediaCodec in Android or AVAssetReader in iOS). Specifically, one of the following strategies can be adopted: uniform sampling or content-adaptive sampling. Uniform sampling involves extracting frames at fixed intervals (e.g., extracting 1 frame every 2 frames from an original video of 60 frames / second to obtain a 30 frames / second sequence). Content-adaptive sampling dynamically adjusts the sampling rate based on features such as the motion amplitude of objects in the video and the complexity of scene changes, maximizing the amount of compressed data while retaining key motion information.
[0038] In this embodiment, by downsampling at the electronic device, transmission optimization is achieved in the end-to-cloud collaborative architecture. Data compression reduces the dependence on uplink bandwidth, improves transmission rate, and reduces traffic consumption. It also achieves computational load balancing, with the cloud server receiving the optimized video sequence, avoiding the processing of redundant frames, and significantly improving the execution efficiency of core algorithms such as motion analysis and background modeling. Furthermore, it achieves key information preservation, with a content-based sampling strategy ensuring the integrity of representative motion frames, providing high-quality input for accurate separation of dynamic and static elements in the cloud.
[0039] In some implementation scenarios, after the electronic device completes frame downsampling, the optimized video sequence is uploaded to a cloud server for subsequent computationally intensive processing (such as motion vector field analysis and dynamic region mask generation). The resulting dynamic low-light image is then transmitted back to the electronic device for display. This collaborative architecture fully leverages the advantages of cloud computing power while effectively overcoming the limitations of electronic devices in terms of computing resources and transmission conditions, providing a feasible solution for the real-time generation of high-quality dynamic low-light images.
[0040] In this embodiment, a motion vector field is used to quantify the motion state of each pixel in a video sequence between consecutive frames, forming a pixel-level motion information description. Specifically, motion estimation algorithms such as optical flow can be used to calculate and generate displacement vectors for each pixel by analyzing the changes in brightness or color features of pixels between consecutive frames. The motion vector field formed by these displacement vectors can accurately characterize the motion direction and amplitude of each pixel in the video scene, thereby completely describing the microscopic motion morphology of the object.
[0041] In this embodiment, the pixel-level motion analysis method used, compared with motion analysis techniques based on regions or the whole image, can accurately identify complex motion patterns and non-rigid deformation features of object boundary regions in terms of motion detail capture, providing high-precision underlying motion data support for the generation of dynamic region masks, thereby effectively improving the accuracy of distinguishing between dynamic pixels and static pixels.
[0042] In some embodiments, step 101 may specifically include the following step: step 1011.
[0043] In step 1011, for consecutive frames in the video sequence, the motion vector field is obtained by calculating the motion vector of each pixel based on the constant brightness constraint of the optical flow method.
[0044] In this embodiment, the video sequence consists of temporally consecutive image frames, each containing scene visual information at a specific moment. Pixel motion estimation is performed using the optical flow method, which is based on the assumption of constant brightness constraint, that is, assuming that the brightness values of the surface pixels of an object remain stable as it moves between consecutive frames.
[0045] In practice, two adjacent frames are selected from the video sequence. For each pixel in the first frame, the target position corresponding to its brightness characteristics is found in the next frame using an established mathematical model and optimization algorithm. The spatial displacement between the two points constitutes the motion vector of that pixel, containing horizontal and vertical motion components. By performing this calculation on all pixels in the image, a motion vector field that accurately describes the pixel-level motion patterns between frames is finally generated.
[0046] For example, the Lucas-Kanade algorithm is used to calculate optical flow in the current frame. and the next frame Convert to grayscale image, for a given pixel. The brightness value in the frame at time t In the next frame (time) The corresponding pixel position in ) brightness value Still the same, we get the following formula (1): (1) By approximating with Taylor expansion, we obtain the following formula (2): (2) in, , , Let x, y, and time be the partial derivatives of the image, respectively. The optical flow equation is as follows: (3) (3) Assuming that all pixels within a small local window (e.g., 5×5 pixels) have the same optical flow (u, v), then for each pixel within the window, an optical flow constraint equation of formula (3) can be written. A 5×5 window would have 25 equations, but they all share the same u and v, forming an overdetermined system of equations (the number of equations exceeds the number of unknowns). Optimization techniques such as least squares can be used to solve for the most suitable (u, v), minimizing the overall error of the assumption that the brightness of all pixels within the window is constant.
[0047] The final calculated vector (u, v) is the motion vector, and the magnitude of the motion vector is the standard used to determine whether the pixel is moving.
[0048] In this embodiment, an optical flow method based on constant brightness constraints is used to calculate the motion vector of each pixel, achieving pixel-level precise analysis of object motion. Compared to motion analysis methods based on regions or the entire image, this method can more accurately identify complex motion patterns such as object boundary motion and local deformation. For example, in video sequences containing rotating objects, traditional region analysis methods can only estimate the overall rotation direction and approximate angle of the object, while this embodiment can accurately depict the motion trajectory of each pixel, thus fully restoring the rotational details of the object. Furthermore, in complex scenes involving multiple objects occluding each other or high-speed motion, the method can still effectively calculate the motion vector of each pixel. Even if the object is partially occluded during motion, the optical flow method can effectively estimate the motion vector of the occluded area using motion information from the area before and around the occlusion, thereby ensuring the continuity and integrity of the motion vector field.
[0049] In this embodiment, the constant brightness constraint can adapt to different lighting conditions to a certain extent. In real-world scenarios, lighting may change, but as long as the change is not drastic enough to cause a sudden change in pixel brightness, the optical flow method can calculate the pixel's motion vector based on the constant brightness constraint. This allows for effective calculation of the motion vector field under different lighting environments, such as changes in indoor lighting or day-night cycles outdoors, demonstrating good robustness. Furthermore, the optical flow method can calculate pixel motion for both rigid objects (such as cars and robots) and non-rigid objects (such as people and animals). For rigid objects, the motion vector field can reflect the object's overall translation and rotation; for non-rigid objects, such as the limb movements of a person, the optical flow method can capture the independent movements of different parts of the limb, thus providing accurate motion information for different types of video scenes.
[0050] In this embodiment, the amplitude information in the motion vector field reflects the intensity of pixel motion. Based on this amplitude information, a dynamic region mask can be generated to distinguish between dynamic and static pixels in a video sequence. An accurate motion vector field enables the dynamic region mask to more precisely identify dynamic regions, avoiding misclassification of static pixels as dynamic pixels or omission of dynamic pixels, thus providing a reliable foundation for subsequently determining static backgrounds and dynamic elements from the video sequence. Furthermore, when compositing dynamic elements with a static background, the motion vector field can provide the motion trajectory information of the dynamic elements. With this information, dynamic elements can be more naturally composited into the static background, making the motion of dynamic elements in the composite dynamic low-light image smoother and more realistic, enhancing the image imaging effect.
[0051] In step 102, a dynamic region mask is generated based on the amplitude information of the motion vector field; wherein, the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence.
[0052] In this embodiment, a dynamic region mask is a binary image used to identify dynamic and static regions in a video sequence. In the fields of image processing and computer vision, it highlights the dynamic regions that are in motion by dividing the image into a set of dynamic pixels (i.e., dynamic regions) and a set of static pixels (i.e., static regions) in the form of a mask (usually represented by 0 and 1, where 0 represents static regions and 1 represents dynamic regions).
[0053] In this embodiment, each motion vector in the motion vector field describes the positional change of a pixel or region in an image between different frames. In a video sequence, when an object moves, the position of its corresponding pixel changes in consecutive frames. By calculating the positional difference of corresponding pixels between adjacent frames, the motion vector can be obtained. For example, in a video of a vehicle moving, the pixels on the vehicle will move in the direction of the vehicle's movement in the next frame, and the calculated motion vector reflects the direction and distance of this movement.
[0054] In this embodiment, each motion vector in the motion vector field has a corresponding amplitude, which reflects the intensity of pixel motion. The motion of dynamic objects (such as moving vehicles or walking people) in a video results in a larger motion vector amplitude (hereinafter referred to as "motion amplitude") for their pixels; while static objects (such as buildings or trees in the background) have pixel motion vector amplitudes close to zero due to the lack of significant positional changes. Therefore, by analyzing the amplitude information of the motion vector field, it is possible to determine which regions in the image contain dynamic elements and which regions are static, thereby generating a binary dynamic region mask. This mask acts like a "mask," clearly identifying which regions in the video are dynamically changing and which regions are relatively static.
[0055] In this embodiment, the dynamic region mask simplifies a complex video sequence into two distinct regions: dynamic and static. This distinction facilitates targeted processing of dynamic elements and static backgrounds, improving the efficiency and quality of image synthesis. For example, when synthesizing dynamic low-light images, different processing strategies can be applied to the dynamic and static regions to achieve better imaging results.
[0056] In the edge-cloud collaborative implementation plan, this step can be deployed on a cloud server, utilizing its computing resources to complete parallel amplitude analysis and threshold determination for a large number of pixels, ensuring the efficiency and accuracy of mask generation. The generated mask data will serve as a key intermediate result, supporting subsequent static background reconstruction and dynamic element extraction processes.
[0057] In step 103, a static background and at least one dynamic element are determined from the video sequence based on a dynamic region mask.
[0058] In this embodiment, the dynamic element is usually a loopable dynamic element, which will be referred to as "loopable dynamic element" for ease of description.
[0059] In this embodiment, after generating a dynamic region mask for multiple frames of images in a video sequence, the pixel regions marked as 0 (static) in the mask are observed. If certain pixel regions are consistently marked as 0 across multiple consecutive frames, it indicates that these regions have not undergone significant motion in the video, and can be preliminarily identified as static background regions. Furthermore, the preliminarily identified static background regions can be further optimized, for example, by using background modeling or statistical calculations, to obtain the final static background.
[0060] In this embodiment, for pixel regions marked as 1 (dynamic) in the dynamic region mask, the motion trajectory of these regions is tracked in consecutive frames. If the motion trajectory of a dynamic element exhibits periodic changes, that is, it repeats the same motion path and shape after a certain time interval, then the element can be determined to be a cyclical dynamic element. For example, in a video sequence of a rotating fan, the motion trajectory of the fan blades can be tracked using the dynamic region mask, and it can be found that the blades return to the same position and angle at certain intervals, thus determining that the fan blades are a cyclical dynamic element.
[0061] In addition, features such as shape, color, and texture of dynamic elements can be extracted and feature matching performed across consecutive frames. If the features of a dynamic element match highly across different frames and its motion pattern exhibits periodicity, then the element is determined to be a cyclical dynamic element. For example, in an animated video, the shape and color features of a running puppy remain largely unchanged across different frames, and its running motion exhibits periodic repetition. Feature matching can determine that the puppy is a cyclical dynamic element.
[0062] In some embodiments, step 103 may specifically include the following steps: step 1031 and step 1032.
[0063] In step 1031, the pixel values of the pixels marked as static pixels by the dynamic region mask in multiple frames of the video sequence are statistically calculated to generate a static background.
[0064] In this embodiment, a dynamic region mask can be used to accurately locate pixel regions that maintain a stable relative position and shape in a video sequence. These regions typically correspond to static backgrounds in the video, such as walls and fixed furniture in indoor scenes, and buildings and mountains in outdoor scenes.
[0065] In some embodiments, the above statistical calculation may be an average value.
[0066] In this embodiment, the pixel values of corresponding static pixels in multiple frames of images can be added together and divided by the number of frames. For example, in five consecutive frames of images, if the pixel values of a certain static pixel position in different frames are 100, 102, 98, 101, and 99 respectively, then the average value of this static pixel is (100+102+98+101+99) / 5=100. By performing such an average calculation on a large number of static pixels, a relatively smooth set of pixel values that can represent the characteristics of a static background can be obtained, thereby generating a static background.
[0067] In some embodiments, the above statistical calculation may be the median value.
[0068] In this embodiment, the pixel values of corresponding static pixels in multiple frames of images can be arranged in ascending order, and the median value can be taken as the representative value of the static pixel. For example, if the above 5 pixel values are sorted as 98, 99, 100, 101, and 102, then the median value is 100. Using the median value can avoid the influence of some extreme pixel values (such as excessively high or low pixel values due to noise) on the generation of static background, making the generated static background more stable.
[0069] As can be seen, in this embodiment of the application, statistical calculations of multiple frames of images can reduce the impact of noise, interference, and other factors that may exist in a single frame image on the static background. For example, in a single frame image, the pixel values of some static pixels may be abnormal due to sudden changes in light or sensor noise. By taking the average or median value, the influence of these abnormal values can be weakened, thereby generating a more accurate and stable static background.
[0070] In step 1032, the starting frame and loop termination frame for the first appearance of the dynamic pixel are determined based on the change information of the dynamic region mask on the time axis; based on the starting frame and loop termination frame, video segments are extracted from the video sequence as dynamic elements.
[0071] In this embodiment, the dynamic region mask changes across different frames, reflecting the appearance, movement, and disappearance of dynamic elements in the video. By analyzing the changes in the dynamic region mask over time, the frame in which the dynamic element first appears (the start frame) can be identified; and the frame in which the dynamic element completes a relatively complete motion cycle and seamlessly connects with other parts (the loop termination frame). For example, in a video of a rotating fan, the frame in which the fan blades begin to rotate and return to a similar position and angle after completing one rotation can be used as the loop termination frame.
[0072] In this embodiment, based on the start frame and the loop termination frame, a corresponding video segment is extracted from the video sequence. This extracted video segment contains loopable dynamic elements, such as the video segment of the fan blades rotating once, which can be looped when needed without causing obvious visual discontinuity.
[0073] In this embodiment, after extracting a video segment containing recurring dynamic elements, only this relatively short video segment needs to be stored. It can then be played in a loop to display longer dynamic effects when needed. This saves storage space compared to storing the entire video containing repetitive dynamic content.
[0074] In this embodiment, by reasonably looping these dynamic elements, a unique visual rhythm and atmosphere can be created. For example, in animation production, looping dynamic elements of natural phenomena, such as falling snowflakes or flowing streams, can enhance the realism and interest of the animation.
[0075] In this embodiment of the application, during the video editing process, the loopable dynamic elements can be easily reused and adjusted. Users can insert, delete, or modify the loop count and playback speed of these dynamic elements at any time as needed, without having to reprocess a large amount of video data.
[0076] As can be seen, in this embodiment, a dynamic region mask can be used to filter static pixels from a video sequence. A static background is generated by statistically analyzing (e.g., taking the average or median value) multiple frames of static pixels. Simultaneously, based on the changes in the dynamic region mask along the time axis, the starting frame and loop termination frame of the first appearance of dynamic elements are determined, thereby extracting recurring dynamic elements from the video sequence. These dynamic elements can be repetitive actions, object movements, etc., in the video. The determined static background and recurring dynamic elements provide important material for subsequent image synthesis. The static background, as the foundation of the image, provides a stable display environment for the dynamic elements; while the recurring dynamic elements can be reused, reducing data volume and improving the flexibility and efficiency of dynamic low-light image generation.
[0077] In step 104, the dynamic elements and the static background are combined to generate a dynamic low-light image.
[0078] In this embodiment, the regions of dynamic elements that are masked as dynamic pixels can be superimposed onto the corresponding positions of the static background according to certain rules (such as transparency blending). The transparency blending can adjust the transparency of the dynamic elements as needed to make the transition between them and the static background smoother, enhancing the realism of the image. Furthermore, the output format of the dynamic low-light image can be set, such as generating a seamless looping dynamic image in GIF or WebP format.
[0079] In this embodiment, dynamic low-light images are generated through synthesis processing, combining the stability of a static background with the vividness of dynamic elements to create artistically compelling dynamic low-light images that blend stillness and motion. These dynamic low-light images can be viewed directly by users and applied to various scenarios, such as virtual reality, game development, and film special effects, providing users with a more immersive and vivid visual experience. Furthermore, this synthesis method offers high flexibility and scalability, allowing for adjustments to the combination of dynamic elements and static backgrounds to generate diverse image effects according to different needs.
[0080] As can be seen from the above embodiments, in this embodiment, a motion vector field is calculated based on the video sequence; wherein, the motion vector field contains pixel-level motion information of each image frame in the video sequence; a dynamic region mask is generated based on the amplitude information of the motion vector field; wherein, the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence; based on the dynamic region mask, a static background and at least one dynamic element are determined from the video sequence; the dynamic element and the static background are synthesized to generate a dynamic low-light image.
[0081] As can be seen, in this embodiment, the introduction of pixel-level motion vector field analysis automates and enhances the precision of the dynamic low-light image generation process. Specifically, by calculating the amplitude information of the motion vector field and automatically generating dynamic region masks, dynamic and static elements in a video sequence can be separated efficiently and accurately. This eliminates the need for traditional manual frame-by-frame editing and masking processes, significantly reducing the technical threshold and operational complexity, allowing ordinary users to conveniently and efficiently generate high-quality dynamic low-light images. Furthermore, the motion vector-based analysis method possesses excellent local motion capture capabilities, effectively identifying subtle dynamic changes such as flowing water and drifting leaves, thereby presenting high-quality cyclical dynamic effects in local areas while ensuring the static stability of the main image subject.
[0082] Figure 3 This is a second flowchart of an image generation method provided in an embodiment of this application, such as... Figure 3 As shown, the method may include the following steps: step 301, step 302, step 303 and step 304.
[0083] In step 301, a motion vector field is calculated based on the video sequence; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence.
[0084] The content of step 301 in this embodiment of the application is the same as... Figure 1 The content of step 101 in the illustrated embodiment is similar and will not be repeated here.
[0085] In step 302, the motion amplitude of each pixel in the video sequence is determined based on the amplitude information of the motion vector field; pixels with motion amplitude greater than or equal to the motion threshold are marked as dynamic pixels, and pixels with motion amplitude less than the motion threshold are marked as static pixels; a binary dynamic region mask is generated based on the marking results of all pixels; wherein, the motion threshold is determined by optimizing the motion detection effect on the video dataset.
[0086] In this embodiment, the motion vector field describes the motion of objects or pixels in a video sequence between different frames. In video processing, each frame can be considered as being composed of numerous pixels, and the motion vector reflects the direction and magnitude of the positional change of each pixel from the current frame to the next frame. The amplitude of the motion vector is an indicator used to quantify the magnitude of this positional change; the larger the amplitude value, the more intense the pixel's motion between frames; the smaller the amplitude value, the weaker or essentially stationary the pixel's motion.
[0087] In this embodiment, the amplitude of each motion vector in the motion vector field is examined and compared with a pre-set motion threshold. This motion threshold is not arbitrarily determined, but rather determined through repeated optimization of motion detection performance on a large number of video datasets.
[0088] If the motion amplitude of a pixel is greater than or equal to the motion threshold, it indicates that the object or region corresponding to that pixel has significant motion between frames, and it is marked as a dynamic pixel; conversely, if the motion amplitude is less than the motion threshold, it indicates that the object or region corresponding to that pixel does not move significantly or is in a static state, and it is marked as a static pixel.
[0089] In this embodiment, a binary dynamic region mask is constructed based on the labeling results of all pixels. In this mask, dynamic pixels and static pixels are represented by different values, for example, 1 represents a dynamic pixel and 0 represents a static pixel. In this way, the entire video frame is transformed into a binary image with only two values, 0 and 1, where the region containing 1 corresponds to the dynamic part of the video and the region containing 0 corresponds to the static part, thus clearly distinguishing the dynamic and static regions in the video.
[0090] In this embodiment, by setting a reasonable motion threshold and comparing it with the motion amplitude of pixels, it is possible to accurately distinguish between moving pixels and relatively stationary pixels in a video. For example, in a video of a street scene, the pixels corresponding to moving vehicles have a larger motion amplitude and will be marked as dynamic pixels; while the pixels corresponding to stationary objects such as buildings and trees on both sides of the road have a smaller motion amplitude and will be marked as static pixels. This avoids misjudging some minor, non-target motions (such as pixel changes caused by slight camera shake) as dynamic areas, thus improving the accuracy of motion detection.
[0091] Furthermore, since the motion threshold is optimized on a video dataset, it takes into account the characteristics of different video scenes. For example, the speed and amplitude of object movement differ greatly in fast-moving sports videos and slowly changing surveillance videos. The optimized motion threshold can automatically adapt to different scenes, ensuring relatively accurate detection of dynamic regions in various types of videos.
[0092] In this embodiment, considering the inevitable impact of various noises during actual video acquisition, such as sensor noise and environmental interference, these noises may cause minor pixel changes. However, through motion thresholding, only pixels whose motion amplitude truly exceeds the threshold are marked as dynamic pixels. This effectively filters out false motion caused by noise, enhancing the algorithm's robustness to noise. Furthermore, videos from different sources may vary significantly in quality, such as resolution, frame rate, and sharpness. The motion threshold, optimized on various video datasets, can adapt to videos of different qualities. Even in low-quality, noisy videos, it can accurately detect dynamic regions, ensuring the algorithm's stability and reliability.
[0093] In step 303, a static background and at least one dynamic element are determined from the video sequence based on a dynamic region mask.
[0094] In step 304, the dynamic elements and the static background are composited to generate a dynamic low-light image. The content of steps 303 and 304 in the embodiments of this application is the same as that of... Figure 1 Steps 103 and 104 in the illustrated embodiment are similar and will not be repeated here.
[0095] As can be seen, in this embodiment, by calculating pixel-level motion vector fields and automatically generating dynamic region masks based on their amplitude information, precise and automated separation of moving and static elements in a video sequence is achieved. This eliminates the need for manual frame-by-frame editing and mask drawing, reducing the technical threshold and operational complexity of creating dynamic low-light images and enabling convenient and efficient generation. Simultaneously, the motion vector-based analysis method effectively captures subtle, localized movements, ensuring that the generated dynamic low-light image maintains the static main subject while presenting high-quality cyclical dynamic effects in local areas, thus improving the imaging effect of the dynamic low-light image. Furthermore, the process of generating the dynamic region mask is based solely on its amplitude information, making the calculation process relatively simple, with lower hardware resource requirements and computation time, thus meeting real-time requirements.
[0096] Figure 4 This is the third flowchart of an image generation method provided in the embodiments of this application, such as... Figure 4 As shown, the method may include the following steps: step 401, step 402, step 403, step 404, step 405 and step 406.
[0097] In step 401, a motion vector field is calculated based on the video sequence; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence.
[0098] The content of step 401 in this embodiment of the application is the same as... Figure 1 The content of step 101 in the illustrated embodiment is similar and will not be repeated here.
[0099] In step 402, an initial background reference image is generated; wherein, the pixel value of each pixel position in the initial background reference image is the average color value of the corresponding pixel position in the first N frames of the video sequence, and N is an integer greater than 1.
[0100] In this embodiment, the background reference image is reference data used to characterize the static baseline of the scene. The initial background reference image is generated as follows: select the initial N consecutive frames (N>1) of the video sequence, calculate the arithmetic mean of the color values of each pixel position in the N frames, and assign the average value as the pixel value of the corresponding position in the background reference image.
[0101] For example, the initial background reference image can be represented by the following formula (4): (4) in, Let be the color value at pixel position (x, y) in the i-th frame, where 1 ≤ i ≤ N.
[0102] In an edge-cloud collaborative architecture, this step can be flexibly deployed on either the electronic device or the cloud: initial calculations on the electronic device reduce data transmission, while execution in the cloud leverages powerful computing capabilities for more complex background modeling and optimization. This method of establishing an initial background reference image provides a stable and reliable static benchmark for the entire dynamic low-light image generation process, and is a key technological foundation for ensuring the final synthesis quality.
[0103] In step 403, for each frame after the first N frames in the video sequence, the following steps are performed: calculate the color difference value of each pixel between the current frame and the background reference image at the previous moment; determine the motion amplitude of each pixel in the current frame based on the motion vector field of the current frame; mark pixels with color difference values less than the difference threshold and motion amplitudes less than the motion threshold as static pixels, and mark the remaining pixels as dynamic pixels; wherein, the difference threshold is determined by optimizing the background modeling effect on the video dataset; the motion threshold is determined by optimizing the motion detection effect on the video dataset.
[0104] In this embodiment, dynamic pixel recognition is performed on each video frame after the first N frames. A dual criterion verification mechanism of color difference test and motion intensity analysis is used to achieve accurate separation of dynamic and static pixels.
[0105] In this embodiment, pixel-level color consistency is determined through color difference testing. Specifically, the difference value between each pixel in the current frame and the background reference image at the previous moment in the color space is calculated (preferably using Euclidean distance to measure the vector distance in color spaces such as RGB). When the difference value is lower than the difference threshold determined by optimization through a large-scale video dataset, the current pixel is determined to meet the color consistency condition with the background reference image.
[0106] In this embodiment, the color difference detection mechanism has three technical advantages: it uses a vector space distance metric, which can accurately reflect the comprehensive differences in multidimensional color features; the threshold setting based on dataset optimization can ensure adaptability to different scene lighting conditions; and it forms a complementary criterion with the motion amplitude detection, which can jointly construct a robust motion-static separation benchmark.
[0107] In this embodiment, pixel-level motion state analysis is achieved through motion amplitude testing. Specifically, the motion amplitude characteristics of each pixel are analyzed based on the motion vector field of the current frame. When the amplitude value is lower than the motion threshold determined by optimization using a multi-scene video dataset, the pixel is determined to meet the static condition.
[0108] In this embodiment, the motion detection mechanism also has three technical advantages: based on direct measurement of physical motion characteristics, it can effectively distinguish between real motion and apparent changes; through dataset-driven threshold optimization, it can adapt to the motion sensitivity requirements of different scenarios; and by forming an orthogonal criterion with the color detection mechanism, it can construct a multi-dimensional decision space.
[0109] In this embodiment, the final classification logic employs a dual-criteria fusion strategy: pixels that pass both the color difference test and the motion amplitude test are classified as static pixels, while pixels that fail either test are classified as dynamic pixels. This decision mechanism significantly improves the accuracy and robustness of static / dynamic region segmentation through complementary verification of color and motion features.
[0110] In this embodiment, the difference threshold and motion threshold are determined through optimization on diverse video datasets to ensure optimal performance in the following scenarios: eliminating interference from minor environmental changes (such as lighting fluctuations and sensor noise), preserving the complete motion characteristics of real dynamic targets, and balancing the dialectical relationship between false positive and false negative rates. For example, in a video monitoring a warehouse, there may be some minor color changes (such as dust drifting) and minor movements (such as slight shaking of goods). By setting reasonable thresholds, these situations can be avoided from being misjudged as important dynamic events.
[0111] In edge-cloud collaborative processing scenarios, the aforementioned verification and analysis steps can be flexibly configured on electronic devices or in the cloud: local execution on electronic devices can reduce data transmission, while cloud processing can utilize rich scene data to achieve more accurate threshold adaptation, improve accuracy, and provide a reliable motion perception basis for the generation of dynamic low-light images.
[0112] In this embodiment, a robust dynamic pixel recognition mechanism is constructed by fusing color and motion information, effectively overcoming the limitations of single feature detection. For apparent changes such as sudden changes in illumination and shadow interference, motion feature verification effectively avoids misjudgment. For example, in an indoor scene with flickering lights, although the color features change significantly, motion amplitude verification can confirm that the area is actually stationary, thus maintaining a correct background judgment. In dynamic scenes with multiple intertwined targets (such as traffic intersections), dual criteria can achieve accurate target separation. For example, for real moving targets such as vehicles and pedestrians, color and motion features are activated simultaneously. For semi-static elements such as swaying trees, motion features are moderately triggered while color features remain stable. For static backgrounds such as roads and buildings, both features remain stationary. This fine distinction ensures the integrity of dynamic element extraction.
[0113] As can be seen, in the embodiments of this application, the complementary characteristics of color and motion features in response to noise can be utilized: sensor noise mainly affects the color space, but the motion vector remains stable; instantaneous motion interference may disturb the local motion field, but the color features remain continuous; through the dual verification mechanism, most random noise interference can be filtered out, significantly improving the signal-to-noise ratio of the output.
[0114] In step 404, a binary dynamic region mask is generated based on the labeling results of all pixels.
[0115] In this embodiment, dynamic pixels and static pixels in the dynamic region mask are represented by different values (e.g., 1 represents dynamic pixels and 0 represents static pixels), thereby clearly dividing the dynamic and static regions in the video.
[0116] In step 405, a static background and at least one dynamic element are determined from the video sequence based on a dynamic region mask.
[0117] In step 406, the dynamic elements and the static background are combined to generate a dynamic low-light image.
[0118] The content of steps 405 and 406 in the embodiments of this application is the same as that of... Figure 1 Steps 103 and 104 in the illustrated embodiment are similar and will not be repeated here.
[0119] As can be seen, in this embodiment, by calculating pixel-level motion vector fields and automatically generating dynamic region masks based on their amplitude information and background reference images, precise and automated separation of moving and static elements in video sequences is achieved. This eliminates the need for manual frame-by-frame editing and mask drawing, reducing the production threshold and operational complexity of dynamic low-light images, and enabling convenient and efficient generation of dynamic low-light images. Simultaneously, the motion vector-based analysis method effectively captures subtle and localized movements, ensuring that the generated dynamic low-light image maintains the static main subject while presenting high-quality cyclical dynamic effects in local areas, thus improving the imaging effect of the dynamic low-light image.
[0120] In some embodiments provided in this application, the provided image generation method does not treat the background reference image as fixed, but allows it to be continuously adjusted and optimized as the video sequence progresses, so as to better adapt to the slow changes that may occur in the background in the video scene. Accordingly, after step 403 above, the following step can be added: step 407.
[0121] In step 407, the pixel value of each static pixel in the current frame is weighted and fused with the pixel value of the corresponding pixel position in the background reference image at the previous time step; based on the weighted fusion result, an updated background reference image is generated; wherein, the fusion weight of the pixel value of the static pixel is 'a', and the fusion weight of the pixel value of the corresponding pixel position in the background reference image at the previous time step is (1-a), where a is the learning rate and 0. <a<1。
[0122] In this embodiment of the application, the updated background reference image can be generated using the following formula (5): (5) in, The updated background reference image for the current time / frame (i.e., time t) is a new background representation obtained by fusing the background reference image of the previous time (i.e., time t-1) with the static pixel values of the current frame. The pixel value of the current frame (i.e., time t) specifically refers to the set of pixel values marked as static pixels in the current frame. That is, only the color values and other information of the pixels that were determined to be static in step 403 are selected to participate in the update of the background reference image. This is the background reference image from the previous time step (i.e., time t-1). It serves as the basis for this update and carries the accumulated knowledge of the background from all previous frames. 'a' represents the learning rate, a key control parameter whose value is limited to between 0 and 1. It is used to adjust the contribution of the static pixel values of the current frame to the update of the background reference image. Increasing the value of 'a' will increase the weight of the current frame data in the background update, thereby accelerating the adaptive adjustment of the background reference image; conversely, a smaller value of 'a' will enhance the persistence of the historical background reference image, making the background reference image update process smoother.
[0123] To address the common phenomenon of gradually changing illumination in real-world video scenarios (such as changes in outdoor sunlight duration or attenuation of indoor light sources), this embodiment employs the aforementioned update mechanism to dynamically calibrate the background reference image. When ambient light intensifies, the corresponding pixel values in the background reference image gradually adjust as the brightness of the current frame increases, thereby maintaining illumination consistency between the background reference image and the actual scene. Furthermore, for microscopic movements or deformations of background objects (such as swaying vegetation or thermal deformation of building materials), the background reference image can be continuously optimized based on static pixel information representing such changes in the current frame. Taking swaying tree branches as an example, subtle changes in their position and shape are gradually integrated into the background reference image through iterative updates, ensuring real-time consistency between the background reference image and the scene state.
[0124] If the background reference image lacks an update mechanism, the difference between the initial background reference image and the actual scene will accumulate over time, causing areas that should belong to the background to be misidentified as dynamic targets due to slow changes. Therefore, in this embodiment, the gradual updating of the background reference image effectively suppresses false detections caused by such slowly changing factors.
[0125] Considering that an accurate background reference image is a prerequisite for achieving precise dynamic region segmentation, in this embodiment, the adaptively updated background reference image can effectively distinguish between the inherent changes of the real moving target and the background when compared with each image frame. Taking indoor person detection as an example, when the background reference image has been adapted to the slight movement of the curtains, the system can accurately separate the person's movement area, avoiding the dynamic background being mistakenly classified as the foreground target, thereby improving the accuracy of dynamic region segmentation.
[0126] To adapt to the different characteristics of various scenarios, this embodiment allows for flexible configuration of the background reference image update strategy by adjusting the learning rate 'a'. Specifically, in highly dynamic scenarios (such as traffic monitoring), appropriately increasing the value of 'a' (e.g., setting a=0.05) improves the response speed of the background reference image to environmental changes, ensuring that the background can quickly recover to its true state after the moving target leaves. In relatively stable scenarios (such as indoor security), decreasing the value of 'a' (e.g., setting a=0.005) enhances the historical continuity of the background reference image, effectively suppressing background fluctuations caused by brief interference. Simultaneously, this update mechanism has inherent noise suppression characteristics: when a single frame image experiences pixel value anomalies due to sensor noise, a smaller 'a' value ensures that the historical background reference image dominates the weighted fusion, achieving smooth filtering of abnormal fluctuations and significantly improving the anti-interference capability of the background reference image. This adaptive update strategy based on learning rate adjustment enables the background reference image to maintain the optimal balance of accuracy, stability, and real-time performance in different application scenarios.
[0127] In summary, in this embodiment of the application, by fusing motion vector amplitude analysis and background difference detection as dual criteria, and combining them with a motion-weighted adaptive background update strategy, the accuracy and robustness of dynamic element extraction in complex scenes such as gradual illumination changes and micro-motions are improved.
[0128] The image generation method provided in this application can be executed by an image generation device. This application uses an image generation device executing the image generation method as an example to illustrate the image generation device provided in this application.
[0129] Figure 5 This is a structural block diagram of an image generation apparatus provided in an embodiment of this application, such as... Figure 5As shown, the image generation device 500 may include: a calculation module 501, a first generation module 502, a determination module 503, and a synthesis module 504; The calculation module 501 is used to calculate a motion vector field based on a video sequence; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence; The first generation module 502 is used to generate a dynamic region mask based on the amplitude information of the motion vector field; wherein, the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence; The determining module 503 is used to determine a static background and at least one dynamic element from the video sequence based on the dynamic region mask. The compositing module 504 is used to perform compositing processing on the dynamic elements and the static background to generate a dynamic low-light image.
[0130] As can be seen from the above embodiments, this embodiment achieves automation and precision in the dynamic low-light image generation process by introducing pixel-level motion vector field analysis. Specifically, by calculating the amplitude information of the motion vector field and automatically generating dynamic region masks, dynamic and static elements in a video sequence can be separated efficiently and accurately. This eliminates the need for traditional manual frame-by-frame editing and mask drawing processes, significantly reducing the technical threshold and operational complexity, allowing ordinary users to conveniently and efficiently generate high-quality dynamic low-light images. Simultaneously, the motion vector-based analysis method possesses excellent local motion capture capabilities, effectively identifying subtle dynamic changes such as flowing water and drifting leaves, thereby presenting high-quality cyclical dynamic effects in local areas while ensuring the static stability of the main image subject.
[0131] Optionally, as an embodiment, the first generation module 502 is specifically used to determine the motion amplitude of each pixel in the video sequence based on the amplitude information of the motion vector field; mark pixels with motion amplitude greater than or equal to a motion threshold as dynamic pixels, and mark pixels with motion amplitude less than the motion threshold as static pixels; wherein, the motion threshold is determined by optimizing the motion detection effect on the video dataset; and generate a binary dynamic region mask based on the marking results of all pixels.
[0132] Optionally, as an embodiment, the image generation apparatus 500 may further include: The second generation module is used to generate an initial background reference image; wherein, the pixel value of each pixel position in the initial background reference image is the average color value of the corresponding pixel position in the first N frames of the video sequence, and N is an integer greater than 1; The first generation module 502 is specifically configured to perform the following steps for each frame in the video sequence after the first N frames: calculate the color difference value of each pixel between the current frame and the background reference image at the previous moment; determine the motion amplitude of each pixel in the current frame based on the motion vector field of the current frame; mark pixels whose color difference value is less than a difference threshold and whose motion amplitude is less than a motion threshold as static pixels, and mark the remaining pixels as dynamic pixels; wherein, the difference threshold is determined by optimizing the background modeling effect on the video dataset; the motion threshold is determined by optimizing the motion detection effect on the video dataset; and generate a binary dynamic region mask based on the marking results of all pixels.
[0133] Optionally, as an embodiment, the image generation apparatus 500 may further include: The fusion module is used to perform weighted fusion of the pixel value of each static pixel in the current frame with the pixel value of the corresponding pixel position in the background reference image of the previous time step; wherein, the fusion weight of the pixel value of the static pixel is 'a', and the fusion weight of the pixel value of the corresponding pixel position in the background reference image of the previous time step is (1-a), where a is the learning rate and 0. <a<1; The update module is used to generate an updated background reference image based on the weighted fusion result.
[0134] Optionally, as an embodiment, the determining module 503 is specifically used to perform statistical calculations on the pixel values of pixels marked as static pixels by the dynamic region mask within multiple frames of the video sequence, so as to generate a static background.
[0135] Optionally, as an embodiment, the determining module 503 is specifically used to determine the starting frame and the loop termination frame of the first appearance of the dynamic pixel based on the change information of the dynamic region mask on the time axis; and to extract video segments from the video sequence as dynamic elements based on the starting frame and the loop termination frame.
[0136] Optionally, as an embodiment, the compositing module 504 is specifically used to superimpose the regions of the dynamic elements that are marked as dynamic pixels by the dynamic region mask onto the corresponding positions of the static background.
[0137] Optionally, as an embodiment, the calculation module 501 is specifically used to calculate the motion vector field of each pixel for consecutive frames in a video sequence based on the constant brightness constraint of the optical flow method.
[0138] The image generation device in this application embodiment can be a device, or a component, integrated circuit, or chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This application embodiment does not impose specific limitations.
[0139] The image generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.
[0140] The image generation apparatus provided in this application embodiment can achieve... Figure 1 , Figure 3 or Figure 4 To avoid repetition, the various processes implemented in the method embodiment shown will not be described again here.
[0141] Optionally, such as Figure 6 As shown, this application embodiment also provides an electronic device 600, including a processor 601, a memory 602, and a program or instructions stored in the memory 602 and executable on the processor 601. When the program or instructions are executed by the processor 601, they implement the various processes of the above-described image generation method embodiment and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0142] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0143] Figure 7 This is a schematic diagram of the hardware structure of an electronic device that implements an embodiment of this application.
[0144] The electronic device 700 includes, but is not limited to, components such as: radio frequency unit 701, network module 702, audio output unit 703, input unit 704, sensor 705, display unit 706, user input unit 707, interface unit 708, memory 709, and processor 710.
[0145] Those skilled in the art will understand that the electronic device 700 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 710 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0146] The processor 710 is configured to: calculate a motion vector field based on a video sequence; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence; generate a dynamic region mask based on the amplitude information of the motion vector field; wherein the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence; determine a static background and at least one dynamic element from the video sequence based on the dynamic region mask; and perform a composite processing on the dynamic element and the static background to generate a dynamic low-light image.
[0147] As can be seen, in this embodiment, the introduction of pixel-level motion vector field analysis automates and enhances the precision of the dynamic low-light image generation process. Specifically, by calculating the amplitude information of the motion vector field and automatically generating dynamic region masks, dynamic and static elements in a video sequence can be separated efficiently and accurately. This eliminates the need for traditional manual frame-by-frame editing and masking processes, significantly reducing the technical threshold and operational complexity, allowing ordinary users to conveniently and efficiently generate high-quality dynamic low-light images. Furthermore, the motion vector-based analysis method possesses excellent local motion capture capabilities, effectively identifying subtle dynamic changes such as flowing water and drifting leaves, thereby presenting high-quality cyclical dynamic effects in local areas while ensuring the static stability of the main image subject.
[0148] Optionally, as an embodiment, the processor 710 is specifically configured to determine the motion amplitude of each pixel in the video sequence based on the amplitude information of the motion vector field; mark pixels with motion amplitude greater than or equal to a motion threshold as dynamic pixels, and mark pixels with motion amplitude less than the motion threshold as static pixels; wherein the motion threshold is determined by optimizing the motion detection effect on the video dataset; and generate a binary dynamic region mask based on the marking results of all pixels.
[0149] Optionally, as an embodiment, the processor 710 is specifically configured to generate an initial background reference image; wherein, the pixel value at each pixel position in the initial background reference image is the color average value of the corresponding pixel positions in the first N frames of the video sequence, and N is an integer greater than 1; for each frame in the video sequence after the first N frames, the following steps are performed: calculate the color difference value of each pixel between the current frame and the background reference image at the previous moment; determine the motion amplitude of each pixel in the current frame according to the motion vector field of the current frame; mark the pixels with the color difference value less than the difference threshold and the motion amplitude less than the motion threshold as static pixels, and mark the remaining pixels as dynamic pixels; wherein, the difference threshold is determined by tuning the background modeling effect on the video dataset; the motion threshold is determined by tuning the motion detection effect on the video dataset; based on the marking results of all pixels, generate a binary dynamic region mask.
[0150] Optionally, as an embodiment, the processor 710 is further configured to perform weighted fusion on the pixel values of each static pixel in the current frame and the pixel values at the corresponding pixel positions in the background reference image at the previous moment; wherein, the fusion weight of the pixel value of the static pixel is a, and the fusion weight of the pixel value at the corresponding pixel position in the background reference image at the previous moment is (1 - a), a is the learning rate and 0 < a < 1; generate an updated background reference image according to the weighted fusion result.
[0151] Optionally, as an embodiment, the processor 710 is specifically configured to perform statistical calculation on the pixel values of the pixels marked as static pixels within multiple frames of the video sequence by the dynamic region mask to generate a static background.
[0152] Optionally, as an embodiment, the processor 710 is specifically configured to determine the starting frame and the loop termination frame when the dynamic pixels first appear according to the change information of the dynamic region mask on the time axis; intercept a video segment from the video sequence as a dynamic element based on the starting frame and the loop termination frame.
[0153] Optionally, as an embodiment, the processor 710 is specifically configured to superimpose the region marked as a dynamic pixel in the dynamic element to the corresponding position of the static background.
[0154] Optionally, as an embodiment, the processor 710 is specifically configured to calculate the motion vector of each pixel based on the brightness constancy constraint of the optical flow method to obtain a motion vector field for consecutive frames in the video sequence.
[0155] It should be understood that, in this embodiment, the input unit 704 may include a graphics processing unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes image data of still images or videos obtained by an image acquisition device (such as a camera) in video acquisition mode or image acquisition mode. The display unit 706 may include a display panel 7061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 707 includes a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include a touch detection device and a touch controller. Other input devices 7072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here. The memory 709 can be used to store software programs and various types of data, including but not limited to applications and operating systems. Processor 710 can integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 710.
[0156] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described image generation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0157] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0158] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described image generation method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0159] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0160] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and each step may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0162] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An image generation method characterized by, The method comprises: calculating a motion vector field based on a video sequence; wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence; generating a dynamic region mask based on amplitude information of the motion vector field; wherein the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence; determining a static background and at least one dynamic element from the video sequence based on the dynamic region mask; performing a synthesis process on the dynamic element and the static background to generate a dynamic glint image.
2. The method of claim 1, wherein, The generating of the dynamic region mask based on the amplitude information of the motion vector field comprises: determining the motion amplitude of each pixel in the video sequence according to the amplitude information of the motion vector field; marking the pixels with a motion amplitude greater than or equal to a motion threshold as dynamic pixels, and marking the pixels with a motion amplitude less than the motion threshold as static pixels; wherein the motion threshold is determined by optimizing the motion detection effect on a video dataset; generating a binary dynamic region mask based on the marking results of all pixels.
3. The method of claim 1, wherein, Before the generating of the dynamic region mask based on the amplitude information of the motion vector field, the method further comprises: generating an initial background reference image; wherein the pixel value of each pixel position in the initial background reference image is the color average value of the corresponding pixel position in the first N frames of the video sequence, and N is an integer greater than 1; The generating of the dynamic region mask based on the amplitude information of the motion vector field comprises: for each frame of the video sequence after the first N frames, the following steps are performed: calculating the color difference value of each pixel between the current frame and the background reference image at the last time; determining the motion amplitude of each pixel in the current frame according to the motion vector field of the current frame; marking the pixels with a color difference value less than a difference threshold and a motion amplitude less than a motion threshold as static pixels, and marking the remaining pixels as dynamic pixels; wherein the difference threshold is determined by optimizing the background modeling effect on a video dataset; the motion threshold is determined by optimizing the motion detection effect on a video dataset; generating a binary dynamic region mask based on the marking results of all pixels.
4. The method of claim 3, wherein, After the marking of the pixels with a color difference value less than a difference threshold and a motion amplitude less than a motion threshold as static pixels, the method further comprises: performing weighted fusion of the pixel value of each static pixel in the current frame and the pixel value of the corresponding pixel position in the background reference image at the last time; wherein the fusion weight of the pixel value of the static pixel is a, and the fusion weight of the pixel value of the corresponding pixel position in the background reference image at the last time is (1-a), a is a learning rate and 0 < a < 1; generating an updated background reference image according to the weighted fusion result.
5. The method of claim 1, wherein, The determining of the static background from the video sequence based on the dynamic region mask comprises: performing statistical calculation on the pixel values of the pixels marked as static pixels in the dynamic region mask in multiple image frames in the video sequence to generate a static background.
6. The method of claim 1, wherein, determining at least one dynamic element from the video sequence based on the dynamic region mask, comprising: determining a start frame and a loop end frame of a first appearance of a dynamic pixel based on variation information of the dynamic region mask on a time axis; cutting a video clip from the video sequence as the dynamic element based on the start frame and the loop end frame.
7. The method of claim 1, wherein, the synthesizing the dynamic element and the static background, comprising: superimposing a region in the dynamic element marked as the dynamic pixel by the dynamic region mask to a corresponding position of the static background.
8. The method of claim 1, wherein, the calculating the motion vector field based on the video sequence, comprising: calculating a motion vector field by calculating a motion vector of each pixel based on a brightness constant constraint of an optical flow method for consecutive frames in the video sequence.
9. An image generation apparatus characterized by comprising: the apparatus, comprising: a calculating module configured to calculate a motion vector field based on a video sequence, wherein the motion vector field contains pixel-level motion information of each image frame in the video sequence; a first generating module configured to generate a dynamic region mask based on amplitude information of the motion vector field, wherein the dynamic region mask is used to identify dynamic pixels and static pixels in the video sequence; a determining module configured to determine a static background and at least one dynamic element from the video sequence based on the dynamic region mask; a synthesizing module configured to synthesize the dynamic element and the static background to generate a dynamic glint image.
10. An electronic device, comprising: The electronic device includes a processor, a memory, and a program or instructions stored on the memory and executable on the processor, and the program or instructions are executed by the processor to implement the steps of the image generation method according to any one of claims 1-8.
11. A readable storage medium, characterized by, The readable storage medium stores a program or instructions, and the program or instructions are executed by the processor to implement the steps of the image generation method according to any one of claims 1-8.
Citation Information
Cited By
Artificial intelligence-based animation video frame rate adjustment method and system
CN122265482A