Image processing method and photographing device

By acquiring multiple video frames to be processed and using deep neural networks for style feature extraction and consistency processing, the problem of inconsistent video frame tones that cannot be achieved by the automatic color correction function in editing software is solved, thus improving video production efficiency and quality.

WO2026152342A1PCT designated stage Publication Date: 2026-07-23ARASHI VISION INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ARASHI VISION INC
Filing Date
2025-01-16
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

The automatic color correction function in existing editing software cannot achieve color tone consistency between video frames, resulting in a decrease in video production efficiency and quality.

Method used

By acquiring multiple video frames to be processed, a reference image frame is determined in response to the user's selection operation. A deep neural network is used for style feature extraction and consistency processing to ensure that all video frames have the same tonal characteristics.

Benefits of technology

It achieves tonal consistency between video frames, improves the efficiency and quality of video production, simplifies user operation, and lowers the technical threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025072833_23072026_PF_FP_ABST
    Figure CN2025072833_23072026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an image processing method and a photographing device. The image processing method comprises: acquiring a plurality of video frames to be processed, the plurality of video frames having different tonal characteristics; in response to a selection operation by a user, determining a reference image frame, the reference image frame having a tonal characteristic desired by the user; and, on the basis of the reference image frame, performing consistency processing on the plurality of video frames, such that the plurality of processed video frames have a same tonal characteristic.
Need to check novelty before this filing date? Find Prior Art

Description

An image processing method and imaging device Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method and an image capturing device. Background Technology

[0002] Currently, users adjust images or videos using the "automatic color correction" function provided in editing software. This function fine-tunes images or videos according to the editing software's algorithm logic. Although this "automatic color correction" method only requires one click to turn on and off, its purpose is merely to correct image parameters, and it makes fixed adjustments to those parameters. Under this adjustment method, multiple video frames still have different tonal characteristics after adjustment, failing to achieve consistent video tonal processing. Summary of the Invention

[0003] This application provides an image processing method and an image capturing device.

[0004] In a first aspect, embodiments of this application provide an image processing method, the method comprising:

[0005] Multiple video frames to be processed are acquired, and the multiple video frames to be processed have different tonal characteristics;

[0006] In response to a user's selection action, a reference image frame is determined, the reference image frame having the tonal characteristics desired by the user;

[0007] The multiple video frames to be processed are subjected to consistency processing based on the reference image frame, so that the multiple video frames to be processed after processing have the same tone characteristics.

[0008] Secondly, embodiments of this application provide a shooting device, including:

[0009] A lens, which is detachably connected to the shooting device;

[0010] A human-computer interaction interface, wherein the human-computer interaction interface is used to receive control commands from the user;

[0011] The imaging device also includes a memory and a processor;

[0012] The memory is used to store instructions; the processor calls the instructions stored in the memory to perform the following operations:

[0013] Multiple video frames to be processed are acquired, and the multiple video frames to be processed have different tonal characteristics;

[0014] In response to a user's selection action, a reference image frame is determined, the reference image frame having the tonal characteristics desired by the user;

[0015] The multiple video frames to be processed are subjected to consistency processing based on the reference image frame, so that the multiple video frames to be processed after processing have the same tone characteristics.

[0016] Thirdly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the image processing method in the first aspect.

[0017] This application provides an image processing method and a shooting device. The image processing method includes: acquiring multiple video frames to be processed, which have different tonal characteristics; determining a reference image frame in response to a user's selection operation, which has tonal characteristics desired by the user; and performing consistency processing on the multiple video frames to be processed based on the reference image frame, so that the processed video frames have the same tonal characteristics. This application provides a flexible stylized reference based on the user-selected reference image frame, performs consistency processing on multiple video frames to be processed, unifies the style of the video frames, achieves tonal consistency of multiple video frames to be processed, and improves the efficiency and quality of video production. Attached Figure Description

[0018] Figure 1 is a schematic flowchart of an image processing method provided in an embodiment of this application;

[0019] Figure 2 is a schematic diagram of the human-computer interaction interface of a shooting device provided in an embodiment of this application;

[0020] Figure 3 is a schematic diagram of the human-computer interaction interface of a shooting device provided in an embodiment of this application;

[0021] Figure 4 is a schematic diagram of the human-computer interaction interface of a shooting device provided in an embodiment of this application;

[0022] Figure 5 is a schematic diagram of the human-computer interaction interface of a shooting device provided in an embodiment of this application;

[0023] Figure 6 is a schematic diagram of the human-computer interaction interface of a shooting device provided in an embodiment of this application;

[0024] Figure 7 is a schematic diagram of a process for constructing a set of sample images according to an embodiment of this application;

[0025] Figure 8 is a key-value pair diagram of a sample image set provided in an embodiment of this application;

[0026] Figure 9 is a schematic diagram of a processing flow in an editing scenario provided by an embodiment of this application;

[0027] Figure 10 is a schematic diagram of an image processing method using a deep neural feature extraction network according to an embodiment of this application;

[0028] Figure 11 is a schematic diagram of an automatic color matching process for video provided in an embodiment of this application;

[0029] Figure 12 is a schematic diagram of the structure of a shooting device provided in an embodiment of this application. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0032] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0033] It should be noted that the terms "first, second, and third" used in the embodiments of this application are merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0034] This application provides an image processing method. In practical applications, the hardware for implementing the image processing method can be either a shooting device or an image processing device.

[0035] For example, the shooting device may include a lens, a human-computer interaction interface, an image sensor, a motion sensor, a processor, a memory, and a housing assembly that carries one or more of the above-mentioned devices.

[0036] For example, when the hardware is an image processing device, the user needs to input the footage recorded by the camera into the image processing device for processing. The image processing device may include a tablet computer, a laptop computer, a personal digital assistant, or a smartphone.

[0037] As shown in Figure 1, in one embodiment, the image processing method includes:

[0038] Step 101: Obtain multiple video frames to be processed, which have different tonal characteristics.

[0039] In practical applications, multiple video frames to be processed originate from one or more video materials.

[0040] In practical applications, the application scenarios of this application embodiment may include: when video materials are obtained, during post-editing or automatic video generation, user operations are received through a human-computer interaction interface to achieve consistent processing of multiple video frames to be processed, resulting in multiple video frames to be processed with the same tonal characteristics. Here, before consistent processing of the multiple video frames to be processed, there are differences in at least one aspect such as scene, artistic style, and tonal color among the multiple video frames to be processed. After consistent processing of the multiple video frames to be processed, the scene, artistic style, and tonal color are unified among the multiple video frames to be processed.

[0041] In practical applications, multiple video frames to be processed include, but are not limited to, panoramic video frames.

[0042] Step 102: In response to the user's selection operation, determine a reference image frame with the tonal characteristics desired by the user.

[0043] In practical applications, the reference image frame serves as the basis for consistent processing of multiple video frames to be processed. It is selected by the user through the human-computer interaction interface, and the reference image frame selected by the user has the tonal characteristics desired by the user.

[0044] In practical applications, the reference image frame is derived from multiple video frames to be processed, or from a set of sample images. That is, the reference image frame in this application can be selected from multiple video frames to be processed, for example, selecting one frame from multiple video frames to be processed as the reference image frame; or, the reference image frame in this application can be selected from a set of sample images, which can be a user-constructed image set.

[0045] Step 103: Perform consistency processing on multiple video frames to be processed based on the reference image frame so that the processed multiple video frames to be processed have the same tone characteristics.

[0046] In practical applications, a flexible stylization reference is provided based on the reference image frame, and multiple video frames to be processed are processed in a consistent manner so that the processed video frames have the same tonal characteristics, which improves the uniformity of video frame style and improves the efficiency and quality of video production.

[0047] In a feasible scenario, as shown in Figure 2, which is a schematic diagram of the human-computer interaction interface of the shooting device, multiple video frames to be processed are displayed on the interface, including: video frame 1, video frame 2, video frame 3, and video frame 4. The interface also displays a reference image frame selected by the user. When the user clicks the consistency processing button on the interface, the shooting device performs consistency processing on the four video frames to be processed based on the reference image frame, so that the four processed video frames have the same tonal characteristics.

[0048] In some embodiments, step 103 involves performing consistency processing on multiple view frames to be processed based on a reference image frame, including:

[0049] Obtain the style vector of the reference image frame;

[0050] The multiple video frames to be processed are subjected to consistency processing based on the style vector of the reference image frame so that the processed video frames have the same tonal characteristics.

[0051] In practical applications, obtaining the style vector of the reference image frame includes: inputting the reference image frame into the deep neural encoder of the deep neural color-dithering network, extracting style features from the reference image frame to obtain the style vector, using the style vector of the reference image frame to perform consistency processing on multiple video frames to be processed to achieve style transfer, and combining the style of the reference image frame with the content of multiple video frames to be processed so that the processed multiple video frames to be processed have the same tonal features.

[0052] In some embodiments, the reference image frame is derived from multiple video frames to be processed, and step 102, in response to a user's selection operation, determines the reference image frame, including:

[0053] Display multiple video frames to be processed;

[0054] In response to the user's selection operation for multiple video frames to be processed, one of the multiple video frames to be processed is determined as the reference image frame.

[0055] In practical applications, when the reference image frame comes from multiple video frames to be processed, multiple video frames to be processed are first displayed. Then, the user selects one frame from the multiple video frames to be processed as the reference image frame through the human-computer interaction interface, so as to flexibly select the reference style and achieve the desired style effect. In related technologies, the reference image frame comes from a stylized image, such as media content for users to compare or refer to, and the style is fixed.

[0056] In a feasible scenario, as shown in Figure 3, a cycling video plays on the human-computer interaction interface. As the video plays, multiple video frames to be processed are displayed on the interface. When the user clicks the pause button (the dotted finger gesture in Figure 3), the camera determines the current frame as the reference image frame. Alternatively, when the user clicks the pause button (the dotted finger gesture in Figure 3), the video playback pauses. Clicking the current frame selection button again (the solid finger gesture in Figure 3) then determines the current frame as the reference image frame. This allows for flexible selection of the reference style during video playback, achieving the user's desired style effect.

[0057] In some embodiments, the reference image frame is derived from a set of sample images, and step 103, in response to a user's selection operation, determines the reference image frame, including:

[0058] In response to the user's selection operation for multiple video frames to be processed, one of the multiple video frames to be processed is determined as the keyframe.

[0059] Feature extraction is performed on keyframes to obtain the content features of the keyframes, which include one or more of the following: image color, image texture, image shape, and image structure displayed in the keyframe.

[0060] Determine one or more target sample images from the sample image set based on the content features of keyframes;

[0061] In response to a user's selection action, a reference image frame is determined from one or more target sample images.

[0062] In practical applications, when the reference image frame comes from the sample image set, the user first selects one frame from multiple video frames to be processed as the key frame through the human-computer interaction interface. Then, by utilizing the content features of the key frame, the reference image frame is determined from the sample image set, thus achieving another flexible selection of reference style and realizing the user's desired style effect.

[0063] For example, taking a cycling video with multiple video frames to be processed as an example, after the user selects a keyframe from multiple cycling frames, the shooting device extracts features from the keyframe. The content features of the keyframe include the color, texture, shape, and structure of multiple objects such as trees, coastal roads, and the sea. The shooting device then selects one or more target sample images with similar style sources from the sample image set based on the color, texture, shape, and structure of the multiple objects such as trees, coastal roads, and the sea for the user to choose the desired style image. As shown in Figure 4, the human-computer interaction interface in Figure 4 displays both the keyframe selected by the user and multiple target sample images selected by the shooting device from the sample image set based on the content features of the keyframe, including target sample image 1, target sample image 2, target sample image 3, and target sample image 4. The user combines the keyframe and the recommended multiple target sample images and selects target sample image 2 as the reference image frame. In this way, the reference image frame selected in this application can better achieve the style effect expected by the user.

[0064] In some embodiments, the above-described determination of one or more target sample images in the sample image set based on the content features of keyframes includes:

[0065] Based on the content features of the keyframes, a similarity comparison is made with the content features of each sample image in the sample image set, and one or more sample images with a similarity greater than a preset threshold are identified as target sample images.

[0066] In practical applications, when selecting target sample images from a set of sample images, the selection is based on the similarity between content features. Specifically, the similarity between the content features of the keyframe and the content features of each sample image in the set of sample images is compared. One or more sample images with a similarity greater than a preset threshold are identified as target sample images. Finally, reference image frames are determined from the selected target sample images, thus achieving accurate selection of reference image frames.

[0067] In some embodiments, the above-described response to a user selection operation to determine a reference image frame from one or more target sample images includes:

[0068] The keyframes are diffried based on one or more target sample images to obtain the diffried keyframes corresponding to each target sample image.

[0069] Display at least one keyframe after dithering;

[0070] In response to the user's selection of a dithered keyframe, the target sample image corresponding to the user-selected dithered keyframe is determined as the reference image frame.

[0071] In practical applications, one or more target sample images are selected from the sample image set based on the content features of the keyframes. The color style of one or more target sample images is then transferred to the keyframes to complete the color grading. At least one color-graded keyframe is then output and displayed for the user to preview and select. The user selects the desired effect from the color grading recommendation preview through the human-computer interaction interface, and the target sample image corresponding to the desired effect is determined as the reference image frame.

[0072] Referring to Figure 5, the human-computer interaction interface in Figure 5 displays not only the keyframe selected by the user, but also the keyframes after color grading for each style of sample image obtained by color grading the keyframes based on the above four target sample images. These include the first style keyframe corresponding to the style of target sample image 1, the second style keyframe corresponding to target sample image 2, the third style keyframe corresponding to target sample image 3, and the fourth style keyframe corresponding to target sample image 4. The user combines the keyframes displayed on the human-computer interaction interface with the keyframes of multiple styles after color grading and selects the third style keyframe. At this time, the target sample image 3 corresponding to the third style keyframe is captured and confirmed as the reference image frame.

[0073] This application first performs diffraction processing on keyframes, allowing the user to select the desired effect before applying it to all video frames. This method significantly improves processing efficiency compared to performing diffraction processing on all video frames from the outset, and then repeating the process multiple times if the user's expectations are not met. Furthermore, the diffraction recommendation preview method provided in this application simplifies the process, eliminating the need for users to master color theory and editing skills. Users can simply select the desired effect from the recommended color style transfer images, simplifying the user operation and making the color grading process easy and straightforward.

[0074] In some embodiments, the above-described display of at least one dithered keyframe includes:

[0075] Show the first subset of dithered keyframes among multiple dithered keyframes;

[0076] In response to a toggle operation on the dithered keyframes contained in the first subset, display the dithered keyframes contained in the second subset of the multiple dithered keyframes.

[0077] In practical applications, during the recommended preview process, this application can achieve infinite updates of the dithered recommended preview. If the first subset is not selected by the user, in response to the switching operation, a new batch can be updated to display the dithered keyframes contained in the second subset. In this way, the color style transfer effect of the dithered keyframes from different batches is presented to the user, which simplifies the difficulty of video color grading while providing multiple optional effects for the user to choose from.

[0078] Referring to Figure 6, the human-computer interaction interface in Figure 6 displays not only the keyframes selected by the user, but also the keyframes obtained after color-dithering based on the four target sample images mentioned above: the first style keyframe corresponding to the style of target sample image 1, and the second style keyframe corresponding to target sample image 2. When the user clicks the switch button, the display area of ​​the dithered keyframes on the human-computer interaction interface displays the third style keyframe corresponding to target sample image 3 and the fourth style keyframe corresponding to target sample image 4. In this way, the color style transfer effect of different batches of dithered keyframes can be presented to the user by switching the button.

[0079] In some embodiments, the above-described response to a user's selection operation for a dithered keyframe, determining the target sample image corresponding to the user-selected dithered keyframe as a reference image frame, includes:

[0080] In response to the user's selection operation for one of the dithered keyframes contained in the second subset, the target sample image corresponding to the dithered keyframe is determined as the reference image frame.

[0081] Referring to Figure 6, after the user clicks the switch button, the third style keyframe is selected by combining the keyframes displayed on the human-computer interaction interface and the keyframes of multiple styles after dithering. At this time, the target sample image 3 corresponding to the third style keyframe is captured as the reference image frame.

[0082] In practical applications, users switch between recommended previews and ultimately select the target sample image corresponding to the desired color style transfer effect as the reference image frame.

[0083] In some embodiments, the above-described determination of one or more target sample images in the sample image set based on the content features of keyframes includes:

[0084] Based on the key-value pairs of the constructed sample image set, find one or more target sample images in the sample image set whose values ​​are indicated by the content features of the keyframes.

[0085] In practical applications, when the reference image frame originates from a set of sample images, this application provides a key-value pair lookup method to quickly locate one or more target sample images. This application uses content features as keys and sample images as values. Given a keyframe, it searches the set of sample images for one or more target sample images indicated by the value corresponding to the content feature of the keyframe, achieving fast and accurate retrieval.

[0086] In some embodiments, this application can construct key-value pairs for a set of sample images through the following steps:

[0087] Obtain a set of sample images;

[0088] A deep neural feature extraction network is used to extract features from each sample image in the sample image set to obtain the content features of each sample image.

[0089] Each sample image is color-corrected using different color correction methods to obtain a reference sample image under each color correction method;

[0090] A deep neural encoder is used to extract style features from the reference sample image to obtain a set of style vectors for the reference sample image.

[0091] Use the content features of each sample image as the key and the style vector set of the reference sample image as the value to construct key-value pairs for the sample image set.

[0092] In practical applications, the sample image set, also known as a resource library, is a collection of reference images with aesthetic value that is stored and managed. Designers and creators can retrieve and use resources from the resource library to support their editing and color grading tasks. This application uses a pre-made resource library and a scene matching algorithm to automatically recommend reference images that match the source material in terms of scene and content. Similar reference materials can more accurately transfer styles. Reference materials are media content provided for comparison or reference, used to guide or influence the creative process. For example, in the style transfer process, reference materials are images or videos that provide the style to be transferred.

[0093] In practical applications, the color grading process described above is also known as video color grading. It can be seen as a step in the video post-production process. By adjusting parameters such as color balance, contrast, and brightness, the visual effects and emotional atmosphere of the video can be changed. Color grading can enhance visual appeal, unify the color tone of the shots, and convey a specific artistic style.

[0094] In some embodiments, the reference sample image for each color tone has different parameter values ​​than each sample image in one or more dimensions of color balance, contrast, and brightness.

[0095] In a scenario where it is feasible to construct a set of key-value pairs for sample images, see Figure 7.

[0096] Reference image i represents any sample image in the above sample image set, and content features i are extracted using a deep neural feature extraction network; i is a positive integer greater than 1 and less than or equal to N;

[0097] The reference image i is processed with different color adjustment methods to obtain reference sample images under each color adjustment method, forming multiple sets of reference images with different styles; among them, the different color adjustment methods include but are not limited to: three-dimensional color lookup table (3DLUT), filter, manual color adjustment, etc.

[0098] Style features are extracted from reference image i using a deep neural encoder to obtain the style vector set i of reference image i;

[0099] Using the content features of reference image i as the key and the style vector set of reference image i as the value, construct key-value pairs for the sample image set, as shown in Figure 8.

[0100] In some embodiments, step 103 performs consistency processing on multiple video frames to be processed based on a reference image frame, so that the processed multiple video frames to be processed have the same tonal characteristics, including:

[0101] The deep neural encoder of the deep neural color grading network extracts the content features of each video frame in multiple video frames to be processed, and obtains the fade vector of each video frame.

[0102] Each video frame is processed using the fade network of a deep neural color-grading network and the fade vector of each video frame to obtain the faded image of each video frame.

[0103] A deep neural colorimetric network is used to perform consistency processing on the faded image and reference image frame of each video frame, so that the processed video frames have the same tonal characteristics.

[0104] In practical applications, colorization networks perform consistency processing by adjusting the colors of an image using algorithms to match the colors of another reference image. This application adjusts the colors of multiple video frames to be processed to match the colors of reference image frames. This process typically involves color model conversion and color mapping to achieve similarity in color distribution and hue.

[0105] In practical applications, the coloring network of a deep neural color-dithering network is used to perform consistency processing on the faded image and the reference image frame of each video frame. The resulting multiple video frames to be processed after the above consistency processing do not change the structure of the video frames to be processed, but only transfer the tone information of the reference image frame. This allows users who do not have color theory and editing skills to obtain a color-dithering image that imitates the tone of the desired image.

[0106] In a feasible editing scenario, see Figure 9.

[0107] Step 401: The user uploads the video or image to the editing software of the image processing device;

[0108] Step 402: The editing software of the image processing device uses a deep neural encoder to extract the fade vector of each frame;

[0109] Step 403: Input the fade vector corresponding to each frame and the frame into the fade network to obtain a high-dimensional faded image;

[0110] Step 404: Feed the high-dimensional faded image and style vector into the coloring network and output the final consistency processing result.

[0111] In some embodiments, the image processing method described above further includes:

[0112] Multiple video frames to be processed after consistency processing are displayed on the human-computer interaction interface. The human-computer interaction interface displays at least one parameter adjustment control, which is used to adjust at least one of the following: color intensity, color warmth, and color brightness.

[0113] In response to the operation of the parameter adjustment control, the multiple video frames to be processed after consistency processing are adjusted to obtain the adjusted video frame set.

[0114] In practical applications, when multiple video frames to be processed after consistency processing are displayed on the human-computer interaction interface, users can further adjust the multiple video frames to be processed after consistency processing through at least one parameter adjustment control displayed on the human-computer interaction interface to obtain the final adjusted set of video frames.

[0115] In some embodiments, the deep neural encoder using a deep neural color-grading network extracts content features from each of the multiple video frames to be processed, obtaining a fade vector for each video frame, including:

[0116] Each video frame in a plurality of video frames to be processed is downsampled to obtain each first downsampled video frame;

[0117] A deep neural encoder is used to extract content features from each first downsampled video frame to obtain each fade vector.

[0118] In practical applications, multiple video frames to be processed are used as source images. This application downsamples these images before feeding them into the deep neural encoder to reduce the computational load on key processing nodes. Simultaneously, the deep neural color-matching network can also unify the source images to the same size during this process to accommodate inputs at different resolutions.

[0119] In some embodiments, the deep neural encoder using a deep neural color-grading network extracts style features from a reference image frame to obtain a style vector, including:

[0120] The reference image frame is downsampled to obtain the second downsampled video frame;

[0121] Style features were extracted from the second downsampled video frame using a deep neural encoder to obtain a style vector.

[0122] In practical applications, the reference image is downsampled before being fed into the depth neural encoder to reduce the computational load of key computing nodes.

[0123] In a feasible scenario of automatic video color matching, the process is explained using the deep neural feature extraction network shown in Figure 10 and the automatic video color matching workflow shown in Figure 11.

[0124] The deep neural feature extraction network provided in this application comprises three modules: a deep neural encoder 501, a fading network 502, and a coloring network 503 (also known as a diffusing network). The deep neural encoder 501 takes a low-resolution image as input and outputs a fading vector and a style vector of the image. The fading network 502 takes the fading vector and the source image as input and outputs a fading image. The coloring network 503 takes the style vector and the fading image as input and outputs a diffusing image with that style.

[0125] This application acquires multiple video frames to be processed (also known as source images). The source images are downsampled to (width, height) = (256 pixels, 256 pixels) to obtain a low-resolution source image.

[0126] The source low-resolution image is fed into the depth neural encoder 501 to obtain the fade vector;

[0127] The source image and the fade vector are fed into the fade network 502 to obtain the faded image;

[0128] The reference image is downsampled to (width, height) = (256 pixels, 256 pixels) to obtain a low-resolution reference image;

[0129] The reference low-resolution image is fed into the deep neural encoder 501 to obtain the style vector;

[0130] The faded image and style vector are fed into the coloring network 503 to obtain the diffusing result image.

[0131] Step 601: The user uploads the video or image to the editing software and completes the editing.

[0132] Here, the user enables the "Dithering Recommendation" function, selecting a specific frame from a video clip as a keyframe.

[0133] Step 602: Input keyframes to a deep neural feature extraction network to extract the content features F of the keyframes.

[0134] Step 603: Calculate the distance between the content feature F and all Keys in the key-value pairs of the constructed sample image set, and regard the k content features with the smallest distance as recommended reference content.

[0135] Step 604: Extract the K smallest content features and their corresponding m style vectors, and feed the m style vectors and keyframes into the deep neural color-sculpting network.

[0136] Step 605: The deep neural color dithering network performs color dithering, transferring the color style to the source material to complete the color dithering, and outputs m style vector transfer results as color dithering recommendation previews for users to choose from.

[0137] Step 606: Input the style vector corresponding to the style reference map selected by the user and the complete video frame into the deep neural color sculpting network for color sculpting.

[0138] Step 607: Render the complete video after dithering.

[0139] Among them, source material for coloring refers to the original media content that has not been processed or modified in any way, such as photos and video clips, which are the basic materials for creating and modifying works.

[0140] As described above, this application recommends style materials to users by constructing an aesthetically valuable material library combined with scene matching algorithms. It utilizes style transfer technology, based on a deep neural network-based automatic faded color-grading network, to automate video color grading, thereby lowering the technical barrier to video color grading. Current video color grading requires not only mastery of color theory and editing skills but also a foundation in aesthetics, making it a complex and time-consuming task. This application can transfer the artistic style and color scheme from reference images to source video footage, simplifying the color grading process. This not only saves time and effort but also allows beginners to easily create aesthetically pleasing video effects, improving the efficiency and quality of video production.

[0141] As shown in Figure 12, in one embodiment, the imaging device 700 includes:

[0142] Lens 701, the lens is detachably connected to the shooting equipment;

[0143] Human-computer interaction interface 702, the human-computer interaction interface is used to receive control commands from the user;

[0144] The imaging device 700 also includes a memory 703 and a processor 704;

[0145] Memory 703 is used to store instructions; processor 704 calls the instructions stored in memory 703 to perform the following operations:

[0146] Multiple video frames to be processed are acquired, and these multiple video frames have different tonal characteristics;

[0147] In response to the user's selection action, a reference image frame is determined, which has the tonal characteristics desired by the user;

[0148] Multiple video frames to be processed are subjected to consistency processing based on a reference image frame so that the processed video frames have the same tonal characteristics.

[0149] In some embodiments, the reference image frame is derived from a plurality of video frames to be processed, or the reference image frame is derived from a set of sample images.

[0150] In some embodiments, multiple video frames to be processed originate from one or more video clips.

[0151] In some embodiments, when performing consistency processing on multiple video frames to be processed based on a reference image frame, the processor 704 calls the instructions stored in memory 703 to perform the following operations:

[0152] Obtain the style vector of the reference image frame;

[0153] The multiple video frames to be processed are subjected to consistency processing based on the style vector of the reference image frame so that the processed video frames have the same tonal characteristics.

[0154] In some embodiments, when obtaining the style vector of a reference image frame, processor 704 calls the instructions of memory storage 703 to perform the following operations:

[0155] Style vectors are obtained by extracting style features from reference image frames using a deep neural encoder of a deep neural colorimetric network.

[0156] In some embodiments, the reference image frame is derived from multiple video frames to be processed, and the human-computer interaction interface 702 displays multiple video frames to be processed in response to the user's selection operation to determine the reference image frame.

[0157] Processor 704 calls the instructions stored in memory 703 to perform the following operations: in response to a user's selection operation for multiple video frames to be processed, one of the multiple video frames to be processed is determined as a reference image frame.

[0158] In some embodiments, the reference image frame is derived from a set of sample images. In response to a user's selection operation to determine the reference image frame, processor 704 invokes the instructions stored in memory 703 to perform the following operations:

[0159] In response to the user's selection operation for multiple video frames to be processed, one of the multiple video frames to be processed is determined as the keyframe.

[0160] Feature extraction is performed on keyframes to obtain the content features of the keyframes, which include one or more of the following: image color, image texture, image shape, and image structure displayed in the keyframe.

[0161] Determine one or more target sample images from the sample image set based on the content features of keyframes;

[0162] In response to a user's selection action, a reference image frame is determined from one or more target sample images.

[0163] In some embodiments, when determining one or more target sample images in the sample image set based on the content features of keyframes, processor 704 calls the instructions of memory storage 703 to perform the following operations:

[0164] Based on the content features of the keyframes, a similarity comparison is made with the content features of each sample image in the sample image set, and one or more sample images with a similarity greater than a preset threshold are identified as target sample images.

[0165] In some embodiments, in response to a user's selection operation to determine a reference image frame from one or more target sample images, processor 704 invokes the instructions stored in memory 703 to perform the following operations:

[0166] The keyframes are diffried based on one or more target sample images to obtain the diffried keyframes corresponding to each target sample image.

[0167] The human-computer interaction interface 702 displays at least one keyframe after dithering;

[0168] The processor 704 calls the instructions stored in memory 703 to perform the following operations: in response to the user's selection operation for the dithered keyframe, the target sample image corresponding to the user-selected dithered keyframe is determined as the reference image frame.

[0169] In some embodiments, when at least one dithered keyframe is displayed, the human-computer interaction interface 702 displays the dithered keyframes contained in a first subset of a plurality of dithered keyframes.

[0170] The human-computer interaction interface 702 responds to a switching operation on the dithered keyframes contained in the first subset by displaying the dithered keyframes contained in the second subset of multiple dithered keyframes.

[0171] In some embodiments, in response to a user's selection operation for a dithered keyframe, when the target sample image corresponding to the user-selected dithered keyframe is determined as a reference image frame, the processor 704 calls the instructions stored in memory 703 to perform the following operations:

[0172] In response to the user's selection operation for one of the dithered keyframes contained in the second subset, the target sample image corresponding to the dithered keyframe is determined as the reference image frame.

[0173] In some embodiments, when determining one or more target sample images in the sample image set based on the content features of keyframes, processor 704 calls the instructions of memory storage 703 to perform the following operations:

[0174] Based on the key-value pairs of the constructed sample image set, find one or more target sample images in the sample image set whose values ​​are indicated by the content features of the keyframes.

[0175] In some embodiments, processor 704 invokes the instructions of memory storage 703 to perform the following operations:

[0176] Obtain a set of sample images;

[0177] A deep neural feature extraction network is used to extract features from each sample image in the sample image set to obtain the content features of each sample image.

[0178] Each sample image is color-corrected using different color correction methods to obtain a reference sample image under each color correction method;

[0179] A deep neural encoder is used to extract style features from the reference sample image to obtain a set of style vectors for the reference sample image.

[0180] Use the content features of each sample image as the key and the style vector set of the reference sample image as the value to construct key-value pairs for the sample image set.

[0181] In some embodiments, the reference sample image for each color tone has different parameter values ​​than each sample image in one or more dimensions of color balance, contrast, and brightness.

[0182] In some embodiments, when multiple video frames to be processed are subjected to consistency processing based on a reference image frame so that the processed video frames to be processed have the same tonal characteristics, the processor 704 calls the instructions stored in memory 703 to perform the following operations:

[0183] The deep neural encoder of the deep neural color grading network extracts the content features of each video frame in multiple video frames to be processed, and obtains the fade vector of each video frame.

[0184] Each video frame is processed using the fade network of a deep neural color-grading network and the fade vector of each video frame to obtain the faded image of each video frame.

[0185] A deep neural colorimetric network is used to perform consistency processing on the faded image and reference image frame of each video frame, so that the processed video frames have the same tonal characteristics.

[0186] In some embodiments, multiple video frames to be processed after consistency processing are displayed on the human-computer interaction interface 702. The human-computer interaction interface 702 displays at least one parameter adjustment control, which is used to adjust at least one of the following: color intensity, color warmth, and color brightness.

[0187] Processor 704 calls the instructions stored in memory 703 to perform the following operations: in response to the operation of the parameter adjustment control, adjusts multiple video frames to be processed after consistency processing to obtain an adjusted set of video frames.

[0188] In some embodiments, when the deep neural encoder of the deep neural color-dimming network extracts content features from each of the multiple video frames to be processed to obtain the fade vector of each video frame, the processor 704 calls the instructions stored in memory 703 to perform the following operations:

[0189] Each video frame in a plurality of video frames to be processed is downsampled to obtain each first downsampled video frame;

[0190] A deep neural encoder is used to extract content features from each first downsampled video frame to obtain each fade vector.

[0191] In some embodiments, when the deep neural encoder of the deep neural color-dimming network extracts style features from the reference image frame to obtain the style vector, the processor 704 calls the instructions stored in memory 703 to perform the following operations:

[0192] The reference image frame is downsampled to obtain the second downsampled video frame;

[0193] Style features were extracted from the second downsampled video frame using a deep neural encoder to obtain a style vector.

[0194] The aforementioned processor is likely an integrated circuit chip with signal processing capabilities.

[0195] It should be noted that the specific implementation process of the steps executed by the processor in this embodiment can be referred to the implementation process in the method provided in the embodiment corresponding to Figure 1, and will not be repeated here.

[0196] It should be noted that the specific implementation process of the steps executed by the processor in this embodiment can be referred to the implementation process in the method provided in the embodiment corresponding to Figure 1, and will not be repeated here.

[0197] The embodiments of this application provide a computer product, including a computer program that can be executed by one or more processors to implement the method provided in the embodiment corresponding to FIG1, which will not be described in detail here.

[0198] It should be understood that the terms "an embodiment," "an embodiment," "an embodiment of this application," "the foregoing embodiment," "some embodiments," or "some implementations" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, the phrases "an embodiment," "an embodiment," "an embodiment of this application," "the foregoing embodiment," "some embodiments," or "some implementations" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0199] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image processing method, comprising: Multiple video frames to be processed are acquired, and the multiple video frames to be processed have different tonal characteristics; In response to a user's selection action, a reference image frame is determined, the reference image frame having the tonal characteristics desired by the user; The multiple video frames to be processed are subjected to consistency processing based on the reference image frame, so that the multiple video frames to be processed after processing have the same tone characteristics.

2. The method according to claim 1, wherein: The reference image frame is derived from multiple video frames to be processed, or the reference image frame is derived from a set of sample images.

3. The method according to claim 1, wherein: The multiple video frames to be processed originate from one or more video materials.

4. The method according to claim 1, wherein the consistency processing of the plurality of video frames to be processed based on the reference image frame comprises: Obtain the style vector of the reference image frame; The multiple video frames to be processed are subjected to consistency processing based on the style vector of the reference image frame, so that the multiple video frames to be processed after processing have the same tonal characteristics.

5. The method according to claim 4, wherein obtaining the style vector of the reference image frame comprises: The style vector is obtained by extracting style features from the reference image frame using a deep neural encoder of a deep neural color-grading network.

6. The method according to claim 2, wherein the reference image frame is derived from a plurality of video frames to be processed, and the step of determining the reference image frame in response to a user selection operation includes: Display multiple video frames to be processed; In response to a user's selection operation for a plurality of video frames to be processed, one of the plurality of video frames to be processed is determined as the reference image frame.

7. The method of claim 2, wherein the reference image frame is derived from the sample image set, and the step of determining the reference image frame in response to a user's selection operation includes: In response to the user's selection operation for a plurality of video frames to be processed, one of the plurality of video frames to be processed is determined as a key frame. Feature extraction is performed on the keyframe to obtain the content features of the keyframe, wherein the content features include one or more of the image color, image texture, image shape, and image structure displayed by the keyframe; One or more target sample images in the sample image set are determined based on the content features of the keyframes; The reference image frame is determined from one or more of the target sample images in response to a user's selection action.

8. The method according to claim 7, wherein determining one or more target sample images in the sample image set based on the content features of the keyframe comprises: Based on the content features of the keyframe, a similarity comparison is performed between the content features of each sample image in the sample image set, and one or more sample images with a similarity greater than a preset threshold are determined as the target sample images.

9. The method of claim 8, wherein determining the reference image frame from one or more of the target sample images in response to a user selection operation comprises: The keyframes are diffraction processed based on one or more of the target sample images to obtain the diffraction keyframes corresponding to each target sample image. Display at least one of the diffraction keyframes; In response to the user's selection operation for the dithered keyframe, the target sample image corresponding to the user-selected dithered keyframe is determined as the reference image frame.

10. The method of claim 9, wherein displaying at least one of the diffraction keyframes comprises: Display the first subset of the multiple dithered keyframes; In response to a toggle operation on the dithered keyframes contained in the first subset, the dithered keyframes contained in the second subset of the plurality of dithered keyframes are displayed.

11. The method of claim 10, wherein determining the target sample image corresponding to the user-selected dithered keyframe as the reference image frame in response to the user's selection operation for the dithered keyframe comprises: In response to the user's selection operation for one of the dithered keyframes included in the second subset, the target sample image corresponding to the dithered keyframe is determined as the reference image frame.

12. The method according to claim 7, wherein determining one or more target sample images in the sample image set based on the content features of the keyframe comprises: Based on the key-value pairs of the constructed sample image set, find one or more target sample images from the sample image set whose values ​​are indicated by the content features of the keyframe.

13. The method according to claim 12, further comprising: Obtain a set of sample images; A deep neural feature extraction network is used to extract features from each sample image in the sample image set to obtain the content features of each sample image. Each of the sample images is color-corrected using different color correction methods to obtain a reference sample image under each color correction method. The style features of the reference sample image are extracted using a deep neural encoder to obtain a set of style vectors for the reference sample image. Using the content features of each sample image as the key and the style vector set of the reference sample image as the value, key-value pairs are constructed for the sample image set.

14. The method of claim 13, wherein the reference sample image under each color adjustment mode has different parameter values ​​from each sample image in one or more dimensions of color balance, contrast, and brightness.

15. The method according to any one of claims 2 to 14, wherein performing consistency processing on a plurality of video frames to be processed based on the reference image frame, so that the processed plurality of video frames to be processed have the same tonal characteristics, comprises: The deep neural encoder of the deep neural color grading network extracts the content features of each video frame in the multiple video frames to be processed, and obtains the fade vector of each video frame. Each video frame is processed using the fade network of a deep neural color-grading network and the fade vector of each video frame to obtain a faded image of each video frame. A coloring network using a deep neural colorimetric network is used to perform consistency processing on the faded image of each video frame and the reference image frame, so that the processed multiple video frames to be processed have the same tonal characteristics.

16. The method according to claim 1, further comprising: Multiple video frames to be processed after the consistency processing are displayed on the human-computer interaction interface. The human-computer interaction interface displays at least one parameter adjustment control, which is used to adjust at least one of the following: color intensity, color warmth, and color brightness. In response to the operation of the parameter adjustment control, the multiple video frames to be processed after the consistency processing are adjusted to obtain an adjusted set of video frames.

17. The method according to claim 15, wherein the step of extracting content features from each of the plurality of video frames to be processed using a deep neural encoder of a deep neural color-grading network to obtain a fade vector for each of the video frames includes: Each of the multiple video frames to be processed is downsampled to obtain each first downsampled video frame. The deep neural encoder is used to extract content features from each of the first downsampled video frames to obtain each fade vector.

18. The method according to claim 5, wherein the step of extracting style features from the reference image frame using a deep neural encoder of a deep neural color-grading network to obtain the style vector comprises: The reference image frame is downsampled to obtain a second downsampled video frame; The style vector is obtained by extracting style features from the second downsampled video frame using the deep neural encoder.

19. A shooting device, comprising: A lens, which is detachably connected to the shooting device; A human-computer interaction interface, wherein the human-computer interaction interface is used to receive control commands from the user; The imaging device also includes a memory and a processor; The memory is used to store instructions; the processor calls the instructions stored in the memory to perform the following operations: Multiple video frames to be processed are acquired, and the multiple video frames to be processed have different tonal characteristics; In response to a user's selection action, a reference image frame is determined, the reference image frame having the tonal characteristics desired by the user; The multiple video frames to be processed are subjected to consistency processing based on the reference image frame, so that the multiple video frames to be processed after processing have the same tone characteristics.

20. The shooting device according to claim 19, wherein, The reference image frame is derived from multiple video frames to be processed, or the reference image frame is derived from a set of sample images.

21. The shooting device according to claim 19, wherein, The multiple video frames to be processed originate from one or more video materials.

22. The shooting device according to claim 19, wherein when performing consistency processing on a plurality of video frames to be processed based on the reference image frame, the processor invokes instructions stored in the memory to perform the following operations: Obtain the style vector of the reference image frame; The multiple video frames to be processed are subjected to consistency processing based on the style vector of the reference image frame, so that the multiple video frames to be processed after processing have the same tonal characteristics.

23. The imaging device according to claim 22, wherein when acquiring the style vector of the reference image frame, the processor invokes the instructions stored in the memory to perform the following operations: The style vector is obtained by extracting style features from the reference image frame using a deep neural encoder of a deep neural color-grading network.

24. The shooting device according to claim 20, wherein the reference image frame is derived from a plurality of video frames to be processed, and when the reference image frame is determined in response to a user's selection operation, the human-computer interaction interface displays a plurality of video frames to be processed; The processor invokes instructions stored in the memory to perform the following operations: in response to a user's selection operation for a plurality of video frames to be processed, one of the plurality of video frames to be processed is determined as the reference image frame.

25. The imaging device according to claim 20, wherein the reference image frame is derived from the sample image set, and when determining the reference image frame in response to a user's selection operation, the processor invokes instructions stored in the memory to perform the following operations: In response to the user's selection operation for a plurality of video frames to be processed, one of the plurality of video frames to be processed is determined as a key frame. Feature extraction is performed on the keyframe to obtain the content features of the keyframe, wherein the content features include one or more of the image color, image texture, image shape, and image structure displayed by the keyframe; One or more target sample images in the sample image set are determined based on the content features of the keyframes; The reference image frame is determined from one or more of the target sample images in response to a user's selection action.

26. The shooting device according to claim 25, when determining one or more target sample images in the sample image set based on the content features of the keyframe, the processor calls the instructions stored in the memory to perform the following operations: Based on the content features of the keyframe, a similarity comparison is performed between the content features of each sample image in the sample image set, and one or more sample images with a similarity greater than a preset threshold are determined as the target sample images.

27. The imaging device of claim 26, wherein when responding to a user's selection operation to determine the reference image frame from one or more of the target sample images, the processor invokes instructions stored in the memory to perform the following operations: The keyframes are diffraction processed based on one or more of the target sample images to obtain the diffraction keyframes corresponding to each target sample image. The human-computer interaction interface displays at least one of the diffraction keyframes; The processor invokes instructions stored in the memory to perform the following operations: in response to a user's selection operation for the dithered keyframe, the target sample image corresponding to the user-selected dithered keyframe is determined as the reference image frame.

28. The shooting device according to claim 27, wherein when displaying at least one of the dithered keyframes, the human-computer interaction interface displays dithered keyframes included in a first subset of the plurality of dithered keyframes; The human-computer interaction interface responds to the switching operation of the dithered keyframes contained in the first subset by displaying the dithered keyframes contained in the second subset of the plurality of dithered keyframes.

29. The shooting device according to claim 28, wherein when the target sample image corresponding to the user-selected dithered keyframe is determined as the reference image frame in response to the user's selection operation for the dithered keyframe, the processor calls the instructions stored in the memory to perform the following operations: In response to the user's selection operation for one of the dithered keyframes included in the second subset, the target sample image corresponding to the dithered keyframe is determined as the reference image frame.

30. The shooting device according to claim 25, wherein when determining one or more target sample images in the sample image set based on the content features of the keyframe, the processor invokes instructions stored in the memory to perform the following operations: Based on the key-value pairs of the constructed sample image set, find one or more target sample images from the sample image set whose values ​​are indicated by the content features of the keyframe.

31. The imaging device according to claim 30, wherein the processor invokes instructions stored in the memory to perform the following operations: Obtain a set of sample images; A deep neural feature extraction network is used to extract features from each sample image in the sample image set to obtain the content features of each sample image. Each of the sample images is color-corrected using different color correction methods to obtain a reference sample image under each color correction method. The style features of the reference sample image are extracted using a deep neural encoder to obtain a set of style vectors for the reference sample image. Using the content features of each sample image as the key and the style vector set of the reference sample image as the value, key-value pairs are constructed for the sample image set.

32. The shooting device according to claim 31, wherein the reference sample image under each color grading mode has different parameter values ​​from each sample image in one or more dimensions of color balance, contrast, and brightness.

33. The shooting device according to any one of claims 20 to 32, wherein when the plurality of video frames to be processed are subjected to consistency processing based on the reference image frame so that the processed plurality of video frames to be processed have the same tonal characteristics, the processor invokes the instructions stored in the memory to perform the following operations: The deep neural encoder of the deep neural color grading network extracts the content features of each video frame in the multiple video frames to be processed, and obtains the fade vector of each video frame. Each video frame is processed using the fade network of a deep neural color-grading network and the fade vector of each video frame to obtain a faded image of each video frame. A coloring network using a deep neural colorimetric network is used to perform consistency processing on the faded image of each video frame and the reference image frame, so that the processed multiple video frames to be processed have the same tonal characteristics.

34. The shooting device according to claim 19, wherein a plurality of video frames to be processed after consistency processing are displayed on the human-computer interaction interface, and at least one parameter adjustment control is displayed in the human-computer interaction interface, the parameter adjustment control being used to adjust at least one of color intensity, color warmth and coolness, and color brightness; The processor invokes instructions stored in the memory to perform the following operations: in response to the operation of the parameter adjustment control, adjusts the plurality of video frames to be processed after the consistency processing to obtain an adjusted set of video frames.

35. The shooting device according to claim 33, wherein when the deep neural encoder using a deep neural color-grading network extracts content features from each of the plurality of video frames to be processed to obtain a fade vector for each of the video frames, the processor calls the instructions stored in the memory to perform the following operations: Each of the multiple video frames to be processed is downsampled to obtain each first downsampled video frame. The deep neural encoder is used to extract content features from each of the first downsampled video frames to obtain each fade vector.

36. The imaging device according to claim 23, wherein when the deep neural encoder using a deep neural colorimetric network extracts style features from the reference image frame to obtain the style vector, the processor calls the instructions stored in the memory to perform the following operations: The reference image frame is downsampled to obtain a second downsampled video frame; The style vector is obtained by extracting style features from the second downsampled video frame using the deep neural encoder.