Video processing method, electronic device, storage medium, chip system and computer program product
Through multi-camera collaborative recording and image registration technology, the combination of cameras with short and long exposure times, combined with the cross-attention mechanism and deep learning model optimization, the problem of poor video noise reduction effect is solved, and the video quality in low-light scenes is improved and the applicability of resource-constrained devices is achieved.
Patent Information
- Application Number
- CN202411719015.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2044-11-28
AI Technical Summary
The video noise reduction effect in the existing technology is poor, especially in low-light scenes where the noise has a greater impact, making it difficult to improve video quality while balancing the frame rate.
It adopts multi-camera collaborative recording technology, using one camera to balance the frame rate with short exposure time and the other camera to improve the signal-to-noise ratio with long exposure time, and improves the video noise reduction effect through image registration and feature fusion methods, including cross-attention mechanism and deep learning model optimization.
It significantly improves the video noise reduction effect while balancing the frame rate, simplifies the low-light recording steps, improves the convenience of low-light recording, and is suitable for resource-constrained terminal devices.
Smart Images

Figure CN119211464B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to a video processing method, an electronic device, a storage medium, a chip system, and a computer program product. Background Art
[0002] With the continuous development of electronic devices, their functions are becoming increasingly diverse. For example, video recording functions provided by electronic devices can be used to record videos in various scenarios. During the video recording process, noise is generated, which can affect video quality. Videos in different scenarios are affected by noise to varying degrees. For example, videos captured in low-light scenes are significantly affected by noise.
[0003] In order to improve the video quality, video noise reduction can be performed. However, in the related art, the video noise reduction effect is poor. Summary of the Invention
[0004] The embodiments of the present application provide a video processing method, an electronic device, a storage medium, a chip system, and a computer program product, which are applied to the field of terminal technology and can improve the video noise reduction effect while balancing the frame rate.
[0005] In a first aspect, an embodiment of the present application proposes a video processing method. The method may include: performing video recording in response to a first operation. During the video recording process, reducing the noise of the first image based on the second image to obtain a video frame. The video frame may be a frame in the recorded video. The first image and the second image may be taken using multiple cameras. The shooting time of the first image and the shooting time of the second image may overlap, and the lens direction of the first camera and the lens direction of the second camera are the same, indicating that the contents of the first image and the second image are similar. The first camera may be a camera that shoots the first image, and the second camera may be a camera that shoots the second image. The exposure time of the first image may be less than the exposure time of the second image. The first image may be used as a reference image. The second image may be used as an auxiliary image.
[0006] The video processing method provided in the embodiment of the present application is as follows: in the case of video recording, the first image is captured by a camera that uses a shorter exposure time to balance the frame rate, and the second image is captured by a camera that uses a longer exposure time to improve the signal-to-noise ratio, and the shooting time of the first image and the shooting time of the second image overlap, which indicates that the contents of the first image and the second image are similar. Therefore, the second image is used to reduce the noise of the first image to obtain the video frame in the video, thereby achieving an improved video noise reduction effect while balancing the frame rate.
[0007] In one possible implementation, recording a video in response to a first operation may include: displaying a camera application interface in response to a second operation. The camera application interface may include a low-light recording control. In response to a third operation on the low-light recording control, displaying a low-light recording interface. The low-light recording interface may include a start control. In response to the first operation on the start control, recording a video.
[0008] Since the camera application can provide a low-light recording mode, low-light video recording can be performed using the camera application.
[0009] In another possible implementation, recording the video in response to the first operation may include: displaying a low-light recording interface in response to detecting a low-light scene. The low-light recording interface may include a start control. In response to the first operation on the start control, recording the video.
[0010] When a dark-light scene is detected, the dark-light recording interface is displayed, and the dark-light recording mode is automatically started. This simplifies the steps for starting dark-light recording and improves the convenience of dark-light recording.
[0011] In another possible implementation, performing video recording in response to the first operation may include: displaying a low-light recording interface in response to the second operation. The low-light recording interface may include a start control. In response to the first operation on the start control, performing video recording.
[0012] When the second operation is detected, the low-light recording interface is displayed, enabling low-light recording mode to be enabled by default when the camera app is launched. Low-light recording mode can be the default launch mode of the camera app or the last launch mode of low-light recording mode, thereby simplifying the process of enabling low-light recording and improving the convenience of low-light recording.
[0013] In one possible implementation, the clarity of the first image may be higher than that of the second image. And / or the exposure time of the second image may be greater than the frame interval of the first image. The exposure time of the first image may be less than or equal to the frame interval.
[0014] By setting the exposure time of the reference image to be less than or equal to the frame interval of the reference image, a balanced frame rate is achieved, and by setting the exposure time of the auxiliary image to be greater than the frame interval of the reference image, an improved signal-to-noise ratio is achieved. Thus, the video noise reduction effect is improved while balancing the frame rate.
[0015] By making the clarity of the first image higher than that of the second image, the video noise reduction effect is further improved.
[0016] In a possible implementation, the first image may be a short-frame image, and the second image may be a long-frame image.
[0017] Since the first image is a short-frame image, it can balance the frame rate, preserve the scene's texture edges and highlight information, and have high definition and a low signal-to-noise ratio. The second image is a long-frame image, which has a high signal-to-noise ratio, low definition, richer dark-light information, and richer color information. This demonstrates that long and short-frame images can complement each other to a certain extent. For example, the high-noise sharp texture of the short-frame image and the low-noise blurred texture of the long-frame image can constrain each other, facilitating detail recovery. The highlight information of the short-frame image can repair the overexposed cutoff areas of the long-frame image. Furthermore, the color information of the long-frame image can accurately restore the scene's colors. Therefore, short-frame images offer advantages in balancing frame rate and / or high definition, while long-frame images offer advantages in improving signal-to-noise ratio. By using long-frame images to complement short-frame images, the quality of the short-frame images is improved, thereby achieving enhanced video noise reduction while balancing frame rate.
[0018] In a possible implementation, the exposure duration of the second image may be greater than or equal to 100 ms.
[0019] Since the exposure time of the second image is greater than or equal to 100 ms, the exposure time of the second image can improve the signal-to-noise ratio, thereby improving the video noise reduction effect.
[0020] In one possible implementation, the first image and the second image may have respective timestamps. The timestamps can be used as one basis for determining whether the first image and the second image have similar content, and further, as a basis for determining whether the second image can be used to reduce noise in the first image. In one possible implementation, if the timestamps of the first image and the second image are determined to be within a predetermined time range, and if the lens orientation of the first camera is the same as that of the second camera, it can be determined that the content of the first image and the second image is similar, and thus the second image can be used to reduce noise in the first image.
[0021] In a possible implementation, reducing the noise of the first image according to the second image to obtain the video frame may include: fusing the first image and the second image to obtain the video frame.
[0022] Since the video frame is obtained by fusing the first image and the second image, the first image is taken by a camera that uses a shorter exposure time to balance the frame rate, and the second image is taken by a camera that uses a longer exposure time to improve the signal-to-noise ratio. Therefore, the video noise reduction effect is improved while balancing the frame rate.
[0023] In a possible implementation, fusing the first image and the second image to obtain the video frame may include: fusing the first image and a registered second image to obtain the video frame. The registered second image is obtained by registering the second image with the first image.
[0024] Because the video frame is created by fusing the first image with the registered second image, and the registered second image is created by registering the second image with the first image, image-level noise reduction is achieved. Image-level noise reduction fully utilizes pixel information, preserving image details as much as possible, and offers good interpretability, thereby improving video noise reduction effectiveness.
[0025] In one possible implementation, the registered second image is obtained by registering the second image with the first image, which may include: the registered second image may be obtained by registering the first image and the second image using an optical flow method. For example, the registered second image may be obtained by registering the second image and the first image using optical flow information. The optical flow information may be pixel-by-pixel relative displacement information between the first image and the second image.
[0026] By utilizing the optical flow method to realize image registration between the second image and the first image, the efficiency of image registration is improved.
[0027] In one possible implementation, fusing the first image and the second image to obtain the video frame may include fusing the first image with a histogram-converted image to obtain the video frame. The histogram-converted image may be obtained by transferring a color distribution of the second image to the first image based on a histogram conversion function.
[0028] Since there is no spatial misalignment between the image obtained by histogram conversion and the first image, the probability of ghosting caused by camera or object motion is reduced. In addition, since the image registration step is not performed, the video noise reduction process is simplified.
[0029] In another possible implementation, fusing the first image and the second image to obtain the video frame may include fusing tenth feature data and eleventh feature data to obtain the video frame. The tenth feature data and the eleventh feature data may be obtained based on the first feature data of the first image and the second feature data of the second image.
[0030] In one possible implementation, the tenth feature data and the eleventh feature data may be obtained based on the first feature data of the first image and the second feature data of the second image, which may include: the tenth feature data may be obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism. The first feature data may be used for query. The second feature data may be used for key and value. The eleventh feature data may be obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism. The second feature data may be used for query. The first feature data may be used for key and value.
[0031] In one possible implementation, the tenth feature data may be obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism, and may include: the tenth feature data may be obtained based on the seventh attention weight matrix and the sixth value matrix. The seventh attention weight matrix may be obtained based on the seventh query matrix and the seventh key matrix. The seventh query matrix may be obtained by performing a fourteenth linear transformation on the first feature data of the first image. The seventh key matrix may be obtained by performing a fifteenth linear transformation on the second feature data of the second image. The seventh attention weight matrix may be used to indicate the similarity between the second feature and the second feature data.
[0032] In one possible implementation, the eleventh feature data may be obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism, and may include: the eleventh feature data may be obtained based on the eighth attention weight matrix and the seventh value matrix. The eighth attention weight matrix may be obtained based on the eighth query matrix and the eighth key matrix. The eighth query matrix may be obtained by performing the sixteenth linear transformation on the second feature data of the second image. The eighth key matrix may be obtained by performing the seventeenth linear transformation on the first feature data of the first image. The eighth attention weight matrix may be used to indicate the similarity between the first feature and the second feature data.
[0033] Since the tenth feature data is obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism, the first feature data is used for query and the second feature data is used for key and value. Therefore, the tenth feature in the tenth feature data is obtained by fusing the second feature data with the first feature data. The eleventh feature data is obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism. The second feature data is used for query and the first feature data can be used for key and value. Therefore, the eleventh feature in the eleventh feature data is obtained by fusing the first feature data with the second feature data. This achieves more effective utilization of the first and second feature data. On this basis, the video frame is obtained by fusing the tenth and eleventh feature data, thereby improving the video noise reduction effect.
[0034] In another possible implementation, fusing the first image and the second image to obtain the video frame may include: fusing first feature data of the first image and registered second feature data to obtain the video frame. The registered second feature data is obtained by registering the second feature data of the second image with the first feature data of the first image.
[0035] Since the first feature data abstracts the first image while retaining the detail information of the first image, and the second feature data abstracts the second image while retaining the detail information of the second image, the first feature data and the second feature data have higher robustness and discriminability. On this basis, the video frame is obtained by fusing the first feature data of the first image and the aligned second feature data, and the aligned second feature data is obtained by aligning the second feature data of the second image and the first feature data of the first image. Therefore, the reliability and robustness of video denoising are improved, and the video denoising effect is improved.
[0036] In one possible implementation, the registered second feature data is obtained by registering the second feature data of the second image with the first feature data of the first image, which may include: the registered second feature data may be obtained by registering the second feature data of the second image with the first feature data of the first image using a cross-attention mechanism. The first feature data may be used for querying. The second feature data may be used for keys and values. The first feature data may include multiple first features. The second feature data may include multiple second features.
[0037] Since the cross-attention mechanism can mine the association between the second feature data and the first feature data, it can achieve interaction and matching between the first image and the second image at the feature level, thereby capturing the similarities and differences between the features, focusing on similarities with high weights and suppressing differences with low weights, and can adaptively adjust the weights. Therefore, it has higher robustness and scalability. Therefore, using the cross-attention mechanism to align the second feature data of the second image with the first feature data of the first image can adjust the feature weight of the second feature in the second feature data. For example, a larger weight is assigned to the second feature in the second feature data that has a high similarity with the first feature data, and a smaller weight is assigned to the second feature in the second feature data that has a low similarity with the first feature data. This improves the accuracy of feature alignment and thus improves the video denoising effect.
[0038] In one possible implementation, the cross-attention mechanism may include at least one of the following: a dot-product attention mechanism, a scaled dot-product attention mechanism, an additive attention mechanism, a linear attention mechanism, or a sparse attention mechanism.
[0039] Since the video processing method described in the embodiments of the present application can be applied to resource-constrained electronic devices, such as terminal devices, the computational complexity of the cross-attention mechanism is proportional to the square of the number of feature data included in the feature data. Therefore, by adopting an optimization strategy based on optimizing the cross-attention mechanism, such as a linear attention mechanism and a sparse attention mechanism, the expectation is achieved that the computational complexity can be reduced while maintaining or improving the accuracy of feature registration.
[0040] In one possible implementation, the registered second feature data may be obtained by registering the second feature data of the second image and the first feature data of the first image using a cross-attention mechanism, and may include: the registered second feature data may be obtained based on at least one fourth intermediate matrix. The fourth intermediate matrix may be obtained based on the third attention weight matrix and the fifth value matrix. The third attention weight matrix may be obtained based on the fourth query matrix and the fourth key matrix. The third attention weight matrix may be used to indicate the similarity between the second feature and the first feature data. The fourth query matrix may be obtained by performing a seventh linear transformation on the first feature data. For example, the fourth query matrix may be obtained based on the first feature data and the seventh transformation matrix. The fourth key matrix may be obtained by performing an eighth linear transformation on the second feature data. For example, the fourth key matrix may be obtained based on the second feature data and the eighth transformation matrix. The fifth value matrix may be obtained by performing a ninth linear transformation on the second feature data. For example, the fifth value matrix may be obtained based on the second feature data and the ninth transformation matrix.
[0041] It should be noted that when there is only one fourth intermediate matrix, the above method can be understood as a single-headed cross-attention mechanism. When there are multiple fourth intermediate matrices, the above method can be understood as a multi-headed cross-attention mechanism. The transformation matrix groups corresponding to at least one fourth intermediate matrix can be different. The transformation matrix group can include a seventh transformation matrix, an eighth transformation matrix, and a ninth transformation matrix. The element values in the transformation matrix group can be model parameters of the deep learning model.
[0042] In one possible implementation, the third attention weight matrix may be obtained based on the fourth query matrix and the fourth key matrix, which may include: the third attention weight matrix may be obtained by normalizing the fifth intermediate matrix. For example, the third attention weight matrix may be obtained by processing the fifth intermediate matrix using a softmax function, a sigmoid function, or a tanh function. The fifth intermediate matrix may be obtained based on the dot product between the fourth query matrix and the fourth key matrix.
[0043] In another possible implementation, the third attention weight matrix may be obtained based on the fourth query matrix and the fourth key matrix, and may include: the third attention mechanism may be obtained by normalizing the sixth intermediate matrix. The sixth intermediate matrix may be obtained by performing a tenth linear transformation on the seventh intermediate matrix. For example, the sixth intermediate matrix may be obtained based on the seventh intermediate matrix and the tenth transformation matrix. The element values in the tenth transformation matrix may be model parameters of the deep learning model. The seventh intermediate matrix may be obtained by performing a nonlinear activation on the eighth intermediate matrix. The eighth intermediate matrix may be obtained based on the fourth query matrix and the fourth key matrix.
[0044] In one possible implementation, the eighth intermediate matrix may be obtained based on the fourth query matrix and the fourth key matrix, which may include: the eighth intermediate matrix may be obtained based on the fifth query matrix and the fifth key matrix. The fifth query matrix may be obtained by performing the eleventh linear transformation on the fourth query matrix. For example, the fifth query matrix may be obtained by performing the eleventh linear transformation on the fourth key matrix. For example, the fifth key matrix may be obtained by performing the twelfth linear transformation on the fourth key matrix.
[0045] In another possible implementation, the eighth intermediate matrix may be obtained based on the fourth query matrix and the fourth key matrix, which may include: the eighth intermediate matrix may be obtained by performing a thirteenth linear transformation on the ninth intermediate matrix. For example, the eighth intermediate matrix may be obtained based on the ninth intermediate matrix and the thirteenth transformation matrix. The ninth intermediate matrix may be obtained by concatenating the fourth query matrix and the fourth key matrix.
[0046] In a possible implementation, the optimization strategy may include at least one of the following: a linear strategy, a sparse strategy, a quantization strategy, a pruning strategy, or other strategies.
[0047] By adopting an optimization strategy based on optimizing the cross-attention mechanism, it is possible to reduce the computational complexity while maintaining or improving the accuracy of feature registration, so that the video processing method described in the embodiment of the present application can be better applied to resource-constrained electronic devices, such as terminal devices.
[0048] In one possible implementation, the fourth key matrix may be obtained by performing an eighth linear transformation and sparse selection on the second feature data of the second image. For example, the fourth key matrix may be obtained by performing a sparse selection on the sixth key matrix. The sixth key matrix may be obtained by performing an eighth linear transformation on the second feature data of the second image.
[0049] Since the fourth key matrix is obtained by sparse selection of the sixth key matrix, the data volume of the fourth key matrix is reduced, thereby reducing the consumption of computing resources, so that the video processing method described in the embodiment of the present application can be better applied to resource-constrained electronic devices.
[0050] In one possible implementation, the third attention weight matrix may be obtained based on the fourth attention weight matrix. In one possible implementation, the third attention weight matrix may be obtained by quantizing the twelfth intermediate matrix. Optionally, the third attention weight matrix may be obtained by quantizing the twelfth intermediate matrix to a 4-bit or 8-bit fixed-point integer. The twelfth intermediate matrix may be obtained by pruning the fourth attention weight matrix. Optionally, the twelfth intermediate matrix may be obtained by setting the element values in the fourth attention weight matrix that are less than or equal to the second predetermined threshold to zero. The second predetermined threshold can be configured according to actual business needs and is not limited here. The fourth attention weight matrix may be obtained by normalizing the tenth intermediate matrix. The tenth intermediate matrix may be obtained by inverse quantizing the eleventh intermediate matrix. The eleventh intermediate matrix may be obtained based on the sixth query matrix and the sixth key matrix. The sixth query matrix and the sixth key matrix may be obtained by quantizing the fourth query matrix and the fourth key matrix.
[0051] Since the third attention weight matrix is obtained through quantization and cropping, the consumption of computing resources is reduced, so that the video processing method described in the embodiment of the present application can be better applied to electronic devices with limited resources.
[0052] In another possible implementation, the registered second feature data may be obtained by registering the second feature data of the second image and the first feature data of the first image using a cross-attention mechanism. This may include: the registered second feature data may be obtained based on a first query matrix, a first intermediate matrix, and a first value matrix. The first intermediate matrix may be used for a key. The first query matrix may be used for a query. The first value matrix may be used for a value. The first value matrix may be obtained based on the first intermediate matrix, the first key matrix, and the second value matrix. The first intermediate matrix may be used for a query. The first key matrix may be used for a key. The second value matrix may be used for a value. For example, the first value matrix may be obtained by processing the first intermediate matrix, the first key matrix, and the second value matrix using a cross-attention mechanism. Optionally, the first value matrix may be obtained based on a fifth attention weight matrix and the second value matrix. The fifth attention weight matrix may be obtained based on the first intermediate matrix and the first key matrix. The first query matrix may be obtained by performing a first linear transformation on the first feature data of the first image. For example, the first query matrix may be obtained based on the first feature data of the first image and a first transformation matrix. The element values in the first transformation matrix may be model parameters of a deep learning model. The first key matrix may be obtained by performing a second linear transformation on the second feature data of the second image. For example, the first key matrix can be obtained based on the second feature data and the second transformation matrix of the second image. The element values in the second transformation matrix can be model parameters of the deep learning model. The second value matrix can be obtained by performing a third linear transformation on the second feature data. For example, the second value matrix can be obtained based on the second feature data and the third transformation matrix. The first intermediate matrix can be obtained based on the first query matrix. For example, the first intermediate matrix can be obtained by pooling the first query matrix. Pooling can include maximum pooling, average pooling, or L2 norm pooling, etc. Optionally, the first intermediate matrix can be obtained based on the first query matrix and the fourteenth transformation matrix. The element values in the fourteenth transformation matrix can be model parameters of the deep learning model. Optionally, the first intermediate matrix can be obtained by fusing similar elements in the first query matrix.
[0053] Since the first value matrix is obtained based on the first intermediate matrix, the first key matrix, and the second value matrix, and the first intermediate matrix serves as the query, the first intermediate matrix can be used to aggregate information from the first key matrix and the second value matrix. Since the registered second feature data is obtained based on the first query matrix, the first intermediate matrix, and the first value matrix, and the first query matrix serves as the query and the first intermediate matrix serves as the key, it can be shown that the first intermediate matrix can transmit aggregated information to the first query matrix.
[0054] Since the first intermediate matrix is obtained based on the first query matrix, the data volume of the first intermediate matrix can be smaller than that of the first query matrix. Therefore, in the process of determining the first value matrix, the first value matrix is obtained based on the first intermediate matrix, the first key matrix, and the second value matrix. The first intermediate matrix serves as the query. Therefore, compared to using the first query matrix as the query, the data processing volume and computing resource consumption are reduced. In addition, the first value matrix aggregates the information of the first key matrix and the second value matrix. In the process of determining the aligned second feature data, the aligned second feature data is obtained based on the first value matrix obtained from the first query matrix, the first intermediate matrix, and the first value matrix. The first query matrix serves as the query, the first intermediate matrix serves as the key, and the first value matrix serves as the value. Therefore, the first value matrix can transmit the aggregated information to the first query matrix, thereby achieving global modeling. The first intermediate matrix can be used to aggregate the information of the first key matrix and the second value matrix and transmit the aggregated information to the first query matrix.
[0055] In one possible implementation, the first intermediate matrix may be obtained by fusing similar elements in the first query matrix, which may include: the first intermediate matrix may be obtained based on fused elements and non-fused elements. The fused element may be obtained by fusing a target element and other elements that have an associated relationship. The other element that has an associated relationship with the target element may be an element that is most similar to the target element, determined from other sets. The target element may be any one in the target set. When the target set is the first set, the other set is the second set. When the target set is the second set, the other set is the first set. The first set and the second set may be obtained by dividing the elements included in the first query matrix. Non-fused elements may refer to elements that have no associated relationship with other elements.
[0056] Because the first intermediate matrix can be obtained by fusing similar elements in the first query matrix, the data size of the first intermediate matrix is smaller than that of the first query matrix. Therefore, the first intermediate matrix serves as a query in determining the first value matrix and as a key in determining the registered second feature data, reducing data processing and computing resource consumption.
[0057] In one possible implementation, the registered second feature data may be obtained based on the first query matrix, the first intermediate matrix and the first value matrix, which may include: the registered second feature data may be obtained based on the second intermediate matrix and the third value matrix. For example, the registered second feature data may be obtained by adding the second intermediate matrix and the third value matrix. The second intermediate matrix may be obtained based on the first query matrix, the first intermediate matrix and the first value matrix. For example, the third intermediate matrix may be obtained based on the sixth attention weight matrix and the first value matrix. The sixth attention weight matrix may be obtained based on the first query matrix and the first intermediate matrix. The third value matrix may be obtained by performing a first depthwise separable convolution on the first value matrix.
[0058] Since the third value matrix is obtained by performing a first depth-wise separable convolution on the first value matrix, feature diversity is improved.
[0059] In another possible implementation, the registered second feature data may be obtained based on the first query matrix, the first intermediate matrix, and the first value matrix, which may include: the registered second feature data may be obtained based on a sixth attention weight matrix and the first value matrix. The sixth attention weight matrix may be obtained based on the first query matrix and the first intermediate matrix.
[0060] In one possible implementation, the registered second feature data may be obtained by registering the second feature data of the second image with the first feature data of the first image using a cross-attention mechanism. This may include: the registered second feature data may be obtained based on a third intermediate matrix and a second query matrix. For example, the registered second feature data may be obtained by multiplying the third intermediate matrix and the second query matrix. The third intermediate matrix may be obtained based on a second key matrix and a fourth value matrix. The second query matrix and the second key matrix may be obtained by processing the third query matrix and the third key matrix, respectively, using a focusing function. The focusing function may be used to adjust the proximity between the third query matrix and the third key matrix. That is, the clustering function may be used to bring similar elements in the third query matrix and the third key matrix closer together and dissimilar elements further apart. For example, the third query matrix may be input into a clustering function to obtain a second query matrix. The third key matrix may be input into a clustering function to obtain a second key matrix. The third query matrix may be obtained by performing a fourth linear transformation on the first feature data of the first image. For example, the third query matrix may be obtained based on the first feature data of the first image and the fourth transformation matrix. The third key matrix may be obtained by performing a fifth linear transformation on the second feature data of the second image. For example, the third key matrix may be obtained based on the second feature data of the second image and the fifth transformation matrix. The third intermediate matrix may be obtained based on the second key matrix and the fourth value matrix. For example, the third intermediate matrix may be obtained by multiplying the second key matrix and the fourth value matrix. The fourth value matrix may be obtained by performing a sixth linear transformation on the second feature data of the second image. For example, the fourth matrix may be obtained based on the second feature data of the second image and the sixth transformation matrix.
[0061] Since the aggregation function can be used to adjust the proximity between the third query matrix and the third key matrix, that is, the aggregation function can be used to make similar elements in the third query matrix and the third key matrix closer and dissimilar elements farther apart, thereby achieving the goal of assigning a larger weight to the second feature with a large similarity between the aligned second feature data and the first feature data, and assigning a smaller weight to the second feature in the second feature data with a small similarity to the first feature data, the accuracy of feature alignment is improved, thereby improving the video noise reduction effect. In addition, since the third intermediate matrix is obtained based on the second key matrix and the fourth value matrix, the aligned second feature data is obtained based on the second query matrix and the third intermediate matrix, that is, the key and value are calculated first, and the result is then calculated with the query, thereby reducing the amount of data processing and reducing the consumption of computing resources.
[0062] In one possible implementation, the first feature data may be determined based on the first visual feature data and the first semantic feature data. Alternatively, the first feature data may be determined based on the first visual feature data, the first semantic feature data, and a first position code. The first position code may be used to indicate a pixel position of a pixel in the first image.
[0063] The second feature data may be determined based on the second visual feature data and the second semantic feature data. Alternatively, the second feature data may be determined based on the second visual feature data, the second semantic feature data, and a second position code. The second position code may be used to indicate a pixel position of a pixel in the second image.
[0064] Since position coding can better capture the dependencies and patterns between features in feature data and can focus on the relative positions of features, the reliability of the first feature data and the second feature data is improved, thereby improving the video noise reduction effect.
[0065] In one possible implementation, the first query matrix may be obtained by performing a second depthwise separable convolution on the first feature data of the first image. And / or, the first key matrix may be obtained by performing a third depthwise separable convolution on the second feature data of the second image. And / or, the first value matrix may be obtained by performing a fourth depthwise separable convolution on the second feature data.
[0066] Since the first query matrix, the first key matrix, and the first value matrix can be obtained by using a depthwise separable convolution method, the depthwise separable convolution can reduce the amount of data processing, thereby reducing the amount of data processing and reducing computing resource consumption.
[0067] In one possible implementation, fusing the first feature data of the first image and the registered second feature data to obtain the video frame may include decoding the fused feature data to obtain the video frame. The fused feature data may be obtained by fusing the first feature data of the first image and the registered second feature data.
[0068] Because the fused feature data is obtained by fusing the first feature data with the registered second feature data, it can reflect the fusion characteristics of the first and registered second feature data, thereby making the information carried by the fused feature more comprehensive and accurate. On this basis, decoding the fused feature data to obtain video frames improves the video frame quality and thus enhances the video noise reduction effect.
[0069] In one possible implementation, the fused feature data may be obtained by fusing the first feature of the first image and the registered second feature data, which may include: the fused feature data may be obtained by concatenating the first feature data and the registered second feature data. Alternatively, the fused feature data may be obtained by adding the first feature data and the registered second feature data. Alternatively, the fused feature data may be obtained by multiplying the first feature data and the registered second feature data.
[0070] In another possible implementation, the fused feature data may be obtained by fusing the first feature of the first image and the registered second feature data, which may include: the fused feature data may be obtained based on the first feature data and the third feature data. For example, the fused feature data may be obtained by fusing the first feature data and the third feature data. The third feature data may be obtained based on the fourth feature data and the registered second feature data. For example, the third feature data may be obtained by multiplying the fourth feature data and the registered second feature data. The fourth feature data may be obtained based on the first feature data and the registered second feature data. For example, the fourth feature data may be obtained by convolving the first feature data with the registered second feature data.
[0071] Since the fused feature data is obtained based on the first feature data and the third feature data, the third feature data is obtained based on the fourth feature data and the aligned second feature data, and the fourth feature data is obtained based on the first feature data and the aligned second feature data, a more complete fusion of the first feature data and the second feature data is achieved, thereby making the information carried by the fused feature data more comprehensive and accurate.
[0072] In one possible implementation, decoding the fused feature data to obtain a video frame may include: the video frame may be obtained based on the first feature data and the fifth feature data. The fifth feature data may be obtained based on the sixth feature data. The sixth feature data may be obtained based on the seventh feature data and the fused feature data. The seventh feature data may be obtained based on the eighth feature data. The eighth feature data may be obtained based on the fused feature data and the ninth feature data. The ninth feature data may be obtained based on the fused feature data.
[0073] Since the video frame is obtained based on the first feature data and the fifth feature data, the fifth feature data is obtained based on the sixth feature data, the sixth feature data is obtained based on the seventh feature data and the fused feature data, the seventh feature data is obtained based on the eighth feature data, the eighth feature data is obtained based on the fused feature data and the ninth feature data, and the ninth feature data can be obtained based on the fused feature data, therefore, more full utilization of the feature data is achieved, thereby improving the quality of the video frame.
[0074] In a possible implementation manner, the image brightness of the first image and the second image are aligned, and / or the image brightness of the first image and the second image are aligned and the field of view angles are aligned.
[0075] By aligning the brightness of the first and second images, the brightness ranges of the first and second images are made the same, thereby improving the accuracy of image registration and image fusion, and thus enhancing the video noise reduction effect. By aligning the field of view of the first and second images, the accuracy of image registration and image fusion is improved, and thus enhancing the video noise reduction effect.
[0076] In one possible implementation, the image brightness of the second image may be determined based on the original image brightness of the second image, the amount of light entering the second image, and the amount of light entering the first image. The amount of light entering the first image may be determined based on the exposure time and sensitivity of the first image. The amount of light entering the second image may be determined based on the exposure time and sensitivity of the second image. For example, the image brightness of the second image may be determined based on a ratio and the original image brightness of the second image. The ratio may be the ratio between the amount of light entering the second image and the amount of light entering the first image.
[0077] By adjusting the image brightness, the brightness range of multiple images is made the same, that is, the brightness of multiple images is aligned. This improves the accuracy of image registration and image fusion, and further enhances the video noise reduction effect.
[0078] In one possible implementation, the target image may be obtained by image cropping. The target image may be an image captured by a target camera among the first and second images. The target camera may be a camera with a larger field of view among the multiple cameras. Alternatively, the target camera may be a camera with a longer exposure time among the multiple cameras. Alternatively, the target camera may be a camera with a larger field of view and a longer exposure time among the multiple cameras.
[0079] By using image cropping to align the field of view of multiple images, the accuracy of image registration and image fusion is improved, thereby enhancing the video noise reduction effect. In addition, it also reduces the amount of data processing and the consumption of computing resources.
[0080] In a possible implementation, the first image and the second image may be RAW images.
[0081] Since the RAW image has not been processed by the image signal processor, ISP processing will make the noise more complicated. Therefore, the first image and the second image are RAW images, which reduces the difficulty of video noise reduction and improves the video noise reduction effect.
[0082] In a second aspect, an embodiment of the present application provides a video processing device, which may include: a recording module for recording a video in response to a first operation. A processing module for reducing the noise of a first image based on a second image during video recording to obtain a video frame. The video frame may be a frame in the recorded video. The first image and the second image may be captured using multiple cameras. The shooting time of the first image and the shooting time of the second image may overlap. The lens direction of the first camera may be the same as the lens direction of the second camera. The first camera may be a camera that captures the first image. The second camera may be a camera that captures the second image. The exposure time of the first image may be less than the exposure time of the second image.
[0083] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory is used to store computer-executable instructions, and the processor is used to run the computer-executable instructions stored in the memory, so that the electronic device executes the method described in any possible implementation of the first aspect.
[0084] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program or instruction is stored. When the computer program or instruction is run on an electronic device, the electronic device executes the method described in any possible implementation of the first aspect.
[0085] In a fifth aspect, an embodiment of the present application provides a computer program product comprising a computer program, which, when the computer program is run on an electronic device, enables the electronic device to execute the method described in any possible implementation manner of the first aspect.
[0086] In a sixth aspect, the present application provides a chip or chip system, comprising at least one processor and a communication interface, wherein the communication interface and the at least one processor are interconnected via a line, and the at least one processor is configured to execute a computer program or instruction to perform the method described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip may be an input / output interface, a pin, or a circuit.
[0087] In one possible implementation, the chip or chip system described above in this application further includes at least one memory, wherein instructions are stored in the at least one memory. The memory may be a storage unit within the chip, such as a register or cache, or a storage unit of the chip (such as a read-only memory or random access memory).
[0088] It should be understood that the second to sixth aspects of the present application correspond to the technical solutions of the first aspect of the present application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation methods are similar and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] Figure 1 A schematic diagram of an image output mode provided in an embodiment of the present application;
[0090] Figure 2 A schematic diagram of video recording according to an embodiment of the present application;
[0091] Figure 3 A schematic diagram illustrating the principle of the video processing method provided in an embodiment of the present application;
[0092] Figure 4 This is a schematic diagram of the principle of using image registration and image fusion to achieve video noise reduction proposed in an embodiment of the present application;
[0093] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0094] Figure 6 A software structure diagram of the electronic device provided in the embodiment of the present application;
[0095] Figure 7 A flowchart of a video processing method proposed in an embodiment of the present application;
[0096] Figure 8A A schematic diagram of an interface for a dark light recording mode provided in an embodiment of the present application;
[0097] Figure 8B A schematic diagram of another interface of a low-light recording mode provided in an embodiment of the present application;
[0098] Figure 9A A schematic diagram of starting a dark light recording mode provided in an embodiment of the present application;
[0099] Figure 9B A schematic diagram of another method of activating a low-light recording mode according to an embodiment of the present application;
[0100] Figure 9C A schematic diagram of another method of activating a low-light recording mode according to an embodiment of the present application;
[0101] Figure 9D A schematic diagram of a dark light recording mode provided in an embodiment of the present application;
[0102] Figure 10 A schematic diagram of a principle for obtaining registered second feature data according to an embodiment of the present application;
[0103] Figure 11 A schematic diagram of another principle for obtaining the registered second feature data provided in an embodiment of the present application;
[0104] Figure 12 A schematic diagram of the principle of obtaining fused feature data provided in an embodiment of the present application;
[0105] Figure 13 A flowchart of another video processing method provided in an embodiment of the present application;
[0106] Figure 14 This is a structural block diagram of the video processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0107] In the embodiments of this application, terms such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. For example, the terms "first chip" and "second chip" are used solely to distinguish between different chips and do not define their order. Those skilled in the art will understand that terms such as "first" and "second" do not define the quantity or execution order, and do not necessarily define differences.
[0108] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0109] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, a--c, bc, or abc, where a, b, c can be single or plural.
[0110] In order to clearly describe the technical solutions of the embodiments of the present application, some terms or concepts involved in the embodiments of the present application are explained below.
[0111] Image clarity can be a measure of the quality of an image or video frame, which reflects the details and sharpness of the visual content.
[0112] Image brightness can refer to the light intensity or brightness of an image or video frame, which can affect the visual consistency and clarity of the image or video frame.
[0113] A video may refer to a data stream that encodes video frames in a time sequence. A video may include multiple video frames. Video quality may be evaluated by at least one of image clarity, image brightness, or smoothness.
[0114] Frame rate refers to the number of frames recorded or played per second. Frame rate can be measured in FPS (Frames Per Second). The interframe interval (or frame duration) refers to the time interval between consecutive frames in a video, or the duration of each frame. The interframe interval is the reciprocal of the frame rate. For example, for a frame rate of 30 FPS, the interframe interval is 1 / 30 second.
[0115] The frame rate may refer to the shooting frame rate or the playback frame rate, and may be determined based on actual business needs. It should be noted that in the embodiments of the present application, the frame rate may refer to the shooting frame rate. More specifically, the shooting frame rate may refer to the number of frames recorded per second by the camera during shooting. The camera may support at least one shooting frame rate, for example, 30 FPS, 90 FPS, 120 FPS, or 480 FPS. The shooting frame rates supported by different cameras may be the same or different.
[0116] Noise refers to factors that affect the understanding of source information. Source information can include images or videos. Noise not only affects the quality of source information but can also interfere with subsequent tasks based on the source information. For example, subsequent tasks can include at least one of the following: image classification, object detection, image recognition, image segmentation, person re-identification, multimodal matching, image retrieval, or image description.
[0117] The signal-to-noise ratio (SNR) can be used to assess the quality of source information. It represents the ratio of useful information to noise in the source. A higher SNR indicates a higher proportion of useful information and a lower proportion of noise in the source, providing clearer and higher-quality source information.
[0118] Exposure refers to the process by which a camera's photosensitive element receives light. Exposure determines the brightness (or lightness / darkness) of an image. Parameters that influence image brightness include exposure time (ET), ISO sensitivity, and aperture value (AV).
[0119] The shutter controls how long light enters the camera, thus determining the image exposure duration. The longer the shutter is open, the more light enters the camera, resulting in a longer image exposure duration. The shorter the shutter is open, the less light enters the camera, resulting in a shorter image exposure duration.
[0120] Exposure duration (or exposure time, shutter speed) can refer to the duration between the shutter opening and closing. The exposure duration can be less than or equal to the frame time interval. Based on the exposure duration, exposure can be divided into long exposure, short exposure, and medium exposure, thereby obtaining long-frame images, short-frame images, and medium-frame images. Long exposure can have a longer exposure duration. Short exposure can have a shorter exposure duration. The exposure duration of medium exposure is longer than that of short exposure and shorter than that of long exposure. Long-frame images can refer to images obtained under long exposure. Short-frame images can refer to images obtained under short exposure. Medium-frame images can refer to images obtained under medium exposure.
[0121] Sensitivity refers to the sensitivity of the camera's photosensitive element to light, and can be represented by ISO (International Organization for Standardization).
[0122] The aperture is an adjustable opening in a camera lens, the size of which determines the amount of light entering. Light intake refers to the amount of light that enters the camera's photosensitive element. The aperture value is the ratio of the camera lens' focal length to the aperture diameter. A larger aperture value allows more light to enter, while a smaller aperture value reduces the amount of light entering. The amount of light entering can be determined by the image's exposure time and ISO sensitivity.
[0123] A low-light scene may refer to a scene with low ambient light intensity (or ambient illumination) and / or low ambient light brightness. As an implementation method, a low-light scene may refer to a scene with ambient light intensity less than or equal to a predetermined ambient light intensity. As another implementation method, a low-light scene may refer to a scene with ambient light brightness less than or equal to a predetermined ambient light brightness. The predetermined ambient light intensity and predetermined ambient light brightness can be configured according to actual business needs and are not limited here. For example, the predetermined ambient light intensity may be less than or equal to 5 lux. For example, low-light scenes may include outdoor low-light scenes and indoor low-light scenes. Outdoor low-light scenes may include night scenes, cloudy scenes, foggy scenes, tunnel scenes, etc. Ambient light intensity may refer to the luminous flux received per unit area, describing the distribution of light in space. Ambient light brightness may refer to the intensity of light emitted by a light source or the surface of the object being photographed.
[0124] Field of View (FOV) refers to the range of view (or content) that a camera can cover. A larger FOV means a larger range of view. A smaller FOV means a smaller range of view.
[0125] Image registration is the process of transforming different images of the same scene into the same coordinate system. Image registration is a critical task used to align images acquired at different times, viewing angles, or sensors for subsequent analysis and comparison. Image registration can be used to ensure pixel consistency across multiple images.
[0126] Image registration refers to the process of matching and aligning multiple images of the same scene captured under different conditions (e.g., at different times, from different perspectives, or with different cameras). For example, the images could be taken at the same time from different perspectives, at different times from the same perspective, or at different times from different perspectives.
[0127] Video noise reduction (or denoising) is used to suppress noise in videos while preserving texture details as much as possible, restoring the real scene and improving video reliability and usability. The demand for video noise reduction is widespread, covering a wide range of fields, such as security, autonomous driving, and audio and video entertainment.
[0128] The Cross-Attention mechanism enables interaction and matching of different images at the feature level, thereby capturing the similarities and differences between features, focusing on similarities with high weights and suppressing differences with low weights, and can adaptively adjust weights. Therefore, it has higher robustness and scalability.
[0129] The RAW domain (or RAW format) is an unprocessed format. RAW image can also refer to an unprocessed image.
[0130] The image output modes may include a high-definition mode, a low signal-to-noise ratio mode, and a low-definition, high signal-to-noise ratio mode.
[0131] For example, a high-definition, low signal-to-noise ratio mode can include a pixel-binning mode. Alternatively, the pixel-binning mode can include a Quadr mode. Quadr mode can combine the charges induced by adjacent pixels in low light conditions and read them out as a single pixel. Quadr mode can also read out images in brighter light conditions as a single pixel, improving image resolution.
[0132] For example, a low-definition, high signal-to-noise ratio image output mode may include a pixel merging mode, etc. Optionally, the pixel merging mode may include a Binning mode. The Binning mode may add together the charges induced by adjacent pixels and read them out in a one-pixel mode. The Binning mode may reduce the image resolution, increase the photosensitive area, and improve the signal-to-noise ratio while maintaining the field of view angle. The Binning mode may include a horizontal Binning mode or a vertical Binning mode. Horizontal Binning is to add together the charges induced by pixels in adjacent rows and read them out. Vertical Binning is to add together the charges induced by pixels in adjacent columns and read them out.
[0133] The following combination Figure 1 Describes pixel-merging modes (e.g., Quadr mode) and pixel-merging modes (e.g., Binning mode).
[0134] Figure 1 A schematic diagram of the image output mode provided in an embodiment of the present application. Figure 1 In the figure, "R" indicates a red pixel, "G" indicates a green pixel, and "B" indicates a blue pixel.
[0135] like Figure 1 As shown, the pixel unit may include one red sub-pixel unit, two green sub-pixel units, and one blue sub-pixel unit. The red sub-pixel unit may include four red pixels, the green sub-pixel unit may include four green pixels, and the blue sub-pixel unit may include four blue pixels.
[0136] It should be noted that the Quadr mode in the embodiment of the present application can be read out in a single pixel mode under low light conditions, for example, Figure 1 The Binning mode in the embodiment of the present application can adopt, for example, Figure 1 Pixel binning mode in . And, Figure 1 This is merely an illustrative example and does not constitute a limitation on the embodiments of the present application.
[0137] The inventive concept of the embodiments of the present application is described below with reference to the accompanying drawings.
[0138] In scenarios where video quality is significantly affected by noise, video noise reduction is required to improve video quality. For example, low-light scenes may be a scenario where video quality is significantly affected by noise.
[0139] In related technologies, video noise reduction can be achieved by extending the exposure time. This is because extending the exposure time can increase the amount of light entering, which can reduce noise and improve the signal-to-noise ratio. For example, during video recording, a single camera can be used to shoot with a longer exposure time to obtain a noise-reduced video frame.
[0140] We found that to achieve good video noise reduction, the exposure time needs to be greater than or equal to 100ms. However, in video recording, the exposure time is limited by the frame rate, that is, the exposure time needs to be less than or equal to the frame interval. For example, if the frame interval is 30ms, even a longer exposure time will not meet the requirement. Therefore, when balancing the frame rate, the above-mentioned video noise reduction method is difficult to achieve good noise reduction effect. Therefore, it is necessary to improve the video noise reduction effect.
[0141] To improve video noise reduction, it was found that long-frame images have a high signal-to-noise ratio, low definition, richer dark and light area information, and richer color information. Short-frame images can balance the frame rate, preserve the scene's texture edges and highlight area information, and have high definition and a low signal-to-noise ratio. This shows that long-frame images and short-frame images can complement each other to a certain extent. For example, the high-noise sharp texture of the short-frame image and the low-noise blurred texture of the long-frame image can constrain each other, facilitating detail recovery. The highlight area information of the short-frame image can repair the overexposed cutoff area of the long-frame image. In addition, the color information of the long-frame image can correctly restore the color of the scene. Therefore, the embodiments of the present application take into account the advantages of short-frame images in balancing frame rate and / or high definition, and the advantages of long-frame images in improving signal-to-noise ratio, and propose that long-frame images can be used as a supplement to short-frame images to improve the quality of short-frame images, thereby improving the video noise reduction effect while balancing the frame rate. The short-frame image can be used as a reference image, and the long-frame image can be used as an auxiliary image.
[0142] Next, we need to consider how to obtain long-frame images and short-frame images, and we find that there are single-camera and multi-camera methods.
[0143] The single-camera approach uses a single camera to capture both long- and short-frame images. However, as mentioned above, the exposure time of the long-frame images obtained with this approach is limited by the frame rate. Therefore, this approach makes it difficult to obtain long-frame images that meet the requirements for video noise reduction.
[0144] Based on a multi-camera approach, long-frame and short-frame images are captured using different cameras. Since long-frame and short-frame images are captured by different cameras, the exposure duration of the long-frame image can exceed the frame rate limit if the exposure duration of the short-frame image can match the frame rate requirement. Therefore, this approach can produce long-frame images that meet video noise reduction requirements.
[0145] Next, we need to consider how to meet the needs of multiple cameras. We have found that with the continuous development of electronic devices, the functions of electronic devices have been enriched and improved. For example, multiple cameras can provide different shooting functions. The field of view of different cameras can be different. Therefore, electronic devices can be equipped with multiple cameras to provide multi-camera functions. Multiple cameras can include at least two of the following: main camera, telephoto camera, wide-angle camera, ultra-wide-angle camera, periscope telephoto camera, macro camera, fisheye camera, infrared camera, depth camera, black and white camera, color camera, etc.
[0146] Based on this, an embodiment of the present application proposes that the multi-camera function provided by the electronic device can be used to obtain images with different exposure times to achieve video noise reduction. For example, in the case of video recording, multiple cameras are used to capture multiple images of the subject to obtain video frames in the video. In addition to cameras that use shorter exposure times to balance the frame rate, there are also cameras that use longer exposure times to improve the signal-to-noise ratio. Thus, an image with a shorter exposure time (for example, a short-frame image) among the multiple images can be used as a reference image, and an image with a longer exposure time (for example, a long-frame image) among the multiple images can be used as an auxiliary image. The auxiliary image is used to reduce the noise of the reference image to obtain a noise-reduced reference image. The noise-reduced reference image is a video frame in the video.
[0147] In the case of video recording, the reference image is obtained by shooting the subject with a camera that uses a shorter exposure time to balance the frame rate, and the auxiliary image is obtained by shooting the subject with a camera that uses a longer exposure time to improve the signal-to-noise ratio. Therefore, using the auxiliary image to reduce the noise of the reference image to obtain the video frame in the video can improve the video noise reduction effect while balancing the frame rate.
[0148] It should be noted that the multiple images may include at least one auxiliary image. Different auxiliary images may have different exposure times. Different auxiliary images may be captured by different cameras. The exposure time of the auxiliary image may be longer than that of the reference image. For example, a mid-frame image may serve as an auxiliary image. Thus, the denoised reference image may be obtained by denoising the reference image using at least one auxiliary image.
[0149] It should also be noted that the image output modes of different cameras can be the same or different, and this is not limited here. For example, the first camera can adopt a high-definition, low signal-to-noise ratio image output mode. The second camera can adopt a low-definition, high signal-to-noise ratio image output mode. Optionally, both the first camera and the second camera can adopt a high-definition, low signal-to-noise ratio image output mode. Optionally, both the first camera and the second camera can adopt a low-definition, high signal-to-noise ratio image output mode. The first camera may refer to a camera that uses a shorter exposure time to balance the frame rate. The second camera may refer to a camera that uses a longer exposure time to improve the signal-to-noise ratio.
[0150] As an implementation method, the exposure duration of the reference image can be set to be less than or equal to the frame interval of the reference image to balance the frame rate. The exposure duration of the auxiliary image can be set to be greater than the frame interval of the reference image to improve the signal-to-noise ratio. For example, the frame interval of the reference image can be 30ms. The exposure duration of the reference image can be less than or equal to 30ms. Optionally, the exposure duration of the reference image can be 20ms. The exposure duration of the auxiliary image can be greater than or equal to 100ms. Optionally, the exposure duration of the auxiliary image can be 150ms.
[0151] As an implementation method, the shooting time of the reference image and the shooting time of the auxiliary image may overlap, and the lens orientation of the first camera and the lens orientation of the second camera may be the same, so that the content of the reference image and the auxiliary image are similar. It should be noted that in the embodiment of the present application, the lens orientation of the first camera and the lens of the second camera are the same, but not absolutely the same, and the two lens orientations can be roughly the same. The two lens orientations can make the content of the reference image and the auxiliary image similar. For example, the first camera and the second camera are both front cameras. Optionally, the first camera and the second camera are both rear cameras.
[0152] As an implementation, the reference image may be obtained by photographing the subject with the first camera at intervals of a first predetermined duration. The auxiliary image may be obtained by photographing the subject with the second camera at intervals of a second predetermined duration. The first predetermined duration is greater than or equal to the exposure duration of the reference image and less than the second predetermined duration. The second predetermined duration is greater than or equal to the exposure duration of the auxiliary image. For example, the first predetermined duration may be 30 ms, and the second predetermined duration may be 150 ms.
[0153] In order to better understand the video recording process of the embodiment of the present application, the following takes multiple cameras including a first camera and a second camera as an example. Figure 2 Provide explanation.
[0154] Figure 2 A schematic diagram of video recording provided according to an embodiment of the present application.
[0155] like Figure 2 As shown, the first camera can capture reference images 1, 2, 3, 4, and 5 at intervals of a first predetermined duration. The second camera can capture auxiliary image 1 at intervals of a second predetermined duration. Each time an auxiliary image is captured, five reference images can be obtained.
[0156] The exposure duration of each reference image is the first exposure duration. The exposure duration of the auxiliary image is the second exposure duration. The first exposure duration is less than the first predetermined duration. The second exposure duration is equal to the second predetermined duration. The first exposure duration is less than the second exposure duration.
[0157] It should be noted that Figure 2 This is merely an illustrative example and does not constitute a limitation on the embodiments of the present application.
[0158] In order to better understand the inventive concept of the embodiment of the present application, Figure 3 Further explanation is provided.
[0159] Figure 3 This is a schematic diagram of the principle of the video processing method provided in an embodiment of the present application. The method can be applied to electronic devices.
[0160] like Figure 3 As shown, the electronic device may include N cameras, namely, a first camera, ..., an nth camera, ..., an Nth camera. N may be an integer greater than 1. n may be an integer greater than 1 and less than or equal to N. The lenses of the N cameras may have the same orientation. The exposure duration corresponding to the nth camera may be the nth exposure duration. The nth exposure duration is greater than the (n-1)th exposure duration. The first camera may use a shorter exposure duration to balance the frame rate. The nth to Nth cameras may use longer exposure durations to improve the signal-to-noise ratio. The image captured by the first camera may serve as a reference image. The images captured by the nth to Nth cameras may serve as auxiliary images.
[0161] In the case of video recording, multiple cameras can capture the subject according to their respective exposure times to obtain at least one image corresponding to each of the multiple cameras.
[0162] Specifically, the first camera may capture the object according to the first exposure time to obtain first images 1, ..., first images i, .... i may be an integer greater than or equal to 1.
[0163] The nth camera may capture the subject according to the nth exposure time to obtain the nth image 1, ..., nth image j, .... j may be an integer greater than or equal to 1 and less than i.
[0164] The Nth camera may capture the subject according to the Nth exposure duration to obtain the Nth image 1, ..., the Nth image k, .... k may be an integer greater than or equal to 1 and less than or equal to j.
[0165] The following uses the first image 1, ..., the nth image 1, ..., and the Nth image 1 as examples to illustrate how to implement video noise reduction. The first image 1 serves as the reference image. The nth through Nth images 1 serve as auxiliary images. The capture times of the first through Nth images 1 overlap.
[0166] When 2 ≤ n ≤ N, the first image 1 (i.e., the reference image) is denoised using the nth image 1 (i.e., the auxiliary image) to obtain intermediate data. A denoised first image 1 is obtained based on the N-1 intermediate data. Thus, denoising the first image 1 is achieved using the nth to Nth images 1 to obtain the denoised first image 1. The denoised first image 1 is a video frame in the video.
[0167] As an implementation, the first camera may be a telephoto camera, and the second camera may be a main camera.
[0168] In addition, in order to further improve the video noise reduction effect, the embodiment of the present application proposes that the clarity of the reference image can be made higher than the clarity of the auxiliary image.
[0169] In addition, to further improve the video noise reduction effect, the embodiments of the present application propose that this can be achieved by combining image output modes. The first camera can use a high-definition, low signal-to-noise ratio image output mode. The second camera can use a low-definition, high signal-to-noise ratio image output mode. As a result, the reference image has high definition and low signal-to-noise ratio, while the auxiliary image has low definition and high signal-to-noise ratio, and the reference image has higher definition than the auxiliary image.
[0170] As an implementation, a high-definition, low signal-to-noise ratio image output mode may include a pixel-binning mode. For example, the pixel-binning mode may include a Quadr mode. A low-definition, high signal-to-noise ratio image output mode may include a pixel-binning mode, etc. For example, the pixel-binning mode may include a Binning mode.
[0171] In addition, regarding how to use the auxiliary image to reduce the noise of the reference image to obtain the video frame, the embodiment of the present application proposes that it can be achieved by using image registration and image fusion. Figure 4 Provide explanation.
[0172] Figure 4 This is a schematic diagram of the principle of using image registration and image fusion to achieve video noise reduction proposed in an embodiment of the present application.
[0173] As a way to implement Figure 4 The implicit denoising method shown here refers to performing feature registration based on feature data extracted from the image, fusing the registered feature data, and finally obtaining a denoised video frame based on the fused feature data. Therefore, the implicit denoising method can be understood as feature-level denoising.
[0174] For example, a video frame can be obtained based on fused feature data. The fused feature data can be obtained by fusing the reference feature data of a reference image with the registered auxiliary feature data. The registered auxiliary feature data can be obtained by registering the auxiliary feature data of an auxiliary image with the reference feature data. That is, the auxiliary feature data of the auxiliary image and the reference feature data of the reference image can be registered to obtain the registered auxiliary feature data. Next, the reference feature data and the registered auxiliary feature data are fused to obtain the fused feature data. Finally, a video frame is obtained based on the fused feature data.
[0175] It should be noted that the feature data may include at least one of the following: semantic feature data or visual feature data. Semantic feature data may be used to indicate the semantic information expressed by the image. Semantic feature data may include at least one of the following: local semantic feature data or global semantic feature data. Local semantic feature data may be used to indicate the local association between images. The receptive field of global semantic feature data is larger than that of local semantic feature data. Visual feature data may include at least one of the following: shallow visual feature data or deep visual feature data. Shallow visual feature data may be used to indicate the fine-grained visual features of the image. Fine-grained visual features may include at least one of the following: color features, texture features, edge features, or angular features, etc. Deep visual feature data may be used to indicate the coarse-grained visual features of the image. Coarse-grained visual features may refer to abstract visual features. Abstract visual features may refer to visual features that can express semantic information.
[0176] Since the feature data abstracts the image information while retaining the detail information, it has higher robustness and discriminability, thereby improving the reliability and robustness of video noise reduction and improving the video noise reduction effect.
[0177] As another implementation, Figure 4 The explicit denoising method shown in FIG. Explicit denoising can refer to registering multiple images and obtaining a denoised video frame based on the registered images. Therefore, the explicit denoising method can be understood as image-level denoising.
[0178] For example, a video frame can be obtained by fusing a reference image with a registered auxiliary image. A registered auxiliary image can be obtained by registering the auxiliary image with the reference image. That is, the auxiliary image and the reference image can be registered to obtain a registered auxiliary image. Next, the reference image and the registered auxiliary image are fused to obtain a video frame.
[0179] Since image-level denoising can make full use of pixel information, retain image detail information as much as possible, and has good interpretability, it improves the video denoising effect.
[0180] Furthermore, in order to improve the accuracy of feature registration based on implicit noise reduction, it was found that the association between auxiliary feature data and reference feature data can be mined to make the registered auxiliary feature data as similar as possible to the reference feature data. Therefore, the embodiment of the present application proposes that this can be achieved by adjusting feature weights based on similarity. The auxiliary feature data can include multiple auxiliary features. The reference feature data can include multiple reference features. The registered auxiliary feature data can include multiple registered auxiliary features.
[0181] For example, Figure 4 As shown, the registered auxiliary feature data can be obtained based on the similarity data and the auxiliary feature data. The similarity data may include the similarity corresponding to each of the multiple auxiliary features. The similarity corresponding to the auxiliary feature can be used to indicate the degree of similarity between the auxiliary feature and the reference feature data. For example, the greater the similarity corresponding to the auxiliary feature, the greater the similarity between the auxiliary feature and the reference feature data. The smaller the similarity corresponding to the auxiliary feature, the smaller the similarity between the auxiliary feature and the reference feature data. Therefore, based on the similarity data, each auxiliary feature in the auxiliary feature data can be assigned a weight to reflect the importance of the auxiliary feature in the feature data registration. For example, a larger weight is assigned to the auxiliary feature in the auxiliary feature data with a large similarity to the reference feature data, and a smaller weight is assigned to the auxiliary feature in the auxiliary feature data with a small similarity to the reference feature data. Furthermore, based on the auxiliary features and the weights of the auxiliary features, the registered auxiliary features and the registered auxiliary features are obtained.
[0182] In addition, it was found that the cross-attention mechanism can achieve interaction and matching of different images at the feature level, thereby capturing the similarities and differences between features, focusing on similarities with high weights, suppressing differences with low weights, and adaptively adjusting weights, therefore, having higher robustness and scalability. Therefore, an embodiment of the present application proposes that feature weights can be adjusted in a manner based on a cross-attention mechanism. For example, the aligned auxiliary feature data can be obtained by aligning the auxiliary feature data and reference feature data using a cross-attention mechanism. Reference feature data is used for queries. Auxiliary feature data is used for keys and values.
[0183] The cross-attention mechanism may include at least one of the following: a single-head cross-attention mechanism or a multi-head cross-attention mechanism. The similarity determination method involved in the cross-attention mechanism may include at least one of the following: dot product similarity, scaled dot product similarity, additive similarity or cosine similarity, etc. The cross-attention mechanism based on dot product similarity may be called a dot product attention mechanism. The cross-attention mechanism based on scaled dot product similarity may be called a scaled dot product attention mechanism. The cross-attention mechanism based on additive similarity may be called an additive attention mechanism. The cross-attention mechanism based on cosine similarity may be called a cosine attention mechanism.
[0184] In addition, it is found that since the video processing method described in the embodiment of the present application can be applied to resource-constrained electronic devices, such as terminal devices, the computational complexity of the cross-attention mechanism is proportional to the square of the number of feature data included in the feature data. Therefore, it is expected that the computational complexity can be reduced while maintaining or improving the accuracy of feature registration. To this end, the embodiment of the present application proposes that the cross-attention mechanism can be optimized based on an optimization strategy to achieve this.
[0185] As an implementation, the optimization strategy may include at least one of the following: a linear strategy, a sparse strategy, a quantization strategy, or a pruning strategy. A cross-attention mechanism based on a linear strategy may be referred to as a linear attention mechanism. A cross-attention mechanism based on a sparse strategy may be referred to as a sparse attention mechanism.
[0186] In addition, it is found that there are differences in the field of view, displacement, etc. of multiple images, which increases the difficulty of registration and reduces the accuracy of registration, thereby affecting the video noise reduction effect. To this end, the embodiment of the present application proposes that video noise reduction can be achieved by using a method based on a deep learning model.
[0187] For example, the auxiliary image and the reference image can be input into a video denoising model to obtain a video frame. The video denoising model can be obtained by training a deep learning model using the first sample image, the second sample image, and the third sample image. The third sample image can be a denoised image corresponding to the first sample image. Specifically, the video denoising model can be obtained by adjusting model parameters of the deep learning model based on a loss function value. The loss function value can be obtained by inputting the sample denoised image and the third sample image into the loss function. The sample denoised image can be obtained by inputting the first sample image and the second sample image into the deep learning model.
[0188] The exposure time of the first sample image may be shorter than the exposure time of the second sample image. The clarity of the first sample image may be higher than the clarity of the second sample image. The first sample image, the second sample image, and the third sample image may be captured using multiple cameras. The capture times of the first sample image and the second sample image may overlap. The lens orientation of the camera used to capture the first sample image, the lens orientation of the camera used to capture the second sample image, and the lens orientation of the camera used to capture the second sample image may be the same.
[0189] As an implementation, the exposure duration of the second sample image may be greater than the frame interval of the first sample image. The exposure duration of the first sample image may be less than or equal to the frame interval of the first sample image. For example, the frame interval of the first sample image may be 30 ms. The exposure duration of the first sample image may be less than or equal to 30 ms. Optionally, the exposure duration of the first sample image may be 20 ms. The exposure duration of the second sample image may be greater than or equal to 100 ms. Optionally, the exposure duration of the second sample image may be 150 ms.
[0190] It should be noted that the model structure of the deep learning model can be configured according to actual business needs as long as it can improve the video noise reduction effect, and is not limited here.
[0191] For example, a deep learning model may include a cascaded feature extraction module, a feature registration module, a feature fusion module, and a noise reduction module. The feature extraction module can be used for feature extraction. The feature registration module can be used to implement feature registration using a cross-attention mechanism. The feature fusion module can be used for feature fusion. The noise reduction module can be used to obtain video frames. Based on this, how to use the video noise reduction model to obtain video frames can be achieved through the following methods.
[0192] The auxiliary image and the reference image are input into the feature extraction module to obtain auxiliary feature data and reference feature data. As an implementation method, the feature extraction module may include feature extraction units corresponding to each of the multiple images. The model parameters of the multiple feature extraction units are different, but the model structure is the same. Therefore, feature extraction is achieved based on a non-sharing model parameter method, which can more effectively reflect the characteristics of each of the multiple images. As another implementation method, the auxiliary image and the reference image are processed using the same feature extraction module to obtain auxiliary feature data and reference feature data. Therefore, feature extraction is achieved based on a model parameter sharing method, which can make the space and semantics of the features consistent and reduce the number of model parameters.
[0193] The auxiliary feature data and the reference feature data are input into the feature registration module to obtain the registered auxiliary feature data. The feature registration module can be configured to obtain a third attention weight matrix based on the fourth query matrix and the fourth key matrix. According to the third attention weight matrix and the fifth value matrix, the registered auxiliary feature data are obtained. The fourth query matrix can be obtained based on the reference feature data and the seventh transformation matrix. The fourth key matrix can be obtained based on the auxiliary feature data and the eighth transformation matrix. The fifth value matrix can be obtained based on the auxiliary feature data and the ninth transformation matrix. The element values in the seventh transformation matrix, the eighth transformation matrix and the ninth transformation matrix can be model parameters of the deep learning model.
[0194] The reference feature data and the registered auxiliary feature data are input into the feature fusion module to obtain fused feature data. The fused feature data is input into the noise reduction module to obtain a video frame.
[0195] In addition, it was found that due to the different exposure times of multiple images, the image brightness ranges of the multiple images may be different. The difference in image brightness ranges will affect image registration and image fusion, and thus affect the video noise reduction effect. To this end, the embodiment of the present application proposes that the image brightness adjustment method can be used to make the image brightness ranges of multiple images the same, that is, the image brightness of multiple images are aligned. For example, since the auxiliary image needs to be aligned with the reference image, the image brightness alignment with the reference image can be achieved by adjusting the image brightness of the auxiliary image. As an implementation method, a first ratio between the amount of light entering the auxiliary image and the amount of light entering the reference image can be determined. A second ratio between the original image brightness of the auxiliary image and the first ratio is determined to obtain the image brightness of the auxiliary image. The amount of light entering the auxiliary image can be determined based on the exposure time and sensitivity of the auxiliary image. The amount of light entering the reference image can be determined based on the exposure time and sensitivity of the reference image.
[0196] In addition, it is found that the field of view angles of multiple cameras may be different, which will affect subsequent image fusion and increase computing resource consumption. To this end, the embodiment of the present application proposes that an image cropping method can be used to achieve the alignment of the field of view angles of multiple images. For example, the original image can be cropped to obtain a target image. The target image can be an image captured by a target camera among the multiple images. The target camera can be a camera with a larger field of view angle among the multiple cameras and / or the target camera can be a camera with a longer exposure time among the multiple cameras.
[0197] As an implementation manner, image cropping may include at least one of the following: proportional cropping, center cropping, edge cropping, etc. The embodiment of the present application does not limit the image cropping manner.
[0198] In addition, it is found that RAW images are not processed by an image signal processor (ISP), and ISP processing will complicate the noise. Therefore, the embodiment of the present application proposes that the multiple images can be RAW images.
[0199] The electronic device applicable to the video processing method provided in the embodiment of the present application and the specific process of the method are described below in conjunction with the embodiment of the present application.
[0200] The video processing method provided in the embodiments of the present application can be used in electronic devices, which may include handheld devices, vehicle-mounted devices, etc., that have video processing capabilities. For example, some electronic devices are mobile phones, tablet computers, PDAs, laptop computers, mobile internet devices (MIDs), virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control (i.e., industrial control), wireless terminals in self-driving (i.e., self-driving), wireless terminals in remote medical surgery, wireless terminals in smart grids (i.e., smart grids), wireless terminals in transportation safety (i.e., transportation safety), wireless terminals in smart cities (i.e., smart homes), cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to wireless modems, in-vehicle devices, terminal devices in 5G networks or future evolved public land mobile communication networks (Public Land Mobile The terminal equipment in the Network, PLMN, etc. is not limited to this in the embodiments of the present application.
[0201] As an example and not a limitation, in the embodiments of the present application, the electronic device may also be a wearable device. Wearable devices may also be referred to as wearable smart devices, which are a general term for wearable devices that are intelligently designed and developed using wearable technology for daily wear, such as glasses, gloves, watches, clothing, and shoes. A wearable device is a portable device that is worn directly on the body or integrated into the user's clothes or accessories. Wearable devices are not only hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are fully functional, large in size, and can achieve complete or partial functions without relying on smartphones, such as smart watches or smart glasses, as well as those that only focus on a certain type of application function and need to be used in conjunction with other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.
[0202] The electronic device in the embodiments of the present application may also be referred to as terminal equipment, user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, etc.
[0203] It should be noted that the electronic device described in the embodiments of the present application may be configured with a video processing device.
[0204] As an implementation, the video processing device may include a recording module and a processing module. The recording module may be used for video recording. For example, the recording module may be used to capture a subject using multiple cameras to obtain multiple images. The multiple cameras may include cameras that use a shorter exposure time to balance the frame rate and cameras that use a longer exposure time to improve the signal-to-noise ratio.
[0205] The processing module can be configured to reduce noise from a reference image using the auxiliary image to obtain a reduced noise reference image. The reduced noise reference image is a video frame in a video. The reference image can be an image with a shorter exposure time among the multiple images. The auxiliary image can be an image with a longer exposure time among the multiple images.
[0206] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0207] The electronic device 500 may include a processor 510, an external memory interface 520, an internal memory 521, a Universal Serial Bus (USB) interface 530, a charging management module 540, a power management module 541, a battery 542, an antenna 1, an antenna 2, a mobile communication module 550, a wireless communication module 560, an audio module 570, a speaker 570A, a receiver 570B, a microphone 570C, an earphone interface 570D, a sensor module 580, a button 590, a motor 591, an indicator 592, a camera 593, a display screen 594, and a Subscriber Identification Module (SIM) card interface 595, etc. Among them, the sensor module 580 may include a pressure sensor 580A, a gyroscope sensor 580B, an air pressure sensor 580C, a magnetic sensor 580D, an acceleration sensor 580E, a distance sensor 580F, a proximity light sensor 580G, a fingerprint sensor 580H, a temperature sensor 580J, a touch sensor 580K, an ambient light sensor 580L, a bone conduction sensor 580M, etc.
[0208] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 500. In other embodiments of the present application, the electronic device 500 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0209] The processor 510 may include one or more processing units. For example, the processor 510 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor, a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0210] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.
[0211] Processor 510 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 510 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 510. If processor 510 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of processor 510, and thus improves system efficiency.
[0212] In some instances, the memory of the embodiment of the present application may store instructions and data for implementing the video processing method described in the embodiment of the present application, and when the processor is running, the video processing method described in the embodiment of the present application may be implemented.
[0213] The internal memory 521 can be used to store computer executable program code, which includes instructions. The internal memory 521 may include a program storage area and a data storage area. The program storage area may store an operating system, an application required for at least one function (for example, a video recording function, etc.), etc. The data storage area may store data created during the use of the electronic device 500, etc. In addition, the internal memory 521 may include a high-speed random access memory, and may also include a non-volatile memory, for example, at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 510 executes various functional applications and data processing of the electronic device 500 by running instructions stored in the internal memory 521 and / or instructions stored in a memory provided in the processor 510. In an embodiment of the present application, the internal memory 521 may be used to implement the instructions and data of the video processing method described in the embodiment of the present application.
[0214] The electronic device 500 can implement a video recording function through an ISP, a camera 593 , a video codec, a GPU, a display screen 594 , and an AP, etc. In some embodiments, the ISP can be set in the camera 593 .
[0215] The camera 593 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, or other format. In some embodiments, the electronic device 500 may include one or N cameras 593, where N may be an integer greater than 1. For example, in an embodiment of the present application, the camera 593 may be used for shooting in the context of video recording.
[0216] The touch sensor 580K is also called a "touch device". The touch sensor 580K can be set on the display screen 594, and the touch sensor 580K and the display screen 594 form a touch screen, also called a "touch screen". The touch sensor 580K is used to detect touch operations acting on or near it. The touch sensor 580K can pass the detected touch operation to the application processor to determine the type of touch event. For example, in an embodiment of the present application, the touch sensor 580K can be used to pass the detected first operation, second operation, and third operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 594. In other embodiments, the touch sensor 580K can also be set on the surface of the electronic device 500, which is different from the position of the display screen 594.
[0217] The software system of the electronic device 500 may adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. The present application embodiment takes the Android system of the layered architecture as an example to exemplify the software structure of the electronic device 500.
[0218] Figure 6 This is a software structure block diagram of the electronic device provided in the embodiment of the present application.
[0219] A layered architecture divides software into multiple layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: from top to bottom: the application layer, the application framework layer (i.e., the Framework Layer), the hardware abstraction layer (HAL), the driver layer (i.e., the driver layer), and the hardware layer.
[0220] The application layer may include a series of application packages. For example, an application package may include a camera application. When an electronic device runs a camera application, the electronic device may activate a camera and use the camera to capture images. In an embodiment of the present application, the electronic device may include multiple cameras.
[0221] The application framework layer provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer includes some predefined functions. In embodiments of the present application, the application framework layer may include a camera access interface. The camera access interface may include camera management and camera device functions. The camera access interface can be used to provide an API and programming framework for camera applications.
[0222] The hardware abstraction layer (HAL) is an interface layer located between the application framework layer and the driver layer. It can be an encapsulation of the hardware driver, providing a unified interface for upper-layer applications to call. In embodiments of the present application, the HAL can include a camera HAL and a camera algorithm library. The camera HAL can call algorithms in the camera algorithm library.
[0223] The camera hardware abstraction layer can provide virtual hardware for multiple camera devices, for example, camera device 1, camera device 2, or more camera devices. The camera device can correspond to a camera configured on the electronic device 500. The electronic device 500 can call the camera through the camera device. The camera algorithm library can include the video processing method provided in the embodiment of the present application, so that the video processing method in the camera algorithm library can be called by the camera hardware abstraction layer to execute the video processing method described in the embodiment of the present application.
[0224] The driver layer can be a layer between hardware and software that can be used to drive the hardware. The driver layer can include multiple drivers for each hardware component. For example, the driver layer can include a camera device driver, a digital signal processor driver, a graphics processor driver, and a central processing unit driver. The camera device driver can be used to drive the camera to record video and the image signal processor to process images. The digital signal processor driver can be used to drive the digital signal processor to process images. The graphics processor driver can be used to drive the graphics processor to process images. The central processing unit driver can be used to drive the central processing unit to process images.
[0225] The hardware layer may include multiple cameras (eg, camera 1, camera 2, and other cameras), an image signal processor, a digital signal processor, a graphics processor, and a central processing unit, etc. The camera may be used for video recording.
[0226] It should be noted that the embodiments of the present application are only illustrated using the Android system as an example. In other operating systems, if the functions implemented by the various functional modules are similar to those in the embodiments of the present application, the video processing method described in the embodiments of the present application can also be implemented. A camera application is an application installed on an electronic device that can call a camera to provide a video recording function. The application is not limited to camera applications, and other applications that can call a camera to provide a video recording function can be installed on the electronic device.
[0227] For ease of understanding, the following examples of this application will be described with Figure 5 and Figure 6 Taking the electronic device shown as an example, the video processing method provided by the embodiment of the present application is specifically described in combination with the accompanying drawings and application scenarios. It should be noted that the embodiments of the present application can be implemented independently or in combination with each other, and the same or similar concepts or processes will not be repeated in some embodiments.
[0228] It should be noted that the sequence numbers of the steps in the following method are only used to indicate the steps for the purpose of description and should not be considered as indicating the order in which the steps should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.
[0229] Figure 7 This is a flow chart of a video processing method proposed in an embodiment of the present application. This method can be applied to Figure 5 and Figure 6 The electronic device shown in FIG. The electronic device may include a first camera and a second camera. The field of view of the second camera may be greater than that of the first camera. The first camera may be used to balance the frame rate, whereby the first image captured by the first camera may be used as a reference image. The second camera may be used to improve the signal-to-noise ratio, whereby the second image captured by the second camera may be used as an auxiliary image. Based on this, the second image may be used to reduce noise in the first image.
[0230] As an implementation, the first camera may adopt a high-definition, low signal-to-noise ratio image output mode. The second camera may adopt a low-definition, high signal-to-noise ratio image output mode. Optionally, the high-definition, low signal-to-noise ratio image output mode may include a pixel-binning mode. For example, the pixel-binning mode may include a Quadr mode. The low-definition, high signal-to-noise ratio image output mode may include a pixel-binning mode. For example, the pixel-binning mode may include a Binning mode.
[0231] It should be noted that the first shooting parameters of the first camera may include an interval duration of a first predetermined duration and an exposure duration of the first exposure duration, thereby, the exposure duration of the first image captured by the first camera is the first exposure duration. The second shooting parameters of the second camera may include an interval duration of a second predetermined duration and an exposure duration of the second exposure duration, thereby, the exposure duration of the second image captured by the second camera is the second exposure duration. The first exposure duration may be less than the second exposure duration. The first predetermined duration may be greater than or equal to the first exposure duration and less than the second predetermined duration. The second predetermined duration may be greater than or equal to the second exposure duration. The capture time of the first image and the capture time of the second image may overlap, and the lens orientation of the first camera and the lens orientation of the second camera may be the same, so that the content of the first image and the second image are similar.
[0232] As an implementation, the first exposure duration may be less than or equal to the frame interval of the first image to balance the frame rate. The second exposure duration may be greater than the frame interval of the first image to improve the signal-to-noise ratio. For example, the frame interval of the first image may be 30 ms. The first exposure duration may be less than or equal to 30 ms. Optionally, the first exposure duration may be 20 ms. The first predetermined duration may be 30 ms. The second exposure duration may be greater than or equal to 100 ms. Optionally, the second exposure duration may be 150 ms. The second predetermined duration may be 150 ms.
[0233] As an implementation manner, the first image may be a short-frame image, and the second image may be a long-frame image.
[0234] As an implementation, the first image and the second image may have separate timestamps. The timestamps can be used as one basis for determining whether the first image and the second image have similar content, and thus as a basis for determining whether the second image can be used to reduce noise in the first image. As an implementation, if the timestamps of the first image and the second image are determined to be within a predetermined time range, and if the lens orientation of the first camera is the same as that of the second camera, it can be determined that the content of the first image and the second image is similar, and therefore, the second image can be used to reduce noise in the first image.
[0235] It should also be noted that, since the video processing method described in the embodiment of the present application can be applied to scenes where the video quality is greatly affected by noise, for example, dark light scenes, the video recording mode provided in the embodiment of the present application can be characterized by a dark light recording mode, that is, the dark light recording mode can refer to video recording for scenes where the video quality is greatly affected by noise. The shooting modes currently provided by the camera application may include at least one of the following: recording mode, photo mode, night scene mode, portrait mode, movie mode, short video mode or time-lapse photography mode, etc. The camera application interface may include controls corresponding to the shooting mode, thereby, the camera application interface may include recording controls, photo controls, night scene controls, portrait controls, movie controls, short video controls or time-lapse photography controls, etc.
[0236] Regarding how to combine the dark light recording mode provided by the embodiment of the present application with the camera application, the following is combined with Figure 8A and Figure 8B To explain. It should be noted that, Figure 8A and Figure 8B The possible interface styles of the dark light recording mode are only schematically shown and should not constitute a limitation on the embodiments of the present application.
[0237] Figure 8A A schematic diagram of the interface of a dark light recording mode provided in an embodiment of the present application.
[0238] As an implementation method, low-light recording mode can be used as a shooting mode in parallel with other shooting modes in the camera application. For example, low-light recording mode can be placed alongside other shooting modes such as night scene mode, portrait mode, photo mode, recording mode, short video mode, etc. Figure 8A , that is, in the camera application interface, the low-light recording control is placed side by side with the controls of other shooting modes such as night scene controls, portrait controls, photo controls, recording controls, and short video controls.
[0239] Figure 8B A schematic diagram of the interface of another dark light recording mode provided in an embodiment of the present application.
[0240] As another implementation method, the low-light recording mode can be used as an optional sub-mode in the recording mode of the camera application. Figure 8B , that is, in the camera application interface, the low-light recording control may be a control displayed when the recording control is triggered.
[0241] like Figure 7 As shown, the method includes S701-S710. S701-S702 are the video recording initiation steps. S703-S705 are the video recording steps. S706-S710 are the video noise reduction steps during the video recording process. S706 is the field of view angle alignment step. S707 is the brightness alignment step. S708 is the image registration step. S709 is the image fusion step. This method will be described below with reference to the accompanying drawings.
[0242] At S701 , the camera application generates a recording instruction in response to a first operation.
[0243] According to an embodiment of the present application, the first operation may refer to an operation for initiating recording. For example, the first operation may include at least one of the following: a touch operation, a key operation, a gesture operation, or a voice operation. The touch operation may include at least one of the following: a click operation or a slide operation. The recording instruction may be used to instruct the first camera and the second camera to record a video.
[0244] Regarding how to start video recording, the following Figure 9A 、 Figure 9B 、 Figure 9C and Figure 9D To explain. It should be noted that, Figure 9A 、 Figure 9B and Figure 9C The possible startup methods are only schematically shown and should not constitute a limitation on the embodiments of the present application.
[0245] As an implementation method, the active activation method can refer to the user actively triggering the low-light recording mode. Depending on how the low-light recording mode is integrated with the camera application, the active activation method can include the following implementation methods.
[0246] Regarding the low-light recording mode, which is a shooting mode listed alongside other shooting modes in the camera app, as one implementation, in response to a second operation, a first camera app interface is displayed. The first camera app interface may include low-light recording controls. In response to a third operation on the low-light recording controls, the low-light recording interface is displayed.
[0247] Specifically, the camera application can, in response to the second operation, call the camera access interface of the application framework layer to start the camera application. When the low-light recording control included in the first camera application interface is displayed, the camera application can, in response to the third operation on the low-light recording control, call the camera access interface of the application framework layer to start low-light recording mode.
[0248] As another implementation, in response to the second operation, a dark light recording interface is displayed. In this case, when the camera application is started, the dark light recording mode is started by default. Thus, the dark light recording mode can be the default startup mode of the camera application or the last time the dark light recording mode was started. In some embodiments, if the user used the dark light recording mode of the camera application last time, the camera application can start the dark light recording mode by default when the camera application is started again. In other embodiments, the camera application can set the dark light recording mode as the default startup mode. Specifically, the camera application can call the camera access interface of the application framework layer in response to the second operation to start the camera application. When the camera application is started, the dark light recording interface is displayed.
[0249] It should be noted that the second operation may refer to an operation for starting a camera application. For other descriptions of the second and third operations, please refer to the description of the first operation above, which will not be repeated here.
[0250] The following combination Figure 9A To explain. Figure 9A A schematic diagram of starting a dark light recording mode provided in an embodiment of the present application.
[0251] like Figure 9A As shown, in response to a second operation on first camera application control 901, first camera application interface 902 is displayed. First camera application interface 902 may include low-light recording control 903. In response to a third operation on low-light recording control 903, low-light recording interface 904 is displayed, thereby initiating low-light recording mode. The second and third operations may be touch operations.
[0252] Regarding low-light recording mode, which is a selectable sub-mode within the recording mode in the camera application, as one implementation, in response to a fourth operation, a second camera application interface is displayed. The second camera application interface may include recording controls. In response to a fifth operation on the recording controls, a first recording interface is displayed. The first recording interface may include low-light recording controls. In response to a third operation on the low-light recording controls, the low-light recording interface is displayed.
[0253] Specifically, the camera application may, in response to the fourth operation, call the camera access interface of the application framework layer to start the camera application. When displaying the recording control included in the second camera application interface, the camera application may, in response to the fifth operation on the recording control, call the camera access interface of the application framework layer to display the first recording interface. When displaying the low-light recording control included in the first recording interface, the camera application may, in response to the third operation on the low-light recording control, call the camera access interface of the application framework layer to start low-light recording mode.
[0254] As another implementation, in response to the fourth operation, the first recording interface is displayed. In response to the third operation for the dark-light recording control, the dark-light recording interface is displayed. In this case, when the camera application is started, the recording mode is started by default. Thus, the recording mode can be the default startup mode of the camera application or the recording mode started last time. In some embodiments, if the user used the recording mode of the camera application last time, the camera application can start the recording mode by default when the camera application is started again. In other embodiments, the camera application can set the recording mode as the default startup mode. Specifically, the camera application can call the camera access interface of the application framework layer in response to the fourth operation to start the camera application. When the camera application is started, the first recording interface is displayed. When the dark-light recording control included in the first recording interface is displayed, the camera application can call the camera access interface of the application framework layer in response to the third operation for the dark-light recording control to start the dark-light recording mode.
[0255] It should be noted that the fourth operation may refer to an operation for starting a camera application. For other descriptions of the fourth and fifth operations, please refer to the description of the first operation above, which will not be repeated here. Figure 9B To explain. Figure 9B A schematic diagram of another method of starting a dark light recording mode provided in an embodiment of the present application.
[0256] like Figure 9BAs shown, in response to the fourth operation on the second camera application control 905, the second camera application interface 906 is displayed. The second camera application interface 906 may include a recording control 907. In response to the fifth operation on the recording control 907, the first recording interface 908 is displayed. The first recording interface 908 may include a dark light recording control 909. In response to the third operation on the dark light recording control 909, the dark light recording interface is displayed. The dark light recording interface may be Figure 9A The dark light recording interface 904 in the image processing unit is displayed, thereby starting the dark light recording mode. The fourth operation, the fifth operation and the third operation can be touch operations.
[0257] As another implementation, namely, an automatic start mode, the automatic start mode may refer to automatically starting the low-light recording mode in response to detecting a scene where the video quality is significantly affected by noise (eg, a low-light scene) when the recording mode is started.
[0258] In response to detecting that the scene is low-light, a low-light recording interface is displayed, thereby starting the low-light recording mode.
[0259] How to detect a dark scene can be achieved in the following ways.
[0260] As an implementation, a low-light scene is determined when the ambient light intensity is less than or equal to a predetermined ambient light intensity. Ambient light intensity refers to the luminous flux received per unit area, describing the distribution of light in space. The predetermined ambient light intensity can be configured based on actual business needs and is not limited here. For example, the predetermined ambient light intensity can be less than or equal to 5 lux.
[0261] As another implementation method, when the ambient light brightness is less than or equal to the predetermined ambient light brightness, it is determined to be in a dark light scene. Ambient light brightness may refer to the intensity of light emitted by a light source or the surface of a subject. The predetermined ambient light intensity may be configured according to actual business needs and is not limited here. The ambient light brightness may be perceived by an ambient light sensor of an electronic device. Optionally, the ambient light brightness may be determined based on the aperture value, the image brightness of a predetermined image, the exposure time of a predetermined image, and the sensitivity of a predetermined image. Optionally, the image brightness of a predetermined image may be used to indicate the ambient light brightness. The image brightness may be determined according to an average brightness method or a weighted mean method, etc. The average brightness method may refer to determining the average brightness of the entire image. The predetermined image may be an image obtained in a recording mode.
[0262] In response to detecting a low-light scene, displaying the low-light recording interface can be achieved in the following manner.
[0263] As one implementation, in response to the fourth operation, a second camera application interface is displayed. The second camera application interface may include a recording control. In response to a fifth operation on the recording control, a second recording interface is displayed. Alternatively, in response to the fourth operation, the second recording interface is displayed. While the second recording interface is displayed, in response to detecting a low-light scene, a low-light recording interface is displayed. The fourth operation may be an operation for launching the camera application.
[0264] As another implementation, in response to the fourth operation, a second recording interface is displayed. While the second recording interface is displayed, in response to detecting a low-light scene, the low-light recording interface is displayed. In this case, when the camera application is launched, the recording mode is enabled by default. Thus, the recording mode can be the default launch mode of the camera application or the last launch mode of the recording mode.
[0265] As another implementation, in response to detecting a low-light scene, a prompt message is generated. The prompt message may be used to prompt whether to activate low-light recording mode. In response to a sixth operation in response to the prompt message, a low-light recording interface is displayed. The sixth operation may be an operation for activating low-light recording mode. For further details on the sixth operation, please refer to the description of the first operation above and will not be repeated here.
[0266] As another implementation, in response to the fourth operation, a second camera application interface is displayed. The second camera application interface may include recording controls. In response to a fifth operation on the recording controls, a second recording interface is displayed. Alternatively, in response to the fourth operation, the second recording interface is displayed. While the second recording interface is displayed, a prompt message is generated in response to detecting a low-light scene. The prompt message may indicate whether to enable low-light recording mode. In response to a sixth operation on the prompt message, the low-light recording interface is displayed.
[0267] As another implementation, in response to the fourth operation, a second recording interface is displayed. While the second recording interface is displayed, a prompt message is generated in response to detecting a low-light scene. In response to a sixth operation in response to the prompt message, a low-light recording interface is displayed. In this case, when the camera application is launched, recording mode is enabled by default. Thus, the recording mode can be the default launch mode of the camera application or the last launch mode of recording mode.
[0268] The following combination Figure 9C To explain. Figure 9C A schematic diagram of another method of starting a dark light recording mode provided in an embodiment of the present application.
[0269] like Figure 9CAs shown, in response to the fourth operation on the second camera application control 905, the second camera application interface 906 is displayed. The second camera application interface 906 may include a recording control 907. In response to the fifth operation on the recording control 907, the second recording interface 910 is displayed. When the second recording interface 910 is displayed, in response to detecting that it is in a dark light scene, prompt information 911 is generated. Prompt information 911 can be used to prompt whether to start the dark light recording mode. In response to the sixth operation on the prompt information 911, the dark light recording interface is displayed. The dark light recording interface can be Figure 8A The dark light recording interface 904 in the image processing unit is displayed, thereby starting the dark light recording mode. The fourth operation, the fifth operation and the sixth operation can be touch operations.
[0270] The following describes how to record a video with low-light recording mode enabled.
[0271] The low-light recording interface may include a start control. In response to a first operation on the start control, video recording is performed. Specifically, when the start control included in the low-light recording interface is displayed, the camera application may, in response to the first operation on the start control, send a recording instruction to the first camera and the second camera via the camera access interface, the camera hardware abstraction layer, and the camera device driver. The recording instruction may be used to instruct the first camera and the second camera to perform video recording.
[0272] The following combination Figure 9D To explain. Figure 9D A schematic diagram of the dark light recording mode provided in an embodiment of the present application.
[0273] like Figure 9D As shown, the low-light recording interface 904 may include a start control 912. In response to a first operation on the start control 912, video recording is performed. Figure 9D The right side of the dotted line schematically shows the recording screen 913 of the first camera. Figure 9D The left side of the middle dotted line schematically shows the recording screen 914 of the second camera. Figure 9D It can be seen that since the field of view of the second camera is greater than the field of view of the first camera, the recording screen 914 shows more content than the recording screen 913.
[0274] In response to a second operation on the camera application, the camera application can call a camera access interface of the application framework layer to launch the camera application. If a low-light recording control included in the camera application interface is displayed, the camera application can call a camera access interface of the application framework layer to launch a low-light scene recording mode in response to a first operation on the low-light recording control. If a start control included in the low-light recording interface is displayed, the camera application can send a recording request to the application framework layer in response to a third operation on the start control. The recording request can be used to request the camera to record video.
[0275] In response to receiving a recording request, the application framework layer sends a recording request to multiple camera device drivers in the camera hardware abstraction layer. In response to receiving the recording request, the camera device driver drives the camera corresponding to the camera device to capture the subject and obtain an image. The multiple cameras may include a first camera and a second camera. The first camera may be a camera that uses a shorter exposure time to balance the frame rate. The second camera may be a camera that uses a longer exposure time to improve the signal-to-noise ratio. The multiple images may include a reference image captured by the first camera and an auxiliary image captured by the second camera. The exposure time of the reference image is shorter than the exposure time of the auxiliary image. The camera may send the image to the camera hardware abstraction layer via the camera device driver layer. The camera hardware abstraction layer may store the image in a memory buffer.
[0276] At S702 , the camera application transmits a recording instruction to the first camera and the second camera through the camera access interface, the camera hardware abstraction layer, and the camera device driver.
[0277] According to an embodiment of the present application, in response to receiving a recording instruction, the camera access interface transmits the recording instruction to the camera hardware abstraction layer and the camera device driver. In response to receiving the recording instruction, the camera device driver transmits the recording instruction to the first camera and the second camera to drive the first camera and the second camera to capture the subject to obtain the first image and the second image.
[0278] In S703 , in response to receiving the recording instruction, the first camera shoots according to the first shooting parameter to obtain a first image, and in response to receiving the recording instruction, the second camera shoots according to the second shooting parameter to obtain a second image.
[0279] According to an embodiment of the present application, the first image may include at least one .The second image may include at least one .
[0280] In S704 , the first camera transmits a first image to the camera hardware abstraction layer through the camera device driver, and the second camera transmits a second image to the camera hardware abstraction layer through the camera device driver.
[0281] At S705 , the camera hardware abstraction layer transmits the first image and the second image to the camera algorithm library.
[0282] According to embodiments of the present application, the camera hardware abstraction layer (HAL) can invoke video processing methods in a camera algorithm library to implement video noise reduction using these methods. For example, the HAL can utilize support from an image signal processor, a digital signal processor, and a graphics processor to reduce noise from a first image using a second image, thereby generating a noise-reduced first image. The noise-reduced first image is a video frame in a video.
[0283] At S706 , the camera algorithm library crops the second image to obtain a cropped second image.
[0284] According to an embodiment of the present application, the camera hardware abstraction layer can call a video processing method in the camera algorithm library to use the video processing method to achieve alignment of the field of view angles of the first image and the second image. For example, because the field of view angle of the second camera is larger than the field of view angle of the first camera, the second image can be cropped so that the cropped second image can be aligned with the field of view angle of the first image. Image cropping can include at least one of the following: proportional cropping, center cropping, or edge cropping. The embodiments of the present application do not limit the image cropping method.
[0285] In S707, the camera algorithm library adjusts the original image brightness of the cropped second image according to the first exposure time and sensitivity of the first image, and the second exposure time, sensitivity and original image brightness of the cropped second image to obtain a second image with adjusted brightness.
[0286] According to an embodiment of the present application, the camera hardware abstraction layer can call a video processing method in the camera algorithm library to use the video processing method to achieve brightness alignment between the first image and the cropped second image. For example, a brightness-adjusted second image is obtained based on a ratio and the original image brightness of the cropped second image. The ratio can be determined based on a second amount of light entering the second image and a first amount of light entering the first image. The first amount of light entering can be determined based on a first exposure time and sensitivity of the first image. The second amount of light entering can be determined based on a second exposure time and sensitivity of the cropped second image.
[0287] As an implementation manner, the original image brightness of the cropped second image may be adjusted according to the following formulas (1)-(2).
[0288] (1)
[0289] (2)
[0290] in, The image brightness of the brightness adjusted second image may be characterized. The original image brightness of the cropped second image may be characterized. Can represent ratios. The first exposure duration can be represented. The sensitivity of the first image can be represented. The second exposure time can be represented. The sensitivity of the second image can be represented.
[0291] In S708 , the camera algorithm library uses a cross-attention mechanism to register the second feature data of the second image whose brightness has been adjusted with the first feature data of the first image to obtain registered second feature data.
[0292] According to the embodiments of the present application, it should be noted that the brightness-adjusted second image will be referred to as the second image below. The camera hardware abstraction layer can call video processing methods in the camera algorithm library to implement image registration using these video processing methods. For example, a cross-attention mechanism can be used to align the second feature data of the second image with the first feature data of the first image to obtain registered second feature data. The first feature data can be used for the query (i.e., Q). The second feature data can be used for the key (i.e., K) and value (i.e., V).
[0293] According to an embodiment of the present application, the first feature data may include multiple first features. The second feature data may include multiple second features. The cross-attention mechanism may include at least one of the following: a single-head cross-attention mechanism or a multi-head cross-attention mechanism. The similarity determination method involved in the cross-attention mechanism may include at least one of the following: dot product similarity, scaled dot product similarity, additive similarity, or cosine similarity, etc.
[0294] Regarding how to use the cross-attention mechanism to achieve registration, it can be achieved in the following ways.
[0295] As an implementation, the registered second feature data may be obtained based on at least one fourth intermediate matrix. The fourth intermediate matrix may be obtained based on the third attention weight matrix and the fifth value matrix. The third attention weight matrix may be obtained based on the fourth query matrix and the fourth key matrix. The third attention weight matrix may be used to indicate the similarity between the second feature and the first feature data.
[0296] The fourth query matrix may be obtained by performing the seventh linear transformation on the first feature data. For example, the fourth query matrix may be obtained based on the first feature data and the seventh transformation matrix. The fourth key matrix may be obtained by performing the eighth linear transformation on the second feature data. For example, the fourth key matrix may be obtained based on the second feature data and the eighth transformation matrix. The fifth value matrix may be obtained by performing the ninth linear transformation on the second feature data. For example, the fifth value matrix may be obtained based on the second feature data and the ninth transformation matrix.
[0297] It should be noted that when there is only one fourth intermediate matrix, the above method can be understood as a single-headed cross-attention mechanism. When there are multiple fourth intermediate matrices, the above method can be understood as a multi-headed cross-attention mechanism. The transformation matrix groups corresponding to at least one fourth intermediate matrix can be different. The transformation matrix group can include a seventh transformation matrix, an eighth transformation matrix, and a ninth transformation matrix. The element values in the transformation matrix group can be model parameters of the deep learning model.
[0298] Regarding how to obtain the third attention weight matrix, it can be achieved in the following way.
[0299] As an implementation method, based on dot product similarity or scaled dot product similarity. For example, the third attention weight matrix can be obtained by normalizing the fifth intermediate matrix. For example, the third attention weight matrix can be obtained by processing the fifth intermediate matrix using a Softmax function, a Sigmoid function, or a Tanh function. The fifth intermediate matrix can be obtained based on the dot product between the fourth query matrix and the fourth key matrix.
[0300] As another implementation, based on additive similarity, for example, the third attention mechanism can be obtained by normalizing the sixth intermediate matrix. The sixth intermediate matrix can be obtained by performing the tenth linear transformation on the seventh intermediate matrix. For example, the sixth intermediate matrix can be obtained based on the seventh intermediate matrix and the tenth transformation matrix. The element values in the tenth transformation matrix can be model parameters of the deep learning model. The seventh intermediate matrix can be obtained by performing a nonlinear activation on the eighth intermediate matrix. The eighth intermediate matrix can be obtained based on the fourth query matrix and the fourth key matrix.
[0301] Regarding how to obtain the eighth intermediate matrix, it can be achieved in the following manner.
[0302] As an implementation, the eighth intermediate matrix may be obtained based on the fifth query matrix and the fifth key matrix. The fifth query matrix may be obtained by performing the eleventh linear transformation on the fourth query matrix. For example, the fifth query matrix may be obtained based on the fourth query matrix and the eleventh transformation matrix. The fifth key matrix may be obtained by performing the twelfth linear transformation on the fourth key matrix. For example, the fifth key matrix may be obtained based on the fourth key matrix and the twelfth transformation matrix.
[0303] As another implementation, the eighth intermediate matrix may be obtained by performing a thirteenth linear transformation on the ninth intermediate matrix. For example, the eighth intermediate matrix may be obtained based on the ninth intermediate matrix and the thirteenth transformation matrix. The ninth intermediate matrix may be obtained by concatenating the fourth query matrix and the fourth key matrix.
[0304] It should be noted that the element values in the eleventh transformation matrix, the twelfth transformation matrix, and the thirteenth transformation matrix can be model parameters of the deep learning model.
[0305] In order to reduce computational complexity while maintaining or improving the accuracy of feature registration, so as to better apply the video processing method described in the embodiments of the present application to resource-constrained electronic devices, such as terminal devices, an optimization strategy based on the cross-attention mechanism can be adopted to optimize this.
[0306] As an implementation manner, the optimization strategy may include at least one of the following: a linear strategy, a sparse strategy, a quantization strategy, a pruning strategy, or other strategies.
[0307] Regarding how to optimize the cross-attention mechanism based on the sparse strategy, it can be achieved in the following ways.
[0308] As an implementation method, the fourth key matrix can be obtained by performing an eighth linear transformation and sparse selection on the second feature data of the second image. For example, the fourth key matrix can be obtained by performing a sparse selection on the sixth key matrix. The sixth key matrix can be obtained by performing an eighth linear transformation on the second feature data of the second image. Sparse selection can refer to selecting element values that meet predetermined conditions from the sixth key matrix so that the fourth key matrix is composed of these element values, that is, the element values in the fourth key matrix meet the predetermined conditions. The element values that meet the predetermined conditions can refer to element values in the sixth key matrix whose similarity with the elements in the fourth query matrix is greater than or equal to a first predetermined threshold.
[0309] Since the fourth key matrix is obtained by sparse selection of the sixth key matrix, the data volume of the fourth key matrix is reduced, thereby reducing the consumption of computing resources, so that the video processing method described in the embodiment of the present application can be better applied to resource-constrained electronic devices.
[0310] Regarding how to optimize the cross-attention mechanism based on quantization strategy and pruning strategy, it can be achieved in the following ways.
[0311] As an implementation method, the third attention weight matrix can be obtained based on the fourth attention weight matrix. For example, the third attention weight matrix can be obtained by quantizing the twelfth intermediate matrix. Optionally, the third attention weight matrix can be obtained by quantizing the twelfth intermediate matrix to a 4-bit or 8-bit fixed-point integer. The twelfth intermediate matrix can be obtained by pruning the fourth attention weight matrix. Optionally, the twelfth intermediate matrix can be obtained by setting the element values in the fourth attention weight matrix that are less than or equal to the second predetermined threshold to zero. The second predetermined threshold can be configured according to actual business needs and is not limited here.
[0312] The fourth attention weight matrix may be obtained by normalizing the tenth intermediate matrix. The tenth intermediate matrix may be obtained by dequantizing the eleventh intermediate matrix. The eleventh intermediate matrix may be obtained based on the sixth query matrix and the sixth key matrix. The sixth query matrix and the sixth key matrix may be obtained by quantizing the fourth query matrix and the fourth key matrix.
[0313] Regarding how to optimize the cross attention mechanism based on other strategies, it can be achieved in the following ways. For ease of understanding, the following is combined with Figure 10 and Figure 11 To explain.
[0314] As a way to implement it, Figure 10 The figure is a schematic diagram of a principle for obtaining registered second feature data according to an embodiment of the present application.
[0315] like Figure 10 As shown, the registered second feature data can be obtained based on the first query matrix, the first intermediate matrix, and the first value matrix. The first intermediate matrix can be used for K (i.e., key). The first query matrix can be used for Q (i.e., query). The first value matrix can be used for V (i.e., value).
[0316] The first value matrix can be obtained based on the first intermediate matrix, the first key matrix, and the second value matrix. The first intermediate matrix can be used for Q. The first key matrix can be used for K. The second value matrix can be used for V. For example, the first value matrix can be obtained by processing the first intermediate matrix, the first key matrix, and the second value matrix using a cross-attention mechanism. Optionally, the first value matrix can be obtained based on the fifth attention weight matrix and the second value matrix. The fifth attention weight matrix can be obtained based on the first intermediate matrix and the first key matrix.
[0317] Since the first value matrix is obtained based on the first intermediate matrix, the first key matrix, and the second value matrix, and the first intermediate matrix serves as the query, the first intermediate matrix can be used to aggregate information from the first key matrix and the second value matrix. Since the registered second feature data is obtained based on the first query matrix, the first intermediate matrix, and the first value matrix, and the first query matrix serves as the query and the first intermediate matrix serves as the key, it can be shown that the first intermediate matrix can transmit aggregated information to the first query matrix.
[0318] The first query matrix may be obtained by performing a first linear transformation on the first feature data of the first image. For example, the first query matrix may be obtained based on the first feature data of the first image and the first transformation matrix. The element values in the first transformation matrix may be model parameters of the deep learning model.
[0319] The first key matrix may be obtained by performing a second linear transformation on the second feature data of the second image. For example, the first key matrix may be obtained based on the second feature data of the second image and the second transformation matrix. The element values in the second transformation matrix may be model parameters of the deep learning model.
[0320] The second value matrix may be obtained by performing a third linear transformation on the second feature data. For example, the second value matrix may be obtained based on the second feature data and the third transformation matrix.
[0321] The first intermediate matrix can be obtained based on the first query matrix. For example, the first intermediate matrix can be obtained by pooling the first query matrix. Pooling can include maximum pooling, average pooling, or L2 norm pooling. Optionally, the first intermediate matrix can be obtained based on the first query matrix and a fourteenth transformation matrix. The element values in the fourteenth transformation matrix can be model parameters of the deep learning model. Optionally, the first intermediate matrix can be obtained by fusing similar elements in the first query matrix.
[0322] Regarding how to fuse similar elements in the first query matrix, this can be achieved in the following manner.
[0323] As an implementation, the first intermediate matrix may be obtained based on fused elements and non-fused elements. A fused element may be obtained by fusing a target element and other elements that have an association relationship. The other elements that have an association relationship with the target element may be elements that are most similar to the target element, as determined from other sets. The target element may be any one in the target set. When the target set is the first set, the other set is the second set. When the target set is the second set, the other set is the first set. The first set and the second set may be obtained by partitioning the elements included in the first query matrix. Non-fused elements may refer to elements that do not have an association relationship with other elements.
[0324] Regarding how to obtain the registered second feature data according to the first query matrix, the first intermediate matrix and the first value matrix, it can be achieved in the following manner.
[0325] As an implementation manner, the registered second feature data may be obtained based on the second intermediate matrix and the third value matrix. For example, the registered second feature data may be obtained by adding the second intermediate matrix and the third value matrix.
[0326] The second intermediate matrix can be obtained based on the first query matrix, the first intermediate matrix and the first value matrix. For example, the third intermediate matrix can be obtained based on the sixth attention weight matrix and the first value matrix. The sixth attention weight matrix can be obtained based on the first query matrix and the first intermediate matrix.
[0327] The third value matrix may be obtained by performing a first depthwise separable convolution on the first value matrix.
[0328] Since the third value matrix is obtained by performing a first depth-wise separable convolution on the first value matrix, feature diversity is improved.
[0329] As another implementation, the registered second feature data may be obtained according to the sixth attention weight matrix and the first value matrix. The sixth attention weight matrix may be obtained according to the first query matrix and the first intermediate matrix.
[0330] Since the first intermediate matrix is obtained based on the first query matrix, the data volume of the first intermediate matrix can be smaller than that of the first query matrix. Therefore, in the process of determining the first value matrix, the first value matrix is obtained based on the first intermediate matrix, the first key matrix, and the second value matrix. The first intermediate matrix serves as the query. Therefore, compared to using the first query matrix as the query, the data processing volume and computing resource consumption are reduced. In addition, the first value matrix aggregates the information of the first key matrix and the second value matrix. In the process of determining the aligned second feature data, the aligned second feature data is obtained based on the first value matrix obtained from the first query matrix, the first intermediate matrix, and the first value matrix. The first query matrix serves as the query, the first intermediate matrix serves as the key, and the first value matrix serves as the value. Therefore, the first value matrix can transmit the aggregated information to the first query matrix, thereby achieving global modeling. The first intermediate matrix can be used to aggregate the information of the first key matrix and the second value matrix and transmit the aggregated information to the first query matrix.
[0331] As another implementation, Figure 11 A schematic diagram of another principle for obtaining registered second feature data provided in an embodiment of the present application.
[0332] like Figure 11As shown, the registered second feature data may be obtained according to the third intermediate matrix and the second query matrix. For example, the registered second feature data may be obtained by multiplying the third intermediate matrix and the second query matrix.
[0333] The third intermediate matrix can be obtained according to the second key matrix and the fourth value matrix.
[0334] The second query matrix and the second key matrix can be obtained by processing the third query matrix and the third key matrix respectively using a focusing function. The focusing function can be used to adjust the proximity between the third query matrix and the third key matrix. That is, the aggregation function can be used to bring similar elements in the third query matrix and the third key matrix closer together and dissimilar elements further apart. For example, the third query matrix can be input into the aggregation function to obtain the second query matrix. The third key matrix can be input into the aggregation function to obtain the second key matrix.
[0335] As an implementation manner, the second query matrix and the second key matrix can be determined according to the following formulas (3)-(4).
[0336] (3)
[0337] (4)
[0338] in, Aggregation functions can be characterized. It can be the third query matrix or the fourth query matrix. It can be the second query matrix or the second key matrix. Can be characterized Model. Can be characterized of Power. Can be greater than 1.
[0339] The third query matrix may be obtained by performing a fourth linear transformation on the first feature data of the first image. For example, the third query matrix may be obtained based on the first feature data of the first image and the fourth transformation matrix.
[0340] The third key matrix may be obtained by performing the fifth linear transformation on the second feature data of the second image. For example, the third key matrix may be obtained based on the second feature data of the second image and the fifth transformation matrix.
[0341] The third intermediate matrix may be obtained according to the second key matrix and the fourth value matrix. For example, the third intermediate matrix may be obtained by multiplying the second key matrix and the fourth value matrix.
[0342] The fourth value matrix may be obtained by performing a sixth linear transformation on the second feature data of the second image. For example, the fourth matrix may be obtained based on the second feature data of the second image and the sixth transformation matrix.
[0343] It should be noted that the element values in the fourth transformation matrix, the fifth transformation matrix and the sixth transformation matrix can be model parameters of the deep learning model.
[0344] Since the aggregation function can be used to adjust the proximity between the third query matrix and the third key matrix, that is, the aggregation function can be used to make similar elements in the third query matrix and the third key matrix closer and dissimilar elements farther apart, thereby achieving the goal of assigning a larger weight to the second feature with a large similarity between the aligned second feature data and the first feature data, and assigning a smaller weight to the second feature in the second feature data with a small similarity to the first feature data, the accuracy of feature alignment is improved, thereby improving the video noise reduction effect. In addition, since the third intermediate matrix is obtained based on the second key matrix and the fourth value matrix, the aligned second feature data is obtained based on the second query matrix and the third intermediate matrix, that is, the key and value are calculated first, and the result is then calculated with the query, thereby reducing the amount of data processing and reducing the consumption of computing resources.
[0345] As an implementation, the first feature data may be determined based on the first visual feature data and the first semantic feature data. Alternatively, the first feature data may be determined based on the first visual feature data, the first semantic feature data, and a first position code. The first position code may be used to indicate a pixel position of a pixel in the first image.
[0346] The second feature data may be determined based on the second visual feature data and the second semantic feature data. Alternatively, the second feature data may be determined based on the second visual feature data, the second semantic feature data, and a second position code. The second position code may be used to indicate a pixel position of a pixel in the second image.
[0347] Since position coding can better capture the dependencies and patterns between features in feature data and can focus on the relative positions of features, the reliability of the first feature data and the second feature data is improved, thereby improving the video noise reduction effect.
[0348] As an implementation, the first query matrix may be obtained by performing a second depthwise separable convolution on the first feature data of the first image. And / or, the first key matrix may be obtained by performing a third depthwise separable convolution on the second feature data of the second image. And / or, the first value matrix may be obtained by performing a fourth depthwise separable convolution on the second feature data.
[0349] Since the first query matrix, the first key matrix, and the first value matrix can be obtained by using a depthwise separable convolution method, the depthwise separable convolution can reduce the amount of data processing, thereby reducing the amount of data processing and reducing computing resource consumption.
[0350] In S709 , the camera algorithm library fuses the first feature data and the registered second feature data to obtain fused feature data.
[0351] According to the embodiment of the present application, the camera hardware abstraction layer can call the video processing method in the camera algorithm library to realize image fusion by using the video processing method. How to realize image fusion can be realized by the following methods. Figure 12 To explain.
[0352] As an implementation method, the first feature data and the registered second feature data are concatenated to obtain fused feature data. Alternatively, the first feature data and the registered second feature data are added to obtain fused feature data. Alternatively, the first feature data and the registered second feature data are multiplied to obtain fused feature data.
[0353] As another implementation, Figure 12 A schematic diagram of the principle of obtaining fused feature data provided in an embodiment of the present application.
[0354] like Figure 12 As shown, the fused feature data may be obtained based on the first feature data and the third feature data. For example, the fused feature data may be obtained by fusing the first feature data and the third feature data. The third feature data may be obtained based on the fourth feature data and the registered second feature data. For example, the third feature data may be obtained by multiplying the fourth feature data and the registered second feature data. The fourth feature data may be obtained based on the first feature data and the registered second feature data. For example, the fourth feature data may be obtained by convolving the first feature data with the registered second feature data.
[0355] Since the fused feature data is obtained based on the first feature data and the third feature data, the third feature data is obtained based on the fourth feature data and the aligned second feature data, and the fourth feature data is obtained based on the first feature data and the aligned second feature data, a more complete fusion of the first feature data and the second feature data is achieved, thereby making the information carried by the fused feature data more comprehensive and accurate.
[0356] At S710 , the camera algorithm library obtains a video frame based on the fused feature data.
[0357] According to an embodiment of the present application, the camera hardware abstraction layer can call the video processing method in the camera algorithm library to obtain a video frame using the video processing method. For example, the fused feature data can be decoded to obtain a video frame.
[0358] Because the fused feature data is obtained by fusing the first feature data with the registered second feature data, it can reflect the fusion characteristics of the first and registered second feature data, thereby making the information carried by the fused feature more comprehensive and accurate. On this basis, decoding the fused feature data to obtain video frames improves the video frame quality and thus enhances the video noise reduction effect.
[0359] Regarding how to decode the fused feature data to obtain the video frame, it can be achieved in the following way.
[0360] As an implementation, a video frame may be obtained based on the first feature data and the fifth feature data. The fifth feature data may be obtained based on the sixth feature data. The sixth feature data may be obtained based on the seventh feature data and the fused feature data. The seventh feature data may be obtained based on the eighth feature data. The eighth feature data may be obtained based on the fused feature data and the ninth feature data. The ninth feature data may be obtained based on the fused feature data.
[0361] Since the video frame is obtained based on the first feature data and the fifth feature data, the fifth feature data is obtained based on the sixth feature data, the sixth feature data is obtained based on the seventh feature data and the fused feature data, the seventh feature data is obtained based on the eighth feature data, the eighth feature data is obtained based on the fused feature data and the ninth feature data, and the ninth feature data can be obtained based on the fused feature data, therefore, more full utilization of the feature data is achieved, thereby improving the quality of the video frame.
[0362] It should be noted that the second image may be used to reduce the noise of the first image during the video recording process, and the second image may also be used to reduce the noise of the first image after the video recording is finished.
[0363] It should also be noted that Figure 7 It is only schematically shown that the field angle alignment step S706 can be performed first, and then the brightness alignment step S707 can be performed. In addition, the brightness alignment step S707 can also be performed first, and then the field angle alignment step S706 can be performed. The embodiment of the present application does not limit the execution order of the field angle alignment step and the brightness alignment step. In addition, Figure 7 S706 and S707 may be optional steps.
[0364] Figure 13A flowchart of another video processing method provided in an embodiment of the present application.
[0365] like Figure 13 As shown, the method includes S1310-S1320.
[0366] At S1310 , in response to a first operation, video recording is performed.
[0367] As one implementation, in response to the second operation, a camera application interface is displayed. The camera application interface may include a low-light recording control. In response to a third operation on the low-light recording control, a low-light recording interface is displayed. The low-light recording interface may include a start control. In response to the first operation on the start control, video recording is performed.
[0368] Since the camera application can provide a low-light recording mode, low-light video recording can be performed using the camera application.
[0369] As another implementation, in response to detecting that the scene is dark, a dark recording interface is displayed. In response to a first operation on a start control, video recording is performed.
[0370] When a dark-light scene is detected, the dark-light recording interface is displayed, and the dark-light recording mode is automatically started. This simplifies the steps for starting dark-light recording and improves the convenience of dark-light recording.
[0371] As another implementation, in response to the second operation, the low-light recording interface is displayed. In response to the first operation on the start control, video recording is performed.
[0372] When the second operation is detected, the low-light recording interface is displayed, enabling low-light recording mode to be enabled by default when the camera app is launched. Low-light recording mode can be the default launch mode of the camera app or the last launch mode of low-light recording mode, thereby simplifying the process of enabling low-light recording and improving the convenience of low-light recording.
[0373] It should be noted that for the description of the above implementation method, please refer to the above Figure 7 、 Figure 8A 、 Figure 8B 、 Figure 9A 、 Figure 9B 、 Figure 9C and Figure 9D The corresponding parts will not be repeated here.
[0374] At S1320 , during video recording, the first image is denoised according to the second image to obtain a video frame.
[0375] According to an embodiment of the present application, a video frame may be a frame in a recorded video. The first image and the second image may be captured using multiple cameras. The capture time of the first image and the capture time of the second image may overlap. The lens orientation of the first camera may be the same as the lens orientation of the second camera. The first camera may be the camera that captured the first image. The second camera may be the camera that captured the second image. The exposure time of the first image may be shorter than the exposure time of the second image.
[0376] As an implementation manner, the definition of the first image may be higher than the definition of the second image.
[0377] By making the clarity of the first image higher than that of the second image, the video noise reduction effect is further improved.
[0378] As another implementation, the exposure time of the second image may be greater than the frame interval of the first image, and the exposure time of the first image may be less than or equal to the frame interval of the first image.
[0379] By setting the exposure time of the reference image to be less than or equal to the frame interval of the reference image, a balanced frame rate is achieved, and by setting the exposure time of the auxiliary image to be greater than the frame interval of the reference image, an improved signal-to-noise ratio is achieved. Thus, the video noise reduction effect is improved while balancing the frame rate.
[0380] As another implementation, the clarity of the first image may be higher than that of the second image. Furthermore, the exposure time of the second image may be greater than the frame interval of the first image. The exposure time of the first image may be less than or equal to the frame interval of the first image.
[0381] As an implementation manner, the first image may be a short-frame image, and the second image may be a long-frame image.
[0382] Since the first image is a short-frame image, it can balance the frame rate, preserve the scene's texture edges and highlight information, and have high definition and a low signal-to-noise ratio. The second image is a long-frame image, which has a high signal-to-noise ratio, low definition, richer dark-light information, and richer color information. This demonstrates that long and short-frame images can complement each other to a certain extent. For example, the high-noise sharp texture of the short-frame image and the low-noise blurred texture of the long-frame image can constrain each other, facilitating detail recovery. The highlight information of the short-frame image can repair the overexposed cutoff areas of the long-frame image. Furthermore, the color information of the long-frame image can accurately restore the scene's colors. Therefore, short-frame images offer advantages in balancing frame rate and / or high definition, while long-frame images offer advantages in improving signal-to-noise ratio. By using long-frame images to complement short-frame images, the quality of the short-frame images is improved, thereby achieving enhanced video noise reduction while balancing frame rate.
[0383] As an implementation manner, the exposure time of the second image may be greater than or equal to 100 ms.
[0384] Since the exposure time of the second image is greater than or equal to 100 ms, the exposure time of the second image can improve the signal-to-noise ratio, thereby improving the video noise reduction effect.
[0385] It should be noted that for the description of the exposure time of the first image, the exposure time of the second image, the shooting time of the first image, the shooting time of the second image, the frame interval of the first image, the clarity of the first image and the clarity of the second image, please refer to the corresponding parts above and will not be repeated here.
[0386] As an implementation manner, the image brightness of the first image and the second image are aligned, and / or the image brightness of the first image and the second image are aligned and the field of view angles are aligned.
[0387] By aligning the brightness of the first and second images, the brightness ranges of the first and second images are made the same, thereby improving the accuracy of image registration and image fusion, and thus enhancing the video noise reduction effect. By aligning the field of view of the first and second images, the accuracy of image registration and image fusion is improved, and thus enhancing the video noise reduction effect.
[0388] As an implementation, the image brightness of the second image may be determined based on the original image brightness of the second image, the amount of light entering the second image, and the amount of light entering the first image. The amount of light entering the first image may be determined based on the exposure time and sensitivity of the first image. The amount of light entering the second image may be determined based on the exposure time and sensitivity of the second image.
[0389] For example, the image brightness of the second image may be determined based on a ratio and the original image brightness of the second image. The ratio may be a ratio between the amount of light entering the second image and the amount of light entering the first image.
[0390] By adjusting the image brightness, the brightness range of multiple images is made the same, that is, the brightness of multiple images is aligned. This improves the accuracy of image registration and image fusion, and further enhances the video noise reduction effect.
[0391] It should be noted that, for the description of how to determine the image brightness of the second image, please refer to the corresponding part above. For example, the description of how to determine the image brightness of the cropped second image is not repeated here.
[0392] As an implementation, the target image may be obtained by image cropping. The target image may be an image captured by a target camera among the first and second images. The target camera may be a camera with a larger field of view among the multiple cameras. Alternatively, the target camera may be a camera with a longer exposure time among the multiple cameras. Alternatively, the target camera may be a camera with a larger field of view and a longer exposure time among the multiple cameras. For example, the target camera may be the second camera described above.
[0393] By using image cropping to align the field of view of multiple images, the accuracy of image registration and image fusion is improved, thereby enhancing the video noise reduction effect. In addition, it also reduces the amount of data processing and the consumption of computing resources.
[0394] It should be noted that for the description of image cropping, please refer to the corresponding part above and will not be repeated here.
[0395] As an implementation manner, the first image and the second image may be RAW images.
[0396] Since the RAW image has not been processed by the image signal processor, ISP processing will make the noise more complicated. Therefore, the first image and the second image are RAW images, which reduces the difficulty of video noise reduction and improves the video noise reduction effect.
[0397] It should be noted that for the description of RAW images, please refer to the corresponding part above and will not be repeated here.
[0398] Regarding how to reduce the noise of the first image based on the second image to obtain a video frame, it can be achieved in the following manner.
[0399] As an implementation manner, the first image and the second image are fused to obtain a video frame.
[0400] Since the video frame is obtained by fusing the first image and the second image, the first image is taken by a camera that uses a shorter exposure time to balance the frame rate, and the second image is taken by a camera that uses a longer exposure time to improve the signal-to-noise ratio. Therefore, the video noise reduction effect is improved while balancing the frame rate.
[0401] Regarding how to fuse the first image and the second image to obtain a video frame, this can be achieved in the following manner.
[0402] As an implementation manner, the first image and the registered second image are fused to obtain a video frame. The registered second image is obtained by registering the second image with the first image.
[0403] Because the video frame is created by fusing the first image with the registered second image, and the registered second image is created by registering the second image with the first image, image-level noise reduction is achieved. Image-level noise reduction fully utilizes pixel information, preserving image details as much as possible, and offers good interpretability, thereby improving video noise reduction effectiveness.
[0404] Regarding how to register the first image and the second image, it can be achieved in the following manner.
[0405] As an implementation, the registered second image may be obtained by registering the first image and the second image using an optical flow method. For example, the registered second image may be obtained by registering the second image and the first image using optical flow information. The optical flow information may be pixel-by-pixel relative displacement information between the first image and the second image.
[0406] By utilizing the optical flow method to realize image registration between the second image and the first image, the efficiency of image registration is improved.
[0407] As another implementation, the first image and the histogram-converted image are fused to obtain a video frame. The histogram-converted image can be obtained by transferring the color distribution of the second image to the first image based on a histogram conversion function.
[0408] Since there is no spatial misalignment between the image obtained by histogram conversion and the first image, the probability of ghosting caused by camera or object motion is reduced. In addition, since the image registration step is not performed, the video noise reduction process is simplified.
[0409] As another implementation, the tenth feature data and the eleventh feature data are fused to obtain a video frame. The tenth feature data and the eleventh feature data can be obtained based on the first feature data of the first image and the second feature data of the second image.
[0410] Regarding how to obtain the tenth characteristic data and the eleventh characteristic data, it can be achieved in the following manner.
[0411] As an implementation, the tenth feature data may be obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism. The first feature data may be used for querying. The second feature data may be used for keys and values. The eleventh feature data may be obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism. The second feature data may be used for querying. The first feature data may be used for keys and values.
[0412] Regarding how to use the cross attention mechanism to obtain the tenth feature data and the eleventh feature data, it can be achieved in the following way.
[0413] As an implementation method, the tenth feature data can be obtained based on the seventh attention weight matrix and the sixth value matrix. The seventh attention weight matrix can be obtained based on the seventh query matrix and the seventh key matrix. The seventh query matrix can be obtained by performing a fourteenth linear transformation on the first feature data of the first image. The seventh key matrix can be obtained by performing a fifteenth linear transformation on the second feature data of the second image. The seventh attention weight matrix can be used to indicate the similarity between the second feature and the second feature data.
[0414] The eleventh feature data may be obtained based on the eighth attention weight matrix and the seventh value matrix. The eighth attention weight matrix may be obtained based on the eighth query matrix and the eighth key matrix. The eighth query matrix may be obtained by performing a sixteenth linear transformation on the second feature data of the second image. The eighth key matrix may be obtained by performing a seventeenth linear transformation on the first feature data of the first image. The eighth attention weight matrix may be used to indicate the similarity between the first feature and the second feature data.
[0415] Since the tenth feature data is obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism, the first feature data is used for query and the second feature data is used for key and value. Therefore, the tenth feature in the tenth feature data is obtained by fusing the second feature data with the first feature data. The eleventh feature data is obtained by processing the first feature data of the first image and the second feature data of the second image using a cross-attention mechanism. The second feature data is used for query and the first feature data can be used for key and value. Therefore, the eleventh feature in the eleventh feature data is obtained by fusing the first feature data with the second feature data. This achieves more effective utilization of the first and second feature data. On this basis, the video frame is obtained by fusing the tenth and eleventh feature data, thereby improving the video noise reduction effect.
[0416] As another implementation, the first feature data of the first image and the registered second feature data are fused to obtain a video frame. The registered second feature data is obtained by registering the second feature data of the second image with the first feature data of the first image.
[0417] Since the first feature data abstracts the first image while retaining the detail information of the first image, and the second feature data abstracts the second image while retaining the detail information of the second image, the first feature data and the second feature data have higher robustness and discriminability. On this basis, the video frame is obtained by fusing the first feature data of the first image and the aligned second feature data, and the aligned second feature data is obtained by aligning the second feature data of the second image and the first feature data of the first image. Therefore, the reliability and robustness of video denoising are improved, and the video denoising effect is improved.
[0418] Regarding how to register the second feature data of the second image and the first feature data of the first image, it can be achieved in the following manner.
[0419] As an implementation, the registered second feature data is obtained by registering the second feature data of the second image with the first feature data of the first image using a cross-attention mechanism. The first feature data can be used for querying, and the second feature data can be used for keys and values.
[0420] Since the cross-attention mechanism can mine the association between the second feature data and the first feature data, it can achieve interaction and matching between the first image and the second image at the feature level, thereby capturing the similarities and differences between the features, focusing on similarities with high weights and suppressing differences with low weights, and can adaptively adjust the weights. Therefore, it has higher robustness and scalability. Therefore, using the cross-attention mechanism to align the second feature data of the second image with the first feature data of the first image can adjust the feature weight of the second feature in the second feature data. For example, a larger weight is assigned to the second feature in the second feature data that has a high similarity with the first feature data, and a smaller weight is assigned to the second feature in the second feature data that has a low similarity with the first feature data. This improves the accuracy of feature alignment and thus improves the video denoising effect.
[0421] As an implementation, the cross-attention mechanism may include at least one of the following: a dot product attention mechanism, a scaled dot product attention mechanism, an additive attention mechanism, a linear attention mechanism, or a sparse attention mechanism.
[0422] Since the video processing method described in the embodiments of the present application can be applied to resource-constrained electronic devices, such as terminal devices, the computational complexity of the cross-attention mechanism is proportional to the square of the number of feature data included in the feature data. Therefore, by adopting an optimization strategy based on optimizing the cross-attention mechanism, such as a linear attention mechanism and a sparse attention mechanism, the expectation is achieved that the computational complexity can be reduced while maintaining or improving the accuracy of feature registration.
[0423] Regarding how to use the cross-attention mechanism to obtain the aligned second feature data, it can be achieved in the following way.
[0424] As an implementation manner, the registered second feature data may be obtained according to the first query matrix, the first intermediate matrix and the first value matrix. The first value matrix may be obtained according to the first intermediate matrix, the first key matrix and the second value matrix.
[0425] The first query matrix may be obtained by performing a first linear transformation on the first feature data of the first image. The first intermediate matrix may be obtained based on the first query matrix. The first key matrix may be obtained by performing a second linear transformation on the second feature data of the second image. The second value matrix may be obtained by performing a third linear transformation on the second feature data.
[0426] Since the first value matrix is obtained based on the first intermediate matrix, the first key matrix, and the second value matrix, and the first intermediate matrix serves as the query, the first intermediate matrix can be used to aggregate information from the first key matrix and the second value matrix. Since the registered second feature data is obtained based on the first query matrix, the first intermediate matrix, and the first value matrix, and the first query matrix serves as the query and the first intermediate matrix serves as the key, it can be shown that the first intermediate matrix can transmit aggregated information to the first query matrix.
[0427] Because the first intermediate matrix is obtained based on the first query matrix, the data volume of the first intermediate matrix can be smaller than that of the first query matrix. Therefore, in the process of determining the first value matrix, the first value matrix is obtained based on the first intermediate matrix, the first key matrix, and the second value matrix. The first intermediate matrix serves as the query. Therefore, compared to using the first query matrix as the query, the data processing amount and computing resource efficiency are reduced. In addition, the first value matrix aggregates the information of the first key matrix and the second value matrix. In the process of determining the aligned second feature data, the aligned second feature data is obtained based on the first value matrix obtained from the first query matrix, the first intermediate matrix, and the first value matrix. The first query matrix serves as the query, the first intermediate matrix serves as the key, and the first value matrix serves as the value. Therefore, the first value matrix can transmit the aggregated information to the first query matrix, thereby achieving global modeling. The first intermediate matrix can be used to aggregate the information of the first key matrix and the second value matrix and transmit the aggregated information to the first query matrix.
[0428] Regarding how to obtain the registered second feature data according to the first query matrix, the first intermediate matrix and the first value matrix, it can be achieved in the following manner.
[0429] As an implementation, the registered second feature data may be obtained based on the second intermediate matrix and the third value matrix. The second intermediate matrix may be obtained based on the first query matrix, the first intermediate matrix, and the first value matrix. The third value matrix may be obtained by performing a depthwise separable convolution on the first value matrix.
[0430] Since the third value matrix is obtained by performing a first depth-wise separable convolution on the first value matrix, feature diversity is improved.
[0431] As another implementation, the registered second feature data may be obtained based on the third intermediate matrix and the second query matrix. The third intermediate matrix may be obtained based on the second key matrix and the fourth value matrix. The second query matrix and the second key matrix may be obtained by processing the third query matrix and the third key matrix respectively using a focusing function. The focusing function may be used to adjust the proximity between the third query matrix and the third key matrix. The third query matrix may be obtained by performing a fourth linear transformation on the first feature data of the first image. The third key matrix may be obtained by performing a fifth linear transformation on the second feature data of the second image. The third intermediate matrix may be obtained based on the second key matrix and the fourth value matrix. The fourth value matrix may be obtained by performing a sixth linear transformation on the second feature data of the second image.
[0432] Since the aggregation function can be used to adjust the proximity between the third query matrix and the third key matrix, that is, the aggregation function can be used to make similar elements in the third query matrix and the third key matrix closer and dissimilar elements farther apart, thereby achieving the goal of assigning a larger weight to the second feature with a large similarity between the aligned second feature data and the first feature data, and assigning a smaller weight to the second feature in the second feature data with a small similarity to the first feature data, the accuracy of feature alignment is improved, thereby improving the video noise reduction effect. In addition, since the third intermediate matrix is obtained based on the second key matrix and the fourth value matrix, the aligned second feature data is obtained based on the second query matrix and the third intermediate matrix, that is, the key and value are calculated first, and the result is then calculated with the query, thereby reducing the amount of data processing and reducing the consumption of computing resources.
[0433] It should be noted that for the description of this part, please refer to the corresponding part above and will not be repeated here.
[0434] Regarding how to fuse the first feature data of the first image and the registered second feature data to obtain a video frame, this can be achieved in the following manner.
[0435] As an implementation, the fused feature data is decoded to obtain a video frame. The fused feature data may be obtained by fusing the first feature of the first image and the registered second feature data.
[0436] Because the fused feature data is obtained by fusing the first feature data with the registered second feature data, it can reflect the fusion characteristics of the first and registered second feature data, thereby making the information carried by the fused feature more comprehensive and accurate. On this basis, decoding the fused feature data to obtain video frames improves the video frame quality and thus enhances the video noise reduction effect.
[0437] Regarding how to obtain fused feature data, it can be achieved in the following ways.
[0438] As an implementation, the fused feature data may be obtained based on the first feature data and the third feature data. The third feature data may be obtained based on the fourth feature data and the registered second feature data. The fourth feature data may be obtained based on the first feature data and the registered second feature data.
[0439] Since the fused feature data is obtained based on the first feature data and the third feature data, the third feature data is obtained based on the fourth feature data and the aligned second feature data, and the fourth feature data is obtained based on the first feature data and the aligned second feature data, a more complete fusion of the first feature data and the second feature data is achieved, thereby making the information carried by the fused feature data more comprehensive and accurate.
[0440] It should be noted that for the description of this part, please refer to the corresponding part above and will not be repeated here.
[0441] According to an embodiment of the present application, since in the case of video recording, the first image is taken by a camera that uses a shorter exposure time to balance the frame rate, and the second image is taken by a camera that uses a longer exposure time to improve the signal-to-noise ratio, and the shooting time of the first image and the shooting time of the second image overlap, it means that the contents of the first image and the second image are similar. Therefore, the second image is used to reduce the noise of the first image to obtain the video frame in the video, thereby achieving the improvement of the video noise reduction effect while balancing the frame rate.
[0442] It should be noted that the module names involved in the embodiments of the present application can be defined as other names as long as the functions of each module can be achieved, and there is no specific restriction on the names of the modules.
[0443] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0444] The video processing method of the embodiment of the present application has been described above. The device for performing the above method provided by the embodiment of the present application is described below. Those skilled in the art will understand that the method and device can be combined and referenced with each other, and the relevant device provided by the embodiment of the present application can perform the steps in the above video processing method.
[0445] Figure 14 This is a structural block diagram of the video processing device provided in an embodiment of the present application.
[0446] like Figure 14 As shown, the video processing device 1400 may include a recording module 1410 and a processing module 1420 .
[0447] The recording module 1410 is configured to perform video recording in response to a first operation.
[0448] Processing module 1420 is configured to, during video recording, reduce noise from the first image based on the second image to obtain a video frame. A video frame is a frame in the recorded video. The first image and the second image are captured using multiple cameras. The capture time of the first image and the capture time of the second image overlap. The lens orientation of the first camera is the same as the lens orientation of the second camera. The first camera is the camera that captured the first image. The second camera is the camera that captured the second image. The exposure time of the first image is less than the exposure time of the second image.
[0449] The video processing method provided in the embodiment of the present application can be applied to electronic devices with communication functions. The electronic devices include terminal devices. The specific device form of the terminal device can refer to the above related descriptions and will not be repeated here.
[0450] An embodiment of the present application provides an electronic device comprising a processor and a memory. The memory stores computer-executable instructions. The processor executes the computer-executable instructions stored in the memory, causing the electronic device to perform the video processing method described in the above embodiment. The implementation principles and technical effects are similar to those of the above-described related embodiments and will not be further elaborated here.
[0451] The present application provides a chip or chip system. The chip or chip system includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a line. The at least one processor is configured to execute a computer program or instruction to perform the video processing method described in the above embodiment. The implementation principles and technical effects are similar to those of the above-described related embodiments and will not be further elaborated here.
[0452] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is run on an electronic device, the electronic device executes the video processing method in the above embodiment. The method described in the above embodiment can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. If implemented in software, the function can be stored as one or more instructions or codes on a computer-readable medium or transmitted on a computer-readable medium. Computer-readable media can include computer storage media and communication media, and can also include any medium that can transfer a computer program from one place to another. The storage medium can be any target medium that can be accessed by a computer.
[0453] In one possible implementation, computer-readable media may include RAM, ROM, Compact Disc Read-Only Memory (CD-ROM) or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium designed to carry or store the desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately termed a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. Disk and disc, as used herein, include optical disc, laser disc, magnetic disk, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of the above are also intended to be included within the scope of computer-readable media.
[0454] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed, the computer executes the above method.
[0455] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to produce a machine, so that the instructions executed by the processing unit of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0456] The above specific implementation methods further explain in detail the purpose, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above are only specific implementation methods of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application should be included in the scope of protection of the embodiments of the present application.
Claims
1. A video processing method, characterized in that: include: In response to the first operation, performing video recording in a low-light scene; During the video recording process, the fused feature data is decoded to obtain a video frame, wherein the fused feature data is obtained by fusing the first feature data of the first image and the registered second feature data, wherein the registered second feature data is obtained by registering the second feature data of the second image and the first feature data of the first image using a cross-attention mechanism, wherein the first feature data is used for query, and the second feature data is used for key and value; the video frame is a frame in the recorded video, the first image and the second image are taken by using multiple cameras, the first image and the second image are RAW images, the shooting time of the first image and the shooting time of the second image overlap, the lens direction of the first camera is the same as the lens direction of the second camera, the first camera is the camera that shoots the first image, and the second camera is the camera that shoots the second image, the exposure time of the first image is less than the exposure time of the second image; the exposure time of the second image is greater than the frame interval of the first image, and the exposure time of the first image is less than or equal to the frame interval; The decoding of the fused feature data to obtain a video frame includes: Obtaining a video frame based on the first feature data and the fifth feature data, wherein the fifth feature data is obtained based on the sixth feature data, the sixth feature data is obtained based on the seventh feature data and the fused feature data, the seventh feature data is obtained based on the eighth feature data, the eighth feature data is obtained based on the fused feature data and the ninth feature data, and the ninth feature data is obtained based on the fused feature data; The registered second feature data is obtained by registering the second feature data of the second image and the first feature data of the first image using a cross-attention mechanism, including: The registered second feature data is obtained based on the second intermediate matrix and the third value matrix, the second intermediate matrix is obtained based on the first query matrix, the first intermediate matrix and the first value matrix, the third value matrix is obtained by performing depthwise separable convolution on the second value matrix, and the first value matrix is obtained based on the first intermediate matrix, the first key matrix and the second value matrix; wherein, the first query matrix is obtained by performing a first linear transformation on the first feature data of the first image, the first intermediate matrix is obtained based on the first query matrix, the first key matrix is obtained by performing a second linear transformation on the second feature data of the second image, and the second value matrix is obtained by performing a third linear transformation on the second feature data; or, The registered second feature data is obtained based on the third intermediate matrix and the second query matrix, the third intermediate matrix is obtained based on the second key matrix and the fourth value matrix, the second query matrix and the second key matrix are obtained by respectively processing the third query matrix and the third key matrix using a focusing function, the focusing function is used to adjust the proximity between the third query matrix and the third key matrix, the third query matrix is obtained by performing a fourth linear transformation on the first feature data of the first image, the third key matrix is obtained by performing a fifth linear transformation on the second feature data of the second image, the third intermediate matrix is obtained based on the second key matrix and the fourth value matrix, and the fourth value matrix is obtained by performing a sixth linear transformation on the second feature data of the second image.
2. The method according to claim 1, characterized in that The step of recording the video in response to the first operation includes: In response to the second operation, displaying a camera application interface, wherein the camera application interface includes a low-light recording control; in response to a third operation on the low-light recording control, displaying a low-light recording interface; or in response to detecting a low-light scene, displaying the low-light recording interface; or in response to the second operation, displaying the low-light recording interface, wherein the low-light recording interface includes a start control; In response to a first operation on the start control, video recording is performed.
3. The method according to claim 1 or 2, characterized in that The definition of the first image is higher than that of the second image.
4. The method according to claim 1, wherein The exposure time of the second image is greater than or equal to 100 ms.
5. The method according to claim 1, wherein The fused feature data is obtained by fusing the first feature of the first image and the registered second feature data, and includes: The fused feature data is obtained based on the first feature data and the third feature data, the third feature data is obtained based on the fourth feature data and the aligned second feature data, and the fourth feature data is obtained based on the first feature data and the aligned second feature data.
6. The method according to claim 1, characterized in that The image brightness of the first image and the second image are aligned.
7. The method according to claim 1, characterized in that The image brightness of the first image and the second image are aligned and the field angles are aligned.
8. The method according to claim 6 or 7, characterized in that The image brightness of the second image is determined based on the original image brightness of the second image, the amount of light entering the second image and the amount of light entering the first image. The amount of light entering the first image is determined based on the exposure time and sensitivity of the first image, and the amount of light entering the second image is determined based on the exposure time and sensitivity of the second image.
9. An electronic device, characterized in that: include: processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the electronic device performs the method according to any one of claims 1 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 8.
11. A chip system, characterized in that: The system comprises at least one processor and a communication interface, wherein the communication interface and the at least one processor are interconnected via a line, and the at least one processor is configured to run a computer program or instruction to execute the method according to any one of claims 1 to 8.
12. A computer program product, characterized in that The method comprises a computer program, which, when running on an electronic device, enables the electronic device to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Image acquiring device, method and terminal and video acquiring method
CN103986875A
Method for improving dynamic range of image and camera
CN111726543A
Multi-source remote sensing optical image registration and fusion method based on space deformation field
CN118941600A