Video noise reduction method and device, equipment, storage medium and computer program product

Through collaborative reinforcement processing and spatial mask-based content distinction processing, the problem of being unable to intelligently distinguish characters and backgrounds in video noise reduction is solved, and the noise reduction effect of retaining important details is achieved, which significantly improves the visual sense of the video.

CN120219224AInactive Publication Date: 2025-06-27CHINA MOBILE M2M +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332515.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot intelligently distinguish the target character from the background during video noise reduction, resulting in character features being mishandled and the loss of important details and texture information.

Method used

By acquiring the frame feature map of the target image to be processed and the reference face feature map is used, collaborative reinforcement processing is performed, and combined with content distinction processing based on spatial masks, important regional details in the image are identified and retained.

Benefits of technology

It realizes intelligently distinguishing the target characters and backgrounds during the noise reduction process, retaining important details and texture information, avoiding excessive smoothing, and significantly improving the overall view of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219224A_ABST
    Figure CN120219224A_ABST
Patent Text Reader

Abstract

The invention provides a video denoising method, device and equipment, a storage medium and a computer program product, and relates to the technical field of image denoising, and the method comprises the steps: obtaining a to-be-processed target image frame feature map based on a to-be-denoised video of a target user; performing collaborative enhancement processing on the to-be-processed target image frame feature map and the reference face feature map to obtain a collaborative enhancement to-be-processed target image frame feature map; performing space mask-based content distinguishing processing on the collaborative enhancement to-be-processed target image frame feature map to obtain collaborative enhancement to-be-processed target image frame significant features; and obtaining a noise-reduced to-be-processed target image frame based on the collaborative enhancement of the significant features of the to-be-processed target image frame. According to the method, the important area in the image can be identified, it is ensured that details of the important area are reserved in the noise reduction process, the target person and the background are intelligently distinguished in the noise reduction process, the important details and texture information in the image are reserved, excessive smoothness is avoided, and therefore the overall impression of the video is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image noise reduction, and particularly to a video noise reduction method, device, equipment, storage medium, and computer program product. Background Art

[0002] In video calls or conferences, in order to improve the quality of experience, some technical means are usually adopted to reduce the interference of non-participants. However, traditionally, the entire video frame is usually uniformly denoised without distinguishing the target person and the background, resulting in the possibility that the person's features may be wrongly treated as noise. In addition, while removing noise, traditional video noise reduction methods will smooth out important details in the image, such as the facial features of a person, thus losing important detail and texture information.

[0003] Therefore, how to intelligently distinguish the target person and the background during the noise reduction process and retain the important details and texture information in the image has become an urgent problem to be solved. Summary of the Invention

[0004] The present invention provides a video noise reduction method, device, equipment, storage medium, and computer program product, which are used to solve the defect that the existing noise reduction methods cannot intelligently distinguish the target person and the background while removing noise and retain the important details and texture information in the image, and realize intelligently distinguishing the target person and the background during the noise reduction process and retaining the important details and texture information in the image, avoiding over-smoothing, thereby significantly improving the overall visual perception of the video.

[0005] The present invention provides a video noise reduction method, including the following steps: Based on the video to be denoised of the target user, obtain the feature map of the target image frame to be processed; the target user is the user who needs to be focused on optimizing and protecting during the video noise reduction process; Perform collaborative enhancement processing on the feature map of the target image frame to be processed and the reference face feature map to obtain a collaboratively enhanced feature map of the target image frame to be processed; Perform content discrimination processing based on a spatial mask on the collaboratively enhanced feature map of the target image frame to be processed to obtain the significant features of the collaboratively enhanced target image frame to be processed; Based on the significant features of the collaboratively enhanced target image frame to be processed, obtain the target image frame to be denoised.

[0006] According to the video noise reduction method provided by the present invention, the step of performing collaborative enhancement processing on the feature map of the target image frame to be processed and the reference face feature map to obtain a collaboratively enhanced feature map of the target image frame to be processed includes: Determine the full-channel target reference feature similarity matrix between the feature map of the target image frame to be processed and the reference face feature map; Normalize the full-channel target reference feature similarity matrix to obtain a normalized full-channel target reference feature similarity matrix; Based on the normalized full-channel target reference feature similarity matrix, determine the collaborative attention weight vector of the target image frame to be processed for the feature map of the target image frame to be processed; Based on the collaborative attention weight vector of the target image frame to be processed, perform a per-channel multiplication on the feature map of the target image frame to be processed to obtain the collaboratively enhanced feature map of the target image frame to be processed.

[0007] According to a video denoising method provided by the present invention, the determination of the full-channel target reference feature similarity matrix between the feature map of the target image frame to be processed and the reference face feature map includes: Reshape the feature map of the target image frame to be processed and the reference face feature map to obtain a feature matrix of the target image frame to be processed and a reference face feature matrix; Calculate the product between the feature matrix of the target image frame to be processed and the transposed matrix of the reference face feature matrix to obtain the full-channel target reference feature similarity matrix.

[0008] According to a video denoising method provided by the present invention, the normalization processing of the full-channel target reference feature similarity matrix to obtain a normalized full-channel target reference feature similarity matrix includes: Taking each eigenvalue in the full-channel target reference feature similarity matrix as the exponent of the natural constant, calculate the exponential function value with the natural constant as the base according to the position to obtain a full-channel target reference feature similarity class support matrix; Calculate the sum of each column vector in the full-channel target reference feature similarity class support matrix to obtain a full-channel target reference feature similarity class support column global vector composed of multiple column vector global sum values; Perform a division operation on each eigenvalue in the full-channel target reference feature similarity class support matrix and the corresponding column global vector sum value in the full-channel target reference feature similarity class support column global vector according to the position to obtain the normalized full-channel target reference feature similarity matrix.

[0009] According to a video denoising method provided by the present invention, the determination of the collaborative attention weight vector of the target image frame to be processed for the feature map of the target image frame to be processed based on the normalized full-channel target reference feature similarity matrix includes: Perform global average pooling on the feature map of the target image frame to be processed along the channel dimension to obtain a full-channel feature vector of the target image frame to be processed; Calculate the matrix product between the normalized full-channel target reference feature similarity matrix and the full-channel feature vector of the target image frame to be processed, to obtain the collaborative attention weight vector of the target image frame to be processed.

[0010] According to a video noise reduction method provided by the present invention, the content discrimination processing based on a spatial mask is performed on the collaboratively enhanced feature map of the target image frame to be processed to obtain the significant feature of the collaboratively enhanced target image frame to be processed, including: Perform feature representation on the collaboratively enhanced feature map of the target image frame to be processed to obtain the represented feature map of the collaboratively enhanced target image frame to be processed; Set the feature values greater than or equal to a predetermined threshold at each position of the represented feature map of the collaboratively enhanced target image frame to 1, and the rest to 0, to obtain the masked feature map of the collaboratively enhanced target image frame to be processed; Perform point-by-point multiplication of the masked feature map of the collaboratively enhanced target image frame and the feature map of the collaboratively enhanced target image frame at each position to obtain the significant feature of the collaboratively enhanced target image frame to be processed.

[0011] According to a video noise reduction method provided by the present invention, the performing feature representation on the collaboratively enhanced feature map of the target image frame to be processed to obtain the represented feature map of the collaboratively enhanced target image frame to be processed includes: Taking the negative of the feature value at each position of the collaboratively enhanced feature map of the target image frame to be processed as the exponent of the natural constant, and calculating the exponential function value with the natural constant as the base at each position to obtain the class support feature map of the collaboratively enhanced target image frame to be processed; Calculate the reciprocal of the sum of the feature value at each position in the class support feature map of the collaboratively enhanced target image frame to be processed and the constant 1 to obtain the represented feature map of the collaboratively enhanced target image frame to be processed.

[0012] The present invention also provides a video noise reduction device, including the following modules: A feature map acquisition module, configured to acquire a feature map of a target image frame to be processed based on a video to be noise-reduced of a target user; the target user is a user who needs to be focused on optimizing and protecting in video noise reduction processing; A collaborative enhancement processing module, configured to perform collaborative enhancement processing on the feature map of the target image frame to be processed and a reference face feature map to obtain a collaboratively enhanced feature map of the target image frame to be processed; A content discrimination module, configured to perform content discrimination processing based on a spatial mask on the collaboratively enhanced feature map of the target image frame to be processed to obtain a significant feature of the collaboratively enhanced target image frame to be processed; A video noise reduction module, configured to obtain a noise-reduced target image frame to be processed based on the significant feature of the collaboratively enhanced target image frame to be processed.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the video noise reduction method described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the video noise reduction method described in any one of the above is implemented.

[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the video noise reduction method described in any one of the above is implemented.

[0016] The video noise reduction method, device, equipment, storage medium, and computer program product provided by the present invention obtain a target image frame feature map to be processed based on a video to be noise-reduced of a target user; perform collaborative enhancement processing on the target image frame feature map to be processed and a reference face feature map to obtain a collaboratively enhanced target image frame feature map to be processed; perform content discrimination processing based on a spatial mask on the collaboratively enhanced target image frame feature map to be processed to obtain a collaboratively enhanced target image frame significant feature; and obtain a noise-reduced target image frame to be processed based on the collaboratively enhanced target image frame significant feature. The present invention helps to identify important regions in an image, ensures that details are retained in important regions during the noise reduction process, intelligently distinguishes a target person from a background during the noise reduction process, and retains important details and texture information in the image, avoiding over-smoothing, thereby significantly improving the overall visual perception of the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic flowchart of the video noise reduction method provided by the present invention.

[0019] Figure 2 is a schematic flowchart of the video directional noise reduction method provided by the present invention.

[0020] Figure 3 is a schematic architecture diagram of the video directional noise reduction method provided by the present invention.

[0021] Figure 4 is a schematic flowchart of the collaborative enhancement processing provided by the present invention.

[0022] Figure 5 It is a schematic flowchart of content discrimination processing based on a spatial mask provided by the present invention.

[0023] Figure 6 It is a schematic structural diagram of a video noise reduction device provided by the present invention.

[0024] Figure 7 It is a schematic structural diagram of an electronic device provided by the present invention. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention fall within the protection scope of the present invention.

[0026] The following will be combined with Figures 1-7 Describe the video noise reduction method, device, equipment, storage medium, and computer program product of the present invention.

[0027] The video noise reduction method provided by the present invention is specifically a video directional noise reduction method. It can be understood that directional noise reduction is an image processing technology that analyzes the importance of different regions in an image and targets the image for noise reduction. For a video, directional noise reduction can effectively remove unimportant background noise in the video while retaining the details and textures of important target regions (such as faces).

[0028] Figure 1 It is a schematic flowchart of the video noise reduction method provided by the present invention. As Figure 1 shown, the method includes the following: Step 101, based on the video to be noise-reduced of the target user, obtain the feature map of the target image frame to be processed.

[0029] Among them, the target user is the user who needs to be optimized and protected with emphasis in video noise reduction processing. Specifically, the target user is usually the main participant or core figure in the video frame. The noise reduction method needs to optimize the processing for this user to retain the facial features and details while reducing the noise interference in the background or other non-target regions. It should be understood that the video to be noise-reduced contains the target user; the feature map of the target image frame to be processed can be understood as the result obtained after feature extraction of the image frame to be processed extracted from the target video (i.e., the video to be noise-reduced).

[0030] First, obtain the reference face image of the target user, and obtain the target video of the target user from the media stream (such as the video captured in real time by the camera). Then, considering that the reference face image reflects the implicit feature information about the face, such as facial features, the shape and position of eyes and mouth, etc.; texture features, such as skin texture and wrinkles, etc. And the dilated convolutional neural network model is a special convolutional neural network, which expands the receptive field by introducing holes (that is, skipping some pixels) in the convolutional kernel, which enables the dilated convolutional neural network model to capture a larger range of context information in the image, so as to extract more representative features. Therefore, in the embodiment of the present invention, the reference face image is input into the face feature extractor based on the dilated convolutional neural network model to capture and extract the representative feature information implicit in the reference face image, so as to obtain a reference face feature map with stronger feature representation ability.

[0031] Further, considering that the target video usually contains a large number of frames, if noise reduction processing is performed on all frames, it will consume a large amount of computing resources. Based on this, in the embodiment of the present invention, the target image frame to be processed is extracted from the target video. It can be understood that the background noise may cover up the details and textures of the target area, and by extracting the target image frame to be processed from the target video, noise reduction processing can be performed only on the important frames. Specifically, the target image frame to be processed usually contains the face area of the target user, and the background area is relatively unimportant, so that resources can be concentrated to perform more refined noise reduction on the target area, thereby improving the noise reduction effect.

[0032] Further, considering that the target image frame to be processed also contains the key feature information about the target face image, these key feature information play an important role and influence on the subsequent noise reduction of the image frame. Therefore, in order to mine and capture the key face feature information from the target image frame to be processed, in the embodiment of the present invention, the target image frame to be processed is input into the image feature extractor based on the dilated convolutional neural network model to obtain the target image frame feature map to be processed. Among them, the dilated convolutional neural network can learn the high-level feature representation of the image, these features are crucial for understanding the image content, and the dilated convolution expands the receptive field by increasing the spacing of the convolutional kernel, enabling the network to capture a larger range of context information in the image, which helps to identify the patterns and structures in the image. That is, through the image feature extractor based on the dilated convolutional neural network model, implicit feature information and key features of different scales can be captured, and support for subsequent video noise reduction can be provided.

[0033] Step 102, perform collaborative enhancement processing on the target image frame feature map to be processed and the reference face feature map to obtain a collaboratively enhanced target image frame feature map.

[0034] Considering that the feature map of the target image frame to be processed contains the facial features and motion information specific to the target user, and the reference facial feature map contains the reference facial features and identity features of the face. At the same time, there is a similarity in facial features between the feature map of the target image frame to be processed and the reference facial feature map. Therefore, in order to enhance the features of the feature map of the target image frame to be processed based on the reference facial feature map, in the embodiments of the present invention, the feature map of the target image frame to be processed and the reference facial feature map are subjected to collaborative enhancement processing. For example, the feature map of the target image frame to be processed and the reference facial feature map are input into a collaborative attention module to obtain a collaboratively enhanced feature map of the target image frame to be processed. It can be understood that the collaborative attention module calculates the similarity information of the features of the target image frame to be processed and the reference facial features in the channels, so that the feature map of the target image frame to be processed can learn the facial feature information of the reference face from the reference facial feature map, and enhance the features of the feature map of the target image frame to be processed based on the similarity information, so as to focus on and highlight the important features related to the reference facial feature map, improve the importance and representation ability of the feature map of the target image frame to be processed, and thus generate a more accurate and comprehensive feature representation, providing a better basis for noise reduction of the image frame.

[0035] Step 103, perform content discrimination processing based on a spatial mask on the collaboratively enhanced feature map of the target image frame to be processed to obtain the collaboratively enhanced significant features of the target image frame to be processed.

[0036] Considering that the collaboratively enhanced features of the target image frame to be processed in each channel of the collaboratively enhanced feature map of the target image frame to be processed have different degrees of influence. That is, some of the collaboratively enhanced features of the target image frame to be processed are more important for noise reduction of the video image frame, and some are not particularly relevant. Therefore, in order to focus on the important content area part in the collaboratively enhanced feature map of the target image frame to be processed and reduce the interference and influence of the background area information. In the embodiments of the present invention, the collaboratively enhanced feature map of the target image frame to be processed is input into a content discrimination module based on a spatial mask for content discrimination processing to obtain a collaboratively enhanced significant feature map of the target image frame to be processed. Specifically, the content area refers to the area containing the face of the target user, and the background area refers to the area not containing the face of the target user. Among them, the spatial mask is a binary image, where the pixel value of the content area is 1 and the pixel value of the background area is 0. And through the content discrimination module based on the spatial mask, background noise and irrelevant features can be suppressed, so as to highlight the important content and significant feature information in the collaboratively enhanced feature map of the target image frame to be processed, and thus obtain a collaboratively enhanced significant feature map of the target image frame to be processed containing the important significant features related to the reference facial feature map in the target image frame to be processed, and use the collaboratively enhanced significant feature map of the target image frame to be processed as the collaboratively enhanced significant features of the target image frame to be processed.

[0037] Step 104: Based on the co-enhanced significant features of the target image frame to be processed, obtain the denoised target image frame to be processed.

[0038] Specifically, input the co-enhanced significant feature map of the target image frame to be processed into the decoder-based denoising generator to obtain the denoised target image frame to be processed. That is, perform classification processing on the co-enhanced significant features of the target image frame to be processed obtained by spatially masking and enhancing the feature map of the target image frame to be processed, so as to intelligently obtain the denoised target image frame to be processed. In this way, it helps to identify important regions in the image, such as faces, etc., ensuring that the details of these regions are retained during the denoising process, avoiding over-smoothing, and thus significantly improving the overall visual experience of the video.

[0039] The video denoising method provided by the embodiments of the present invention includes: obtaining the feature map of the target image frame to be processed based on the video to be denoised of the target user, where the target user is the user who needs to be focused on for optimization and protection during video denoising processing; performing co-enhanced processing on the feature map of the target image frame to be processed and the reference face feature map to obtain the co-enhanced feature map of the target image frame to be processed; performing content differentiation processing based on spatial masking on the co-enhanced feature map of the target image frame to be processed to obtain the co-enhanced significant features of the target image frame to be processed; and obtaining the denoised target image frame to be processed based on the co-enhanced significant features of the target image frame to be processed. The present invention helps to identify important regions in the image, such as faces, etc., ensuring that the details of these important regions are retained during the denoising process, realizing intelligent differentiation between the target person and the background during the denoising process, and retaining important details and texture information in the image, avoiding over-smoothing, and thus significantly improving the overall visual experience of the video.

[0040] Based on the above embodiments, performing co-enhanced processing on the feature map of the target image frame to be processed and the reference face feature map to obtain the co-enhanced feature map of the target image frame to be processed includes: Step 1020: Determine the full-channel target reference feature similarity matrix between the feature map of the target image frame to be processed and the reference face feature map; Step 1021: Perform normalization processing on the full-channel target reference feature similarity matrix to obtain the normalized full-channel target reference feature similarity matrix; Step 1022: Based on the normalized full-channel target reference feature similarity matrix, determine the co-attention weight vector of the target image frame for the feature map of the target image frame to be processed; Step 1023: Based on the co-attention weight vector of the target image frame, perform channel-wise multiplication on the feature map of the target image frame to be processed to obtain the co-enhanced feature map of the target image frame to be processed.

[0041] It can be understood that the full-channel target reference feature similarity matrix is a matrix generated by calculating the channel-level similarity between the feature map of the target image frame to be processed and the reference face feature map. This matrix reflects the similarity between the two feature maps in all channels and is used to measure the similarity between the feature map of the target image frame and the reference face feature map.

[0042] In one embodiment, step 1020 includes: reshaping the feature map of the target image frame to be processed and the reference face feature map to obtain the feature matrix of the target image frame to be processed and the reference face feature matrix; calculating the product between the feature matrix of the target image frame to be processed and the transposed matrix of the reference face feature matrix to obtain the full-channel target reference feature similarity matrix.

[0043] It can be understood that reshaping refers to rearranging the dimensions or shape of the data without changing the data content to adapt to specific calculation or processing requirements. For example, assume that the shape of the feature map of the target image frame to be processed is (C, H, W), where C is the number of channels, and H and W are the spatial dimensions (height and width); assume that the shape of the reference face feature map is the same as that of the target image frame to be processed, i.e., (C, H, W); the purpose of reshaping is to convert the feature map from the shape of (C, H, W) to the matrix form of (C, N), where N = H × W is the total number of spatial positions.

[0044] In the embodiment of the present invention, through matrix multiplication, the calculation of the similarity of high-dimensional feature maps is converted into matrix operations, avoiding complex calculations at each position or each channel, reducing the computational complexity, and at the same time maintaining high precision.

[0045] In one embodiment, step 1021 includes: taking each eigenvalue in the full-channel target reference feature similarity matrix as the exponent of the natural constant, calculating the exponential function value with the natural constant as the base for each position to obtain the full-channel target reference feature similarity class support matrix; calculating the sum of each column vector in the full-channel target reference feature similarity class support matrix to obtain the full-channel target reference feature similarity class support column global vector composed of multiple column vector global sum values; performing a division operation on each eigenvalue in the full-channel target reference feature similarity class support matrix and the corresponding column global vector sum value in the full-channel target reference feature similarity class support column global vector for each position to obtain the normalized full-channel target reference feature similarity matrix.

[0046] It can be understood that calculating the exponential function maps the similarity values to the positive range while amplifying the larger similarity values; calculating the sum of each column vector in the all-channel target reference feature similarity class support matrix calculates the global support degree of each column for subsequent normalization; performing division operation by position (i.e., position-wise division) normalizes the all-channel target reference feature similarity class support matrix so that the sum of elements in each column is 1, facilitating subsequent analysis or comparison.

[0047] In one embodiment, step 1022 includes: performing global average pooling on the feature map of the target image frame to be processed along the channel dimension to obtain the all-channel feature vector of the target image frame to be processed; calculating the matrix product between the normalized all-channel target reference feature similarity matrix and the all-channel feature vector of the target image frame to be processed to obtain the collaborative attention weight vector of the target image frame to be processed.

[0048] It can be understood that global average pooling compresses the spatial information of the feature map into a global feature vector at the channel level, retaining the global information of the channel. Matrix product weights the global feature vector using the normalized similarity matrix to generate a collaborative attention weight vector, reflecting the importance of different channels.

[0049] In one embodiment, the collaborative attention module processes the feature map of the target image frame to be processed and the reference face feature map using the following attention formula to obtain the collaboratively enhanced feature map of the target image frame to be processed. The attention formula is: ; ; ; ; ; ; ; where and respectively represent the feature map of the target image frame to be processed and the reference face feature map, represents shape reshaping processing, and are respectively the feature matrix of the target image frame to be processed and the reference face feature matrix, represents the transpose matrix of, represents the all-channel target reference feature similarity matrix, and both represent the feature values at each position in the all-channel target reference feature similarity matrix, represents the number of row vectors in the all-channel target reference feature similarity matrix, represents the exponential operation with the natural constant e as the base, represents the eigenvalues at each position in the normalized all-channel target reference feature similarity matrix, represents performing global average pooling on the feature map, represents the all-channel feature vector of the target image frame to be processed, represents the normalized all-channel target reference feature similarity matrix, represents the collaborative attention weight vector of the target image frame to be processed, represents the collaboratively enhanced feature map of the target image frame to be processed, represents matrix multiplication.

[0050] In the embodiments of the present invention, by calculating the similarity information of the features of the target image frame to be processed and the reference face features in the channels, the feature map of the target image frame to be processed can learn the facial feature information of the reference face from the reference face feature map, and based on the similarity information, the feature map of the target image frame to be processed is feature-enhanced to focus on and highlight the important features related to the reference face feature map, so as to improve the importance and representation ability of the features of the target image frame to be processed, thereby generating a more accurate and comprehensive feature representation, providing a better basis for noise reduction of the image frame.

[0051] Based on the above embodiments, the content discrimination processing based on the spatial mask is performed on the collaboratively enhanced feature map of the target image frame to be processed to obtain the significant features of the collaboratively enhanced target image frame, including: Step 1030, performing feature representation on the collaboratively enhanced feature map of the target image frame to be processed to obtain the represented feature map of the collaboratively enhanced target image frame; Step 1031, setting the feature values greater than or equal to a predetermined threshold at each position of the represented feature map of the collaboratively enhanced target image frame to 1, and the rest to 0, to obtain the masked feature map of the collaboratively enhanced target image frame; Step 1032, performing element-wise multiplication of the masked feature map of the collaboratively enhanced target image frame and the feature map of the collaboratively enhanced target image frame to obtain the significant features of the collaboratively enhanced target image frame.

[0052] It can be understood that feature representation generally refers to further extracting or transforming the feature map through certain operations (such as convolution, activation functions, etc.) to enhance its feature expression ability. Compare the feature values at each position of the feature map that collaboratively enhances the representation of the target image frame to be processed with a predetermined threshold. If the feature value is greater than or equal to the threshold, set the value at that position to 1; if the feature value is less than the threshold, set the value at that position to 0. Among them, the binary mask feature map is used to indicate which positions have significant features (value is 1) and which positions have insignificant features (value is 0). The role of the mask feature map is to retain significant features (positions with value 1) and suppress insignificant features (positions with value 0). By multiplying element-wise (i.e., multiplying each element), the significant part in the feature map that collaboratively enhances the target image frame to be processed can be extracted.

[0053] The extraction of significant features can be understood as: further extracting or transforming the feature map that collaboratively enhances the target image frame to be processed; generating a binary mask through threshold processing to indicate the positions of significant features; using the mask to extract significant features to obtain the significant features of the target image frame that collaboratively enhances the frame to be processed. Through feature representation, mask generation, and element-wise multiplication, significant features are extracted from the feature map that collaboratively enhances the target image frame to be processed, retaining the target information and suppressing noise, providing a higher-quality feature representation for subsequent tasks.

[0054] In one embodiment, step 1030 includes: taking the negative of the feature values at each position of the feature map that collaboratively enhances the target image frame to be processed as the exponent of the natural constant, calculating the element-wise exponential function value with the natural constant as the base, to obtain the class support feature map of the target image frame that collaboratively enhances the frame to be processed; calculating the reciprocal of the sum of the feature values at each position in the class support feature map of the target image frame that collaboratively enhances the frame to be processed and the constant 1, to obtain the feature map that represents the target image frame that collaboratively enhances the frame to be processed.

[0055] It can be understood that the exponential function maps the feature values to a positive range, and at the same time, the role of the negative sign is to map larger feature values to smaller values and smaller feature values to larger values. The reciprocal operation maps the values of the class support feature map to a smaller range (between 0 and 1), while retaining the relative magnitude relationship of the feature values. Based on this, obtaining the feature map that represents the target image frame that collaboratively enhances the frame to be processed can be understood as: mapping the feature values to a positive range and reversing the magnitude relationship of the feature values through the negative sign; then, normalizing the values of the class support feature map to between 0 and 1 for subsequent processing or analysis.

[0056] In one embodiment, use a content discrimination module based on a spatial mask to process the feature map of the target image frame that collaboratively enhances the frame to be processed with the following saliency extraction formula to obtain the saliency feature map of the target image frame that collaboratively enhances the frame to be processed. Among them, the saliency extraction formula is: ; ; ; Among them, represents the feature value at the position for co-reinforcing the feature map of the target image frame to be processed, represents masking processing, represents the exponential function with the natural constant e as the base, represents the feature value at the position for co-reinforcing the masked feature map of the target image frame to be processed, represents a hyperparameter, represents the co-reinforced masked feature map of the target image frame to be processed, represents the co-reinforced feature map of the target image frame to be processed, represents element-wise multiplication by position, represents the co-reinforced significant feature map of the target image frame to be processed.

[0057] In the embodiment of the present invention, the content discrimination module based on the spatial mask can suppress background noise and irrelevant features, thereby highlighting the important content and significant feature information in the co-reinforced feature map of the target image frame to be processed, and thus obtaining a co-reinforced significant feature map of the target image frame to be processed that contains important significant features related to the reference face feature map in the target image frame to be processed, and using the co-reinforced significant feature map of the target image frame to be processed as the co-reinforced significant feature of the target image frame to be processed.

[0058] To further analyze and explain the video noise reduction method proposed by the present invention, refer to the following embodiments.

[0059] The embodiment of the present invention specifically proposes a video directional noise reduction method. By obtaining a reference face image of a target user and obtaining the target video of the target user from a media stream, and using image and video processing technologies based on deep learning to respectively extract features from the reference face image and the target video of the target user for each image frame, so as to intelligently obtain a noise reduction target image frame to be processed according to the significant features of the target video image frame. In this way, it helps to identify important regions in the image, such as faces, etc., ensuring that the details of these regions are retained during the noise reduction process and avoiding over-smoothing, thereby significantly improving the overall visual perception of the video.

[0060] Refer to Figures 2-3 , according to the video directional noise reduction method of the embodiment of the present invention, it mainly includes the following steps: S110, obtaining a reference face image of a target user; S120, extract face features from the reference face image to obtain a reference face feature map; S130, obtain the target video of the target user from the media stream; S140, extract the target image frame to be processed from the target video; S150, extract image features from the target image frame to be processed to obtain a target image frame feature map to be processed; S160, input the target image frame feature map to be processed and the reference face feature map into the collaborative attention module to obtain a collaboratively enhanced target image frame feature map; S170, input the collaboratively enhanced target image frame feature map into the content discrimination module based on a spatial mask to obtain a collaboratively enhanced target image frame salient feature map as the salient feature of the collaboratively enhanced target image frame; S180, obtain a denoised target image frame to be processed based on the salient feature of the collaboratively enhanced target image frame to be processed.

[0061] In S110, obtain the reference face image of the target user. It should be understood that the reference face image of the target user usually refers to a high-quality and clear face photo, which is used as a benchmark for analysis and recognition. Based on this, in the embodiments of the present invention, obtaining the reference face image of the target user and analyzing and processing it can provide accurate reference information for the face in the video, thereby improving the accuracy of video denoising.

[0062] In S120, input the reference face image into a face feature extractor based on a dilated convolutional neural network model to obtain a reference face feature map. Correspondingly, considering that the reference face image reflects implicit feature information about the face, such as facial features, the shape and position of eyes and mouth; texture features, such as skin texture and wrinkles. And the dilated convolutional neural network model is a special convolutional neural network, which expands the receptive field by introducing holes (i.e., skipping some pixels) in the convolutional kernel, which enables the dilated convolutional neural network model to capture a larger range of context information in the image, thereby extracting more representative features. Therefore, in the embodiments of the present invention, inputting the reference face image into a face feature extractor based on a dilated convolutional neural network model can capture and extract the representative feature information implicit in the reference face image, thereby obtaining a reference face feature map with stronger feature representation ability.

[0063] In S130, obtain the target video of the target user from the media stream. It should be understood that considering that the target video of the target user refers to the actually captured video stream, which may be real-time or recorded and involves the target user. Therefore, in the embodiments of the present invention, the target video of the target user is obtained from the media stream, and video image frames are extracted and analyzed. At the same time, reference comparison is performed based on the reference image to better denoise the video frames.

[0064] In S140, extract the target image frames to be processed from the target video. Correspondingly, considering that the target video usually contains a large number of frames, denoising all frames will consume a large amount of computing resources. Based on this, in the embodiments of the present invention, the target image frames to be processed are extracted from the target video. That is to say, background noise may mask the details and textures of the target area, and by extracting the target image frames to be processed from the target video, only important frames can be denoised. Specifically, the target image frames to be processed usually contain the face area of the target user, while the background area is relatively unimportant, so that resources can be concentrated to perform more refined denoising on the target area, thereby improving the denoising effect.

[0065] In S150, input the target image frames to be processed into an image feature extractor based on the dilated convolutional neural network model to obtain the feature map of the target image frames to be processed. It should be understood that considering that the target image frames to be processed also contain key feature information about the target face image, these key feature information play an important role and have an impact on subsequent denoising of the image frames. Therefore, in the embodiments of the present invention, the target image frames to be processed are input into an image feature extractor based on the dilated convolutional neural network model to obtain the feature map of the target image frames to be processed.

[0066] In S160, considering that the feature map of the target image frames to be processed contains the face features and motion information of a specific target user, and the reference face feature map contains the reference face features and identity features. At the same time, there is a similarity in face features between the feature map of the target image frames to be processed and the reference face feature map. Therefore, in the embodiments of the present invention, in order to enhance the features of the feature map of the target image frames to be processed based on the reference face feature map, the feature map of the target image frames to be processed and the reference face feature map are input into the collaborative attention module to obtain the collaboratively enhanced feature map of the target image frames to be processed.

[0067] As Figure 4 shown, inputting the feature map of the target image frames to be processed and the reference face feature map into the collaborative attention module to obtain the collaboratively enhanced feature map of the target image frames to be processed includes: S210, calculate the full-channel target-reference feature similarity matrix between the feature map of the target image frames to be processed and the reference face feature map; S220, normalize the full-channel target reference feature similarity matrix to obtain a normalized full-channel target reference feature similarity matrix; S230, calculate the co-attention weight vector of the target image frame to be processed based on the normalized full-channel target reference feature similarity matrix; S240, use the co-attention weight vector of the target image frame to be processed as the weight vector, and calculate the per-channel product of the target image frame feature map to be processed to obtain the co-enhanced target image frame feature map to be processed.

[0068] In one embodiment, calculating the full-channel target reference feature similarity matrix between the target image frame feature map to be processed and the reference face feature map includes: reshaping the target image frame feature map to be processed and the reference face feature map to obtain a target image frame feature matrix to be processed and a reference face feature matrix; and calculating the product between the target image frame feature matrix to be processed and the transposed matrix of the reference face feature matrix to obtain the full-channel target reference feature similarity matrix between the target image frame feature map to be processed and the reference face feature map.

[0069] In one embodiment, normalizing the full-channel target reference feature similarity matrix to obtain a normalized full-channel target reference feature similarity matrix includes: taking each eigenvalue in the full-channel target reference feature similarity matrix as the exponent of the natural constant to calculate the exponential function value of the natural constant at each position to obtain a full-channel target reference feature similarity class support matrix; calculating the sum of each column vector in the full-channel target reference feature similarity class support matrix to obtain a full-channel target reference feature similarity class support column global vector composed of multiple column vector global sum values; and calculating the per-position division between each eigenvalue in the full-channel target reference feature similarity class support matrix and the corresponding column global vector sum value in the full-channel target reference feature similarity class support column global vector to obtain the normalized full-channel target reference feature similarity matrix.

[0070] In one embodiment, calculating the co-attention weight vector of the target image frame to be processed based on the normalized full-channel target reference feature similarity matrix includes: performing global average pooling on the target image frame feature map along the channel dimension to obtain a full-channel feature vector of the target image frame to be processed; and calculating the matrix product between the normalized full-channel target reference feature similarity matrix and the full-channel feature vector of the target image frame to be processed to obtain the co-attention weight vector of the target image frame to be processed.

[0071] In S170, it should be understood that considering that the co-reinforcement target image frame features in each channel of the co-reinforcement target image frame feature map to be processed have different degrees of influence, that is, some co-reinforcement target image frame features are more important for video image frame noise reduction, while some are not particularly relevant. Therefore, in order to focus on the important content area part in the co-reinforcement target image frame feature map and reduce the interference and influence of background area information, in the embodiment of the present invention, the co-reinforcement target image frame feature map is input into the content discrimination module based on a spatial mask to obtain the co-reinforcement target image frame salient feature map.

[0072] As Figure 5 shown, inputting the co-reinforcement target image frame feature map into the content discrimination module based on a spatial mask to obtain the co-reinforcement target image frame salient feature map as the co-reinforcement target image frame salient feature includes: S310, performing feature representation on the co-reinforcement target image frame feature map to obtain the co-reinforcement target image frame represented feature map; S320, setting the feature values greater than or equal to a predetermined threshold at each position of the co-reinforcement target image frame represented feature map to one, and the rest to zero to obtain the co-reinforcement target image frame mask feature map; S330, performing element-wise multiplication of the co-reinforcement target image frame mask feature map and the co-reinforcement target image frame feature map to obtain the co-reinforcement target image frame salient feature map.

[0073] In one embodiment, performing feature representation on the co-reinforcement target image frame feature map to obtain the co-reinforcement target image frame represented feature map includes: taking the negative of the feature value at each position of the co-reinforcement target image frame feature map as the exponent of the natural constant to calculate the element-wise exponential function value with the natural constant as the base to obtain the co-reinforcement target image frame class support feature map; and calculating the reciprocal of the sum of the feature value at each position in the co-reinforcement target image frame class support feature map and the constant one to obtain the co-reinforcement target image frame represented feature map.

[0074] In S180, the co-reinforced significant feature map of the target image frame to be processed is input into a decoder-based noise reduction generator to obtain a noise-reduced target image frame to be processed. That is, the co-reinforced significant features obtained by performing spatial mask saliency on the co-reinforced target image frame feature map of the target image frame to be processed are classified to intelligently obtain the noise-reduced target image frame to be processed. In this way, it helps to identify important regions in the image, such as faces, etc., ensuring that the details of these regions are retained during the noise reduction process and avoiding over-smoothing, thereby significantly improving the overall visual experience of the video.

[0075] The video directional noise reduction method provided by the embodiments of the present invention obtains a reference face image of a target user and the target video of the target user from a media stream, and uses deep learning-based image and video processing technologies to extract features from the reference face image and the target video image frames of the target user respectively, so as to intelligently obtain a noise-reduced target image frame to be processed according to the significant features of the target video image frames. In this way, it helps to identify important regions in the image, such as faces, etc., and at the same time ensures that the details of these regions are retained during the noise reduction process and avoids over-smoothing, so as to significantly improve the overall visual experience of the video.

[0076] The video noise reduction device provided by the present invention will be described below. The video noise reduction device described below can be mutually corresponded and referred to the video noise reduction method described above.

[0077] Reference Figure 6 , the video noise reduction device provided by the present invention includes a feature map acquisition module 601, a co-reinforcement processing module 602, a content discrimination module 603, and a video noise reduction module 604.

[0078] The feature map acquisition module 601 is configured to acquire a feature map of a target image frame to be processed based on a video to be noise-reduced of a target user; the target user is a user who needs to be key optimized and protected during video noise reduction processing; The co-reinforcement processing module 602 is configured to perform co-reinforcement processing on the feature map of the target image frame to be processed and the reference face feature map to obtain a co-reinforced feature map of the target image frame to be processed; The content discrimination module 603 is configured to perform content discrimination processing based on a spatial mask on the co-reinforced feature map of the target image frame to be processed to obtain a co-reinforced significant feature of the target image frame to be processed; The video noise reduction module 604 is configured to obtain a noise-reduced target image frame to be processed based on the co-reinforced significant feature of the target image frame to be processed.

[0079] The video noise reduction device provided by the embodiment of the present invention obtains a target image frame feature map to be processed based on the video to be noise-reduced of the target user; the target user is the user who needs to be key optimized and protected in the video noise reduction process; the target image frame feature map to be processed and the reference face feature map are subjected to collaborative enhancement processing to obtain a collaborative enhanced target image frame feature map to be processed; the collaborative enhanced target image frame feature map to be processed is subjected to content differentiation processing based on a spatial mask to obtain a significant feature of the collaborative enhanced target image frame to be processed; based on the significant feature of the collaborative enhanced target image frame to be processed, a noise reduction target image frame to be processed is obtained. The present invention helps to identify important regions in the image, such as the face, etc., ensures that the details of these important regions are retained during the noise reduction process, realizes intelligent differentiation between the target person and the background during the noise reduction process, and retains important details and texture information in the image, avoiding over-smoothing, thereby significantly improving the overall visual experience of the video.

[0080] In one embodiment, the collaborative enhancement processing module 602 is specifically configured to: Determine a full-channel target reference feature similarity matrix between the target image frame feature map to be processed and the reference face feature map; Perform normalization processing on the full-channel target reference feature similarity matrix to obtain a normalized full-channel target reference feature similarity matrix; Based on the normalized full-channel target reference feature similarity matrix, determine a collaborative attention weight vector of the target image frame to be processed for the target image frame feature map to be processed; Based on the collaborative attention weight vector of the target image frame to be processed, perform a per-channel product on the target image frame feature map to be processed to obtain the collaborative enhanced target image frame feature map to be processed.

[0081] In one embodiment, the collaborative enhancement processing module 602 is specifically configured to: Reshape the target image frame feature map to be processed and the reference face feature map to obtain a target image frame feature matrix to be processed and a reference face feature matrix; Calculate the product between the target image frame feature matrix to be processed and the transposed matrix of the reference face feature matrix to obtain the full-channel target reference feature similarity matrix.

[0082] In one embodiment, the collaborative enhancement processing module 602 is specifically configured to: Taking each eigenvalue in the full-channel target reference feature similarity matrix as the exponent of the natural constant, calculate the exponential function value with the natural constant as the base according to the position to obtain a full-channel target reference feature similarity class support matrix; Calculate the sum of each column vector in the full-channel target reference feature similarity class support matrix to obtain a full-channel target reference feature similarity class support column global vector composed of multiple column vector global sum values; Perform a division operation on each eigenvalue in the full-channel target reference feature similarity class support matrix and the corresponding column global vector sum value in the full-channel target reference feature similarity class support column global vector by position to obtain the normalized full-channel target reference feature similarity matrix.

[0083] In one embodiment, the collaborative reinforcement processing module 602 is specifically configured to: Perform global average pooling on the to-be-processed target image frame feature map along the channel dimension to obtain a to-be-processed target image frame full-channel feature vector; Calculate the matrix product between the normalized full-channel target reference feature similarity matrix and the to-be-processed target image frame full-channel feature vector to obtain the to-be-processed target image frame collaborative attention weight vector.

[0084] In one embodiment, the content discrimination module 603 is specifically configured to: Perform feature characterization on the collaboratively reinforced to-be-processed target image frame feature map to obtain a collaboratively reinforced to-be-processed target image frame characterization feature map; Set the feature values greater than or equal to a predetermined threshold at each position in the collaboratively reinforced to-be-processed target image frame characterization feature map to 1, and the rest to 0 to obtain a collaboratively reinforced to-be-processed target image frame mask feature map; Perform a pointwise multiplication of the collaboratively reinforced to-be-processed target image frame mask feature map and the collaboratively reinforced to-be-processed target image frame feature map by position to obtain the collaboratively reinforced to-be-processed target image frame salient feature.

[0085] In one embodiment, the content discrimination module 603 is specifically configured to: Use the negative of the feature value at each position of the collaboratively reinforced to-be-processed target image frame feature map as the exponent of the natural constant, and calculate the exponential function value with the natural constant as the base by position to obtain a collaboratively reinforced to-be-processed target image frame class support feature map; Calculate the reciprocal of the sum of the feature value at each position in the collaboratively reinforced to-be-processed target image frame class support feature map and the constant 1 to obtain the collaboratively reinforced to-be-processed target image frame characterization feature map.

[0086] Figure 7 Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 7As shown in the figure, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 complete communication with each other through the communication bus 740. The processor 710 may call the logical instructions in the memory 730 to execute a video noise reduction method, which includes: obtaining a feature map of a target image frame to be processed based on the video to be noise-reduced of a target user; the target user being a user who needs to be optimized and protected with emphasis in video noise reduction processing; performing collaborative enhancement processing on the feature map of the target image frame to be processed and a reference face feature map to obtain a collaboratively enhanced feature map of the target image frame to be processed; performing content discrimination processing based on a spatial mask on the collaboratively enhanced feature map of the target image frame to be processed to obtain a significant feature of the collaboratively enhanced target image frame to be processed; and obtaining a target image frame to be noise-reduced based on the significant feature of the collaboratively enhanced target image frame to be processed.

[0087] In addition, when the logical instructions in the foregoing memory 730 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0088] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the video noise reduction method provided by each of the above methods. The method includes: obtaining a target image frame feature map to be processed based on the video to be noise-reduced of a target user; the target user is a user who needs to be optimized and protected with emphasis in video noise reduction processing; performing collaborative enhancement processing on the target image frame feature map to be processed and a reference face feature map to obtain a collaboratively enhanced target image frame feature map to be processed; performing content discrimination processing based on a spatial mask on the collaboratively enhanced target image frame feature map to be processed to obtain a significantly featured collaboratively enhanced target image frame to be processed; and obtaining a target image frame to be noise-reduced based on the significantly featured collaboratively enhanced target image frame to be processed.

[0089] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the execution of the video noise reduction method provided by each of the above methods. The method includes: obtaining a target image frame feature map to be processed based on the video to be noise-reduced of a target user; the target user is a user who needs to be optimized and protected with emphasis in video noise reduction processing; performing collaborative enhancement processing on the target image frame feature map to be processed and a reference face feature map to obtain a collaboratively enhanced target image frame feature map to be processed; performing content discrimination processing based on a spatial mask on the collaboratively enhanced target image frame feature map to be processed to obtain a significantly featured collaboratively enhanced target image frame to be processed; and obtaining a target image frame to be noise-reduced based on the significantly featured collaboratively enhanced target image frame to be processed.

[0090] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0091] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0092] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video noise reduction method, characterized in that: include: Based on the target user's video to be denoised, a feature map of the target image frame to be processed is obtained; The target users are users who need to be optimized and protected in the video noise reduction process; Performing collaborative enhancement processing on the target image frame feature map to be processed and the reference face feature map to obtain a collaboratively enhanced target image frame feature map to be processed; Performing content differentiation processing based on a spatial mask on the collaboratively enhanced target image frame feature map to be processed, to obtain significant features of the collaboratively enhanced target image frame to be processed; Based on the collaboratively enhanced salient features of the target image frame to be processed, a denoised target image frame to be processed is obtained.

2. The video noise reduction method according to claim 1, characterized in that: The step of performing collaborative enhancement processing on the target image frame feature map to be processed and the reference face feature map to obtain a collaborative enhanced target image frame feature map to be processed includes: Determine a full-channel target reference feature similarity matrix between the target image frame feature map to be processed and the reference face feature map; Normalizing the full-channel target reference feature similarity matrix to obtain a normalized full-channel target reference feature similarity matrix; Determining a target image frame collaborative attention weight vector for the target image frame feature map to be processed based on the normalized all-channel target reference feature similarity matrix; Based on the collaborative attention weight vector of the target image frame to be processed, the feature map of the target image frame to be processed is multiplied channel by channel to obtain the collaboratively enhanced feature map of the target image frame to be processed.

3. The video noise reduction method according to claim 2, characterized in that: The step of determining a full-channel target reference feature similarity matrix between the target image frame feature map to be processed and the reference face feature map comprises: Reshape the target image frame feature map to be processed and the reference face feature map to obtain a target image frame feature matrix to be processed and a reference face feature matrix; The product between the feature matrix of the target image frame to be processed and the transposed matrix of the reference face feature matrix is ​​calculated to obtain the full-channel target reference feature similarity matrix.

4. The video noise reduction method according to claim 2, characterized in that: The normalizing process of the full-channel target reference feature similarity matrix to obtain a normalized full-channel target reference feature similarity matrix includes: Taking each eigenvalue in the full-channel target reference feature similarity matrix as an exponent of a natural constant, calculating the exponential function value based on the natural constant according to the position, and obtaining a full-channel target reference feature similarity class support matrix; Calculate the sum of each column vector in the full-channel target reference feature similarity class support matrix to obtain a full-channel target reference feature similarity class support column global vector composed of multiple column vector global sum values; Each eigenvalue in the full-channel target reference feature similarity class support matrix is ​​divided by the sum of the corresponding column global vectors in the full-channel target reference feature similarity class support column global vectors by position to obtain the normalized full-channel target reference feature similarity matrix.

5. The video noise reduction method according to claim 2, characterized in that: The step of determining the target image frame collaborative attention weight vector of the target image frame feature map to be processed based on the normalized all-channel target reference feature similarity matrix comprises: Performing global mean pooling along the channel dimension on the feature map of the target image frame to be processed to obtain a full-channel feature vector of the target image frame to be processed; The matrix product between the normalized full-channel target reference feature similarity matrix and the full-channel feature vector of the target image frame to be processed is calculated to obtain the collaborative attention weight vector of the target image frame to be processed.

6. The video noise reduction method according to claim 1, characterized in that: The step of performing content differentiation processing based on a spatial mask on the collaboratively enhanced target image frame feature map to be processed to obtain a collaboratively enhanced target image frame salient feature includes: Characterizing the collaboratively enhanced target image frame to be processed, to obtain a collaboratively enhanced target image frame to be processed characterization feature map; The feature values ​​greater than or equal to the predetermined threshold value in each position of the collaboratively enhanced target image frame to be processed are set to 1, and the rest are set to 0, so as to obtain a collaboratively enhanced target image frame to be processed mask feature map; The collaboratively enhanced target image frame mask feature map to be processed and the collaboratively enhanced target image frame feature map to be processed are multiplied by position to obtain the collaboratively enhanced target image frame salient features to be processed.

7. The video noise reduction method according to claim 6, characterized in that: The step of performing feature characterization on the collaboratively enhanced target image frame feature map to be processed to obtain the collaboratively enhanced target image frame feature map to be processed includes: Taking the negative number of the characteristic value of each position of the collaboratively enhanced target image frame to be processed as the exponent of the natural constant, calculating the exponential function value based on the natural constant according to the position, and obtaining the collaboratively enhanced target image frame to be processed class support characteristic map; The reciprocal of the sum of the characteristic values ​​of each position in the collaboratively enhanced target image frame to be processed class support feature map and the constant 1 is calculated to obtain the collaboratively enhanced target image frame to be processed representation feature map.

8. A video noise reduction device, characterized in that: include: A feature map acquisition module is used to acquire a feature map of a target image frame to be processed based on the target user's video to be denoised; The target users are users who need to be optimized and protected in the video noise reduction process; A collaborative enhancement processing module, used for performing collaborative enhancement processing on the target image frame feature map to be processed and the reference face feature map to obtain a collaborative enhanced target image frame feature map to be processed; A content differentiation module, used for performing content differentiation processing on the collaboratively enhanced target image frame feature map to be processed based on a spatial mask to obtain significant features of the collaboratively enhanced target image frame to be processed; The video noise reduction module is used to obtain a noise-reduced target image frame to be processed based on the collaboratively enhanced salient features of the target image frame to be processed.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the video noise reduction method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the video noise reduction method according to any one of claims 1 to 7 is implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the video noise reduction method according to any one of claims 1 to 7 is implemented.