A video noise reduction method and system of a dorsoventral-like attention mechanism
A video denoising method and system based on a dorsal-ventral attention mechanism utilizes the design of dorsal and ventral features to achieve adaptive optimization of the video denoising system in different scenarios. This solves the problems of poor adaptability and high cost of manual parameter tuning in existing technologies, and improves the denoising effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG XINMAI SILICON CO LTD
- Filing Date
- 2022-11-29
- Publication Date
- 2026-05-22
AI Technical Summary
Existing video noise reduction technologies have poor adaptability to different external conditions, high scene dependence and feature complexity, high cost of manual parameter tuning, and difficulty in achieving adaptive optimization effects.
A dorsal-ventral attention mechanism is introduced, and video denoising is adaptively performed through the design of dorsal and ventral features. Feature control is designed in conjunction with the dorsal-ventral attention mechanism to achieve feature segmentation and adaptive switching, and multiple video image features are integrated to improve the denoising effect.
It improves the adaptability and noise reduction effect of the video noise reduction system, reduces the cost of manual parameter tuning, and achieves adaptive optimization in different scenarios.
Smart Images

Figure CN116309093B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image and video processing technology, and in particular to a video noise reduction method and system that resembles a dorsoventral attention mechanism. Background Technology
[0002] Currently, video noise reduction is mainly divided into two categories: traditional noise reduction and machine learning. Traditional spatiotemporal frequency hybrid noise reduction models are mostly designed for fixed feature detection and noise reduction for a certain type of noise. Although the noise reduction effect is relatively good, the input adaptability is poor and the effect varies greatly under different external conditions. Similarly, common machine learning models also have similar drawbacks, either being too dependent on the scene or having too high feature complexity, requiring a lot of manual parameter tuning. Summary of the Invention
[0003] One of the objectives of this invention is to provide a video denoising method and system with a dorsal-ventral attention mechanism. The method and system introduce dorsal-like features and ventral-like features to design and classify the denoising features of the video denoising system, so that the video denoising method and system can adaptively input the region of interest of the system for denoising. Through the design of the dorsal-like features and ventral-like features, the denoising system has better adaptability to feature segmentation.
[0004] Another objective of this invention is to provide a video noise reduction method and system with a dorsal-ventral attention mechanism. The method and system, through the design of the dorsal and ventral features and the design of feature control using the dorsal-ventral attention interruption mechanism, enable better adaptive switching in different scenarios and reduce the cost of manual parameter tuning.
[0005] Another objective of this invention is to provide a video noise reduction method and system with a dorsal-ventral attention mechanism. The method and system, through the design of the dorsal and ventral features, enable the integration of various video image features in a manner that conforms to the visual mechanism, thereby improving the optimization effect of noise reduction partitioning.
[0006] To achieve at least one of the above-mentioned objectives, the present invention further provides a video noise reduction method simulating a dorsal-ventral attention mechanism, the method comprising:
[0007] The video streams are sequentially subjected to brightness and contrast calibration to obtain a calibration video sequence.
[0008] Based on the benchmark video sequence, video visual spatial features are obtained as back-side features, and visual stimulus features are obtained as ventral features.
[0009] The dorsal-lateral and ventral-lateral features are integrated using adaptive noise reduction feature control based on a dorsal-ventral attention mechanism to obtain integrated features.
[0010] Calculate the feature ratio of the integrated features, and perform adaptive noise reduction to output the video stream based on the feature ratio of the integrated features.
[0011] According to a preferred embodiment of the present invention, the benchmarking data acquisition includes: acquiring a video stream, performing brightness and contrast benchmarking operations on the video stream, wherein the average brightness of each frame of the input video is calculated, a target brightness value is preset, the quotient of the target brightness value and the average brightness of the input video is calculated as a comparison coefficient for video frame brightness benchmarking, and brightness benchmarking operations are performed on each video frame to generate first benchmarking data.
[0012] According to another preferred embodiment of the present invention, after completing the brightness calibration and obtaining the first calibration data, a target contrast parameter is preset, a contrast histogram of the first calibration data is obtained, and a target contrast calibration operation under the first calibration data is performed according to the target contrast parameter and the contrast histogram to obtain the second calibration data. The second calibration data is then converted to a standard calibration video sequence.
[0013] According to another preferred embodiment of the present invention, at least one of the intra-frame pixel gradient features, intra-frame spatial detail distribution features, inter-frame pixel difference features, and inter-frame block orientation features of the benchmark video sequence are extracted as the back-side features; at least one of the local brightness clustering features, local color clustering features, motion light intensity difference clustering features, and motion special angle clustering features of the benchmark video sequence are extracted as the ventral features.
[0014] According to another preferred embodiment of the present invention, the video denoising method includes: constructing and classifying denoising features using the constructed backside-like features and ventral-like features, wherein the ventral-like features are constructed as static-level features and motion-level features, and the static-level feature weights and motion-level feature weights corresponding to the ventral-like features and backside-like features are calculated, and the corresponding weights are calculated using an adaptive weighting method.
[0015] According to another preferred embodiment of the present invention, the proportion is calculated based on the calculated static-level feature weights and motion-level feature weights, and a first proportion threshold is set, wherein the first proportion threshold is greater than 50%. When the proportion of the static-level feature weights is greater than the proportion threshold, noise reduction is performed on the input benchmark video based on the weight proportion, including inter-frame IIR noise reduction method.
[0016] According to another preferred embodiment of the present invention, a second percentage threshold is set based on the calculated static feature weights and motion feature weights, wherein the second percentage threshold is less than 50%. When the percentage of the static feature weights is less than the second percentage threshold, the spatial frequency domain denoising method is used to denoise the input benchmark video based on the weight percentage.
[0017] According to another preferred embodiment of the present invention, the proportion is calculated based on the calculated static feature weights and motion feature weights. When the proportion of the static feature weights is between the first proportion threshold and the second proportion threshold, the inter-frame IIR denoising method or the spatial frequency domain denoising method is used simultaneously to denoise the benchmark video.
[0018] According to another preferred embodiment of the present invention, the feature integration method includes: calculating the basis weights, offset weights, and compensation weights under the static-level features; calculating the static-level weights based on the basis weights, offset weights, and compensation weights under the static-level features; calculating the basis weights, offset weights, and compensation weights under the motion-level features; and calculating the motion-level weights based on the basis weights, offset weights, and compensation weights under the motion-level features.
[0019] According to another preferred embodiment of the present invention, the maximum and minimum values of the intra-frame spatial detail distribution features are obtained. Based on the maximum and minimum values of the intra-frame spatial detail distribution features and the current intra-frame spatial detail distribution feature value, the basis weights of the intra-frame spatial detail distribution features are calculated as the basis weights of the static-level features. The maximum and minimum values of the intra-frame pixel gradient features are also calculated. Based on the maximum and minimum values of the intra-frame pixel gradient features and the current intra-frame pixel gradient feature value, offset weights are calculated and used as the offset weights of the static-level features. Static-level compensation features are calculated based on local color clustering features and local brightness clustering features, and used to calculate the static-level feature weights. The static-level feature weights [w] ij The calculation formula is as follows:
[0020] [w] ij =[w] baseij *4096+[w] ofstij *16+[w] cmpij ;
[0021] The basis weights for the static-level features are [w]. baseij The static-level feature offset weight is [w]. ofstij The compensation weight for static-level features is [w]. cmpij .
[0022] According to another preferred embodiment of the present invention, the motion-level basis weights are calculated based on the inter-frame block orientation features, the offset weights are calculated based on the motion intensity difference clustering features, the compensation weights are calculated based on the motion-specific angle clustering features, and the motion-level feature weights are calculated. The motion-level feature weights are calculated using the following formula: [wm] ij =[wm] basetype *32+[wm] ofsttype *16+[wm] cmpij ;
[0023] The basis weights for the motion-level features are [wm]. basetype The motion-level feature offset weight is [wm]. ofsttype The motion-level feature compensation weight is [wm]. cmpij .
[0024] To achieve at least one of the above-mentioned objectives, the present invention further provides a video noise reduction system with a dorsoventral attention mechanism, wherein the system performs the above-mentioned video noise reduction method with a dorsoventral attention mechanism.
[0025] The present invention further provides a computer-readable storage medium storing a computer program that can be executed by a processor to perform the aforementioned video noise reduction method based on a dorsal-ventral attention mechanism. Attached Figure Description
[0026] Figure 1 The diagram shown is a flowchart of a video noise reduction method based on a dorsal-ventral attention mechanism according to the present invention.
[0027] Figure 2 The diagram shown is a structural schematic of a video noise reduction system based on a dorsal-ventral attention mechanism according to the present invention.
[0028] Figure 3 This diagram illustrates the backside intra-pixel gradient features of this invention.
[0029] Figure 4 This diagram illustrates the orientation features of the back-side inter-frame block in this invention.
[0030] Figure 5 This displays the special angle clustering features of ventral motion in this invention, where the central black block represents the current pixel and the boundary black blocks represent the corresponding angle feature blocks. Detailed Implementation
[0031] The following description is intended to disclose the present invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art. The basic principles of the invention defined in the following description can be applied to other embodiments, modifications, improvements, equivalents, and other technical solutions that do not depart from the spirit and scope of the invention.
[0032] It is understood that the term "a" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element can be one, while in another embodiment, the number of the element can be multiple, and the term "a" should not be understood as a limitation on the number.
[0033] Please combine Figures 1-5 This invention discloses a video denoising method and system based on a dorsoventral attention mechanism. The system mainly includes: a preprocessing subsystem; a dorsoventral attention mechanism denoising feature extraction subsystem; a ventral attention mechanism denoising feature extraction subsystem; a dorsoventral attention mechanism adaptive denoising feature control subsystem; and a denoising and output subsystem. The preprocessing subsystem is used to perform a calibration operation on input video data of the same or different formats, generating a calibration video sequence that meets the requirements for video denoising feature detection. The dorsoventral attention mechanism denoising feature extraction subsystem is used to extract feature information similar to brain space. In this invention, the dorsoventral attention mechanism denoising feature for video information is a video spatial feature, and the video is constructed. The ventral attention mechanism denoising feature extraction subsystem is used to extract visual stimulus features of video information and construct ventral attention mechanism features for the video. The dorsoventral attention mechanism adaptive denoising feature control subsystem is used to integrate the constructed dorsoventral attention mechanism features and ventral attention mechanism features to construct an integrated feature, which is used for subsequent denoising operations. The noise reduction and output subsystem will perform different dominant noise reduction methods according to the weight ratio of the static and motion features in the above-mentioned integrated features.
[0034] Specifically, the preprocessing subsystem performs a preprocessing calibration operation on the input video. Taking an input RAW format video stream as an example, the calibration operation includes: setting a preset target brightness of 'a', calculating the average brightness of each frame of the input RAW video as 'b', and calculating the brightness calibration coefficient k = a / b for each frame. Further, based on the brightness calibration coefficient for each frame, the brightness of each input video frame is multiplied by the corresponding frame's calibration coefficient to complete the calibration operation for each frame. The video data that has undergone the brightness calibration operation is defined as the first calibration data.
[0035] After completing the brightness calibration of the video data, target contrast calibration parameters numl, numh, thrl, and thrh are further preset, where numl and numh represent the set number of pixels in the dark and bright areas, respectively, and thrl and thrh represent the corresponding grayscale thresholds for the dark and bright areas. The histogram distribution of the first calibration data is further statistically analyzed, and a non-linear transformation is applied to the first calibration data so that the data distribution below thrl is close to numl, and the data distribution above thrh is close to numh. The data that has completed contrast calibration is further defined as the second calibration data. The preprocessing also includes format conversion of the second calibration data, and the above calibration operation is performed on each format conversion. For example, if the original second calibration data is in RAW format, it can be converted to RGB format for the above calibration operation, and further converted to YUV data for the above calibration operation. The calibration data completed after data format conversion is defined as a video calibration sequence.
[0036] Furthermore, the back-side noise reduction feature extraction subsystem and the ventral noise reduction feature extraction subsystem are used to extract back-side and ventral features from the video benchmark sequence, respectively. The back-side features include four types: intra-frame pixel gradient features, intra-frame spatial detail distribution features, inter-frame pixel difference features, and inter-frame block orientation features. The ventral features include: local brightness clustering features, local color clustering features, motion intensity difference clustering features, and motion-specific angle clustering features.
[0037] The intra-frame pixel gradient features are the gradient differences in eight directions within the pixel neighborhood, specifically... Figure 3 Taking the display as an example, within a similar neighborhood of size 9*9, the gradient difference in each direction is calculated as the absolute value of the difference between the white pixels on both sides of the black pixel division in each direction. Furthermore, the maximum value of the absolute difference in each of the eight directions is calculated.
[0038] The method for extracting intra-frame spatial detail distribution features includes: taking the detail map of each frame of the benchmark video sequence as statistical input, and the detail map calculation method includes general Gaussian mask, Sobel difference operator, etc., to calculate the variance of the detail map of each pixel in a neighborhood of size such as 31*31, and the obtained variance is recorded as the corresponding intra-frame spatial detail distribution feature; for system environments that support large data processing, multiple neighborhood sizes can be set to obtain multiple intra-frame spatial detail distribution features for subsequent noise reduction adaptive control.
[0039] The inter-frame pixel difference feature extraction method includes: for each pixel in each frame of the benchmark video sequence, calculating the absolute value of the difference between it and the data of the frames before and after it, and taking the larger value as the inter-frame pixel difference feature of the current point.
[0040] The method for extracting inter-frame block orientation features includes: extracting the aforementioned inter-frame pixel difference features, such as... Figure 4 As an example of azimuth, the inter-frame pixel difference features of each azimuth block are summed and statistically analyzed to obtain the inter-frame block azimuth features.
[0041] It should be noted that the above four methods of extracting backside features are only illustrative examples. This invention can set up similar regions of different sizes and different neighborhood directions for different feature extractions. This invention is not limited to the examples described above.
[0042] The ventral visual stimulus features extracted by the ventral noise reduction feature subsystem include: local brightness clustering features, local color clustering features, motion light intensity difference clustering features, and motion special angle clustering features.
[0043] The method for calculating local brightness clustering features includes: performing brightness downsampling on the benchmark video sequence, using block mean downsampling, and setting the downsampling size to resolution adaptive mode, i.e., setting a fixed downsampling factor Ndn according to the input resolution.
[0044] The method for calculating local color clustering features includes: mainly local saturation clustering features, that is, extracting saturation from the UV components of the benchmark data 3, and simultaneously performing adaptive downsampling on the saturation.
[0045] The method for calculating the motion intensity difference clustering feature includes: performing inter-frame difference calculation on the benchmark video sequence, arranging each frame of data in the format [x,y,diffabs], where x and y represent the horizontal and vertical coordinates of the pixel, respectively, and diffabs is the inter-frame difference of the pixel. Using the kmeans clustering algorithm, the motion intensity difference clustering feature of the frame image is statistically analyzed, that is, the centroid of the motion intensity difference distribution [xc, yc].
[0046] The method for calculating the special motion angle clustering features includes: calculating the inter-frame motion direction of the benchmark video sequence, such as... Figure 4 In this embodiment, a neighborhood size of 31*31 is selected for calculating the inter-frame motion direction, with 12 special angles. The motion direction calculation uses SAD block matching with a block size of 7*7. It should be noted that the above-described ventral feature extraction method is merely illustrative, and the embodiments of this invention are not limited to the above examples.
[0047] After extracting the aforementioned dorsal and ventral features, adaptive feature integration is further performed on these features. The feature integration method employs the dorsal-ventral attention mechanism corresponding to the adaptive denoising feature control subsystem, and adaptive denoising is applied to the integrated features. A specific example is provided below:
[0048] The intra-pixel gradient features of the backside frame are denoted as [d]. ij , where ij represents pixel coordinates;
[0049] The spatial detail distribution features within the backside frame are denoted as [σ]. ij , where ij represents pixel coordinates;
[0050] The backside inter-frame pixel difference feature is denoted as [fd]. ij , where ij represents pixel coordinates;
[0051] The inter-frame block orientation feature of the back side is denoted as [fdoa], which is the block orientation identifier;
[0052] The clustering feature of local brightness on the ventral side of the class is denoted as [Lit]. we , where we represents the downsampling coordinates;
[0053] The local color clustering feature on the ventral side of the class is denoted as [Sat]. we , where we represents the downsampling coordinates;
[0054] The clustering feature of the ventral motion intensity difference is denoted as [xc, yc], which is the centroid of the motion intensity difference distribution.
[0055] Clustering features for ventral motion at specific angles are denoted as [Ang]. ij Where ij represents pixel coordinates;
[0056] The feature integration involves constructing features based on the still-level and motion-level features of the benchmark video sequence. The still-level features are static features within the video, and the motion-level features are moving features within the video. This invention applies different adaptive noise reduction schemes to the still-level and motion-level features constructed in different video frames according to their weight ratios.
[0057] The feature weights include base weights, offset weights, and supplementary weights.
[0058] In the described static-level feature construction method, the spatial detail distribution feature [σ] within the backside frame is used. ij Calculate the basis weights [w] baseij With 8-bit precision, the maximum value of the spatial detail distribution feature of the back side of the current frame class is calculated [σ]. max The minimum value of the spatial detail distribution feature in the backside frame [σ] min The formula for calculating the basis weights of the intra-frame spatial detail distribution features of the backside class is as follows:
[0059] In the construction of the static-level features, intra-frame pixel gradient features [d] are used. ij Calculate the offset weight [w] ofstijPrecision 8-bit, statistical analysis of the maximum intra-frame pixel gradient feature in the current frame [d] max ] and minimum intra-frame pixel gradient features [d min The formula for calculating the offset weight of the intra-frame pixel gradient features is as follows:
[0060]
[0061] In the static-level feature construction method, local brightness clustering features [Lit] are used. we And local color clustering features [Sat] we Calculate the compensation weight [w] cmpij The precision is configured to 4 bits, and each downsampled local brightness cluster feature [Lit] is used. we And local color clustering features [Sat] we Find the mean, denoted as [LS]. we The same as above, the largest [LS] max ] and minimum [LS min The compensation weights [w] are calculated using the same method as the base weights and offset weights described above. cmpij .
[0062] The basis weights [w] are constructed from the static-level features. baseij Offset weight [w] ofstij and compensation weight [w] cmpij Then, the static-level feature weights [w] are calculated using the following formula. ij :
[0063] [w] ij =[w] baseij *4096+[w] ofstij *16+[w] cmpij .
[0064] In motion-level feature construction methods, basis weights [wm] are calculated using inter-frame block orientation features. basetype With a precision of 1 bit, the pixel [wm] within the block area is assigned the block orientation. basetype Set to 1, for pixels outside the corresponding block area in the block orientation [wm]. basetype Set to 0.
[0065] In the motion-level feature construction method, the offset weight [wm] is calculated using the motion light intensity difference clustering feature [xc, yc]. ofsttype With a precision of 1 bit, the radiation range is located in the vicinity of [xc, yc], and the pixel size is 1 / Ndn (resolution). Within the radiation range, the pixel size is [wm]. basetype Set to 1 for pixels outside the radiation range [wm] basetypeSet to 0, where Ndn is the downsampling factor set by the above clustering features;
[0066] In motion-level feature construction methods, features are clustered from specific motion perspectives [Ang]. ij Calculate the compensation weights [wm] cmpij With a precision of 4 bits, its value is equal to [Ang]. ij The feature itself.
[0067] Basis weights in motion-level feature acquisition [wm] basetype [wm] ofsttype And compensation weights [wm] cmpij Then, the motion-level feature weights [wm] are calculated using the following formula. ij :
[0068] [wm] ij =[wm] basetype *32+[wm] ofsttype *16+[wm] cmpij .
[0069] It should be noted that the output of the adaptive noise reduction feature control subsystem of the dorsoventral attention mechanism is a static-level feature weight [w]. ij Motion-level feature weights [wm] ij .
[0070] The noise reduction and output subsystem will be based on the static level feature weights [w]. ij Motion-level feature weights [wm] ij Different spatiotemporal frequency mixing noise reduction methods are applied to the input video based on the proportion of different parameters.
[0071] The weighting of static and motion features can be classified using either hard or soft thresholding. For example, a soft thresholding method for weighting can be used as follows: First, a first weighting threshold is set based on the calculated static and motion feature weights. If the first threshold is greater than 50%, and the static feature weight's weight is greater than this threshold, then inter-frame IIR denoising is used to denoise the input benchmark video. Second, a second weighting threshold is set based on the calculated static and motion feature weights. If the second threshold is less than 50%, and the static feature weight's weight is less than this threshold, then spatial frequency domain denoising is used to denoise the input benchmark video. Third, if the static feature weight's weight is between the first and second weighting thresholds, then both inter-frame IIR denoising and spatial frequency domain denoising are used simultaneously to denoise the benchmark video.
[0072] It should be noted that the aforementioned dorsal and ventral features, as well as the noise reduction methods, may include, but are not limited to, the following optional components:
[0073] Backside intra-frame gradient feature design includes, but is not limited to, pixel-level gradients, block-level gradients, and even subject-level gradients.
[0074] As an improvement, the intra-frame distribution features of the back side can be designed in terms of dimensionality, including but not limited to spatial distribution and value range distribution; and in terms of feature design, including but not limited to brightness distribution, detail distribution, and color distribution.
[0075] As an improvement, the design of back-side inter-frame differential features includes, but is not limited to, inter-frame pixel difference and inter-frame optical flow difference;
[0076] As an improvement, the design of back-side inter-frame orientation features includes, but is not limited to, inter-frame block orientation differences;
[0077] For video visual stimulus features:
[0078] As an improvement, the visual stimulus features of the ventral type include, but are not limited to, brightness clustering features, color clustering features, motion difference clustering features, motion direction clustering features, etc.
[0079] As an improvement, the design of ventral brightness clustering features includes, but is not limited to, global brightness clustering, local brightness clustering, and local brightness difference clustering;
[0080] As an improvement, the design of color clustering features for the ventral region includes, but is not limited to, global color clustering, local color clustering, and local color difference clustering;
[0081] As an improvement, the design of strong clustering features based on ventral motion difference includes, but is not limited to, motion light intensity difference clustering and motion projection transformation clustering.
[0082] As an improvement, the design of clustering features for ventral motion direction includes, but is not limited to, motion vector clustering and special angle clustering;
[0083] The adaptive noise reduction feature control subsystem of the dorsal-ventral attention mechanism is mainly used to analyze the dorsal-ventral features of the input video and obtain adaptive noise reduction features that match the input video.
[0084] As an improvement, the dorsoventral feature analysis method of the self-interrupted dorsoventral attention mechanism adaptive denoising feature control subsystem includes, but is not limited to, adaptive weighted analysis and adaptive feature aggregation analysis.
[0085] As an improvement, the adaptive noise reduction feature output of the dorsoventral attention mechanism adaptive noise reduction feature control subsystem includes, but is not limited to, weight distribution, center distribution, etc.
[0086] As an improvement, the noise reduction and output subsystem mainly performs noise reduction processing based on adaptive noise reduction features and outputs the corresponding noise reduction result video stream;
[0087] As an improvement, the noise reduction and output subsystem noise reduction processing includes, but is not limited to, spatial domain noise reduction, frequency domain noise reduction, time domain noise reduction, and combinations thereof;
[0088] As an improvement, the noise reduction and output subsystem outputs video streams including but not limited to common unencoded format types.
[0089] In particular, according to embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs the functions defined in the methods of this application. It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wire segments, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless segments, wire segments, optical fibers, RF, etc., or any suitable combination thereof.
[0090] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0091] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention. The purpose of the present invention has been fully and effectively achieved. The functions and structural principles of the present invention have been shown and explained in the embodiments. Without departing from the stated principles, the implementation of the present invention may have any variations or modifications.
Claims
1. A video noise reduction method based on a dorsoventral attention mechanism, characterized in that, The method includes: The video streams are sequentially subjected to brightness and contrast calibration to obtain a calibration video sequence. Based on the benchmark video sequence, video visual spatial features are obtained as back-side features, and visual stimulus features are obtained as ventral features. The dorsal-lateral and ventral-lateral features are integrated using adaptive noise reduction feature control based on a dorsoventral attention mechanism to obtain integrated features, including: The feature integration method includes: calculating the basis weights, offset weights, and compensation weights under static features, and calculating the static weights based on the basis weights, offset weights, and compensation weights under the static features; calculating the basis weights, offset weights, and compensation weights under motion features, and calculating the motion weights based on the basis weights, offset weights, and compensation weights under the motion features. Calculate the feature proportion of the integrated features, and perform adaptive noise reduction to output the video stream based on the feature proportion of the integrated features, including: The proportions of static-level feature weights and motion-level feature weights are calculated, and a first proportion threshold is set, wherein the first proportion threshold is greater than 50%. When the proportion of static-level feature weights is greater than the proportion threshold, noise reduction of the input benchmark video is performed based on the weight proportion, including inter-frame IIR noise reduction method. The proportions of static and motion feature weights are calculated, and a second proportion threshold is set, wherein the second proportion threshold is less than 50%. When the proportion of static feature weights is less than the second proportion threshold, spatial frequency domain noise reduction is used to denoise the input benchmark video based on the weight proportions. The proportions of the static-level feature weights and motion-level feature weights are calculated. When the proportion of the static-level feature weights is between the first proportion threshold and the second proportion threshold, the inter-frame IIR denoising method or the spatial frequency domain denoising method is used simultaneously to denoise the benchmark video.
2. The video noise reduction method for a dorsoventral attention mechanism according to claim 1, characterized in that, At least one of the following features from the benchmark video sequence—intra-frame pixel gradient features, intra-frame spatial detail distribution features, inter-frame pixel difference features, and inter-frame block orientation features—is extracted as the back-side feature; at least one of the following features from the benchmark video sequence—local brightness clustering features, local color clustering features, motion light intensity difference clustering features, and motion special angle clustering features—is extracted as the ventral feature.
3. The video noise reduction method for a dorsoventral attention mechanism according to claim 1, characterized in that, The video denoising method includes: constructing and classifying denoising features using the constructed backside and ventral features, wherein the ventral features are constructed as static features and motion features, and the static feature weights and motion feature weights corresponding to the ventral and backside features are calculated, and the corresponding weights are calculated using an adaptive weighting method.
4. The video noise reduction method for a dorsoventral attention mechanism according to claim 1, characterized in that, Obtain the maximum and minimum values of the intra-frame spatial detail distribution features. Based on the maximum and minimum values of the intra-frame spatial detail distribution features and the current intra-frame spatial detail distribution feature value, calculate the basis weights of the intra-frame spatial detail distribution features as the basis weights of the static-level features. Also calculate the maximum and minimum values of the intra-frame pixel gradient features. Based on the maximum and minimum values of the intra-frame pixel gradient features and the current intra-frame pixel gradient feature value, calculate the offset weights as the offset weights of the static-level features. Finally, calculate static-level compensation features based on local color clustering features and local brightness clustering features, used to calculate the static-level feature weights [w]. ij The calculation formula is as follows: ; The basis weights for the static-level features are [w]. baseij The static-level feature offset weight is [w]. ofstij The compensation weight for static-level features is [w]. cmpij .
5. The video noise reduction method for a dorsoventral attention mechanism according to claim 1, characterized in that, The motion-level basis weights are calculated based on the inter-frame block orientation features, the offset weights of the motion intensity difference clustering features are calculated, the compensation weights of the motion-specific angle clustering features are calculated, and the motion-level feature weights are calculated. The formula for calculating the motion-level feature weights is as follows: ; The basis weights for the motion-level features are [wm]. basetype The motion-level feature offset weight is [wm]. ofsttype The motion-level feature compensation weight is [wm]. cmpij .
6. A video noise reduction system with a dorsoventral attention mechanism, characterized in that, The system executes a video noise reduction method based on a dorsal-ventral attention mechanism as described in any one of claims 1-5.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed by a processor as a video noise reduction method with a dorsal-ventral attention mechanism as described in any one of claims 1-5.