Image noise processing method and device, electronic equipment and storage medium
By acquiring noise and target image data from i-frame images, and combining human visual effects and preset image attention events, the noise coarseness level is calculated, solving the problem of discrepancy between human visual perception and existing technologies, and achieving image denoising effects that are more in line with human visual perception.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TP-LINK
- Filing Date
- 2023-09-06
- Publication Date
- 2026-04-17
AI Technical Summary
Existing image denoising techniques fail to effectively consider human visual perception, resulting in unsatisfactory denoising effects. The images are far removed from human visual perception, leading to problems such as noise fluctuations or alternating blurriness.
By acquiring the noise image data and target image data of i-frame images, and combining human visual effects and preset image attention events, the target noise coarseness factor and video influence factor are calculated to determine the noise coarseness level and achieve rational noise reduction.
It improves the user experience of image denoising, ensuring that the timing and intensity of denoising are consistent with human visual perception, and reduces noise fluctuations and blurring.
Smart Images

Figure CN117196979B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, and in particular relates to image noise processing methods, apparatus, electronic devices and storage media. Background Technology
[0002] Currently, there are many image noise reduction techniques, but they mainly focus on the objective numerical quantification of a single image, paying less attention to the timing and intensity of noise removal, and failing to evaluate the severity and removal effect of image noise from the perspective of how the human eye perceives image noise. This can easily lead to problems such as noticeable noise fluctuations or blurring due to unreasonable timing or intensity of noise removal, resulting in a significant discrepancy between the final denoised image or video and the human eye's perception, leading to unsatisfactory image noise reduction results. Summary of the Invention
[0003] This application provides an image noise processing method, apparatus, electronic device, and storage medium, which can solve the technical problems that the image result after denoising by existing denoising technology does not meet the subjective evaluation of the human eye and the image denoising effect is not ideal.
[0004] In a first aspect, embodiments of this application provide an image noise processing method, including:
[0005] At the current moment, the audio and video stream to be inspected is acquired. The audio and video stream to be inspected includes i frames of images. The i frames of images include the current frame image corresponding to the current moment, where i is a positive integer greater than 1.
[0006] Obtain the corresponding noise image data from the i-frame images respectively to obtain i noise image data;
[0007] Based on the i noisy image data, obtain the target noise coarsening factor corresponding to the current frame image;
[0008] The corresponding target image data is obtained from the i-frame images respectively, and the target image data is determined according to the preset image attention event;
[0009] Based on the noise image data and the target image data corresponding to the i-frame image, obtain the target video influence factor corresponding to the current frame image;
[0010] The noise coarseness level of the current frame image is determined based on the target noise coarseness factor and the target video influence factor.
[0011] Secondly, embodiments of this application provide an image noise processing apparatus, comprising:
[0012] The audio and video acquisition module is used to acquire the audio and video stream to be inspected at the current time. The audio and video stream to be inspected includes i frames of images, and the i frames of images include the current frame image corresponding to the current time, where i is a positive integer greater than 1.
[0013] The first extraction module is used to obtain the corresponding noise image data from the i-frame images respectively;
[0014] The first determining module is used to obtain the target noise coarsening factor corresponding to the current frame image based on the noise image data corresponding to the i-frame image;
[0015] The second image extraction module is used to obtain corresponding target image data from the i-frame images respectively, wherein the target image data is determined according to a preset image attention event;
[0016] The second determining module is used to obtain the target video influence factor corresponding to the current frame image based on the noise image data and the target image data corresponding to the i-frame image;
[0017] The noise level acquisition module is used to determine the noise coarseness level of the current frame image based on the target noise coarseness factor and the target video influence factor.
[0018] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect.
[0019] Fourthly, embodiments of this application provide a storage medium, which is a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.
[0020] The beneficial effects of the embodiments in this application compared with the prior art are:
[0021] In this embodiment of the application, the audio and video stream to be detected, including i-frame images, is acquired at the current moment in combination with human visual effects. The i-frame images include the current frame image corresponding to the current moment and other frame images involving human visual effects.
[0022] Obtain the corresponding noise image data from the i-frame images respectively, and then obtain the target noise coarsening factor corresponding to the current frame image based on the noise image data corresponding to each frame image, so as to characterize the noise coarsening degree of the current frame image;
[0023] According to the preset image attention event, the corresponding target image data is obtained from the i-frame image respectively. Then, based on the noise image data and target image data corresponding to each frame image, the target video influence factor corresponding to the current frame image is obtained, so as to objectively reflect whether the noise of the current frame image belongs to the image content that the viewing user cares about.
[0024] Finally, the noise coarsening level of the current frame image is determined based on the target noise coarsening factor and the target video influence factor, so that computer electronic devices can perform rational denoising of the current frame image from two dimensions: denoising timing and denoising intensity, according to the noise coarsening level, thereby improving the user experience. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0026] Figure 1 This is a schematic flowchart of an image noise processing method provided in an embodiment of this application;
[0027] Figure 2 This is a schematic diagram of the structure of a noise processing device provided in one embodiment of this application;
[0028] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0030] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0031] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0032] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0033] References to "one embodiment" or "some embodiments" in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized.
[0034] Example 1:
[0035] Currently, there are many image noise reduction techniques, but they mainly focus on the objective numerical quantification of a single image, paying less attention to the timing and intensity of noise removal, and failing to evaluate the severity and removal effect of image noise from the perspective of how the human eye perceives image noise. This can easily lead to problems such as noticeable noise fluctuations or blurring due to unreasonable timing or intensity of noise removal, resulting in a significant discrepancy between the final denoised image or video and the human eye's perception, leading to unsatisfactory image noise reduction results.
[0036] Therefore, this embodiment provides an image noise processing method, which can obtain a noise coarseness level close to human eye perception to guide computer equipment to rationally denoise the image;
[0037] The image noise processing method in this embodiment is executed by an electronic device. This electronic device can be a security monitoring device with a high-definition camera, or a terminal device with an IoT monitoring application installed, such as a mobile terminal (e.g., a mobile phone or tablet computer) or a computer host (e.g., a desktop computer or laptop computer) with an IoT monitoring application installed. The terminal device with the IoT monitoring application can wirelessly communicate with the security monitoring device, and the mobile terminal or computer host can remotely acquire the audio and video streams collected by the security monitoring device.
[0038] The noise processing method provided in the embodiments of this application is described below with reference to the accompanying drawings.
[0039] Figure 1 A flowchart of an image noise processing method provided in an embodiment of this application is shown, which is applied to an electronic device, and is described in detail below:
[0040] Step S11: At the current moment, acquire the audio and video stream to be inspected. The audio and video stream to be inspected includes i frames of images. The i frames of images include the current frame image corresponding to the current moment and i-1 frames of images located before the current moment, where i is a positive integer greater than 1.
[0041] Understandably, in order to take advantage of the persistence characteristics of human vision (0.4 seconds) and hearing (0.02 seconds), electronic devices can cache a sequence of images (including the current image) of a certain length as a sequence to be inspected (the audio and video stream to be inspected) before playing audio and video streams.
[0042] The audio and video stream to be inspected can be a cached sequence of the audio and video stream obtained from the security monitoring equipment. In this embodiment, the audio and video stream to be inspected, including i-frame images, is obtained at the current moment in combination with human visual effects. The i-frame images include the current frame image corresponding to the current moment and other i-1 frame images involving human visual effects.
[0043] In this specific implementation, the standard frame rate of 25fps is used. When watching a video, the first 10 images and the first 0.06 seconds of audio are temporarily retained in the brain. Let the image sequence be denoted as R. t =[I t-i+1 I t-i+2 , ..., I t ], where I t Let i be the image at time t, and i be the length of the image sequence (10 in this embodiment, i.e., i = 10).
[0044] Step S12: Obtain the corresponding noise image data from the i-frame images respectively to obtain i noise image data;
[0045] In a specific implementation, the noisy image data can specifically include noisy images. In this embodiment, a pre-trained image noise extractor (CNN network) can be used to extract noise from each image in the above image sequence, resulting in a total of N noisy images for the i-frame image sequence. t =[n t-i+1 ,n t-i+2 ,…n t ]; where n t This represents the noise image data at time t.
[0046] Step S13: Obtain the target noise coarsening factor corresponding to the current frame image based on the i noisy image data;
[0047] Understandably, due to the persistence of vision in the human eye, the i-1 frames preceding the current frame will also affect the noise coarseness of the current frame. For example, the historical image sequence 0.4 seconds before the current frame will also affect the noise perception of the current image, and this effect is strongly correlated with the order in which the preceding images appear. Therefore, in one embodiment, a pre-trained time attention method can be used to calculate the influence weight of each preceding image in the audio / video stream to be inspected on the noise of the current frame. The obtained influence weight is used as a preset weight (the preset weight is determined by the temporal order between the i frames in the audio / video stream to be inspected). Finally, the image sequence N is processed according to different preset weights. t =[n t-i+1 ,n t-i+2 ,…n t The target noise coarsening factor CI corresponding to the current frame image is obtained by weighted summation of all i-frame noise images. t .
[0048] Step S14: Obtain the corresponding target image data from the i-frame images respectively, wherein the target image data is determined according to a preset image attention event;
[0049] It is understood that the preset image attention event can be understood as an event that occurs on a pre-set object of attention in the software program of the electronic device. The object of attention can be a person and / or a target body, the target body can be a vehicle or other objects that the user pays attention to, etc., and the event can be the behavior of a person / animal, or the driving mode of a vehicle / vehicle, etc.
[0050] Step S15: Based on the noise image data and the target image data corresponding to the i-frame image, obtain the target video impact factor corresponding to the current frame image. The target video impact factor characterizes the degree of influence of noise in the current frame image on the viewing experience.
[0051] In practical implementation, a target image event detector (such as human figures, vehicles, and other target objects of user interest) can be pre-trained, and the R of the image sequence can be processed accordingly. t =[I t-i+1 I t-i+2 , ..., I t For each frame of the image, target image data (which may be the target image event region) is extracted. The image feature overlap (which may be pixel feature overlap) between the noise image data and the target image data is calculated. For example, it may be the ratio between the number of noise pixels falling within the target image event region and the total number of pixels in the target image event region. Finally, the average value of the image feature overlap corresponding to each frame is calculated, and this average value is used as the target video impact factor EI corresponding to the current frame. t .
[0052] Step S16: Determine the noise coarseness level of the current frame image based on the target noise coarseness factor and the target video influence factor.
[0053] In one embodiment, an audio-visual dataset (referred to as a self-built audio-visual dataset) can be pre-established in the software program of the electronic device. This self-built audio-visual dataset includes various typical audio-visual data to comprehensively list all possible image noise conditions (including corresponding audio-visual segments). A human visual rating of these image noise conditions (typical audio-visual data) is obtained using MOS (Mean Opinion Score, a pre-rating system for users based on their experience). The typical audio-visual data in the self-built audio-visual dataset can be dynamically expanded; users can add, modify, and replace existing audio-visual data according to their needs.
[0054] For ease of explanation, this embodiment denotes the human eye rating as y. Each typical audio and video data corresponds to a human eye rating y, and each human eye rating y is strongly correlated with at least the image noise coarsening factor CI (noise content instance) and the video impact factor EI (image content instance). The software program of this embodiment pre-maps the feature factors (CI and EI) that are strongly correlated with the image noise rating to different human eye ratings. That is, by fitting, a mapping function relationship f is found such that y = f(CI, EI) holds true for the vast majority of data (typical audio and video data) in the self-built audio and video dataset. In other words, the mapping function relationship is related to each typical audio and video data in the self-built audio and video dataset.
[0055] The electronic device combines the target noise coarsening factor and the target video influence factor of the current frame image with the pre-fitted mapping function relationship f, compares the combined undetermined rating result with different human eye ratings in the self-built audio and video dataset, and takes the human eye rating that is closest to the undetermined rating result as the target human eye rating, which represents the noise coarsening level of the current frame image.
[0056] Understandably, the electronic device will use the target noise coarsening factor CI of the current frame image in step S16. t and the target video impact factor EI t Substituting the mapping function relation f(CI,EI), since CI=CI t EI = EI t The final rating result was y. t =f(CI) t EI t The pending rating result is y. tThe results are compared with different human eye ratings in the self-built audio and video dataset, and the result that matches the pending rating is selected. t The closest human eye rating is used as the target human eye rating, which represents the noise coarseness level of the current frame image, i.e., the noise coarseness level that this application ultimately aims to obtain.
[0057] In this embodiment: The audio / video stream to be inspected, including an i-frame image, is acquired at the current moment, taking into account human visual effects. The i-frame image includes the current frame image corresponding to the current moment and other frame images involving human visual effects. Corresponding noise image data is obtained from each of the i-frame images (each frame image). Then, a target noise coarsening factor corresponding to the current frame image is obtained based on the noise image data corresponding to each frame image to characterize the noise coarsening degree of the current frame image. According to a preset image attention event, corresponding target image data is obtained from each of the i-frame images. Then, a target video impact factor corresponding to the current frame image is obtained based on the noise image data and target image data corresponding to each frame image to objectively reflect whether the noise in the current frame image belongs to the image content that the viewing user cares about. Finally, the noise coarsening level of the current frame image is determined based on the target noise coarsening factor and the target video impact factor, so that computer electronic devices can rationally denoise the current frame image according to the noise coarsening level from two dimensions: denoising timing and denoising intensity, thereby improving the user experience.
[0058] Example 2:
[0059] In some embodiments, the noise image data specifically includes a noise image and a background image;
[0060] Specifically, in this embodiment, a pre-trained image noise extractor (CNN network) can be used to extract noise from each image in the above image sequence to obtain a total of i frames of image sequences with noise images and background images (the background images can be obtained from the original images after removing the noise images).
[0061] Wherein, the noisy image sequence is represented by N. t =[n t-i+1 ,n t-i+2 ,…n t ] indicates that, where n t The background image sequence is represented by B. t =[b t-i+1 ,b t-i+2 ,…b t ] indicates that, among them, B t This represents the background image at time t.
[0062] Understandably, the content of the background image can affect the human eye's perception of noise (for example, noise on a large monochrome area will be particularly noticeable). Therefore, this embodiment considers that the background image can help to more accurately determine the noise coarseness of the current frame image.
[0063] Accordingly, the image noise processing method further includes:
[0064] A1, obtain the pixel data corresponding to each noise point in the noise image;
[0065] A2, assign corresponding noise weights to each noise point in the noisy image according to the type of pixel data;
[0066] A3, update the pixel data corresponding to each noise point in the noise image according to the noise weight, and obtain a new noise image with new noise point pixel values;
[0067] Specifically, the pixel data in step A1 can be the L (luminance), a (chrominance), and b (chrominance) information of the noisy image;
[0068] For noisy images at the same time (same frame), based on the L, a, b information of the noisy image, a pre-trained spatial attention method is used to calculate the corresponding single-point saliency weight according to the type of noise pixel data of the noisy image (that is, to assign corresponding noise weights to each noise point of the noisy image according to the type of pixel data). The type of noise pixel data can include at least the specific value of brightness L (black and white depth, i.e. brightness), the specific value of chroma a (red and green chroma), and the specific value of chroma b (blue and yellow chroma).
[0069] Simultaneously, to facilitate calculation, the pixel data at this point is converted from L, a, b information into RGB information of noise pixel values; and the obtained noise weights are multiplied by the corresponding RGB information of noise pixel values to update the pixel data of each noisy image, obtaining new noise pixel values, and updating the original noisy image sequence to N`. t =[n` t-i+1 ,n` t-i+2 ,…n` t This yields a new noise image with new noise point pixel values;
[0070] A4, obtain the brightness information of the background image;
[0071] For a background image corresponding to a noise image at the same time, obtain the L channel information (which may be the brightness information of the background image) for each background image.
[0072] A5 assigns corresponding human eye noise perception weights to different background images according to the type of brightness information;
[0073] A6. Obtain the noise coarsening factor corresponding to the new noise image based on the human eye noise perception weight;
[0074] Specifically, this embodiment can calculate the human eye noise perception weight of each pixel in a single background image based on the L-channel information (brightness information) of each background image through a spatial attention mechanism. The calculated human eye noise perception weight is then multiplied by the new noise image updated at that moment (obtained in step A3) to obtain the noise coarsening factor sequence C corresponding to the sequence of the new noise image. t =[C t-i+1 C t-i+2 ,…C t ], where C t This represents the noise coarsening factor of the noisy image at time t.
[0075] Accordingly, step S13, "obtaining the target noise coarsening factor corresponding to the current frame image based on the noise image data corresponding to each frame image," further includes the following sub-steps:
[0076] A7, obtain the preset weights corresponding to each frame image in the audio and video stream to be inspected, wherein the preset weights are determined by the temporal order between the i-frame images;
[0077] A8, according to the preset weights, the noise coarsening factors corresponding to each new noise image are weighted and summed to obtain the target noise coarsening factor corresponding to the current frame image.
[0078] Understandably, due to the persistence of vision in the human eye, the i-1 frames preceding the current frame will also affect the noise coarsness of the current frame. For example, the historical image sequence 0.4 seconds before the current frame will also affect the noise perception of the current image, and this effect is strongly correlated with the order in which the preceding images appear. Therefore, in this embodiment, a pre-trained time attention method is used to calculate the influence weight of each preceding image in the audio-visual stream to be inspected on the noise coarsness factor of the current frame. The obtained influence weight is used as a preset weight (the preset weight is determined by the temporal order between the i frames in the audio-visual stream to be inspected). Finally, based on different preset weights, the noise coarsness factor C of the total i frame image sequence is calculated. t =[C t-i+1 C t-i+2 ,…C t The target noise coarsening factor CI corresponding to the current frame image is obtained by performing a weighted summation. t .
[0079] In this embodiment, the noise images in the video stream under test are updated using pixel data corresponding to each noise point. Background images from the video stream are introduced to determine the noise coarsening factor for each new noise image. Finally, the noise coarsening factors for each new noise image are weighted and summed according to a preset weight determined by the temporal order of the frames to obtain the target noise coarsening factor for the current frame image. This allows for a more accurate determination of the noise coarsening level of the current frame image. Therefore, obtaining the target noise coarsening factor in this embodiment is beneficial for more accurately determining the noise coarsening level of the current frame image.
[0080] Example 3:
[0081] In some embodiments, during the process of determining the target video impact factor corresponding to the current frame image, audio data in the audio-visual stream to be inspected is also taken into consideration. The image noise processing method of this embodiment further includes:
[0082] C1, Obtain audio data corresponding to the time of the current frame image from the audio and video stream to be inspected, the audio data including multiple frame audio segments;
[0083] In this specific implementation, the same as in Embodiment 1 above, this example uses a standard frame rate of 25fps. When the human eye watches a video, the first 10 images and the first 0.06 seconds (including the 0.04-second audio frame corresponding to the current image) are temporarily retained in the brain. Let the image sequence be denoted as R. t ={I t-i+1 I t-i+2 , ..., I t ], where I t Let A be the image at time t, and i be the length of the image sequence (in this embodiment, i = 10). Correspondingly, the audio segment sequence of the audio data corresponding to the image at time t is denoted as A. t ;
[0084] Understandably, assuming the video frame rate is 25fps, the duration of the current frame at the current moment is 40ms. Adding the 20ms of the persistence of hearing, the audio sequence length corresponding to the current frame is 60ms. This audio sequence can be divided into N frames (e.g., 6 frames) of a specific length (e.g., 10ms), denoted as A. t =[a t-N-1 ,a t-N-2 ,…,a t-1, a t ], where t is the current time and N is the number of frames after the audio sequence is split.
[0085] C2, the target sound is obtained from each frame audio segment respectively, and the target sound is determined according to a preset voice attention event;
[0086] C3, based on the target sound in each frame audio segment, obtain the target audio influence factor of the current frame image;
[0087] It should be noted that the preset voice attention events in this embodiment can include sounds made by humans, sounds made by animals, artificial sounds, and other target sounds that the user is interested in; by pre-training a detector for voice attention events, the image sequence R is processed. t The corresponding audio sequence A t Target sound detection is performed, with a result of 1 or 0 (representing the presence or absence of the target sound, respectively). The target sound detection results for the entire audio sequence are multiplied by the corresponding influence weight (predefined), and this value is called the target audio influence factor EA for the current frame image. t ;
[0088] In the specific implementation, A t After processing by the target audio event detector, the audio event sequence is output. ae t,M For audio frame a t The detection result of the Mth predefined audio event takes a value of 1 or 0 (representing the presence or absence of the target sound, respectively). For example, the first predefined audio event is crying, the second is footsteps, and the audio frame a... t If crying is detected but footsteps are not, then ae t,1 =1, ae t,2 The value is 0. Let the predefined audio event influence weight matrix be... Then EA t =WA·AE t (Dot product).
[0089] In one embodiment, the target audio influence factor EA of the current frame image t This can be achieved through the following sub-steps:
[0090] Sub-step (1) constructs the required high-level audio semantic matrix by category and assigns corresponding influence weights. The weights are predefined. Taking the working environment of electronic devices as a security scenario as an example, audio semantics that are of primary concern (e.g., humans) or semantic concern (e.g., illustrations) usually have higher influence weights in security scenarios, while the weights of audio semantic types that are not of concern can be set to 0. For security scenarios, an example of constructing dimensions is as follows:
[0091]
[0092] Sub-step (2) Establish an audio semantic dataset based on the audio semantic matrix. The method for establishing the dataset is as follows: use a specific audio device (e.g., an artificial ear) to collect various types of sounds at different distances and directions, and label the corresponding categories, semantics, and specific sound type labels.
[0093] Sub-step (3) Train the class extractor. Using the class labels of all data in the established dataset, train the class extractor based on the publicly available Transformer model.
[0094] Sub-step (4) Train the high-level semantic extractor. Based on the category extractor trained above, freeze the parameters of the Encoder part of the Transformer, and use the high-level semantic labels of all data in the established dataset to fine-tune the parameters of the Decoder part to obtain the high-level semantic extractor.
[0095] Sub-step (5) Calculate the target audio impact factor. Input the audio sequence to be examined into the high-level semantic extractor to extract the high-level semantic matrix (elements are 1 or 0, representing the presence or absence of the specified semantics, respectively). Then multiply the high-level semantic matrix by the predefined impact weight matrix, and the resulting value is the target audio impact factor EA. t .
[0096] Furthermore, in this embodiment, step S16, which determines the noise coarsening level of the current frame image based on the target noise coarsening factor and the target video influence factor, further includes:
[0097] C4: Determine the noise coarseness level of the current frame image based on the noise coarseness factor, the target video influence factor, and the target audio influence factor.
[0098] In a specific implementation, the electronic device inputs the target noise coarsening factor, the target video influence factor, and the target audio influence factor of the current frame image as a whole into the pre-fitted mapping function relationship f. The output undetermined rating result is compared with different human eye ratings in the self-built audio and video dataset. The human eye rating that is closest to the undetermined rating result is taken as the target human eye rating, which represents the noise coarsening level of the current frame image.
[0099] In one embodiment, an audio and video dataset (referred to as a self-built audio and video dataset) can be pre-established in the software program of the electronic device. The self-built audio and video dataset includes a variety of typical audio and video data, with the aim of listing all possible image noise situations (including corresponding audio and video segments) and obtaining the human eye rating of these image noise situations (typical audio and video data) through MOS (Mean Opinion Score, which is a pre-rated score by the user group based on their own experience).
[0100] Specifically, the human eye rating can be a score given by a user based on their viewing experience after watching a typical audio-visual data set as a whole; this score represents a human eye rating. In another embodiment, each typical audio-visual data set includes noise content examples, image content examples, and audio event examples. The user pre-scores the noise content examples to obtain a noise example opinion score, the user pre-scores the image content examples to obtain a video example opinion score, and the user pre-scores the audio event examples to obtain an audio example opinion score. The noise example opinion score, the video example opinion score, and the audio example opinion score are combined to form a human eye rating.
[0101] Understandably, each typical audio and video data corresponds to a human eye rating y, and each human eye rating y is at least strongly correlated with the image noise coarsening factor CI (noise content instance), video impact factor EI (image content instance), and audio impact factor EA (audio event instance). The software program in this embodiment pre-maps the feature factors (CI, EI, and EA) that are strongly correlated with the image noise rating to different human eye ratings, that is, it finds a mapping function relationship f by fitting, such that y = f(CI, EI, EA) holds true for the vast majority of data (typical audio and video data) in the self-built audio and video dataset.
[0102] In a specific implementation, the electronic device will use the target noise coarsening factor CI of the current frame image in step C4. t The target video impact factor EI t Target audio impact factor EA t Substituting the mapping function relation f(CI,EI,EA), since CI=CI t EI = EI t EA = EA t The final rating result was y. t =f(CI) t EI t EA t The pending rating result is y. t The results are compared with different human eye ratings in the self-built audio and video dataset, and the result that matches the pending rating is selected. t The closest human eye rating is used as the target human eye rating, which represents the noise coarseness level of the current frame image, i.e., the noise coarseness level that this application ultimately aims to obtain.
[0103] The image noise processing method in this embodiment considers both the human eye's perception characteristics of the current frame image and the semantic information of the corresponding audio segment. By fusing the semantic information of the video and audio content, the noise coarseness level of the current frame image is finally determined to represent a noise level that is closer to the human eye's perception of the current frame image.
[0104] Example 4:
[0105] In some embodiments, considering that the existing conventional technology for classifying image noise intensity is limited to a single scale of a single image, it does not take into account the multiple image scales that users may use when perceiving image noise in security monitoring scenarios (different viewing devices, viewing distances, etc. cause different sensory perceptions);
[0106] Taking security monitoring scenarios as an example, users may use devices such as mobile phones / tablets, central control screens, PC monitors, and displays (large TVs or PC monitor arrays) to view surveillance footage. These viewing devices vary in size, and users also have different viewing habits. For example:
[0107] Mobile phones / tablets: In typical home settings, the main viewer is the camera owner, who usually views the camera from both a distance and at close range.
[0108] Central control screen: In smart home scenarios, cameras are usually connected to the central control screen, allowing indoor users to view them at a relatively close distance;
[0109] PC monitors: typically used in home or small commercial settings, primarily viewed by camera owners or managers (e.g., security guards), usually sitting at close range or standing at a relatively close range;
[0110] Large TV or PC monitor arrays: Commercial scenarios, where the main viewers are the camera owner's clients, such as restaurants (open kitchen projects), schools, construction sites, and other scenarios where surveillance cameras need to be mounted on walls, typically requiring a long viewing distance.
[0111] In view of the above factors, the image noise processing method proposed in this embodiment can calculate the image noise coarseness level and smoothness level by introducing an attention mechanism based on the multi-scale background, noise and audio-visual semantic information of a past image sequence of a specific length, determine the degree of image noise removal based on the noise coarseness level, and determine whether to remove image noise based on the noise smoothness level.
[0112] Specifically, before step S12, the method further includes:
[0113] D1: Scaling each frame of the i-frame image according to different scaling sizes, so that each frame of the audio and video stream to be detected has n different display sizes;
[0114] Specifically, according to the preset scale X = [X1, ... X... j ,…,X n (Magnification or reduction, simulating different viewing scales of the human eye) for image sequence R t The images are scaled and generated as n (the number of images at a preset scale). A pre-trained image noise extractor (CNN network) is used to extract noise from each image in the sequence, resulting in a noise image and a background image at each preset scale. The preset scale X is then used to calculate the noise level. j For example, let N be the noise image sequence and the background image sequence. t,j =[n t-i+1,j ,n t-i+2,j ,…n t,j ] and B t,j =[b t-i+1,j ,b t-i+2,j ,…b t,j ];n t,j b t,j This represents the noise image and background image at scale j at time t.
[0115] Step S12 further includes:
[0116] D2: Obtain the corresponding noise image data from the i-th frame image of the nth display size;
[0117] Step S13 further includes:
[0118] D3: Based on each noise image data, obtain the target noise coarsening factor of the current frame image under the nth display size;
[0119] Step S14 further includes:
[0120] D4: Obtain the corresponding target image data from the i-frame images of the nth display size respectively;
[0121] Step S15 further includes:
[0122] D5: Based on the noise image data and the target image data corresponding to each frame image of the nth display size, obtain the target video influence factor of the current frame image under the nth display size.
[0123] Understandably, before extracting the noise-related features of the video to be inspected, the security monitoring equipment needs to first enlarge and reduce the image sequence (video) to be inspected according to the possible viewing scenarios of the security monitoring equipment's camera, in order to obtain the noise feature factors perceived by "different viewing devices (i.e., the nth display size)" (target noise coarseness factor under the nth display size and target video influence factor under the nth display size).
[0124] Step S16 further includes:
[0125] D61: Obtain the playback size of the device playing the audio / video stream to be tested, wherein the device playback size is the nth display size;
[0126] D62: Determine the typical audio and video data set corresponding to the nth display size from the preset audio and video library;
[0127] D63: The target noise coarsening factor of the current frame image at the nth display size, the target video influence factor of the current frame image at the nth display size, and / or the target audio influence factor are treated as a whole and matched with various audio and video typical data in the audio and video typical data set corresponding to the nth display size. If the match is successful, the target human eye rating corresponding to the successfully matched target audio and video typical data is obtained. The target human eye rating represents the noise coarsening level of the current frame image at the nth display size.
[0128] Understandably, when constructing a pre-built library of typical audio and video data, the typical audio and video data can be assigned a score according to the display size of the playback device. Considering the differences in human visual perception under different viewing devices, the same typical audio and video data in the pre-built library may have different scores under different display sizes.
[0129] The image noise processing method in this embodiment takes into account the characteristics of human eye perception. It integrates semantic information of video content on the basis of multi-scale image noise intensity features, which is closer to the human eye's perception of the current image noise.
[0130] Example 5:
[0131] Further, after step S16, the procedure further includes:
[0132] D7: Select the corresponding processing strategy based on the noise coarsness level of the current frame image to perform noise reduction processing on the audio and video stream to be inspected.
[0133] It is understandable that the execution subject of steps S16 (steps D61, D62, and D63) and D7 can be a security monitoring device that collects the audio and video stream to be inspected, or a user-side terminal device that plays the audio and video stream to be inspected.
[0134] In one application scenario embodiment, if the executing entity is a security monitoring device, the security monitoring device will obtain the device playback size of the user-side terminal device connected to it, and perform noise reduction processing on the audio and video stream to be tested on the security monitoring device side. After denoising the audio and video stream according to the noise coarseness level of the current frame image under the nth display size, it will be sent to the user-side terminal device for playback.
[0135] In another application scenario embodiment, if the executing entity is a user-side terminal device, the user-side terminal device will receive the audio and video stream to be inspected transmitted by the security monitoring device. Then, the user-side terminal device will obtain its own device playback size, and play the audio and video stream to be inspected after denoising according to the noise coarseness level of the current frame image under the nth display size.
[0136] Specifically, the noise coarseness level can be divided into 5 levels. If the noise coarseness level of the current frame image reaches a certain level, denoising processing is performed.
[0137] In some embodiments, if the noise coarsening level of the current frame image is greater than a preset level (e.g., greater than 2), then pre-denoising with a corresponding intensity (the higher the noise coarsening level, the stronger the denoising) is performed on the current frame image according to the preset level, and the noise coarsening factor is recalculated as the noise coarsening level of the current image after pre-denoising. The noise coarsening levels of the previous N seconds' images (the value of N can be adjusted according to the specific application scenario, and N also includes the current image) are read, and a pre-trained temporal attention method is used to calculate the noise smoothing level. If the noise smoothing level of this image is less than a preset level (e.g., 3), then denoising is performed. Otherwise, denoising is not performed. The final noise coarsening level of the image is stored for subsequent calculation of the noise smoothing level.
[0138] Understandably, since subsequent steps require the noise coarseness levels of the previous N seconds of the current image when calculating the noise smoothness level, in order to avoid redundant calculations, the noise coarseness level of the current frame image is stored immediately after it is determined, so as to avoid recalculating the noise coarseness levels of all previous images when calculating the noise smoothness level of the next frame image.
[0139] The image noise processing method in this embodiment can perform image noise removal more smoothly and naturally based on human visual perception.
[0140] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0141] Example 6:
[0142] Corresponding to the noise processing method described in the above embodiments, Figure 2 A structural block diagram of an image noise processing apparatus provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiments of this application are shown.
[0143] Reference Figure 2 This image noise processing is applied to electronic devices, including:
[0144] The audio and video acquisition module 10 is used to acquire the audio and video stream to be inspected at the current time. The audio and video stream to be inspected includes i frames of images, and the i frames of images include the current frame image corresponding to the current time, where i is a positive integer greater than 1.
[0145] The first extraction module 20 is used to obtain corresponding noise image data from the i-frame images respectively;
[0146] The first determining module 30 is used to obtain the target noise coarsening factor corresponding to the current frame image based on the noise image data corresponding to the i-frame image;
[0147] The second image extraction module 40 is used to obtain corresponding target image data from the i-frame images respectively, wherein the target image data is determined according to a preset image attention event;
[0148] The second determining module 50 is used to obtain the target video influence factor corresponding to the current frame image based on the noise image data and the target image data corresponding to the i-frame image;
[0149] The noise level acquisition module 60 is used to determine the noise coarseness level of the current frame image based on the target noise coarseness factor and the target video influence factor.
[0150] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0151] Example 7:
[0152] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3As shown, the electronic device 3 of this embodiment includes: at least one processor 31 ( Figure 3 The diagram shows only one processor, memory 32, and computer program 33 stored in the memory 32 and executable on the at least one processor 31, which, when executed, implements the steps in any of the above method embodiments.
[0153] The electronic device 3 can be a security monitoring device with a high-definition camera, or a terminal device with an IoT monitoring program installed, such as a mobile terminal (such as a mobile phone or tablet computer) or a computer host (such as a desktop computer or laptop computer) with an IoT monitoring application installed.
[0154] The electronic device may include, but is not limited to, a processor 31 and a memory 32. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0155] The processor 31 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0156] In some embodiments, the memory 32 may be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. In other embodiments, the memory 32 may be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. Furthermore, the memory 32 may include both internal and external storage units of the electronic device 3. The memory 32 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory 32 can also be used to temporarily store data that has been output or will be output.
[0157] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0158] This application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor executes the computer program to implement the steps in any of the above method embodiments.
[0159] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0160] This application provides a computer program product that, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.
[0161] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying computer program code to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0162] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0163] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0164] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0165] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0166] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. An image noise processing method, characterized in that, include: At the current moment, the audio and video stream to be inspected is acquired. The audio and video stream to be inspected includes i frames of images. The i frames of images include the current frame image corresponding to the current moment, where i is a positive integer greater than 1. Obtain the corresponding noise image data from the i-frame images respectively to obtain i noise image data; Based on the i noisy image data, obtain the target noise coarsening factor corresponding to the current frame image; The corresponding target image data is obtained from the i-frame images respectively, and the target image data is determined according to the preset image attention event; Based on the noise image data and the target image data corresponding to the i-frame image, obtain the target video influence factor corresponding to the current frame image; The noise coarseness level of the current frame image is determined based on the target noise coarseness factor and the target video influence factor; The method further includes: Calculate the image feature overlap between the noisy image data and the target image data; The step of obtaining the target video impact factor corresponding to the current frame image based on the noise image data and the target image data corresponding to each frame image includes: Calculate the average value of the image feature overlap corresponding to each frame image to obtain the target video influence factor corresponding to the current frame image; The noise image data includes a noise image and a background image; The image noise processing method further includes: Obtain the pixel data corresponding to each noise point in the noise image; Assign corresponding noise weights to each noise point in the noisy image according to the type of pixel data. The pixel data corresponding to each noise point in the noise image is updated according to the noise weight to obtain a new noise image with new noise point pixel values. Obtain the brightness information of the background image; Assign corresponding human eye noise perception weights to different background images according to the type of brightness information; The noise coarsening factor corresponding to the new noise image is obtained based on the human eye noise perception weight; The step of obtaining the target noise coarsening factor corresponding to the current frame image based on the i noisy image data includes: Obtain the preset weights corresponding to i-frame images in the audio / video stream to be inspected, wherein the preset weights are determined by the temporal order between the i-frame images; The noise coarsening factors corresponding to each new noise image are weighted and summed according to the preset weights to obtain the target noise coarsening factor corresponding to the current frame image.
2. The image noise processing method as described in claim 1, characterized in that, The image noise processing method further includes: obtaining audio data corresponding to the time of the current frame image from the audio-visual stream to be detected, wherein the audio data includes multiple frame audio segments; The target sound is obtained from each frame of audio segment, and the target sound is determined according to a preset speech attention event; The target audio influence factor of the current frame image is obtained based on the target sound in each frame audio segment; Determining the noise coarsness level of the current frame image based on the target noise coarsness factor and the target video impact factor includes: The noise coarseness level of the current frame image is determined based on the noise coarseness factor, the target video impact factor, and the target audio impact factor.
3. The image noise processing method as described in claim 1, characterized in that, After determining the noise coarseness level of the current frame image based on the target noise coarseness factor and the target video impact factor, the method further includes: Based on the noise coarsness level of the current frame image, a corresponding processing strategy is selected to perform denoising processing on the audio and video stream to be inspected.
4. The image noise processing method according to any one of claims 1-3, characterized in that, Determining the noise coarsness level of the current frame image based on the target noise coarsness factor and the target video impact factor includes: Obtain the mapping function relationship between each typical audio and video data in the self-built audio and video dataset; The target noise coarsening factor, the target video influence factor, and / or the target audio influence factor are input as a whole into the mapping function relationship. The output undetermined rating result is compared with different human eye ratings in the self-built audio and video dataset. The human eye rating that is closest to the undetermined rating result is taken as the target human eye rating. The target human eye rating represents the noise coarsening level of the current frame image.
5. The image noise processing method according to any one of claims 1-4, characterized in that, Before obtaining the corresponding noise image data from the i-frame images, the method further includes: Each frame of the i-frame image is scaled according to different scaling sizes, so that each frame of the audio and video stream to be detected has n different display sizes; The step of obtaining the corresponding noise image data from the i-frame images includes: Obtain the corresponding noise image data from the i-th frame image of the nth display size; The step of obtaining the target noise coarsening factor corresponding to the current frame image based on the noise image data corresponding to each frame image includes: The target noise coarsening factor of the current frame image at the nth display size is obtained based on each noise image data. The step of obtaining the corresponding target image data from the i-frame images includes: Obtain the corresponding target image data from the i-frame images of the nth display size respectively; The step of obtaining the target video impact factor corresponding to the current frame image based on the noise image data and the target image data corresponding to each frame image includes: Based on the noise image data and the target image data corresponding to each frame image of the nth display size, the target video influence factor of the current frame image under the nth display size is obtained.
6. The image noise processing method as described in claim 5, characterized in that, The step of determining typical target audio and video data from a preset audio and video library based on the target noise coarsening factor, the target video impact factor, and / or the target audio impact factor includes: Obtain the playback size of the device playing the audio / video stream to be tested, wherein the device playback size is the nth display size; Determine the typical audio and video data set corresponding to the nth display size from the preset audio and video library; Based on the target noise coarsening factor of the current frame image at the nth display size, the target video influence factor of the current frame image at the nth display size, and / or the target audio influence factor, target audio and video typical data are determined from the audio and video typical data set corresponding to the nth display size.
7. An image noise processing apparatus, characterized in that, include: The audio and video acquisition module is used to acquire the audio and video stream to be inspected at the current time. The audio and video stream to be inspected includes i frames of images, and the i frames of images include the current frame image corresponding to the current time, where i is a positive integer greater than 1. The first extraction module is used to obtain the corresponding noise image data from the i-frame images respectively; The first determining module is used to obtain a target noise coarsening factor corresponding to the current frame image based on the noise image data corresponding to the i-frame image; wherein, the noise image data includes a noise image and a background image, and the target noise coarsening factor is obtained by: obtaining pixel data corresponding to each noise point in the noise image; assigning corresponding noise weights to each noise point in the noise image according to the type of pixel data; updating the pixel data corresponding to each noise point in the noise image according to the noise weights to obtain a new noise image with new noise point pixel values; obtaining the brightness information of the background image; assigning corresponding human eye noise perception weights to different background images according to the type of brightness information; obtaining the noise coarsening factor corresponding to the new noise image according to the human eye noise perception weights; obtaining a preset weight corresponding to the i-frame image in the audio-visual stream to be inspected, wherein the preset weights are determined by the time order between the i-frame images; and performing a weighted summation of the noise coarsening factors corresponding to each new noise image according to the preset weights to obtain the target noise coarsening factor corresponding to the current frame image; The second image extraction module is used to obtain corresponding target image data from the i-frame images respectively, wherein the target image data is determined according to a preset image attention event; The second determining module is used to obtain the target video impact factor corresponding to the current frame image based on the noise image data and the target image data corresponding to the i-frame image; wherein, the target video impact factor is obtained by: calculating the image feature overlap degree between the noise image data and the target image data; calculating the average value of the image feature overlap degree corresponding to each frame image to obtain the target video impact factor corresponding to the current frame image; The noise level acquisition module is used to determine the noise coarseness level of the current frame image based on the target noise coarseness factor and the target video influence factor.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 6.
9. A storage medium, said storage medium being a computer-readable storage medium, said computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image processing method and device and computer readable storage medium
CN110796614A
Video real-time noise reduction method and device based on motion estimation and noise estimation
CN116437024A