Audio recovery methods, devices, equipment and media

By selecting reference frames in visual microphone technology to calculate pixel-level sharpness and statistically analyzing block-average sharpness, a block-level binary mask and spatial weight map are generated. Multi-scale and multi-directional splitting and weighted aggregation are then performed, solving the problem of difficult identification of high-confidence regions under the global shutter system and improving the quality and stability of audio restoration.

CN122493867APending Publication Date: 2026-07-31BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING INFORMATION SCI & TECH UNIV
Filing Date
2026-04-27
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing visual microphone technology struggles to effectively identify high-confidence areas under a global shutter mechanism, resulting in significant noise interference and poor audio quality and stability.

Method used

By selecting reference frames from the video sequence of the target object, calculating pixel-level sharpness and calculating the block mean sharpness, generating a block-level binary mask, constructing a spatial weight map, and performing multi-scale and multi-directional splitting and weighted aggregation, noise interference in low-reliability areas is suppressed.

Benefits of technology

It significantly improves the signal purity and speech intelligibility of the recovered audio, achieves active suppression of interference in unreliable regions in the phase domain, and improves the quality and stability of audio recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493867A_ABST
    Figure CN122493867A_ABST
Patent Text Reader

Abstract

This application relates to an audio restoration method, apparatus, device, and medium, comprising: calculating the pixel-level sharpness of a reference frame image based on each pixel of the reference frame; dividing the reference frame image into multiple local blocks; calculating the block mean sharpness based on the pixel set corresponding to each local block; generating a corresponding block-level binary mask based on the mean sharpness value of all local blocks of the reference frame and a preset candidate threshold; constructing a spatial weight map based on the block-level binary mask and the square of the pixel-level sharpness; splitting the reference frame and the current frame to generate multiple sub-bands; calculating the local phase change values ​​of the corresponding sub-bands of the reference frame and the current frame; weighting and aggregating the local phase change values ​​according to the spatial weights to generate the time response of each sub-band; and fusing the time responses of each sub-band to generate restored audio. This solves the problems in related technologies, such as the difficulty in effectively identifying high-reliability regions and the large noise interference in low-reliability regions under the global shutter system, leading to poor audio quality and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to an audio recovery method, apparatus, device, and medium. Background Technology

[0002] Visual microphone technology is a non-contact acoustic sensing method that retrieves audio signals by analyzing the minute deformations of an object's surface caused by sound waves. This technology requires no dedicated microphone and can achieve remote sound reconstruction via video imaging equipment, showing broad application prospects in fields such as security monitoring, industrial inspection, and voice enhancement.

[0003] In related technologies, visual microphone methods typically recover audio signals by analyzing the minute vibrations of a target object in a video caused by sound wave excitation. However, for high-speed imaging scenarios with global shutter speeds, the reliability of phase signals extracted from different spatial locations within a video frame varies significantly due to factors such as differences in object surface texture, depth of field variations, edge blurring, and reflections. Directly performing global averaging or weighting based solely on amplitude can easily introduce noise and spurious signals from low-reliability areas into the recovery result, leading to decreased audio quality and impacting speech intelligibility. Summary of the Invention

[0004] This application provides an audio recovery method, apparatus, device, and medium to solve the problems in related technologies, such as the difficulty in effectively identifying high-reliability areas and the large noise interference in low-reliability areas under the global shutter system, resulting in poor audio quality and stability.

[0005] The first aspect of this application provides an audio restoration method, comprising the following steps: selecting a reference frame from a video sequence of a target object; calculating the pixel-level sharpness of the reference frame image based on each pixel of the reference frame, dividing the reference frame image into multiple local blocks, and calculating the block mean sharpness based on the pixel set corresponding to each local block; generating a corresponding block-level binary mask based on the mean sharpness value of all local blocks of the reference frame and a preset candidate threshold, and constructing a spatial weight map based on the block-level binary mask and the square of the pixel-level sharpness of the reference frame; splitting the reference frame and the current frame into multiple sub-bands in multiple scales and directions, calculating the local phase change values ​​of the corresponding sub-bands of the reference frame and the current frame, mapping the spatial weight map to the spatial weights of the corresponding sub-bands, weighting and aggregating the local phase change values ​​to generate the time response of each sub-band, and fusing the time responses of each sub-band to generate restored audio.

[0006] Optionally, after generating the restored audio by fusing the time responses of each sub-band, the method further includes: identifying candidate restored audio generated based on all preset candidate thresholds; calculating the evaluation score of the candidate restored audio; and selecting the candidate restored audio with the highest evaluation score to generate the final restored audio.

[0007] Optionally, calculating the pixel-level sharpness of the reference frame image based on each pixel of the reference frame includes: obtaining the horizontal gradient value and the vertical gradient value of each pixel of the reference frame; and calculating the pixel-level sharpness of the reference frame image based on the horizontal gradient value and the vertical gradient value.

[0008] Optionally, the formula for calculating the block mean sharpness based on the pixel set corresponding to each local block is as follows: ; in, These are the column and row coordinates of the pixel, respectively. Let p be the block mean sharpness of the p-th local block. pixel coordinates Pixel-level sharpness at that location Represents the pixel position within the p-th local block. Perform a traversal. Let be the set of coordinates of all pixels covered by the p-th local block.

[0009] Optionally, generating the corresponding block-level binary mask based on the mean sharpness value of all local blocks in the reference frame and a preset candidate threshold includes: if the mean sharpness value of a local block is greater than or equal to the preset candidate threshold, then the block-level binary mask value of the local block is a first target value; if the mean sharpness value of a local block is less than the preset candidate threshold, then the block-level binary mask value of the local block is a second target value, wherein the second target value is less than the first target value.

[0010] Optionally, calculating the local phase change value of the corresponding sub-band of the reference frame and the current frame includes: extracting the phase values ​​of the reference frame and the current frame at the target pixel position in the target scale and target direction; and calculating the local phase change value based on the difference between the phase values ​​of the current frame and the reference frame at the target pixel position.

[0011] Optionally, based on the spatial weights mapped to the corresponding sub-bands by the spatial weight map, the local phase change values ​​are weighted and aggregated to generate the calculation formula for the time response of each sub-band: ; in, Let the scale be r, and the direction be r. The sub-band's time response at time t, where t is the time of the current frame. These are the spatial coordinates of pixels within the sub-band. The weight values ​​are the sub-band weights after the spatial weight map is mapped. For scale r, direction In the sub-band, pixels At time t relative to the reference frame The local phase change value.

[0012] A second aspect of this application provides an audio restoration apparatus, comprising: a selection module for selecting a reference frame from a global shutter video sequence of a target object; a calculation module for calculating the pixel-level sharpness of the reference frame image based on each pixel of the reference frame, dividing the reference frame image into multiple local blocks, and calculating the block mean sharpness based on the pixel set corresponding to each local block; a generation module for generating a corresponding block-level binary mask based on the mean sharpness value of all local blocks of the reference frame and a preset candidate threshold, and constructing a spatial weight map based on the block-level binary mask and the square of the pixel-level sharpness of the reference frame; and a splitting module for splitting the reference frame and the current frame into multiple sub-bands at multiple scales and in multiple directions, calculating the local phase change values ​​of the corresponding sub-bands of the reference frame and the current frame, mapping the spatial weight map to the spatial weights of the corresponding sub-bands, weighting and aggregating the local phase change values ​​to generate the time response of each sub-band, and fusing the time responses of each sub-band to generate restored audio.

[0013] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to perform the audio recovery method as described in the above embodiments.

[0014] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to perform the audio recovery method as described in the above embodiments.

[0015] A fifth aspect of this application provides a computer program product, including a computer program or instructions, which, when executed, implement the audio recovery method as described in the above embodiments.

[0016] Therefore, this application has at least the following beneficial effects: This application embodiment calculates pixel-level sharpness from a reference frame and obtains block-average sharpness through block-by-block statistics. Then, a block-level binary mask is generated using a preset candidate threshold to forcibly remove low-reliability regions such as low-texture, defocused, or reflective areas. Within the retained regions, the square of the pixel-level sharpness is used as a continuous spatial weight, thereby constructing a spatial weight map that can distinguish the reliability of different regions. Based on this, multi-scale, multi-directional complex decomposition is performed on the reference frame and the current frame to extract local phase changes, and the spatial weights are mapped to each sub-band for weighted aggregation. This ensures that phase changes in high-sharpness, high-texture regions dominate the aggregation, while noise contributions from low-reliability regions are effectively suppressed. Finally, the recovered audio is generated by fusing the time responses of each sub-band, achieving the technical effect of actively suppressing interference from unreliable regions in the phase domain, significantly improving the signal purity and speech intelligibility of the recovered audio.

[0017] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0018] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of an audio recovery method provided according to an embodiment of this application; Figure 2 This is a schematic diagram illustrating the reference frame block-level sharpness distribution, binary mask, and spatial weight construction according to embodiments of this application. Figure 3 This is a schematic diagram of a single recovery process for local phase change extraction and spatial weighted aggregation according to an embodiment of this application; Figure 4 This is a flowchart of the adaptive spatially weighted audio recovery process provided according to an embodiment of this application; Figure 5 This is a block diagram of an audio restoration device provided according to an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation

[0019] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0020] Existing visual microphone technology is based on methods for detecting minute motions in video and extracting phase-based vibrations. Early methods often employed optical flow estimation, feature point tracking, or Euler amplification to detect minute motions in video. Subsequently, phase-based video motion analysis methods based on multi-scale, multi-directional decomposition were further developed, characterizing sub-pixel displacement through local phase changes and improving the ability to depict subtle vibrations. Building on this foundation, a visual microphone is proposed, and it has been systematically demonstrated that speech and music can be recovered from minute vibrations generated on the surface of a target in high-speed video by acoustic excitation.

[0021] Building upon visual microphones, subsequent research has primarily focused on local region selection, local signal aggregation, input video quality optimization, and imaging method expansion. One approach, exemplified by local visual microphones, divides the image into multiple local blocks, recovers local sound from each block, and then filters and combines them based on local quality to reduce destructive interference caused by inconsistent vibration characteristics in different regions. Other studies have attempted to reduce the computational complexity of visual microphones through methods such as singular value decomposition. Another approach addresses input video quality, improving recovery conditions through frame-by-frame denoising, color space conversion, or visual enhancement. Further research has focused on improving visual microphone systems based on imaging structures such as rolling shutter and single-pixel imaging. Regarding spatial region selection and weighting design, a focal length index has been proposed, identifying focal length based on the image's sharpness along the rolling shutter direction. This is further enhanced by eliminating out-of-focus areas, weighting phase changes in the focused area, or combining both methods to improve recovered sound quality.

[0022] Existing visual microphone methods still have shortcomings in audio restoration under high-speed imaging conditions. While some visual microphone methods have demonstrated the feasibility of video-based sound restoration, their restoration process relies heavily on global fusion, lacking pre-control over reliability differences at different spatial locations. Unreliable regions, once included in the unified aggregation, can easily affect the final restoration result. Local visual microphones, although mitigating the destructive interference of simple averaging through local block selection, still rely on first restoring local sound and then selecting and combining based on posterior quality, failing to directly constrain low-reliability regions in the front-end phase retrieval stage. Preprocessing methods primarily focus on improving input image quality, but existing research indicates that such operations may weaken or even distort subtle phase changes, thus affecting subsequent sound restoration. Furthermore, their effectiveness is closely related to specific scenes and evaluation metrics, making it difficult to establish a stable and universal preprocessing scheme. While methods based on rolling shutter focusing areas or focal lengths demonstrate the help of prioritizing sharp regions, these methods depend on rolling shutter imaging conditions, and the region organization and weighting methods employed are inconsistent with the more common two-dimensional local block differences under global shutter conditions. Therefore, existing technologies still lack an audio restoration method that is oriented towards a global shutter visual microphone, can identify high-confidence regions based on the local sharpness distribution of the reference frame, and adaptively constructs a spatial weighting strategy to directly act on the phase aggregation process.

[0023] This application aims to solve the problems of difficulty in effectively identifying high-reliability regions and lack of adaptive adjustment capability in spatial weighting under the global shutter system, thereby improving the quality and stability of the recovered audio while suppressing background noise and unreliable signal interference.

[0024] The audio recovery method, apparatus, device, and medium of this application are described below with reference to the accompanying drawings.

[0025] Specifically, Figure 1 This is a schematic flowchart of an audio recovery method provided in an embodiment of this application.

[0026] like Figure 1 As shown, the audio recovery method includes the following steps: In step S101, a reference frame is selected from the global shutter video sequence of the target object. It is understood that, in the embodiments of this application, a reference frame can be selected from the global shutter video sequence of the target object in order to facilitate the subsequent calculation of the pixel-level sharpness of the reference frame image.

[0027] It should be noted that the global shutter video sequence of the target object is acquired by a global shutter high-speed camera. This camera can capture minute vibrations on the surface of the target object at a high frame rate (typically hundreds to thousands of frames per second), and all pixels are exposed simultaneously, avoiding the time shift caused by line-by-line exposure with a rolling shutter. This is suitable for subsequent analysis of minute motions based on phase changes. The reference frame is usually selected from the target object video sequence acquired by the global shutter camera. Specific methods for determining the reference frame include, but are not limited to, the following: using the first frame of the video sequence as the reference frame, and calculating the phase difference between all subsequent current frames and this reference frame; this method is simple to implement and suitable for scenarios where the camera and the target object are relatively stable; calculating the global sharpness of each frame image (such as the mean or median of all pixel-level sharpness), and selecting the frame with the highest sharpness as the reference frame. High sharpness means clear texture and rich edges, providing a more reliable phase reference, which is beneficial for the subsequent construction of spatial weights.

[0028] In step S102, the pixel-level sharpness of the reference frame image is calculated based on each pixel of the reference frame, the reference frame image is divided into multiple local blocks, and the block mean sharpness is calculated based on the pixel set corresponding to each local block. It is understood that the embodiments of this application can calculate the pixel-level sharpness of each pixel in the reference frame, which can quantify the local edge strength and texture clarity at various locations on the target surface, providing a fine-grained measure for spatial reliability. However, single-pixel sharpness is easily affected by noise, isolated bright spots, or fragmented structures, and its direct use is not stable enough. To address this, the reference frame is divided into multiple local blocks and the mean sharpness of each block is calculated. By summing and averaging the sharpness of all pixels within a block, random noise and local outliers are effectively suppressed, resulting in a more robust region-level sharpness representation. The block mean sharpness can stably reflect the overall texture quality and focus of different local regions, providing a reliable and interference-resistant statistical basis for subsequent threshold-based screening of high-reliability regions. This ensures that only regions with overall clarity and rich texture are retained for audio restoration, thereby fundamentally avoiding noise contamination in low-reliability regions.

[0029] It should be noted that for each pixel in the reference frame, the gradient value is calculated using the Sobel operator in the horizontal and vertical directions respectively, and then the square and square root are taken to obtain the pixel-level sharpness. The sharpness value reflects the degree of brightness change in the local neighborhood of the pixel, that is, the edge strength or texture clarity. The richer the texture and the sharper the edge of the area, the higher the sharpness value. The smoother, defocused or reflective areas have lower sharpness values. Since the small vibrations caused by sound waves are reflected as local phase changes in the image, and the phase change signal-to-noise ratio of high texture areas is higher, pixel-level sharpness can be used as a fine-grained measure of spatial reliability.

[0030] Because single-pixel sharpness is easily affected by noise, isolated bright spots, or minor texture fluctuations, directly using it for decision-making is not stable enough. Therefore, the reference frame is divided into multiple local blocks (e.g., 16×16 pixels). For each block, the sharpness values ​​of all pixels within it are summed and divided by the total number of pixels to obtain the block-average sharpness. This block-level statistical process is equivalent to performing spatial low-pass filtering on pixel-level sharpness, effectively suppressing random noise and local outliers, and obtaining a more robust region-level sharpness representation. The block-average sharpness stably reflects the overall texture quality and focus of each local region.

[0031] In this embodiment of the application, the pixel-level sharpness of the reference frame image is calculated based on each pixel of the reference frame, including: obtaining the horizontal gradient value and the vertical gradient value of each pixel of the reference frame; and calculating the pixel-level sharpness of the reference frame image based on the horizontal gradient value and the vertical gradient value. It is understood that the embodiments of this application can obtain the horizontal and vertical gradient values ​​of each pixel in the reference frame and calculate the pixel-level sharpness based on the two. This can accurately quantify the edge strength and texture clarity of each local location on the target surface, providing a fine-grained, high-resolution reliable prior for subsequent region selection and spatial weighting. This ensures that high-texture regions receive greater weight in phase aggregation, thereby effectively improving the purity of the recovered audio.

[0032] It should be noted that the horizontal and vertical gradients reflect the rate of change of brightness in the image along the horizontal and vertical directions, respectively. The square root of the sum of the squares of the two gradients yields the gradient magnitude. This magnitude is significantly higher in areas with rich edges or fine textures, and lower in smooth, defocused, or reflective areas. This makes pixel-level sharpness a sensitive indicator of spatial reliability; high-sharpness pixels correspond to higher phase signal-to-noise ratios, while low-sharpness pixels are more susceptible to noise interference.

[0033] Specifically, such as Figure 2 and Figure 4As shown, a reference frame is used as the baseline image for spatial reliability analysis. Local sharpness is employed to describe the clarity and texture intensity at different spatial locations on the target surface, and this forms the basis for subsequent region selection and spatial weighting. Unlike methods that directly aggregate the entire image or perform posterior selection based on local audio quality after restoration, this approach first establishes a spatial reliability distribution on the reference frame and then introduces it into the phase retrieval process.

[0034] The local sharpness of the reference frame is represented using Tenengrad sharpness based on the Sobel gradient. Let the reference frame image be... Its gradient responses in the horizontal and vertical directions are as follows:

[0035]

[0036] in, This refers to the grayscale value (or brightness value) of the reference frame image at pixel coordinates. Column coordinates (horizontal direction) The coordinates are row coordinates (vertical direction). and These are the Sobel operators for the horizontal and vertical directions, respectively. This is the convolution operator. For reference frame in pixels The horizontal gradient value at that point, The vertical gradient value at pixel (x,y) of the reference frame; in,

[0037]

[0038] This gives the position of the reference frame. Pixel-level sharpness value at:

[0039] in, For reference frame in pixels The horizontal gradient value at that point, Let be the vertical gradient value at pixel (x, y) of the reference frame. For pixels The pixel-level sharpness value at that location.

[0040] In this embodiment of the application, the formula for calculating the block mean sharpness based on the pixel set corresponding to each local block is as follows: ; in, These are the column and row coordinates of the pixel, respectively. Let p be the block mean sharpness of the p-th local block. pixel coordinates Pixel-level sharpness at that location Represents the pixel position within the p-th local block. Perform a traversal. Let be the set of coordinates of all pixels covered by the p-th local block.

[0041] It is understood that the embodiments of this application transform the spatial distribution of single-pixel sharpness into the statistical characteristics of local blocks, effectively suppressing pixel-level fluctuations caused by noise, isolated bright spots, or fragmented structures. This allows the block mean sharpness to stably and reliably reflect the overall texture richness and focus clarity of the region, avoiding misjudgment of the reliability of the entire region due to individual pixel anomalies. It ensures that only high-confidence regions with overall clarity and uniform texture are preserved, thereby improving the accuracy and robustness of spatial weighted aggregation during audio restoration.

[0042] It should be noted that, as Figure 2 As shown, pixel-level sharpness can reflect the intensity of local edges and the clarity of textures, but the response of a single pixel is easily affected by noise, isolated bright spots, and local fragmented structures, so it is not directly used as a basis for judging the reliability of a region. To obtain a more stable spatial representation, block-level statistics are further performed on the sharpness map, and the block mean sharpness is used to characterize the overall clarity of local areas.

[0043] Let the first The set of pixels corresponding to each local block is Then its block mean sharpness is defined as: ; Among them, among them, These are the column and row coordinates of the pixel, respectively. Let p be the block mean sharpness of the p-th local block. pixel coordinates Pixel-level sharpness at that location Represents the pixel position within the p-th local block. Perform a traversal. Let be the set of coordinates of all pixels covered by the p-th local block.

[0044] While pixel-level sharpness can accurately reflect the local texture clarity of each pixel, single-pixel values ​​are easily affected by random noise, isolated bright spots, or extremely subtle texture fluctuations, leading to instability when directly used for region reliability determination. By using block-level summation and averaging, a spatial low-pass filter is applied to the pixel-level sharpness map, effectively suppressing high-frequency random noise and local outliers, resulting in more robust and interference-resistant region-level statistical characteristics. The value of the block-mean sharpness directly characterizes the overall clarity and texture richness of the p-th local region: a larger value indicates clearer overall texture and sharper edges, with a higher corresponding phase signal-to-noise ratio; a smaller value indicates a smooth, defocused region or the presence of reflection interference, with noise dominating the phase signal.

[0045] In step S103, a corresponding block-level binary mask is generated based on the mean sharpness value of all local blocks in the reference frame and a preset candidate threshold. The corresponding spatial weights are determined based on the block-level binary mask, and a spatial weight map is constructed based on the block-level binary mask and the squared pixel-level sharpness of the reference frame. The preset candidate threshold can be set according to actual needs without specific limitations.

[0046] It is understood that the embodiments of this application can generate a block-level binary mask by comparing the mean sharpness of each local block with a preset candidate threshold. If the block mean sharpness is greater than or equal to the threshold, it is marked as 1; otherwise, it is marked as 0. The mask hard-kills low-reliability areas such as low texture, defocus, or reflection, and only retains high-resolution, high-texture, high-reliability areas. Therefore, for the retained blocks with a mask value of 1, the square of the pixel-level sharpness of each pixel in the block is further used as the continuous spatial weight, while the weight of the area with a mask value of 0 is set to zero. This generates a spatial weight map. The progressive process from block-level binary screening to pixel-level continuous weighting not only suppresses the interference of noise and local outliers through block-level statistics, but also amplifies the contribution of high-texture pixels in the retained area by using the square of sharpness. This enables the spatial weight map to accurately distinguish the reliability differences of different spatial locations, providing active constraints at the front end for subsequent phase aggregation, and fundamentally preventing noise from low-reliability areas from mixing into the recovery results.

[0047] It should be noted that by using block-level statistics and threshold screening, unreliable parts are first eliminated at the regional level, avoiding misjudgments based on single-pixel noise. Then, within the retained regions, continuous weighting is achieved through the square of sharpness, allowing the clearest and most reliable pixels to play a dominant role in subsequent phase aggregation. The combination of these two methods ensures both the robustness of the screening and the preservation of fine-grained differentiated contribution capabilities.

[0048] like Figure 2As shown, the first sub-image is the reference frame, the second sub-image is the corresponding block-level sharpness distribution map, reflecting the differences in local sharpness and texture intensity at different locations on the target surface; the third sub-image is the block-level binary mask obtained based on the threshold, where the white area represents the high-confidence area participating in the subsequent phase calculation, and the black area represents the low-confidence area that is removed; the fourth sub-image is the final spatial weight map, showing the spatial distribution result after further continuous weighting based on the block-level removal.

[0049] In this embodiment of the application, generating a corresponding block-level binary mask based on the mean sharpness value of all local blocks in the reference frame and a preset candidate threshold includes: if the mean sharpness value of a local block is greater than or equal to the preset candidate threshold, then the block-level binary mask value of the local block is a first target value; if the mean sharpness value of a local block is less than the preset candidate threshold, then the block-level binary mask value of the local block is a second target value, wherein the second target value is less than the first target value. The first target value can be 1, and the second target value can be 0, without any specific restrictions.

[0050] It is understood that the embodiments of this application can generate a block-level binary mask by comparing the mean sharpness value of each local block with a preset candidate threshold. If the mean sharpness is greater than or equal to the threshold, a first target value (e.g., 1) is assigned; otherwise, a smaller second target value (e.g., 0) is assigned. Based on the principle that the mean sharpness of the block can stably represent the overall sharpness of the local area, the threshold is used to classify low-reliability areas such as low texture, defocus, or reflection into high-reliability areas with high clarity and high texture. The higher the mean sharpness of the block, the clearer the area is, and the higher the corresponding phase signal signal-to-noise ratio is. Therefore, a higher mask value is assigned to retain it; otherwise, it is discarded. This ensures that only areas with overall quality meet the standard and enter the subsequent phase aggregation, fundamentally preventing noise and pseudo-signal interference from low-reliability areas, and providing reliable regional constraints for the construction of the subsequent spatial weight map.

[0051] It should be noted that, as Figure 3 and Figure 4 As shown, after obtaining the block-level sharpness distribution, region filtering is performed based on the block mean, and continuous spatial weights are constructed within the retained regions. Multiple candidate threshold ratios are set to filter the block mean sharpness distribution, generating corresponding block-level hard masks and spatial weights. The resulting region filtering range and spatial weight distribution are no longer fixed, but dynamically adjusted according to changes in the local sharpness distribution of the reference frame and the candidate thresholds.

[0052] Let the sharpness threshold be... The block-level hard mask is then: ; in, The sharpness threshold. Let p be the block mean sharpness of the p-th local block. is the block-level hard mask value for the p-th local block.

[0053] When the block mean sharpness of a local block is lower than the threshold, the local block is identified as a low-reliability region and is removed in subsequent phase aggregation.

[0054] Building upon block-level filtering, a continuous spatial weighting structure is further constructed. Let the pixel position... The corresponding block-level mask value is Then the spatial weights are: ; in, pixel position Block-level mask value at the location, pixel coordinates Pixel-level sharpness at that location pixel position The final spatial weight at that location.

[0055] In this process, a block-level hard mask is used to control whether local regions participate in aggregation, and the squared sharpness is used to distinguish the contribution levels within the preserved regions. After this processing, low-reliability regions are first suppressed through block-level culling, while high-sharpness regions take on greater weight in subsequent phase retrieval.

[0056] In step S104, the reference frame and the current frame are split into multiple sub-bands in multiple scales and directions. The local phase change values ​​of the corresponding sub-bands of the reference frame and the current frame are calculated. The local phase change values ​​are weighted and aggregated according to the spatial weight map mapped to the spatial weight of the corresponding sub-band to generate the time response of each sub-band. The time responses of each sub-band are fused to generate the restored audio.

[0057] Specifically, based on the spatial weights mapped to the corresponding sub-bands from the spatial weight map, the local phase change values ​​are weighted and aggregated to generate the time response calculation formula for each sub-band: ; in, Let the scale be r, and the direction be r. The sub-band's time response at time t, where t is the time of the current frame. These are the spatial coordinates of pixels within the sub-band. The weight values ​​are the sub-band weights after the spatial weight map is mapped. For scale r, direction In the sub-band, pixels At time t relative to the reference frame The local phase change value.

[0058] It is understood that the embodiments of this application can divide the image into multiple sub-bands of different scales and directions by performing multi-scale, multi-directional complex controllable pyramid decomposition on the reference frame and the current frame. The complex response in each sub-band contains amplitude and phase information. The phase difference between the current frame and the reference frame at the same sub-band and the same pixel position is extracted and processed by winding to obtain the local phase change value. This value is proportional to the sub-pixel displacement generated by the target surface under the excitation of sound waves. After mapping the pre-constructed spatial weight map to the resolution of each sub-band, the local phase change value of all pixels in each sub-band is weighted and averaged. This makes the vibration signal of high texture and high sharpness regions dominate in the aggregation, while the noise contribution of low reliability regions is effectively suppressed, thereby generating a high signal-to-noise ratio time response for each sub-band. The time responses of all sub-bands are time-aligned and summed, and the vibration information of different frequency bands and directions is fused to recover the broadband audio signal. This realizes the end-to-end optimization from the spatial domain to the phase domain and then to the time domain, effectively improving the purity and speech intelligibility of the recovered audio.

[0059] It should be noted that, as Figure 3 As shown, complex controllable pyramid decomposition is performed on both the reference frame and the current frame. This is a multi-scale, multi-directional image transformation method. The decomposition results in multiple sub-bands, each corresponding to a specific spatial scale (from coarse to fine, reflecting different frequencies) and a specific direction (such as horizontal, vertical, diagonal, etc.). Each pixel location within each sub-band has a complex response, which consists of two parts: amplitude and phase. The amplitude represents the local energy intensity of the pixel location at the corresponding scale and direction, and the phase represents the local positional offset of the pixel location at the corresponding scale and direction.

[0060] The phase difference between the reference frame and the current frame at the same sub-band and the same pixel location is calculated and then coiled to obtain the local phase change value. This change value is proportional to the sub-pixel displacement of the target object's surface caused by acoustic excitation at that scale / direction. Since acoustic vibrations contain components of different frequencies and directions, multi-scale decomposition can capture low-frequency large-scale motion and high-frequency local jitter respectively, while multi-directional decomposition can capture components of different vibration directions.

[0061] A pre-constructed spatial weight map (based on the sharpness of the reference frame) is mapped to the resolution of each sub-band, resulting in a sub-band weight map that matches the sub-band size. For each sub-band, a weighted average of the local phase change values ​​of all pixels within the current frame is calculated: the numerator is the sum of the weight of each pixel multiplied by its phase change value, and the denominator is the sum of all weights. This weighted aggregation process ensures that the phase changes of high-sharpness, high-texture regions (pixels with large weights) dominate the aggregation result, while the phase changes of low-reliability regions (pixels with zero or very small weights) are effectively suppressed. This results in a high signal-to-noise ratio time response for the sub-band in the current frame. This time response is a scalar sequence that varies with the frame number, representing the overall vibration waveform captured by the sub-band.

[0062] Because filters of different scales and orientations have different group delays, the same physical vibration will exhibit slight time shifts in the time responses of different sub-bands. Therefore, it is necessary to first align the time responses of each sub-band. After alignment, the time responses of all sub-bands are directly added together to fuse vibration information from different frequencies and orientations, resulting in a broadband recovered audio signal. This signal is then post-processed, including drift suppression (high-pass filtering to remove extremely low-frequency components), initial segment smoothing, and normalization, to output the final audible audio.

[0063] In summary, as Figure 3 As shown, the input image sequence selects reference frames and constructs a local sharpness distribution. A block-level hard mask is obtained through block-level statistics and thresholding, and then combined with pixel-level sharpness to form a spatial weight map. On the other hand, the input sequence undergoes complex controllable pyramid decomposition to extract the local phase changes of each sub-band. Subsequently, the spatial weights are embedded into the phase aggregation process to perform weighted aggregation of the phase changes of each sub-band, and then temporal alignment and multi-scale fusion are performed to obtain the restored audio.

[0064] In this embodiment of the application, calculating the local phase change value of the corresponding sub-band of the reference frame and the current frame includes: extracting the phase values ​​of the reference frame and the current frame at the target pixel position in the target scale and target direction; and calculating the local phase change value based on the difference between the phase values ​​of the current frame and the reference frame at the target pixel position.

[0065] It is understood that, in the embodiments of this application, the original phase difference can be obtained by extracting the phase values ​​of the reference frame and the current frame at the same pixel position in the target scale and target direction, and calculating the difference between the two. This difference is proportional to the sub-pixel displacement generated by the target surface under acoustic excitation. It directly utilizes the linear relationship between phase and displacement after complex controllable pyramid decomposition. Without optical flow estimation or feature point tracking, it can accurately capture the phase change caused by small vibrations from the video frame. By calculating the phase difference value sub-band and pixel by pixel, multi-band and multi-directional vibration observation signals are obtained, suppressing the interference of non-real vibration components on the subsequent audio restoration process, thereby improving the accuracy and noise resistance of the restored signal.

[0066] Specifically, setting a scale ,direction The complex response is as follows: ; in, This serves as a spatial scale index, representing the scale hierarchy of the complex controllable pyramid decomposition. The direction index indicates the direction of the sub-band. Here, represents the spatial coordinates of the pixels within the sub-band, and t represents the time index. The amplitude of the complex response, The phase of the complex response.

[0067] Current frame Relative to reference frame The local phase change is ; Where t is the time of the current frame. The time of the reference frame The reference phase is the same sub-band and the same pixel position of the reference frame, and wrap is the phase wrapping function. This represents the local phase change value after winding.

[0068] Mapping spatial weights to the corresponding sub-band resolutions at the specified scale and orientation, denoted as The local phase changes are then weighted and aggregated to obtain the time response at this scale and in this direction: ; in, Let the scale be r, and the direction be r. The sub-band's time response at time t, where t is the time of the current frame. These are the spatial coordinates of pixels within the sub-band. The weight values ​​are the sub-band weights after the spatial weight map is mapped. For scale r, direction In the sub-band, pixels At time t relative to the reference frame The local phase change value.

[0069] Local phase changes are used to characterize the minute vibration response of the target surface to acoustic excitation, while spatial weights are used to control the contribution of different regions to the polymerization results.

[0070] In this embodiment of the application, after fusing the time responses of each sub-band to generate the restored audio, the method further includes: identifying candidate restored audio generated based on all preset candidate thresholds; calculating the evaluation score of the candidate restored audio; and selecting the candidate restored audio with the highest evaluation score to generate the final restored audio.

[0071] It is understood that the embodiments of this application can obtain a set of candidate audio recordings by performing a complete spatial weight construction, phase extraction, weighted aggregation and fusion process on each preset candidate threshold. Each candidate audio recording corresponds to a different region screening intensity. An objective score is calculated for each candidate audio recording. This score is based on a reference signal or a no-reference quality metric and can quantify the purity, intelligibility and noise suppression level of the audio. The candidate audio recording with the highest evaluation score is selected as the final recovery output. This achieves data-driven adaptive threshold optimization, avoiding the subjectivity and scene limitations caused by manually fixing thresholds. It can automatically select the most suitable spatial weighting configuration according to the actual imaging conditions such as the texture distribution, defocusing degree and reflection interference of the target object, so that the noise in the low reliability region is suppressed to the greatest extent, while retaining sufficient high reliability signal energy, significantly improving the overall quality and robustness of the recovered audio recording.

[0072] Specifically, such as Figure 3 and Figure 4 As shown, after obtaining the time response at various scales and in various directions, it is temporally aligned and fused to obtain the recovered signal:

[0073] in, To complete the timing-aligned subband response, For the finally recovered audio signal, For spatial scale indexing. t is the direction index, and t is the time index.

[0074] After further processing the recovered signal by drift suppression, initial segment smoothing and normalization, the final recovered audio is obtained.

[0075] To adapt to different materials, texture conditions, and imaging scenarios, a candidate threshold set is used. The corresponding block-level hard masks and spatial weights are generated respectively, and multiple candidate results are obtained under the unified recovery framework. Let the evaluation function be... The optimal threshold is:

[0076] Where Λ is the set of candidate thresholds, For candidate thresholds, For the evaluation function, This means selecting the threshold from the candidate threshold set Λ that maximizes the evaluation function. , The optimal threshold is the threshold selected through the above maximization operation.

[0077] After determining the optimal threshold, the corresponding spatial weights are selected to participate in the final recovery output. In this way, the region selection range, spatial weight distribution, and phase aggregation results are all adaptively adjusted according to the current data, rather than relying on a single fixed parameter.

[0078] This application calculates the local sharpness of the reference frame and extracts the local phase change of the input sequence after inputting the image sequence. Then, it generates corresponding spatial weights under the current candidate threshold and maps these weights to the resolution of each sub-band. It performs spatial weighted aggregation of the local phase changes and obtains candidate audio through multi-scale fusion. The candidate audio is then evaluated, and it is determined whether the candidate threshold has been traversed. If not, it continues to generate spatial weights under the next candidate threshold and repeats the recovery process. If the traversal is complete, the optimal recovered audio is selected and output based on the evaluation results.

[0079] In summary, this application addresses the issue of noise interference easily introduced by defocused, low-texture, and reflective areas under global shutter conditions. It introduces local sharpness as a spatial reliability characterization into the audio restoration process and combines block-level filtering and continuous spatial weighting to constrain the phase contribution of different regions. Through this technical solution, the influence of low-reliability regions on phase convergence is suppressed, while the contribution of local phase changes in high-reliability regions is highlighted, thereby improving the quality, purity, and stability of the restored audio.

[0080] Meanwhile, in response to the differences in response, mechanical delay, and vibration consistency at different spatial locations on the target surface under acoustic excitation, this application sets a candidate threshold set and combines it with an evaluation function to filter the results. This allows the region selection range and spatial weight distribution to be adjusted according to the current data characteristics, thereby reducing the misalignment and aliasing caused by asynchronous vibration regions and improving the envelope clarity and overall robustness of the recovery results.

[0081] According to the audio restoration method proposed in this application, a reference frame is selected from the video sequence and pixel-level sharpness is calculated. Then, the block mean sharpness is obtained through block-by-block statistics. A block-level binary mask is generated using a preset candidate threshold to remove low-reliability areas such as low texture, defocus, or reflection. Within the retained area, the square of the pixel-level sharpness is used as a continuous spatial weight to construct a spatial weight map that can distinguish the reliability of different areas. The reference frame and the current frame are subjected to multi-scale and multi-directional complex decomposition to extract local phase changes. The spatial weights are mapped to each sub-band for weighted aggregation, so that the phase changes of high-sharpness and high-texture areas dominate in the aggregation, and the noise contribution of low-reliability areas is effectively suppressed. Finally, the restored audio is generated by fusing the time response of each sub-band. Active spatial filtering is implemented in the phase domain. Local sharpness is used to characterize the reliability of the area, and weighted aggregation is used to suppress the interference of low-reliability areas on the restoration result, thereby improving the signal purity and speech intelligibility of the restored audio.

[0082] Next, the audio recovery device proposed according to the embodiments of this application is described with reference to the accompanying drawings.

[0083] Figure 5 This is a block diagram of an audio restoration device according to an embodiment of this application.

[0084] like Figure 5 As shown, the audio restoration device 10 includes: a selection module 100, a calculation module 200, a generation module 300, and a splitting module 400.

[0085] The selection module 100 is used to select a reference frame from the video sequence of the target object; the calculation module 200 is used to calculate the pixel-level sharpness of the reference frame image based on each pixel of the reference frame, divide the reference frame image into multiple local blocks, and calculate the block mean sharpness based on the pixel set corresponding to each local block; the generation module 300 is used to generate a corresponding block-level binary mask based on the mean sharpness value of all local blocks of the reference frame and a preset candidate threshold, determine the corresponding spatial weight based on the block-level binary mask, and construct a spatial weight map based on the block-level binary mask and the square of the pixel-level sharpness of the reference frame; the splitting module 400 is used to split the reference frame and the current frame into multiple sub-bands in multiple scales and directions, calculate the local phase change value of the corresponding sub-band of the reference frame and the current frame, map the spatial weight map to the spatial weight of the corresponding sub-band, perform weighted aggregation of the local phase change value, generate the time response of each sub-band, and fuse the time responses of each sub-band to generate restored audio.

[0086] According to the audio restoration apparatus proposed in this application, a reference frame is selected from a video sequence and pixel-level sharpness is calculated. Then, the block mean sharpness is obtained through block-by-block statistics. A block-level binary mask is generated using a preset candidate threshold to remove low-reliability areas such as low texture, defocus, or reflection. Within the retained areas, the square of the pixel-level sharpness is used as a continuous spatial weight to construct a spatial weight map that can distinguish the reliability of different areas. The reference frame and the current frame are subjected to multi-scale and multi-directional complex decomposition to extract local phase changes. The spatial weights are mapped to each sub-band for weighted aggregation, so that the phase changes of high-sharpness and high-texture areas dominate in the aggregation, and the noise contribution of low-reliability areas is effectively suppressed. Finally, the restored audio is generated by fusing the time responses of each sub-band. Active spatial filtering is implemented in the phase domain, local sharpness is used to characterize the reliability of the area, and weighted aggregation is used to suppress the interference of low-reliability areas on the restoration result, thereby improving the signal purity and speech intelligibility of the restored audio.

[0087] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.

[0088] When the processor 602 executes the program, it implements the audio recovery method provided in the above embodiments.

[0089] Furthermore, electronic devices also include: Communication interface 603 is used for communication between memory 601 and processor 602.

[0090] The memory 601 is used to store computer programs that can run on the processor 602.

[0091] The memory 601 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0092] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0093] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.

[0094] The processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0095] This application also provides a computer-readable storage medium storing a computer program or instructions thereon, which, when executed by a processor, implements the above-described audio recovery method.

[0096] This application also provides a computer program product, including a computer program or instructions, which, when executed, implement the above-described audio recovery method.

[0097] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0098] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0099] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0100] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0101] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

Claims

1. An audio recovery method, characterized in that, Includes the following steps: Select a reference frame from the global shutter video sequence of the target object; The pixel-level sharpness of the reference frame image is calculated based on each pixel of the reference frame. The reference frame image is divided into multiple local blocks, and the block mean sharpness is calculated based on the pixel set corresponding to each local block. A corresponding block-level binary mask is generated based on the mean sharpness value of all local blocks in the reference frame and a preset candidate threshold, and a spatial weight map is constructed based on the block-level binary mask and the squared pixel-level sharpness of the reference frame. The reference frame and the current frame are split into multiple sub-bands in multiple scales and directions. The local phase change values ​​of the corresponding sub-bands of the reference frame and the current frame are calculated. The local phase change values ​​are weighted and aggregated according to the spatial weight map mapped to the spatial weight of the corresponding sub-band to generate the time response of each sub-band. The time responses of each sub-band are fused to generate the restored audio.

2. The audio recovery method according to claim 1, characterized in that, After generating the restored audio by fusing the time responses of each sub-band, the method further includes: Identify candidate recovered audio based on all preset candidate thresholds; Calculate the evaluation score of the candidate recovered audio; The candidate audio with the highest evaluation score is selected to generate the final restored audio.

3. The audio recovery method according to claim 1, characterized in that, The step of calculating the pixel-level sharpness of the reference frame image based on each pixel of the reference frame includes: Obtain the horizontal and vertical gradient values ​​for each pixel in the reference frame; The pixel-level sharpness of the reference frame image is calculated based on the horizontal gradient value and the vertical gradient value.

4. The audio recovery method according to claim 1, characterized in that, The formula for calculating the block mean sharpness based on the pixel set corresponding to each local block is as follows: ; in, These are the column and row coordinates of the pixel, respectively. Let p be the block mean sharpness of the p-th local block. pixel coordinates Pixel-level sharpness at that location Represents the pixel position within the p-th local block. Perform a traversal. Let be the set of coordinates of all pixels covered by the p-th local block.

5. The audio recovery method according to claim 1, characterized in that, The step of generating a corresponding block-level binary mask based on the mean sharpness value of all local blocks in the reference frame and a preset candidate threshold includes: If the mean sharpness value of a local block is greater than or equal to the preset candidate threshold, then the block-level binary mask value of the local block is the first target value. If the mean sharpness value of a local block is less than a preset candidate threshold, then the block-level binary mask value of the local block is the second target value, wherein the second target value is less than the first target value.

6. The audio recovery method according to claim 1, characterized in that, The calculation of the local phase change values ​​of the corresponding sub-bands of the reference frame and the current frame includes: At the target scale and in the target orientation, extract the phase values ​​of the reference frame and the current frame at the target pixel location; The local phase change value is calculated based on the difference between the phase values ​​of the current frame and the reference frame at the target pixel position.

7. The audio recovery method according to claim 1, characterized in that, Based on the spatial weights mapped to the corresponding sub-bands from the spatial weight map, the local phase change values ​​are weighted and aggregated to generate the calculation formula for the time response of each sub-band: ; in, Let the scale be r, and the direction be r. The sub-band's time response at time t, where t is the time of the current frame. These are the spatial coordinates of pixels within the sub-band. The weight values ​​are the sub-band weights after the spatial weight map is mapped. For scale r, direction In the sub-band, pixels At time t relative to the reference frame The local phase change value.

8. An audio restoration device, characterized in that, include: The selection module is used to select a reference frame from the global shutter video sequence of the target object; The calculation module is used to calculate the pixel-level sharpness of the reference frame image based on each pixel of the reference frame, divide the reference frame image into multiple local blocks, and calculate the block mean sharpness based on the pixel set corresponding to each local block. The generation module is used to generate a corresponding block-level binary mask based on the mean sharpness value of all local blocks in the reference frame and a preset candidate threshold, and to construct a spatial weight map based on the block-level binary mask and the squared pixel-level sharpness of the reference frame. The splitting module is used to split the reference frame and the current frame in multiple scales and directions to generate multiple sub-bands, calculate the local phase change values ​​of the corresponding sub-bands of the reference frame and the current frame, and perform weighted aggregation of the local phase change values ​​according to the spatial weight map mapped to the spatial weights of the corresponding sub-bands to generate the time response of each sub-band, and fuse the time responses of each sub-band to generate the restored audio.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the audio recovery method as described in any one of claims 1-6.

10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they are used to implement the audio recovery method as described in any one of claims 1-6.