Audio Restoration Method Based on Nonlinear Interleaving Mapping and Video Microphone System
By adopting a non-linear interleaving mapping system based on two-dimensional trigonometric functions in the audio recovery method, the problems of poor scale scalability and low detail resolution in the prior art are solved, and higher quality audio recovery is achieved.
Patent Information
- Application Number
- CN202210751918.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-28
AI Technical Summary
The existing audio recovery methods have poor scale scalability and low detail resolution during nonlinear interleaving mapping, so they cannot effectively restore high-quality sound.
A nonlinear interleaving mapping system based on two-dimensional trigonometric functions is adopted. By nonlinear interleaving mapping of video frames and two-dimensional trigonometric function interleaving auxiliary functions, a two-dimensional 0-1 matrix is generated, and dimensionality reduction and filtering and denoising are performed to finally restore the audio.
Improves the scale scalability and detail resolution of audio recovery, improves the recovery quality of sound details, and reduces algorithm complexity.
Smart Images

Figure CN115499759B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing, and more particularly to an audio recovery method and a video microphone system for a non-linear interleaving mapping system based on two-dimensional trigonometric functions. Background Art
[0002] At present, audio is mainly collected through microphones. However, due to factors such as long distance, the collected speech is relatively weak, and it is impossible to completely receive and recognize the speech signal. Therefore, appropriate methods need to be adopted to amplify the extremely weak vibrations caused by speech so as to recover and collect the speech. In the early 20th century, the invention of electron tubes and radio waves promoted the organic combination of speech acoustics and electroacoustics, which amplified the weak sound signals that were difficult to receive to a certain extent, and the electrical microphone followed. The principle of the electrical microphone is to transmit the vibration of sound to the diaphragm of the microphone, push the internal magnet to generate a changing current, and transport the changing current to the sound processing circuit for amplification processing to complete the amplification and recovery of speech. In the 1970s and 1980s, with the in-depth research and application of fiber optic sensing technology, the invention and application of fiber optic microphones marked the realization of long-distance collection of speech using optical fibers. Soquet et al. studied the characteristics of fiber optic microphones that the optical signal will cause changes in parameters such as light intensity and phase when vibrating, and used the corresponding signal detection means and demodulation system to restore the sound signal, thereby completing the conversion of dynamic sound signals to dynamic optical signals. Furstenau et al. proposed a fiber optic microphone based on an external Fabry-Perot (FP) microinterferometer of the optical fiber, coupled to an external membrane for modulating the length of the FP cavity and a low-coherence superluminescent diode as the light source. It can identify the models of takeoff and landing aircraft by detecting and analyzing the characteristic noise spectrum of the aircraft, and is also applied to traffic monitoring and vehicle classification. Konle et al. developed a high-temperature-resistant fiber optic microphone based on a Fabry-Perot interferometer and successfully applied it to the combustion chamber at 1400k. The laser microphone proposed by Rothberg et al. uses a laser beam to detect the vibration of a glass surface or a mirror surface to collect speech. By detecting the phase change of the reflected laser beam, the distance change of the reflection plane can be tracked, and the Doppler frequency shift of the reflected laser beam is detected by LDV to track the speed change of the reflection plane to recover high-quality speech at a long distance. Their work makes the positioning of the receiver more flexible, but still relies on recording the reflected laser.
[0003] In addition, an optical microphone also includes a video microphone that utilizes ordinary light. The video microphone inherits the advantage of the fiber optic microphone in being able to collect long-distance or enclosed audio information. At the same time, it does not require projecting a laser beam or pattern onto the vibrating surface, which means it does not rely on an active light source. The Abe Davis team first proposed the concept of a video microphone, extracting local video motion signals in the dimension of a complex manipulable pyramid. These local signals are aligned and averaged into a single one-dimensional motion signal that captures the overall motion of the object over time, and then filtered and denoised to produce the restored sound. Kim Y J proposed a patch-based visual microphone framework based on the complex manipulable pyramid microphone to solve the sound restoration problem by restoring sound from sub-regions centered on key points in the image. However, the pyramid decomposition algorithm has a high complexity and is insensitive to audio differences, resulting in low audio quality.
[0004] Chinese Patent Application No. 2022104924747 presents an audio restoration method and a video microphone system based on a non-linear dynamic system. This invention can restore the sound in a room through the silent video content of objects in the room captured by a high-frame-rate camera placed outside the room, overcoming the technical problems of excessive application limitations and large product size existing in existing audio restoration methods. However, in the non-linear interleaved mapping process, it uses a logarithmic function as an auxiliary function, resulting in poor scale scalability and low detail resolution of the restored audio. Summary of the Invention
[0005] In view of the technical problems of poor scale scalability and low detail resolution in the above-mentioned existing technologies, an audio restoration method and a video microphone system based on non-linear interleaved mapping are provided. This invention mainly uses a non-linear interleaved mapping system based on two-dimensional trigonometric functions for audio restoration, which has better scale scalability from a global perspective, can improve the overall detail resolution, and helps to improve the restoration quality of sound details.
[0006] The technical means adopted by this invention are as follows:
[0007] An audio restoration method based on non-linear interleaved mapping, comprising:
[0008] S1. Obtain the video to be processed, preprocess the video to be processed to generate a continuous sequence of grayscale images, where the video to be processed includes vibration images of an object affected by surrounding sound sources;
[0009] S2. Perform non-linear interleaved mapping of pixel coordinates and image brightness for each image in the grayscale image sequence with a two-dimensional trigonometric function interleaved auxiliary function, so as to obtain a group of two-dimensional 0-1 matrices corresponding to the current image;
[0010] S3. Perform dimensionality reduction processing on each two-dimensional 0-1 matrix respectively to generate one-dimensional data;
[0011] S4. After filtering and denoising each of the one-dimensional data, the restored audio is obtained.
[0012] Further, perform non-linear interleaving mapping of pixel coordinates and image brightness on each image in the grayscale image sequence with a two-dimensional trigonometric interleaving auxiliary function, so as to obtain a set of two-dimensional 0-1 matrices corresponding to the current image, including:
[0013] S201. Construct an interleaving auxiliary function matrix based on the following two-dimensional trigonometric interleaving auxiliary function:
[0014] f(x,y) = cos(ax) + cos(by) (0 < a, b < 1)
[0015] The interleaving auxiliary function matrix is:
[0016]
[0017] Among them, M and N select the maximum value of image grayscale 256, and floor is rounding down.
[0018] Further, perform non-linear interleaving mapping of pixel coordinates and image brightness on each image in the grayscale image sequence with a two-dimensional trigonometric interleaving auxiliary function, and it also includes:
[0019] S202. Select W points in the matrix iteration range of size from top to bottom and from left to right as the initial value points, where W < 65536::
[0020]
[0021] S203. For each initial value point, perform N interleaving iterations to generate N two-dimensional points, as follows:
[0022]
[0023] S204. Construct an interleaving mapping matrix of size M×N:
[0024]
[0025] According to S203, record the coordinates of the set 1 points in the interleaving mapping matrix generated by N iterations of W initial points for each frame of the target image, set the elements corresponding to these coordinates in the interleaving mapping matrix to 1, and set other elements to 0, that is:
[0026]
[0027] The obtained 0-1 two-dimensional interleaved mapping matrix I.
[0028] Further, after dimensionality reduction processing is performed on each two-dimensional 0-1 matrix respectively, one-dimensional data is generated, including:
[0029] S301. Generate P row vectors of size Q from P groups of two-dimensional chaotic attractors I generated by using non-linear interleaved mapping on P-frame video images P ≤ the number of frames of the target video image, and:
[0030]
[0031] where floor is for rounding down and mod is for taking the remainder;
[0032] S302. Combine the P row vectors vertically in sequence into a matrix S of size P×Q:
[0033]
[0034] S303. Calculate the covariance matrix covS of matrix S;
[0035] S304. Obtain all the eigenvalues of the covariance matrix covS and find the eigenvectors of the covariance matrix;
[0036] S305. Weight the eigenvectors corresponding to the appropriate eigenvalues with matrix covS to generate one-dimensional data H of size P×1 i as the subsequent voice output:
[0037]
[0038] Further, the appropriate eigenvalues are obtained according to the following method:
[0039] Arrange the eigenvalues from largest to smallest, and the eigenvectors corresponding to the top three eigenvalues in proportion are multiplied by the standardized S respectively to generate three one-dimensional data G of size P×1 i , i = 1, 2, 3:
[0040]
[0041] Transpose the three segments of G i and intercept N consecutive data K at the same position, where
[0042] K i = [K i (1), K i (2), …, Ki (N)]
[0043] Find K i Calculate autocorrelation (40 < T < 50):
[0044]
[0045] Normalize and calculate the integral:
[0046]
[0047] Take the eigenvalue corresponding to the maximum r value as the appropriate eigenvalue.
[0048] Furthermore, preprocess the video to be processed to generate a continuous sequence of grayscale images, including:
[0049] S101. Obtain each frame image of the video to be processed in sequence;
[0050] S102. Obtain the grayscale value Gray of each frame image by weighting the three primary colors of the color RGB image:
[0051] Gray = 0.299R + 0.587G + 0.114B
[0052] S103. Crop the grayscale image into a target grayscale image of 256×256, and the starting point (U, V) of the target grayscale image is randomly obtained within the feasible range.
[0053] Furthermore, filter and denoise each of the one-dimensional data, including:
[0054] Pass each of the one-dimensional data through a Butterworth high-pass filter and an IIR Bessel low-pass filter in sequence.
[0055] A video microphone system includes an audio recovery unit, and the audio recovery unit is used to execute the audio recovery method based on non-linear interleaved mapping described in any one of the above.
[0056] Compared with the prior art, the present invention has the following advantages:
[0057] 1. The present invention mainly uses a non-linear interleaved mapping system based on two-dimensional trigonometric functions for audio recovery, which has better scale scalability from a global perspective, can improve the overall detail resolution, and helps to improve the recovery quality of sound details.
[0058] 2. Compared with the existing audio algorithms, the algorithm complexity of the present invention is lower, the extraction of differences is faster and more convenient, and the audio quality is higher.
[0059] 3. The present invention can directly extract differential features without amplifying the vibration amplitude of an object to extract more obvious image differences. It not only reflects the variation differences in the two-dimensional plane translation of the image but also extracts the differences in three-dimensional space scaling.
[0060] 4. The video microphone of the present invention does not rely on active illumination compared to a laser microphone and can achieve its function with only a high-frame camera and any object that can generate vibrations in the room. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0062] Figure 1 It is a flowchart of the audio recovery method based on non-linear interleaved mapping of the present invention.
[0063] Figure 2 It is a schematic diagram of the implementation process of the audio recovery method of non-linear interleaved mapping in the embodiment.
[0064] Figure 3 It is a schematic diagram of the non-linear interleaved mapping process in the embodiment.
[0065] Figure 4a It is a schematic diagram of the differential extraction by the method of the present invention.
[0066] Figure 4b It is a schematic diagram of the differential extraction by the comparative method.
[0067] Figure 5 It is a schematic diagram of the comparison between the audio recovered by the method of the present invention and the original audio, where the upper part is the schematic diagram of the original audio and the lower part is the schematic diagram of the audio recovered by the method of the present invention.
[0068] Figure 6 It is a schematic diagram of the comparison of the processing time between the method of the present invention and the comparative method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0069] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0070] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0071] As Figure 1 shown, this embodiment provides an audio recovery method for a non-linear interleaved mapping system based on two-dimensional trigonometric functions. First, a high-frame-rate video of the resonance between an object and surrounding sound sources is collected. Each frame image of the video is sequentially input into the constructed non-linear interleaved mapping system of two-dimensional trigonometric functions for iterative generation of a two-dimensional interleaved mapping matrix representing the image features. The two-dimensional interleaved mapping matrix generated by each frame image is dimensionally reduced and transformed into one-dimensional information using PCA technology. The one-dimensional information is filtered and denoised to recover the audio.
[0072] The specific steps are as follows:
[0073] 1. Collect high-frequency video. This includes placing a sound source, any object suitable for vibration, and a high-frame-rate camera with a frame rate higher than the object vibration frequency in a relatively enclosed room. When the sound source emits sound (such as playing music or speaking), use the high-frame-rate camera to perform lossless vibration information collection on the object resonating with the audio to obtain a high-frame-rate video of the object vibration. The video frequency is usually in the range of 2 kHz - 20 kHz.
[0074] Furthermore, preprocess the video to be processed. Store each frame of the video as a video picture. To ensure that the iteration of gray level and coordinates does not exceed the boundary during the coordinate-gray level interleaving process, it is necessary to ensure the same value range for both. Therefore, the image is cropped to a size of 256×256 to match the gray level value range of 0 - 256. Take the (U, V) points of the target gray image as the starting point to crop the target image matrix G with a size of M×N (M = N = 256):
[0075]
[0076] Among them, the starting point (U, V) is randomly obtained within the feasible range, that is, U + 256 and V + 256 are any values not exceeding the size of the target image. Gray is the gray value corresponding to the coordinates of each video picture read. The image gray value Gray is obtained by weighting the three primary colors of the RGB image of the color image.
[0077] Gray = 0.299R + 0.587G + 0.114B.
[0078] 2. Construct a non - linear interleaving mapping system of two - dimensional trigonometric functions. It includes performing non - linear interleaving mapping of pixel coordinates and image brightness between each image in the gray - scale image sequence and the two - dimensional trigonometric function interleaving auxiliary function, so as to obtain a set of two - dimensional 0 - 1 matrices corresponding to the current image. Specifically, it includes:
[0079] 2a). Construct a two - dimensional interleaving auxiliary function matrix.
[0080] Since the two - dimensional trigonometric function has more obvious chaotic characteristics and initial - value sensitivity can retain and amplify differences, the two - dimensional trigonometric function f(x, y) is selected as the two - dimensional interleaving auxiliary function of the non - linear interleaving mapping system:
[0081] f(x, y)=cos(ax)+cos(by)(0 < a, b < 1)
[0082] Re - construct the matrix of this two - dimensional interleaving auxiliary function as the interleaving auxiliary function matrix:
[0083]
[0084] Among them, M and N are selected as the maximum gray value 256 of the image, and floor is for rounding down.
[0085] 2b). Iteratively generate the two - dimensional interleaving mapping matrix algorithm. Specifically:
[0086] Select the matrix iteration range W initial value points in the order from top to bottom and from left to right, where W < 65536:
[0087]
[0088] For each initial value point, perform N times of interleaving iterations to generate N two - dimensional points, as follows:
[0089]
[0090] Construct an interleaving mapping matrix of size M×N:
[0091]
[0092] Record each frame of the target image (i.e., the set target image per frame) according to the above steps. After N iterations with W initial points, generate the coordinates of the set 1 points in the interleaved mapping matrix, and set the elements corresponding to these coordinates in the interleaved mapping matrix to 1, and set other elements to 0, that is:
[0093]
[0094] The resulting 0-1 two-dimensional interleaved mapping matrix I Serves as the representation of this target image for this image feature:
[0095]
[0096] S3. Perform dimensionality reduction processing on each two-dimensional 0-1 matrix respectively to generate one-dimensional data. Specifically, it includes:
[0097] 3a) Construct a matrix suitable for processing
[0098] Generate a row vector from each two-dimensional interleaved mapping matrix of size M×N (where Q = M×N)
[0099]
[0100] where floor is rounding down and mod is taking the remainder.
[0101] The P row vectors obtained from P target images (the selected high-frame video contains P frames) Are merged vertically in order into a matrix S of size P×Q:
[0102]
[0103] 3b) Process the matrix using PCA technology
[0104] The matrix S is (where Is a row vector):
[0105]
[0106] According to the covariance formula, the i,j term of the covariance matrix is defined in the following form:
[0107]
[0108] where E is the expectation, and μ i Is the expectation of the i-th element, that is
[0109] The covariance matrix obtained is:
[0110] CovS = E[(S - E[S])(S - E[S]) T
[0111] Find all the eigenvalues λ1, λ2, …, λ of the generated covariance matrix covS k , and find that there exists a number λ i and a non-zero column vector (with the same dimension as covS) such that
[0112]
[0113] holds (λ i is the matrix eigenvalue), then is the corresponding eigenvector. The matrix V is the eigenvector matrix composed of eigenvectors:
[0114]
[0115] Arrange the eigenvalues from largest to smallest, and multiply the eigenvectors corresponding to the top three eigenvalues in proportion (i = 1, 2, 3) by the standardized S respectively to generate three one-dimensional data H of size P×1 i (i = 1, 2, 3):
[0116]
[0117] Transpose the three segments of H i and intercept N consecutive data K at the same position, where
[0118] K i = [K i (1), K i (2), …, K i (N)]
[0119] Find K i Calculate the autocorrelation (40 < T < 50):
[0120]
[0121] Normalize and calculate the integral:
[0122]
[0123] Take the eigenvalue corresponding to the maximum r value as the appropriate eigenvalue, and take the one-dimensional data H corresponding to the maximum r value as the subsequent voice output.
[0124] S4. After filtering and denoising each of the one-dimensional data, the restored audio is obtained.
[0125] 4a) Perform processing using a high-pass filter. Pass through a high-pass filter of order 3
[0126] 4b) Process using a low-pass filter. Design and adopt an IIR Butterworth low-pass filter.
[0127] The following further illustrates the solution and effect of the present invention through a specific application example.
[0128] In this embodiment, a bag of potato chips that can generate vibrations is placed in a relatively enclosed room, and the English children's song "Mary had a little lamb" is played using audio equipment in the room. A high-frame-rate camera is placed at a distance of 0.5 meters to 2 meters from the bag of potato chips, and the vibrations of the bag of potato chips are collected through a soundproof glass at intervals.
[0129] Restore the sound of the collected silent video. As Figures 2-5 shown, the frame rate of the collected video of the bag of potato chips is 2200 hz, the resolution is 700x400 pixels, and 8000 frames in the video are selected as the processed video. Use a two-dimensional trigonometric function interleaving auxiliary function:
[0130]
[0131] Generate an interleaving auxiliary function matrix and construct a non-linear interleaving mapping system of two-dimensional trigonometric functions with the obtained high-frame-rate video frame images. Select 10000 points and iterate 10 times respectively to obtain a two-dimensional interleaving mapping matrix representing its phase characteristics. Then perform dimensionality reduction processing on the two-dimensional interleaving mapping matrix using PCA technology to obtain one-dimensional data. Use a high-pass filter with an order of 3 for preliminary audio processing, and then use a low-pass filter with a passband ripple coefficient rp = 1, a stopband ripple coefficient rs = 20, a stopband frequency Ft = 1000, a passband frequency Fp = 5000, and a sampling frequency Fs = 22000 to restore the audio. From Figure 5 and the restored audio played, it can be obtained that this method can restore certain initial sound source content.
[0132] In summary, the video microphone system of the present invention uses a non-linear interleaving mapping system of two-dimensional trigonometric functions to extract the phase difference of the video when restoring audio. The code is simpler and more modifiable, and different interleaving auxiliary functions can be selected according to different scenarios to construct a non-linear interleaving mapping system of two-dimensional trigonometric functions, with more flexible application. In practical applications, different from a laser microphone that needs to actively irradiate a laser on an object, it only needs to observe the light already existing in the scene. Compared with other algorithms, it also does not need to change the movement of the video object - amplify the vibration amplitude of the object to extract the phase difference.
[0133] The present invention also discloses a video microphone system, including an audio restoration unit, and the audio restoration unit is used to execute the audio restoration method based on non-linear interleaving mapping described in any one of the above.
[0134] For the embodiments of the present invention, since they correspond to those in the above embodiments, the description is relatively simple. For relevant similarities, please refer to the description in the corresponding part of the above embodiments, and details will not be repeated here.
[0135] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio restoration method based on non-linear interleaving mapping, characterized in that, Including: S1. Obtain the video to be processed, preprocess the video to be processed to generate a continuous sequence of grayscale images, where the video to be processed includes vibration images generated by an object affected by surrounding sound sources; S2. Perform non-linear interleaving mapping of pixel coordinates and image brightness for each image in the grayscale image sequence with a two-dimensional trigonometric interleaving auxiliary function, so as to obtain a set of two-dimensional 0-1 matrices corresponding to the current image, including: S201. Construct an interleaving auxiliary function matrix based on the following two-dimensional trigonometric interleaving auxiliary function: f(x,y) = cosax + cos(by) (0 < a, b < 1) The interleaving auxiliary function matrix is: where M and N select the maximum grayscale value of the image 256, and floor is rounding down; S202. Select W points in the order from top to bottom and from left to right in the matrix iteration range of size as the initial value points, where W <65536: S203. For each initial value point, perform N interleaving iterations to generate N two-dimensional points as follows: S204. Construct an interleaving mapping matrix of size M×N: According to S203, record that each frame of the target image is iterated N times by W initial points to generate the coordinates of the set 1 points in the interleaving mapping matrix, set the elements corresponding to these coordinates in the interleaving mapping matrix to 1, and set other elements to 0, that is: Thus, the obtained 0-1 two-dimensional interleaving mapping matrix I; S3. Perform dimensionality reduction processing on each two-dimensional 0-1 matrix respectively to generate one-dimensional data, including: S301. Generate P row vectors of size Q from the P - group two - dimensional chaotic attractor I generated by non - linearly interleaving and mapping the P - frame video image P ≤ the number of frames of the target video image, and: where floor is rounding down and mod is taking the remainder; S302. Combine P row vectors vertically in sequence into a matrix S of size P×Q: S303. Calculate the covariance matrix covS of matrix S; S304. Obtain all the eigenvalues of the covariance matrix covS, and find the eigenvectors of the covariance matrix; S305. Weight the eigenvector corresponding to the appropriate eigenvalue with the matrix covS to generate a one-dimensional data H of size P×1 i for subsequent voice output: S4. After filtering and denoising each of the one-dimensional data, obtain the restored audio.
2. The audio restoration method based on non-linear interleaving mapping according to claim 1, characterized in that The appropriate eigenvalue is obtained in the following manner: Arrange the eigenvalues from largest to smallest, and the eigenvectors corresponding to the top three eigenvalues in terms of proportion Multiply each of them with the standardized S respectively to generate three one-dimensional data H of size P×1 i , where i = 1, 2, 3: Transpose three segments of H i and intercept N consecutive data K at the same position, where K i = [K i (1), K i (2), …, K i (N)] Find K i Calculate autocorrelation (40 < T < 50): Normalize and calculate the integral: Take the eigenvalue corresponding to the maximum r value as the appropriate eigenvalue.
3. The audio restoration method based on non-linear interleaving mapping according to claim 1, characterized in that Preprocess the video to be processed to generate a continuous sequence of grayscale images, including: S101. Sequentially obtain each frame of the video to be processed; S102. Obtain the grayscale value Gray of each frame of the image by weighting the three primary colors of the color RGB image: Gray = 0.299R + 0.587G + 0.114B S103. Crop the grayscale image into a target grayscale image of 256×256, and the starting point (U, V) of the target grayscale image is randomly obtained within the feasible range.
4. The audio restoration method based on non-linear interleaving mapping according to claim 1, characterized in that Filter and denoise each of the one-dimensional data, including: Pass each of the one-dimensional data through a Butterworth high-pass filter and an IIR Bessel low-pass filter in sequence.
5. A video microphone system, characterized in that, Including an audio restoration unit, and the audio restoration unit is used to execute the audio restoration method based on non-linear interleaving mapping described in any one of claims 1 to 4.