Noise reduction preprocessing method, audio noise reduction method and related device and equipment
By configuring audio input devices in the target location and using calibrated audio signals to obtain audio difference data, a virtual noise reduction array is constructed, which solves the problem of high hardware costs in audio acquisition scenarios with multiple people and achieves efficient audio noise reduction processing.
Patent Information
- Application Number
- CN202511085555.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-04
AI Technical Summary
In audio acquisition scenarios involving multiple people, existing technologies require the addition of specialized directional sound recording equipment to reduce interference, resulting in high costs and poor applicability. They cannot provide reliable noise reduction reference data without increasing additional hardware costs.
By configuring audio input devices in the target location, acquiring audio difference data using calibrated audio signals, determining the reference seat for the target seat, constructing a virtual noise reduction array, and using the recording signal from the reference seat to perform noise reduction processing on the target seat.
Without increasing additional hardware costs, it improves the adaptability and accuracy of audio noise reduction, provides reliable reference data, and enhances the audio quality of the target recording signal.
Smart Images

Figure CN120895046A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech signal processing, in particular to a noise reduction preprocessing method and an audio noise reduction method and related devices and equipment. BACKGROUND
[0002] In an audio acquisition scene where multiple people coexist, such as a conference room, a classroom, a lecture hall and the like, it is often necessary to accurately record the audio of a specific target position. In these scenarios, due to the dense distribution of personnel in the space, the target position is close to other surrounding positions, and the sound produced by the surrounding positions easily interferes with the audio acquisition of the target position.
[0003] In the prior art, a professional directional sound receiving device is usually arranged at each position in the target place to reduce the interference of other positions on the audio acquisition of the target position. However, such devices have high costs, and large-scale replacement of devices will result in high modification costs, leading to poor applicability. Therefore, how to provide as reliable reference data as possible for subsequent noise reduction processing without increasing additional hardware costs has become a problem to be solved. SUMMARY
[0004] The technical problem solved by the present application is to provide a noise reduction preprocessing method and an audio noise reduction method and related devices and equipment, which can provide as reliable reference data as possible for subsequent noise reduction processing without increasing additional hardware costs.
[0005] To solve the above technical problem, the first aspect of the present application provides a noise reduction preprocessing method, comprising: regarding other seats in a target place except a target seat as candidate seats; wherein at least one audio input device is configured at each seat in the target place respectively to meet the recording needs of an object at each seat; obtaining a first audio signal collected by an audio input device at the target seat and a second audio signal collected by an audio input device at each candidate seat based on a calibration audio signal played by an audio output device at the target seat; determining a plurality of reference seats of the target seat based on audio difference data between the first audio signal and each second audio signal; wherein in an audio recording application scenario, the audio input devices at each reference seat of the target seat form a virtual noise reduction array, and the first recording signal collected by the audio input device at the target seat is configured to be reduced in noise by referring to the second recording signal collected by the audio input device at each reference seat.
[0006] To solve the above technical problems, the second aspect of the present application provides an audio noise reduction method, comprising: obtaining a first recording signal collected by an audio input device at a target seat in a target place and a second recording signal collected by an audio input device at each reference seat of the target seat; wherein the reference seats of the target seat are obtained based on the noise reduction preprocessing method of the first aspect; performing a noise reduction operation on the first recording signal based on each second recording signal to obtain a target recording signal of the target seat.
[0007] To solve the above technical problems, the third aspect of the present application provides a noise reduction preprocessing device, comprising: a selection module, an acquisition module and a determination module, the selection module is used to select other seats except the target seat in the target place as candidate seats; wherein at least one audio input device is configured at each seat in the target place respectively to meet the recording needs of the object at each seat; the acquisition module is used to obtain a first audio signal collected by an audio input device at the target seat and a second audio signal collected by an audio input device at each candidate seat based on a calibration audio signal played by an audio output device at the target seat; the determination module is used to determine a plurality of reference seats of the target seat based on audio difference data between the first audio signal and each second audio signal; wherein in the application scenario of audio recording, the audio input devices at each reference seat of the target seat form a virtual noise reduction array, and the first recording signal collected by the audio input device at the target seat is configured to be reduced based on the second recording signal collected by the audio input device at each reference seat.
[0008] To solve the above technical problems, the fourth aspect of the present application provides an audio noise reduction device, comprising: a collection module and a noise reduction module, the collection module is used to obtain a first recording signal collected by an audio input device at a target seat in a target place and a second recording signal collected by an audio input device at each reference seat of the target seat; wherein the reference seats of the target seat are obtained based on the noise reduction preprocessing method of the first aspect; the noise reduction module is used to perform a noise reduction operation on the first recording signal based on each second recording signal to obtain a target recording signal of the target seat.
[0009] To solve the above technical problems, the fifth aspect of the present application provides an electronic device, comprising a memory and a processor coupled to each other, the memory at least stores program instructions, and the processor is used to execute the program instructions to implement the noise reduction preprocessing method of the first aspect or the audio noise reduction method of the second aspect.
[0010] To solve the above technical problems, the sixth aspect of the present application provides a computer readable storage medium, which stores program instructions capable of being run by a processor, and the program instructions are used to implement the noise reduction preprocessing method of the first aspect or the audio noise reduction method of the second aspect.
[0011] The above scheme takes other seats in the target site except the target seat as candidate seats, at least an audio input device is respectively arranged at each seat in the target site, respectively used to meet the recording needs of the object at each seat, based on the calibration audio signal played by the audio output device at the target seat, the first audio signal collected by the audio input device at the target seat and the second audio signal collected by the audio input device at each candidate seat are obtained, based on the audio difference data between the first audio signal and each second audio signal, a plurality of reference seats of the target seat are determined, the first recording signal collected by the audio input device at the target seat in the target site and the second recording signal collected by the audio input device at a plurality of reference seats of the target seat are obtained, and the noise reduction operation is performed on the first recording signal based on each second recording signal to obtain the target recording signal of the target seat. In the noise reduction preprocessing stage, on the one hand, the reference seats of the target seat are established through the calibration audio signal, so as to construct the virtual noise reduction array of the target seat, and therefore the professional equipment does not need to be additionally added, and the adaptability to most application scenarios is improved. On the other hand, the calibration audio played by the audio output device at the target seat is used to obtain the first audio signal and the second audio signal, and the difference data is calculated, so as to improve the accuracy of the determination of the reference seats as much as possible, and then the effectiveness of the reference seats in the subsequent audio noise reduction process is improved. Therefore, in the audio noise reduction process, the first recording signal of the target seat is subjected to the noise reduction processing by using the second recording signal of the reference seats, and the audio quality of the target recording signal can be improved. Therefore, reliable reference data as much as possible can be provided for the subsequent noise reduction processing without increasing the additional hardware cost. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 FIG. 1 is a flowchart of an embodiment of the noise reduction preprocessing method of the present application; Figure 2 FIG. 2 is a flowchart of another embodiment of the noise reduction preprocessing method of the present application; Figure 3 FIG. 3 is a flowchart of an embodiment of the audio noise reduction method of the present application; Figure 4 FIG. 4 is a flowchart of another embodiment of the audio noise reduction method of the present application; Figure 5 FIG. 5 is a framework diagram of an embodiment of the audio noise reduction method of the present application; Figure 6 FIG. 6 is a framework diagram of another embodiment of the audio noise reduction method of the present application; Figure 7 FIG. 7 is a framework diagram of an embodiment of the noise reduction preprocessing device of the present application; Figure 8 FIG. 8 is a framework diagram of an embodiment of the audio noise reduction device of the present application; Figure 9 This is a schematic diagram of the framework of an embodiment of the electronic device of this application; Figure 10 This is a schematic diagram of a framework of an embodiment of the computer-readable storage medium of this application. Detailed Implementation
[0013] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0014] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.
[0015] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the slash " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper indicates two or more objects.
[0016] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the noise reduction preprocessing method of this application. Specifically, it may include the following steps: Step S11: Select other seats in the target location besides the target seat as candidate seats.
[0017] In this embodiment of the disclosure, the target venue has several seats. Different target venues have different requirements for audio noise reduction processing. Specifically, if audio noise reduction optimization is required for all seats in the venue, all seats are designated as target seats. For example, if the target venue has 6 seats, when seat A is designated as the target seat, the candidate seats are seats B, C, D, E, and F. When seat B is designated as the target seat, the candidate seats are seats A, C, D, E, and F, and so on. Each seat is designated as a target seat in turn, and the remaining 5 seats become candidate seats accordingly. Alternatively, only the seats with noise reduction requirements are designated as target seats. For example, if the target venue has 30 seats, seats 1-20 are exam seats, and corresponding recordings need to be saved, thus requiring noise reduction. Seats 21-30 are lecture seats, and no corresponding recordings need to be saved, thus requiring no noise reduction. When seat 1 is designated as the target seat, all seats 2-30 are candidate seats. Similarly, seats 2-20 are processed. Seats 21-30, which do not require noise reduction, are not designated as target seats, and therefore, there is no need to determine candidate seats for them. It should be noted that the method of selecting the target seat in the target location is not limited in this application.
[0018] In one implementation scenario, at least one audio input device is respectively configured at each seat in the target venue to meet the recording needs of the object at each seat. Specifically, the audio input device can be a microphone or the like, which can be fixed near the seat, such as the seat armrest, desktop, or the like, which is not limited in the present application.
[0019] In one specific implementation scenario, the time of the audio input device respectively configured at each seat in the target venue is synchronized, specifically, the terminal clock of the audio input device at each seat in the target venue is synchronized.
[0020] Step S12: Based on the calibration audio signal played by the audio output device at the target seat, the first audio signal collected by the audio input device at the target seat and the second audio signal collected by the audio input device at each candidate seat are obtained.
[0021] In one implementation scenario, the calibration audio signal played by the audio output device at the target seat has a known signal characteristic, such as frequency characteristic, amplitude characteristic, phase characteristic, etc., which can be used as a reference for subsequent audio signal analysis. Specifically, the calibration audio signal is a sweep frequency tone + pseudo-random sequence, and the specific type is not limited.
[0022] In one specific implementation scenario, during the process of playing the calibration audio signal by the audio output device at the target seat, the audio input device arranged at the target seat is triggered to perform signal collection operation, thereby obtaining the first audio signal, which can directly reflect the actual propagation effect and reception of the calibration audio signal at the target seat, and the audio input devices distributed at each candidate seat are also started to collect the second audio signal, which can reflect the change of the audio characteristic of the calibration audio signal from the target seat to the candidate seat.
[0023] In one specific implementation scenario, for each target seat, a corresponding audio input device and an audio output device are configured. It should be noted that the audio input device and the audio output device on the same seat can adopt an integrated design, that is, the two are the same device, which has both the functions of audio signal collection and playing, such as a smart sound box with a built-in microphone and a speaker, a headset with a radio function, etc. The audio input device and the audio output device on the same seat can also adopt a split design, that is, the two are different independent devices, such as a separate microphone and a separate speaker.
[0024] In a specific implementation scenario, the audio input device and the audio output device on the same seat are designed in a split type, and the positions where the audio input device collects signals and the audio output device plays signals are as close as possible to avoid interference on subsequent signal analysis due to position difference, by fixing the microphone on the shell of the loudspeaker, or installing them at the same height and angle on the same support, and the like.
[0025] Step S13: determining several reference seats of the target seat based on the audio difference data between the first audio signal and each second audio signal.
[0026] In an implementation scenario, in the application scenario of audio recording, the audio input devices at the reference seats of the target seat form a virtual noise reduction array. The virtual noise reduction array is not composed of a plurality of audio input devices fixedly installed on the same physical carrier, but is composed of audio input devices distributed at the reference seats of the target seat. These audio input devices are independently deployed at different reference seats, but form a logical array structure through data transmission and cooperative processing mechanism, i.e., the "virtual" array can adapt to different spatial layouts, flexibly utilize the existing seat resources in the scene to deploy audio input devices, and does not need to additionally build a complex physical array structure. The first recording signal collected by the audio input device at the target seat is configured to be reduced by the second recording signal collected by the audio input device at the reference seat. The above scheme, on the one hand, establishes the reference seat of the target seat by calibrating the audio signal, thereby constructing the virtual noise reduction array of the target seat, and thus does not need to additionally add professional equipment, thereby improving the adaptability to most application scenarios. On the other hand, the first audio signal and the second audio signal are obtained through the calibration audio played by the audio output device at the target seat, and the difference data is calculated, thereby improving the accuracy of the determination of the reference seat, and further improving the effectiveness of the reference seat in the subsequent audio noise reduction process.
[0027] In a specific implementation scenario, the audio difference data can be embodied by various acoustic parameters, such as the frequency response difference, amplitude attenuation difference, phase offset, signal-to-noise ratio difference, and the like of the signal. Based on the audio difference data between the first audio signal and each second audio signal, the difference degree between each candidate seat and the target seat is quantified. Specifically, the audio difference data can be obtained by calculating the frequency spectrum similarity, time domain waveform matching degree, sound pressure level difference, and the like of the signal to obtain a comprehensive quantification result. The result directly reflects the deviation degree of the candidate seat and the target seat in the acoustic characteristics. For example, the smaller the quantification result value, the lower the difference degree, the greater the influence of the audio at the candidate seat on the audio collection at the target seat, and vice versa. The greater the difference degree, the smaller the influence of the audio at the candidate seat on the audio collection at the target seat.
[0028] In one implementation scenario, the candidate seats are ranked according to the difference degree represented by the audio difference data, for example, in the order from low to high, with the candidate seat having the smallest difference degree at the front end of the sequence and the candidate seat having the largest difference degree at the tail end of the sequence, or in the order from high to low, with the candidate seat having the largest difference degree at the front end of the sequence and the candidate seat having the smallest difference degree at the tail end of the sequence, and the candidate seat having a difference degree within a preset low position is selected as the reference seat. Correspondingly, if the ranking is in the order from low to high, the preset number of candidate seats are selected as the reference seats from front to back based on the ranking, and if the ranking is in the order from high to low, the preset number of candidate seats are selected as the reference seats from back to front based on the ranking. It should be noted that the specific value of the preset low position is not limited in the present application.
[0029] In one specific implementation scenario, in order to improve the flexibility and adaptability of the selection of the reference seat, the preset low position can be dynamically adjusted according to the needs of the actual application scenario. For example, in the scenario where the position planning of the target site is reasonable, the preset low position can be set to a smaller value to ensure that the selected reference seat is highly consistent and effective in acoustic characteristics with the target seat, thereby improving the effect of audio noise reduction, and in the scenario where the position planning of the target site is not reasonable, the value of the preset low position can be appropriately relaxed to expand the range of the selectable reference seat and improve the fault tolerance and robustness of the system.
[0030] In one specific implementation scenario, the audio difference data includes at least one of time delay data and volume attenuation data. Specifically, the time delay data refers to the difference between the time taken by an audio signal to propagate from the audio output device at the target seat to the audio input device at a certain candidate seat and the time taken by the signal to propagate to the audio input device at the target seat. For example, if the audio signal takes 0.01 seconds longer to propagate to the audio input device at a certain candidate seat than to the audio input device at the target seat, the time delay data is 0.01 seconds. The volume attenuation data refers to the difference between the volume of the calibration audio signal when it propagates to the audio input device at the candidate seat and the volume of the signal received by the audio input device at the target seat, usually measured in sound pressure level. During the propagation of the audio signal, the energy is gradually lost due to factors such as air absorption, obstacle blocking and diffusion, and the farther the distance, the greater the loss, and the more obvious the volume attenuation, thereby forming the volume attenuation data. For example, the sound pressure level of the calibration audio signal received at the target seat is 80 decibels, and the sound pressure level received at a certain candidate seat is 65 decibels due to the long distance and the presence of some obstacles, so the volume attenuation data between the two is 15 decibels.
[0031] In a specific implementation scenario, before the candidate seats are sorted according to the difference degree represented by the audio difference data, when the audio difference data includes the time delay data, the candidate seats with time delay data not greater than a time delay threshold are filtered out, and when the audio difference data includes the volume attenuation data, the candidate seats with volume attenuation data not greater than an attenuation threshold are filtered out. The specific values of the time delay threshold and the attenuation threshold are not limited in the present application, and can be pre-configured or dynamically configured according to the application scenario. The candidate seats are sorted according to the difference degree represented by the audio difference data of the filtered candidate seats. The above scheme only retains the candidate seats with small time delay and volume attenuation as potential interference sources, thereby further narrowing the target range of subsequent processing, eliminating the candidate seats that have little influence on the target seat audio collection due to long distance or obstacles, and improving the accuracy and efficiency of noise reduction processing.
[0032] In a specific implementation scenario, as a possible implementation, a difference data generation model can be pre-trained to generate audio difference data between the first audio signal and each second audio signal. Specifically, the difference data generation model can include but is not limited to an Encoder-Decoder architecture network model, etc. In order to ensure the generation accuracy of the difference data generation model as much as possible, sample first audio and sample second audio can be collected, each sample first audio is labeled with real difference data between each sample second audio, and the difference data generation model is used to obtain predicted difference data between each sample first audio and each sample second audio. Therefore, the network parameters of the difference data generation model can be adjusted based on the difference between the real difference data and the predicted difference data until the difference data generation model converges. That is, the trained difference data generation model can be used to process the first audio signal and the second audio signal to obtain the audio difference data between the first audio signal and each second audio signal. It should be noted that the specific processing process of the difference data generation model can refer to the technical details of the Encoder-Decoder architecture network model, etc., which will not be described here.
[0033] In an implementation scenario, when the seat position in the target place changes or the audio input device is replaced, the noise reduction preprocessing process for the target place can be triggered to update the latest reference seat information corresponding to each target seat. The specific processing steps can be obtained by referring to the steps in any of the above noise reduction preprocessing method embodiments. For brevity, they will not be described here.
[0034] Please refer to Figure 2 , Figure 2 is a flowchart of another embodiment of the noise reduction preprocessing method of the present application. Specifically, it can include the following steps: Step S111: Synchronize the time of the audio input device at the target seat in the target venue.
[0035] For brevity, the steps of the foregoing embodiments will not be repeated here.
[0036] Step S121: The audio output device at the target seat plays the calibration audio signal, and the audio input device at the target seat in the target venue collects signals to obtain the first audio signal and the second audio signal.
[0037] For brevity, the steps of the foregoing embodiments will not be repeated here.
[0038] Step S131: Obtain the audio difference data between the first audio signal and the second audio signal to obtain the time delay data and the volume attenuation data.
[0039] For brevity, the steps of the foregoing embodiments will not be repeated here.
[0040] The above scheme takes other seats in the target venue except the target seat as candidate seats, at least one audio input device is configured at each seat in the target venue, respectively, to meet the recording needs of the objects at each seat, based on the calibration audio signal played by the audio output device at the target seat, the first audio signal collected by the audio input device at the target seat and the second audio signal collected by the audio input device at each candidate seat are obtained, based on the audio difference data between the first audio signal and each second audio signal, a plurality of reference seats of the target seat are determined, in the application scenario of audio recording, the audio input devices at each reference seat of the target seat form a virtual noise reduction array, and the first recording signal collected by the audio input device at the target seat is configured to reduce noise by referring to the second recording signal collected by the audio input device at each reference seat. In the above scheme, in the noise reduction preprocessing stage, on the one hand, the reference seats of the target seat are determined through the calibration audio signal, so as to construct the virtual noise reduction array of the target seat, thus no additional professional equipment is needed, and the adaptability to most application scenarios is improved; on the other hand, the first audio signal and the second audio signal are obtained through the calibration audio played by the audio output device at the target seat, and the difference data is calculated, so as to improve the accuracy of the determination of the reference seats as much as possible, and then improve the effectiveness of the reference seats in the subsequent audio noise reduction process. Therefore, reliable reference data can be provided for subsequent noise reduction processing without increasing the cost of additional hardware.
[0041] Please refer to Figure 3 , Figure 3 is a flowchart of an embodiment of the audio noise reduction method of the present application. Specifically, it can include the following steps: Step S21: obtaining a first recording signal collected by an audio input device at a target seat in a target site and a second recording signal collected by an audio input device at a reference seat of the target seat.
[0042] In the embodiments of the present disclosure, the reference seats of the target seat are obtained based on the steps in any of the above noise reduction preprocessing method embodiments, which will not be described here for brevity.
[0043] In one implementation scenario, the audio input device for collecting the first recording signal at the target seat in the target site is the same as the audio input device for collecting the first audio signal at the target position in any of the above noise reduction preprocessing method embodiments, and the audio input device for collecting the second recording signal at the reference seat in the target site is the same as the audio input device for collecting the second audio signal at the reference position in any of the above noise reduction preprocessing method embodiments.
[0044] It should be noted that in the present scheme, the audio noise reduction and the noise reduction preprocessing are performed on the same target site, and the relative physical positions of the seats in the site remain unchanged. In addition, the physical parameters and deployment of the audio input, output and other hardware devices used when performing audio noise reduction are consistent with those when performing noise reduction preprocessing.
[0045] In one implementation scenario, as a possible implementation, the target site is a spoken language test room, in which a plurality of examinee seats are arranged. During the test, each examinee seat collects the test content recording of the corresponding examinee. Any examinee seat is selected as the target seat, and the reference seat of the corresponding target seat is determined based on any of the above noise reduction preprocessing method embodiments.
[0046] Step S22: performing noise reduction operation on the first recording signal based on each second recording signal to obtain a target recording signal of the target seat.
[0047] In one implementation scenario, a first recording signal collected by an audio input device at a target seat in a target venue and a second recording signal collected by an audio input device at each reference seat of the target seat are obtained, and a noise reduction operation is performed on the first recording signal based on each second recording signal to obtain a target recording signal of the target seat. In the noise reduction preprocessing stage, on the one hand, the reference seats of the target seat are established by calibrating the audio signal, so that a virtual noise reduction array of the target seat is constructed, and therefore, no professional equipment needs to be additionally added, and the adaptability to most application scenarios is improved. On the other hand, the first audio signal and the second audio signal are obtained by playing the calibration audio by the audio output device at the target seat, and the difference data is calculated, so as to improve the accuracy of the determination of the reference seat as much as possible, and then the effectiveness of the reference seat in the subsequent audio noise reduction process is improved. Therefore, in the audio noise reduction process, the first recording signal of the target seat is processed by the second recording signal of the reference seat, and the audio quality of the target recording signal is improved.
[0048] In one implementation scenario, before obtaining the first recording signal collected by the audio input device at the target seat in the target venue and the second recording signal collected by the audio input device at each reference seat of the target seat, the terminal clock of the audio input device at each seat in the target venue is synchronized to add an absolute time stamp to the recording signal of each audio input device at each seat. Specifically, the terminal clock refers to the timing device carried by the audio input device (such as a microphone, a recording terminal, etc.) at each seat in the target venue, which is used to record the time information in the running process of the device. The absolute time stamp refers to the accurate time identification based on a unified time reference attached to each piece of signal data when the audio input device collects the recording signal. The audio difference data between the target seat and the reference seat is obtained, and the audio difference data is obtained by any of the above noise reduction preprocessing methods. The operation of synchronizing the terminal clock ensures that the recording signals collected by the audio input devices at each seat have a consistent reference in the time dimension, which provides key time synchronization information for subsequent audio signal processing and noise reduction, so that the system can accurately compare and analyze the recording signals at different seats. After adding the absolute time stamp, the system can accurately mark the time point of each recording signal. At the same time, the audio difference data between the target seat and the reference seat is obtained based on the noise reduction preprocessing method, which is key information reflecting the difference characteristics between the audio signal at the target seat and the audio signal at the reference seat.
[0049] In one specific implementation scenario, the second recording signal is adjusted based on the audio difference data, the volume data of the second recording signal, and the absolute timestamp of the second recording signal to obtain a reference signal, and a noise reduction operation is performed on the first recording signal based on each reference signal to obtain target audio. The adjustment of the second recording signal may involve amplification, attenuation, filtering, and other processing of the signal, with the purpose of making the second recording signal more similar to the noise to be filtered out in the first recording signal in terms of volume, frequency characteristics, etc., so as to improve the noise reduction effect.
[0050] In one specific implementation scenario, the audio difference data includes time delay data and volume attenuation data, the volume data of the second recording signal is reduced based on the volume attenuation data, and the absolute timestamp of the second recording signal is delayed based on the time delay data, and the reference signal is obtained based on the adjusted second recording signal. This adjustment process ensures that the reference signal is highly matched with the noise component in the first recording signal in terms of volume and timing, thereby providing a more accurate basis for subsequent noise reduction processing. The introduction of time delay data enables the system to accurately simulate and compensate for the time difference in the propagation of audio signals between different seats, and the use of volume attenuation data further enhances the relevance and effectiveness of noise reduction processing. By accurately adjusting the volume of the second recording signal, the system can more effectively filter out the noise component in the first recording signal while preserving important information in the original audio signal. Therefore, the noise reduction preprocessing method can provide as reliable reference data as possible for subsequent noise reduction processing without increasing additional hardware costs.
[0051] In one specific implementation scenario, the noise reduction operation may use various noise reduction algorithms such as adaptive filtering, spectral subtraction, Wiener filtering, etc. These algorithms can intelligently remove noise components in the first recording signal based on the difference between the reference signal and the first recording signal, and retain useful speech information. The final target audio is a clear and high-quality speech signal after noise reduction processing.
[0052] Please refer to Figure 5 , Figure 5is the framework schematic diagram of an embodiment of the audio noise reduction method of the present application. In a specific implementation scenario, the spectrum of the target voice is reserved by subtracting the noise spectrum estimated from the second recording signal from the spectrum of the first recording signal to obtain the target recording signal of the target seat. Specifically, the first recording signal and the second recording signal are subjected to short-time Fourier transform to separate into amplitude spectrum and phase spectrum, the first recording signal retains the phase information for subsequent reconstruction, and the second recording signal is used for noise estimation. The power spectrum of the first recording signal is subtracted from the noise power spectrum obtained from the second recording signal, and then the phase spectrum and the enhanced amplitude spectrum of the first recording signal are used to synthesize the time-domain signal after noise reduction through inverse STFT (Short-Time Fourier Transform) to obtain the target recording signal.
[0053] Please refer to Figure 6 , Figure 6 is the framework schematic diagram of another embodiment of the audio noise reduction method of the present application. In another specific implementation scenario, the time-domain signal of the first recording signal is converted into a frequency-domain signal, the amplitude spectrum and the phase spectrum of the first recording signal are retained, and the time-domain signal of the second recording signal is converted into a frequency-domain signal, only the amplitude spectrum of the second recording signal is retained. The amplitude spectrums of the first recording signal and the second recording signal are spliced in the channel dimension to form a double-channel input, and the double-channel input is processed through a convolutional neural network (CNN). For reference, please refer to the following convolutional neural network structure reference example: # Input shape: (Batch=1, Channels=2, Freq=F, Time=T) Conv2D(in=2, out=64, kernel=(3, 3)) -> ReLU -> BatchNorm -> Conv2D(64 -> 128, kernel=(3, 3)) -> ReLU -> MaxPooling(pool_size=(1, 2)) -> Flatten() -> Dense(-> F x T) -> Reshape -> Sigmoid # Output mask shape (F, T) The mask value in the range of [0, 1] is output through Sigmoid, indicating the reservation ratio of the voice component. The mask is multiplied point by point with the amplitude spectrum of the first recording signal to reserve the target voice and suppress the noise. The phase spectrum of the first recording signal is directly used to avoid phase estimation error. The enhanced amplitude spectrum and the phase spectrum are combined to obtain the time-domain signal after noise reduction through inverse transform to obtain the target recording signal of the target seat.
[0054] In one implementation scenario, time delay compensation is performed on each second recording signal based on the first recording signal. Specifically, the time delay compensation parameter can be obtained based on any of the above noise reduction preprocessing method embodiments. Feature extraction is performed on each second recording signal after time delay compensation to obtain a plurality of audio features. Noise reduction operation is performed on the first recording signal based on each audio feature one by one, and the final audio is output as the target recording signal of the target seat.
[0055] In a specific implementation scenario, based on the time delay compensation processed second recording signals, feature extraction is performed to obtain a plurality of audio features. Specifically, when performing feature extraction on the time delay compensated second recording signals, the time domain and frequency domain analysis methods can be combined. At the time domain level, the amplitude fluctuation law of the audio is captured by calculating the short-time energy, zero-crossing rate and other parameters of the signal. At the frequency domain level, the signal is converted to the frequency domain by Fourier transform to extract the spectral distribution characteristics of the audio. The extracted features are fused to obtain the audio features of the second recording audio. Specifically, when fusing the features, methods such as weighted average, principal component analysis or deep learning can be used to comprehensively consider the information of the time domain and the frequency domain to ensure that the features can fully and accurately reflect the characteristics of the audio signal. Then, based on the audio features, a noise reduction operation is performed on the first recording signal. In the noise reduction process, by comparing the audio features of the first recording signal and the second recording signal, the noise components in the first recording signal are identified and eliminated, while the effective information of the original audio is preserved, thereby outputting a clear and high-quality target recording signal. This method not only improves the noise reduction effect, but also is suitable for audio noise reduction needs in various complex scenarios.
[0056] In a specific implementation scenario, after feature extraction is performed based on the time delay compensated second recording signals to obtain a plurality of audio features, and before performing a noise reduction operation on the first recording signal based on each audio feature to output the final audio as the target recording signal of the target seat, audio difference data between the target seat and each reference seat is obtained, and the audio difference data is obtained by any of the above noise reduction preprocessing method embodiments. Based on the difference degree represented by the audio difference data, the audio features corresponding to each audio difference data are sorted. Based on the sorting result, the operation order of the noise reduction operation is determined. Based on the operation order, a noise reduction operation is performed on the first recording signal based on each audio feature, and the final audio is output as the target recording signal of the target seat.
[0057] In another specific implementation scenario, a noise reduction operation is performed on the first recording signal based on each second recording signal, and the order of selecting the second recording signal during the noise reduction operation is not limited in the present application. The target recording signal of the target seat is obtained.
[0058] Please refer to Figure 4 , Figure 4 is the flowchart of another embodiment of the audio noise reduction method of the present application. Specifically, it can include the following steps: Step S221: Obtain the second recording signal, and delay the absolute timestamp of the second recording signal based on the time delay data.
[0059] Specifically, please refer to the steps of the preceding embodiments. For brevity, they will not be repeated here.
[0060] Step S222: based on the volume attenuation data, reduce the volume data of the second recording signal to obtain a reference signal, and the noise reduction model generates a target recording signal based on the first recording signal and the reference signal.
[0061] For brevity, the steps of the foregoing embodiments will not be repeated here.
[0062] The above scheme takes other seats in the target site except the target seat as candidate seats, at least one audio input device is respectively arranged at each seat in the target site, which is respectively used to meet the recording needs of the object at each seat, based on the calibration audio signal played by the audio output device at the target seat, the first audio signal collected by the audio input device at the target seat and the second audio signal collected by the audio input device at each candidate seat are obtained, based on the audio difference data between the first audio signal and each second audio signal, a plurality of reference seats of the target seat are determined, the first recording signal collected by the audio input device at the target seat and the second recording signal collected by the audio input device at a plurality of reference seats of the target seat are obtained, and the noise reduction operation is performed on the first recording signal based on each second recording signal to obtain the target recording signal of the target seat. In the noise reduction preprocessing stage, on the one hand, the reference seats of the target seat are established through the calibration audio signal, so as to construct the virtual noise reduction array of the target seat, and therefore the professional equipment does not need to be additionally added, and the adaptability to most application scenarios is improved. On the other hand, the first audio signal and the second audio signal are obtained through the calibration audio played by the audio output device at the target seat, and the difference data is calculated, so as to improve the accuracy of the determination of the reference seats as much as possible, and then the effectiveness of the reference seats in the subsequent audio noise reduction process is improved. Therefore, in the audio noise reduction process, the first recording signal of the target seat is processed by the second recording signal of the reference seat, and the audio quality of the target recording signal can be improved.
[0063] Please refer to Figure 7 , Figure 7 is a schematic diagram of the framework of an embodiment of the noise reduction preprocessing device 30 of the present application. As Figure 7As shown, the noise reduction preprocessing device 30 comprises a selection module 31, an acquisition module 32 and a determination module 33. The selection module 31 is configured to select other seats in the target site except the target seat as candidate seats. Each seat in the target site is provided with at least an audio input device, respectively, to meet the recording needs of the object at each seat. The acquisition module 32 is configured to acquire a first audio signal collected by the audio input device at the target seat and a second audio signal collected by the audio input device at each candidate seat based on a calibration audio signal played by the audio output device at the target seat. The determination module 33 is configured to determine a plurality of reference seats of the target seat based on audio difference data between the first audio signal and each second audio signal. In the application scenario of audio recording, the audio input devices at each reference seat of the target seat form a virtual noise reduction array, and the first recording signal collected by the audio input device at the target seat is configured to be reduced in noise by referring to the second recording signal collected by the audio input device at each reference seat.
[0064] Therefore, the noise reduction preprocessing device 30 selects other seats in the target site except the target seat as candidate seats, each seat in the target site is provided with at least an audio input device, respectively, to meet the recording needs of the object at each seat, acquires a first audio signal collected by the audio input device at the target seat and a second audio signal collected by the audio input device at each candidate seat based on a calibration audio signal played by the audio output device at the target seat, and determines a plurality of reference seats of the target seat based on audio difference data between the first audio signal and each second audio signal. In the application scenario of audio recording, the audio input devices at each reference seat of the target seat form a virtual noise reduction array, and the first recording signal collected by the audio input device at the target seat is configured to be reduced in noise by referring to the second recording signal collected by the audio input device at each reference seat. The above scheme, in the noise reduction preprocessing stage, on the one hand, establishes the reference seats of the target seat through the calibration audio signal, thereby constructing the virtual noise reduction array of the target seat, and thus does not need to additionally add professional equipment, thereby improving the adaptability to most application scenarios. On the other hand, the first audio signal and the second audio signal are acquired by the calibration audio played by the audio output device at the target seat, and the difference data is calculated, so as to improve the accuracy of the determination of the reference seats as much as possible, thereby improving the effectiveness of the reference seats in the subsequent audio noise reduction process. Therefore, the reference data as reliable as possible can be provided for the subsequent noise reduction processing without increasing the additional hardware cost.
[0065] In some disclosed embodiments, the determination module 33 further comprises a sorting module (not shown) configured to sort each candidate seat based on the difference degree represented by the audio difference data, and a seat selection module (not shown) configured to select the candidate seat whose difference degree is within a preset low position as a reference seat.
[0066] In some disclosed embodiments, the audio difference data comprises at least one of time delay data and volume attenuation data, and the noise reduction preprocessing device 30 further comprises a screening module (not shown) for screening the candidate seats with time delay data not greater than a time delay threshold, in the case that the audio difference data comprises time delay data, and screening the candidate seats with volume attenuation data not greater than an attenuation threshold, in the case that the audio difference data comprises volume attenuation data, before sorting the candidate seats based on the difference degree characterized by the audio difference data.
[0067] Please refer to Figure 8 , Figure 8 is a frame diagram of an embodiment of the audio noise reduction device 40 of the present application. As shown in Figure 8 , the audio noise reduction device 40 comprises a collection module 41 and a noise reduction module 42, the collection module 41 is configured to obtain a first recording signal collected by an audio input device at a target seat in a target place and a plurality of second recording signals respectively collected by audio input devices at a plurality of reference seats of the target seat, wherein the plurality of reference seats of the target seat are obtained based on the noise reduction preprocessing method of the first aspect described above; the noise reduction module 42 is configured to perform a noise reduction operation on the first recording signal based on each of the second recording signals to obtain a target recording signal of the target seat.
[0068] Therefore, the audio noise reduction device 40 takes other seats except the target seat in the target place as candidate seats, at least one audio input device is arranged at each seat in the target place respectively to meet the recording needs of the object at each seat, based on the calibration audio signal played by the audio output device at the target seat, the first audio signal collected by the audio input device at the target seat and the second audio signal collected by the audio input device at each candidate seat are obtained, the reference seats of the target seat are determined based on the audio difference data between the first audio signal and each second audio signal, the first recording signal collected by the audio input device at the target seat in the target place and the second recording signal collected by the audio input device at each reference seat of the target seat are obtained, and the noise reduction operation is performed on the first recording signal based on each second recording signal to obtain the target recording signal of the target seat. The above scheme, in the noise reduction preprocessing stage, on the one hand, the reference seats of the target seat are established through the calibration audio signal, so as to construct the virtual noise reduction array of the target seat, and therefore, the professional equipment does not need to be additionally added, and the adaptability to most application scenarios is improved. On the other hand, the first audio signal and the second audio signal are obtained through the calibration audio played by the audio output device at the target seat, and the difference data is calculated, so as to improve the accuracy of the determination of the reference seats as much as possible, and then the effectiveness of the reference seats in the subsequent audio noise reduction process is improved. Therefore, in the audio noise reduction process, the first recording signal of the target seat is subjected to the noise reduction processing by using the second recording signal of the reference seat, and the audio quality of the target recording signal can be improved.
[0069] In some disclosed embodiments, before obtaining the first recording signal collected by the audio input device at the target seat in the target place and the second recording signal collected by the audio input device at each reference seat of the target seat, the audio noise reduction device 40 further comprises a timestamp adding module (not shown) for synchronizing the terminal clock of the audio input device at each seat in the target place, adding an absolute timestamp to the recording signal of the audio input device at each seat respectively, and obtaining the audio difference data between the target seat and the reference seat; wherein the audio difference data is obtained by any one of the noise reduction preprocessing method embodiments; the noise reduction module 42 further comprises a reference signal module (not shown) for adjusting the second recording signal based on the audio difference data, the volume data of the second recording signal and the absolute timestamp of the second recording signal to obtain a reference signal; the noise reduction module 42 further comprises a noise reduction sub-module (not shown) for performing the noise reduction operation on the first recording signal based on each reference signal to obtain the target audio.
[0070] In some disclosed embodiments, the audio difference data comprises time delay data and volume attenuation data, the reference signal module (not shown) further comprises a difference processing module (not shown) configured to reduce the volume data of the second recording signal based on the volume attenuation data, and delay the absolute timestamp of the second recording signal based on the time delay data; the reference signal module (not shown) further comprises a reference adjustment module (not shown) configured to obtain the reference signal based on the adjusted second recording signal.
[0071] In some disclosed embodiments, the noise reduction module 42 further comprises a time delay compensation module (not shown) configured to perform time delay compensation on each second recording signal based on the first recording signal; the noise reduction module 42 further comprises a feature extraction module (not shown) configured to perform feature extraction on each time delay compensated second recording signal to obtain a plurality of audio features; the noise reduction module 42 further comprises a recording noise reduction module (not shown) configured to perform noise reduction operation on the first recording signal based on each audio feature one by one, and output the final audio as the target recording signal of the target seat.
[0072] In some disclosed embodiments, after performing feature extraction on each time delay compensated second recording signal to obtain a plurality of audio features, and before performing noise reduction operation on the first recording signal based on each audio feature one by one, and outputting the final audio as the target recording signal of the target seat, the audio noise reduction device 40 further comprises a difference acquisition module (not shown) configured to acquire audio difference data between the target seat and each reference seat; wherein the audio difference data is obtained by any of the above noise reduction preprocessing method embodiments; the audio noise reduction device 40 further comprises a feature sorting module (not shown) configured to sort the audio features corresponding to each audio difference data based on the difference degree represented by the audio difference data; the audio noise reduction device 40 further comprises an operation order determination module (not shown) configured to determine the operation order of performing noise reduction operation based on the sorting result; the recording noise reduction module (not shown) further comprises a sequential noise reduction module (not shown) configured to perform noise reduction operation on the first recording signal based on each audio feature one by one based on the operation order, and output the final audio as the target recording signal of the target seat.
[0073] Please refer to Figure 9 , Figure 9 is a frame schematic diagram of an embodiment of the electronic device 50. The electronic device 50 at least comprises a memory 51 and a processor 52 coupled with each other, the memory 51 at least stores program instructions, and the processor 52 is configured to execute the program instructions to realize the steps in any of the above noise reduction preprocessing method or audio noise reduction method embodiments. For details, please refer to the foregoing disclosed embodiments, which will not be repeated here. It should be noted that the electronic device 50 can include but is not limited to learning machines, office books, smart large screens, etc. The specific type of the electronic device 50 is not limited here.
[0074] Specifically, the processor 52 is configured to control itself and the memory 51 to implement the steps in any of the above-mentioned noise reduction preprocessing methods or audio noise reduction method embodiments. The processor 52 can also be referred to as a CPU (Central Processing Unit). The processor 52 can be an integrated circuit chip with processing capability. The processor 52 can also be a general-purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 52 can be implemented by an integrated circuit chip together.
[0075] The above scheme, the electronic device 50 takes other seats in the target place except the target seat as candidate seats, at least an audio input device is configured at each seat in the target place respectively, which is used to meet the recording needs of the object at each seat, based on the calibration audio signal played by the audio output device at the target seat, the first audio signal collected by the audio input device at the target seat and the second audio signal collected by the audio input device at each candidate seat are obtained, based on the audio difference data between the first audio signal and each second audio signal, a plurality of reference seats of the target seat are determined, the first recording signal collected by the audio input device at the target seat in the target place and the second recording signal collected by the audio input device at a plurality of reference seats of the target seat are obtained, and the noise reduction operation is performed on the first recording signal based on each second recording signal to obtain the target recording signal of the target seat. In the noise reduction preprocessing stage, on the one hand, the reference seats of the target seat are established through the calibration audio signal, so as to construct the virtual noise reduction array of the target seat, thus without additional professional equipment, the adaptability to most application scenarios is improved. On the other hand, the first audio signal and the second audio signal are obtained through the calibration audio played by the audio output device at the target seat, and the difference data is calculated, so as to improve the accuracy of the determination of the reference seats as much as possible, and then the effectiveness of the reference seats in the subsequent audio noise reduction process is improved. Therefore, in the audio noise reduction process, the first recording signal of the target seat is processed by the second recording signal of the reference seats, so as to improve the audio quality of the target recording signal.
[0076] Please refer to Figure 10 , Figure 10is a schematic diagram of a framework of an embodiment of the computer readable storage medium 60 of the present application. The computer readable storage medium 60 stores program instructions 61 capable of being executed by a processor, the program instructions 61 being used to implement the steps in any of the above-described noise reduction preprocessing method or audio noise reduction method embodiments.
[0077] The above scheme, the computer readable storage medium 60 takes other seats in the target site except the target seat as candidate seats, at least respectively configures an audio input device at each seat in the target site, respectively used to meet the recording needs of the object on each seat, based on the calibration audio signal played by the audio output device at the target seat, obtains the first audio signal collected by the audio input device at the target seat and the second audio signal collected by the audio input device at each candidate seat, based on the audio difference data between the first audio signal and each second audio signal, determines a plurality of reference seats of the target seat, obtains the first recording signal collected by the audio input device at the target seat in the target site and the second recording signal collected by the audio input device at a plurality of reference seats of the target seat, respectively, based on each second recording signal, performs noise reduction operation on the first recording signal to obtain the target recording signal of the target seat. The above scheme, in the noise reduction preprocessing stage, on the one hand, the reference seats of the target seat are established through the calibration audio signal, so as to construct a virtual noise reduction array of the target seat, thus without the need of additionally adding professional equipment, the adaptability to most application scenarios is improved; on the other hand, through the calibration audio played by the audio output device at the target seat, the first audio signal and the second audio signal are obtained, and the difference data is calculated, so as to improve the accuracy of the determination of the reference seats as much as possible, and then the effectiveness of the reference seats in the subsequent audio noise reduction process is improved. Therefore, in the audio noise reduction process, the first recording signal of the target seat is processed by the second recording signal of the reference seat, so as to improve the audio quality of the target recording signal.
[0078] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or contains modules which can be used to execute the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be described here.
[0079] The above description of each embodiment tends to emphasize the differences between each embodiment, and the same or similar parts can be mutually referred to. For the sake of brevity, it will not be described here.
[0080] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other manners. For example, the division of the apparatus embodiments described above is merely an example, and the division of the modules or units can be different, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between the units can be indirect coupling or communication connection through some interfaces, and can be electrical, mechanical or other forms.
[0081] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.
[0082] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can be physically present separately, or two or more units can be integrated into one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0083] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (processor) to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0084] If the technical solution of the present application involves personal information, the product applying the technical solution of the present application has clearly informed the personal information processing rules before processing the personal information and obtained the personal independent consent. If the technical solution of the present application involves sensitive personal information, the product applying the technical solution of the present application has obtained the personal independent consent before processing the sensitive personal information and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that it has entered the personal information collection range and will collect personal information. If the individual voluntarily enters the collection range, it is considered to agree to collect personal information. Or on the device for processing personal information, through the pop-up information or by asking the individual to upload his / her personal information, the individual's authorization is obtained under the condition of using obvious mark / information to inform the personal information processing rules. The personal information processing rules can include personal information processor, personal information processing purpose, processing method and personal information type, etc.
Claims
1. A noise reduction preprocessing method, characterized in that, include: Other seats in the target location besides the target seat are selected as candidate seats; wherein, each seat in the target location is equipped with at least one audio input device to meet the recording needs of the person in each seat. Based on the calibration audio signal played by the audio output device at the target seat, the first audio signal collected by the audio input device at the target seat and the second audio signal collected by the audio input devices at each of the candidate seats are obtained. Based on the audio difference data between the first audio signal and each of the second audio signals, a plurality of reference seats for the target seat are determined; wherein, in the application scenario of audio recording, the audio input devices at each of the reference seats of the target seat form a virtual noise reduction array, and the first recording signal collected by the audio input device at the target seat is configured to perform noise reduction with reference to the second recording signals collected by the audio input devices at the plurality of reference seats.
2. The method according to claim 1, characterized in that, The step of determining several reference seats based on the audio difference data between the first audio signal and each of the second audio signals includes: Based on the degree of difference represented by the audio difference data, each candidate seat is ranked. The candidate seats whose difference level is within a preset low range are selected as the reference seats.
3. The method according to claim 2, characterized in that, The audio difference data includes at least one of time delay data and volume attenuation data. Before ranking the candidate seats based on the degree of difference represented by the audio difference data, the method further includes: When the audio difference data includes latency data, candidate seats whose latency data is not greater than a latency threshold are selected; when the audio difference data includes volume attenuation data, candidate seats whose volume attenuation data is not greater than an attenuation threshold are selected. The process of ranking the candidate seats based on the degree of difference represented by the audio difference data includes: The candidate seats are ranked based on the degree of difference represented by the audio difference data of the selected candidate seats.
4. An audio noise reduction method, characterized in that, include: The method involves acquiring a first recording signal collected by an audio input device at a target seat in a target location and second recording signals collected by audio input devices at several reference seats at the target seat; wherein the several reference seats at the target seat are obtained based on the noise reduction preprocessing method described in any one of claims 1 to 3. The first recording signal is subjected to noise reduction operation based on each of the second recording signals to obtain the target recording signal of the target seat.
5. The method according to claim 4, characterized in that, Before acquiring the first recording signal collected by the audio input device at the target seat in the target location and the second recording signals collected by the audio input devices at several reference seats of the target seat, the method further includes: The terminal clocks of the audio input devices at each seat in the target location are synchronized to add absolute timestamps to the recording signals of the audio input devices at each seat, and audio difference data between the target seat and the reference seat is obtained; wherein, the audio difference data is obtained by the noise reduction preprocessing method; The step of performing noise reduction on the first recording signal based on each of the second recording signals to obtain the target recording signal for the target seat includes: Based on the audio difference data, the volume data of the second recording signal, and the absolute timestamp of the second recording signal, the second recording signal is adjusted to obtain a reference signal; The first recording signal is subjected to noise reduction operation based on each of the reference signals to obtain the target audio.
6. The method according to claim 5, characterized in that, The audio difference data includes time delay data and volume attenuation data. The step of adjusting the second recording signal based on the audio difference data, the volume data of the second recording signal, and the absolute timestamp of the second recording signal to obtain a reference signal includes: Based on the volume attenuation data, the volume of the second recording signal is reduced, and based on the delay data, the absolute timestamp of the second recording signal is delayed; The reference signal is obtained based on the adjusted second recording signal.
7. The method according to claim 4, characterized in that, The step of performing noise reduction on the first recording signal based on each of the second recording signals to obtain the target recording signal for the target seat includes: Time delay compensation is performed on each of the second recording signals based on the first recording signal; Based on the time delay compensation of each of the second recording signals, feature extraction is performed to obtain several audio features; The noise reduction operation is performed on the first recording signal based on each of the aforementioned audio features, and the final audio is output as the target recording signal for the target seat.
8. The method according to claim 7, characterized in that, After extracting features from each of the second recording signals based on time delay compensation to obtain several audio features, and before performing noise reduction operations on each of the first recording signals based on each of the audio features to output the final audio as the target recording signal for the target seat, the method further includes: Acquire audio difference data between the target seat and each of the reference seats; wherein the audio difference data is obtained by the noise reduction preprocessing method; Based on the degree of difference represented by the audio difference data, the audio features corresponding to each audio difference data are sorted. Based on the sorting results, the order in which the noise reduction operation is performed is determined; The step of performing noise reduction operations on the first recording signal based on each of the aforementioned audio features, and outputting the final audio as the target recording signal for the target seat, includes: Based on the operation sequence, noise reduction operations are performed on the first recording signal one by one based on each of the audio features, and the final audio is output as the target recording signal of the target seat.
9. A noise reduction preprocessing device, characterized in that, include: The selection module is used to select other seats in the target venue besides the target seat as candidate seats; wherein, each seat in the target venue is equipped with at least one audio input device to meet the recording needs of the object in each seat; The acquisition module is used to acquire, based on the calibration audio signal played by the audio output device at the target seat, a first audio signal collected by the audio input device at the target seat and a second audio signal collected by the audio input devices at each of the candidate seats; The determining module is used to determine a plurality of reference seats for the target seat based on the audio difference data between the first audio signal and each of the second audio signals; wherein, in the application scenario of audio recording, the audio input devices at each of the reference seats of the target seat form a virtual noise reduction array, and the first recording signal collected by the audio input device at the target seat is configured to perform noise reduction with reference to the second recording signals collected by the audio input devices at the plurality of reference seats.
10. An audio noise reduction device, characterized in that, include: The acquisition module is used to acquire a first recording signal acquired by an audio input device at a target seat in a target location and a second recording signal acquired by audio input devices at several reference seats of the target seat respectively; wherein the several reference seats of the target seat are obtained based on the noise reduction preprocessing method according to any one of claims 1 to 3; The noise reduction module is used to perform noise reduction operation on the first recording signal based on each of the second recording signals to obtain the target recording signal of the target seat.
11. An electronic device, characterized in that, The device includes a memory and a processor coupled to each other, wherein the memory stores program instructions and the processor executes the program instructions to implement the noise reduction preprocessing method according to any one of claims 1 to 3, or to implement the audio noise reduction method according to any one of claims 4 to 8.
12. A computer-readable storage medium having program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, they implement the noise reduction preprocessing method according to any one of claims 1 to 3, or the audio noise reduction method according to any one of claims 4 to 8.