A microphone modeling method, apparatus, device, and medium

Through a multi-stage audio data processing workflow, including convolution processing, equalization modeling, transient modeling, nonlinear distortion modeling, and proximity effect simulation, a highly realistic target microphone model is generated, which solves the problem of insufficient flexibility and adaptability in microphone modeling in existing technologies and achieves accurate simulation of microphone timbre and characteristics.

CN122138115APending Publication Date: 2026-06-02SHENZHEN HOLLYLAND TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN HOLLYLAND TECH CO LTD
Filing Date
2026-01-30
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing microphone modeling techniques rely on machine learning and neural networks, resulting in black-box models with poor flexibility and adaptability, making it difficult to meet the needs of accurate simulation of microphone timbre and characteristics in different scenarios.

Method used

A highly realistic target microphone model is generated through a multi-stage audio data processing workflow, including convolution processing, equalization modeling, transient modeling, nonlinear distortion modeling, and proximity effect simulation.

Benefits of technology

It achieves accurate simulation of microphone timbre and characteristics, improves the quality and effect of audio processing, and enhances the flexibility and adaptability of microphone modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122138115A_ABST
    Figure CN122138115A_ABST
Patent Text Reader

Abstract

This invention relates to the field of audio processing technology, and more particularly to a microphone modeling method, apparatus, device, and medium. The method involves acquiring raw audio data collected by a source microphone within a preset time period, performing convolution processing on the raw audio data to obtain first intermediate audio data, performing equalization modeling processing on the first intermediate audio data to obtain second intermediate audio data, performing transient modeling processing on the second intermediate audio data to obtain third intermediate audio data, performing nonlinear distortion modeling processing on the third intermediate audio data to obtain fourth intermediate audio data, performing proximity effect simulation processing on the fourth intermediate audio data to obtain fifth intermediate audio data, and generating a target microphone model based on the fifth intermediate audio data. Therefore, this application, through a multi-stage audio data processing flow, can accurately achieve microphone modeling, thereby generating a highly realistic target microphone model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio processing technology, and in particular to a microphone modeling method, apparatus, device, and medium. Background Technology

[0002] Microphone modeling is an audio processing technique designed to simulate the timbre and characteristics of classic microphones. Classic microphones typically have unique frequency responses, dynamic characteristics, and harmonic distortion, which give them a unique "retro" timbre and are widely used in music production, broadcasting, and live sound.

[0003] However, existing microphone modeling techniques have significant drawbacks. They primarily rely on supervised training of machine learning and neural networks, making them largely black-box models with unclear underlying principles and a limited number of adjustable parameters. This results in poor modeling flexibility and adaptability, making it difficult to accurately simulate microphone timbre and characteristics in different scenarios. Ultimately, this limits the types of ideal microphones available to users. Therefore, how to accurately achieve microphone modeling to meet the precise simulation needs of microphone timbre and characteristics in various scenarios is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] Therefore, in order to address the aforementioned technical problems, this invention provides a microphone modeling method, apparatus, device, and medium that can accurately model microphones to meet the need for precise simulation of microphone timbre and characteristics in different scenarios.

[0005] A first aspect of this application provides a microphone modeling method, the microphone modeling method comprising: Acquire the raw audio data collected by the source microphone within a preset time period; The original audio data is convolved to obtain the first intermediate audio data; The first intermediate audio data is subjected to equalization modeling processing to obtain the second intermediate audio data; The second intermediate audio data is subjected to transient modeling processing to obtain the third intermediate audio data; The third intermediate audio data is subjected to nonlinear distortion modeling processing to obtain the fourth intermediate audio data; The fourth intermediate audio data is subjected to proximity effect simulation processing to obtain the fifth intermediate audio data, and a target microphone model is generated based on the fifth intermediate audio data.

[0006] A second aspect of this application provides a microphone modeling apparatus, the microphone modeling apparatus comprising: The acquisition module is used to acquire the raw audio data collected by the source microphone within a preset time period; The processing module is used to perform convolution processing on the original audio data to obtain first intermediate audio data; perform equalization modeling processing on the first intermediate audio data to obtain second intermediate audio data; perform transient modeling processing on the second intermediate audio data to obtain third intermediate audio data; and perform nonlinear distortion modeling processing on the third intermediate audio data to obtain fourth intermediate audio data. The generation module is used to perform proximity effect simulation processing on the fourth intermediate audio data to obtain the fifth intermediate audio data, and generate a target microphone model based on the fifth intermediate audio data.

[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the microphone modeling method as described in the first aspect.

[0008] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, implements the microphone modeling method as described in the first aspect.

[0009] In summary, this invention provides a microphone modeling method, apparatus, device, and medium. It acquires raw audio data collected by a source microphone within a preset time period, performs convolution processing on the raw audio data to obtain first intermediate audio data, performs equalization modeling processing on the first intermediate audio data to obtain second intermediate audio data, performs transient modeling processing on the second intermediate audio data to obtain third intermediate audio data, performs nonlinear distortion modeling processing on the third intermediate audio data to obtain fourth intermediate audio data, performs proximity effect simulation processing on the fourth intermediate audio data to obtain fifth intermediate audio data, and generates a target microphone model based on the fifth intermediate audio data. As can be seen, this application, through a multi-stage audio data processing flow, can accurately achieve microphone modeling, thereby generating a highly realistic target microphone model, meeting the precise simulation requirements for microphone timbre and characteristics in different scenarios, and further improving the quality and effect of audio processing. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a microphone modeling method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a microphone modeling device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0013] It should be understood that, when used in this specification and the appended claims, terms include indicating the presence of the described feature, integral, step, operation, element and / or component, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0014] It should also be understood that the terms used in this specification and the appended claims refer to any combination of one or more of the associated listed items and all possible combinations, and include such combinations.

[0015] As used in this specification and the appended claims, terms if can be interpreted in context as when... or once or in response to determination. Similarly, the phrase if determined or if matched to [described condition or event] can be interpreted in context as once determined or in response to determination or once matched to [described condition or event] or in response to matching to [described condition or event].

[0016] Furthermore, in the description of this invention and the appended claims, the terms first, second, third, etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0017] References to one or more embodiments described in this specification mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the invention. Therefore, phrases appearing in different parts of this specification as referring to one embodiment, some embodiments, some other embodiments, and others do not necessarily refer to the same embodiment, but rather mean one or more, but not all, embodiments, unless otherwise specifically emphasized. The terms include, comprise, have, and variations thereof mean including but not limited to, unless otherwise specifically emphasized.

[0018] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0019] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0020] See Figure 1 This is a flowchart illustrating a microphone modeling method according to an embodiment of the present invention, as shown below. Figure 1 As shown, this microphone modeling method can be implemented through the following steps.

[0021] S101: Acquire the raw audio data collected by the source microphone within a preset time period.

[0022] In one implementation, an audio acquisition device connected to a source microphone collects all audio signals acquired by the microphone within a preset time period, such as several seconds to several minutes, to ensure a sufficient length of audio samples are collected for subsequent modeling and analysis. During the collection process, the audio acquisition device ensures signal integrity and accuracy, avoiding signal loss or distortion. Simultaneously, the audio acquisition device possesses high-precision sampling capabilities, accurately sampling audio signals at a preset sampling rate to obtain high-quality raw audio data. This raw audio data serves as a crucial foundation for subsequent modeling and analysis. Furthermore, these audio signals encompass sound information of various frequencies and amplitudes. After processing, the raw audio data is preprocessed, including but not limited to noise reduction and filtering, to remove potential noise and interference signals, thereby improving the quality and purity of the audio data and providing a more accurate and reliable data foundation for subsequent modeling and analysis.

[0023] S102: Perform convolution processing on the original audio data to obtain the first intermediate audio data.

[0024] In one implementation, the original audio data is convolved to obtain first intermediate audio data. During convolution, a specific convolution kernel is used to operate on the original audio data. This kernel is carefully designed and selected to extract specific feature information from the original audio data. Through convolution, different frequency components and temporal features in the original audio data can be effectively separated and extracted, resulting in first intermediate audio data containing rich feature information. Compared to the original audio data, this first intermediate audio data is more prominent and explicit in its feature representation, providing more valuable data support for subsequent analysis and processing. Specifically, this convolution processing can use a pre-set unit impulse response of the target microphone to perform convolution operations with the original audio data to simulate the effect of a specific acoustic environment on the sound. For example, a general reverb effect can be selected, its impulse response loaded, and then convolved with the original audio data to replicate the transient response and timbre of the target microphone, thereby introducing a certain sense of space into the first intermediate audio data. As can be seen, the above steps can effectively perform preliminary convolution processing on the original audio data, making the processed audio data more conducive to the subsequent training and modeling of the microphone model, thereby improving the performance and effectiveness of the entire microphone modeling method.

[0025] S103: Perform equalization modeling processing on the first intermediate audio data to obtain the second intermediate audio data.

[0026] In one implementation, second intermediate audio data is obtained by performing equalization modeling on the first intermediate audio data. This equalization modeling process mainly relies on the energy distribution characteristics of the audio signal in different frequency bands, and uses a specific equalization algorithm to dynamically adjust the first intermediate audio data. This adjustment aims to optimize the spectral characteristics of the audio, making the processed second intermediate audio data closer to the actual output characteristics of the target microphone in terms of frequency response curve and timbre. Therefore, through the above steps, potential issues such as missing or overly emphasized frequency bands in the first intermediate audio data can be further eliminated, providing a more accurate and comprehensive data foundation for subsequent microphone model training.

[0027] S104: Perform transient modeling processing on the second intermediate audio data to obtain the third intermediate audio data.

[0028] In one implementation, the third intermediate audio data is obtained by performing transient modeling processing on the second intermediate audio data. This transient modeling process primarily focuses on the transient characteristics of the audio signal, i.e., the changing characteristics of the audio signal over a short period of time. By performing detailed transient analysis on the second intermediate audio data and employing appropriate transient modeling algorithms, transient components in the audio signal can be captured, such as the percussion sounds of percussion instruments or plosive sounds in speech. Furthermore, this transient modeling processing can utilize a compressor with fixed attack and release times to perform dynamic range compression on the second intermediate audio data. For example, a faster attack time can be set to respond to the initial impact of the sound, and a medium release time can be set to smooth the sound decay, thereby simulating the dynamic response characteristics of a microphone in the third intermediate audio data. It is evident that this processing can preserve the original dynamic range of the audio signal while enhancing or optimizing the performance of transient components, making the processed third intermediate audio data more closely resemble the actual output characteristics of the target microphone in terms of transient response.

[0029] S105: Perform nonlinear distortion modeling processing on the third intermediate audio data to obtain the fourth intermediate audio data.

[0030] In one implementation, a fourth intermediate audio data is obtained by performing nonlinear distortion modeling on the third intermediate audio data. This nonlinear distortion modeling aims to simulate the nonlinear distortion characteristics that a microphone may introduce during audio signal processing. This distortion may originate from factors such as the microphone's internal electronic components, circuit design, or material properties. During processing, a specific nonlinear distortion modeling algorithm is used to meticulously analyze and simulate the third intermediate audio data to generate the fourth intermediate audio data that incorporates these nonlinear distortion characteristics. Specifically, an adjustable scheme for odd and even harmonics can be used, combined with tanh distortion (simulating basic tube distortion) to process the third intermediate audio data. This provides users with a more controllable operation and a more retro listening experience. Alternatively, a simple hard or soft clipping algorithm can be used. When the amplitude of the audio signal exceeds a certain threshold, it is clipped, thus introducing harmonic distortion. For example, a fixed clipping threshold can be set, causing signals exceeding this threshold to be truncated, simulating the distortion effect produced by the microphone preamplifier under overload. As can be seen, this processing method can make the final audio output closer to the output characteristics of an actual microphone under the same conditions, thereby improving the accuracy and reliability of microphone modeling.

[0031] S106: Perform proximity effect simulation processing on the fourth intermediate audio data to obtain the fifth intermediate audio data, and generate a target microphone model based on the fifth intermediate audio data.

[0032] In one implementation, the fifth intermediate audio data is obtained by performing proximity effect simulation processing on the fourth intermediate audio data. This proximity effect simulation primarily targets the low-frequency boost phenomenon that occurs when audio data is picked up at close range. Specifically, a specific filter or proximity effect transfer function can be designed to enhance the low-frequency components of the fourth intermediate audio data, simulating the low-frequency response characteristics of an actual microphone during close-range pickup. This can be achieved using a simple low-frequency boost filter that provides a fixed gain boost to the low-frequency components at a fixed frequency point. For example, a low-profile filter that boosts the low-frequency response by 3dB below 100Hz can be designed to simulate the enhanced low-frequency response of a microphone during close-range pickup.

[0033] Furthermore, based on the acquired fifth intermediate audio data, a more accurate and robust target microphone model is generated. This model can more accurately simulate the pickup characteristics of a real microphone in various environments, especially in close-range pickup, where the simulation of low-frequency response is more precise. By continuously optimizing and adjusting the filter parameters in the proximity effect simulation processing, the adaptability and flexibility of the target microphone model can be further improved, enabling it to meet the audio processing needs of different scenarios. Through this processing, the generated fifth intermediate audio data can more closely resemble the microphone pickup effect in real-world scenarios, and the generated target microphone model can thus provide a more accurate reference for actual pickup conditions. This achieves precise simulation of microphone timbre and characteristics, overcoming the limitations of traditional black-box models, such as unclear working principles and insufficient parameter adjustability. This significantly improves the flexibility and adaptability of microphone modeling, meeting the diverse needs for microphone timbre and characteristics in different application scenarios.

[0034] In summary, this invention provides a microphone modeling method, apparatus, device, and medium. It acquires raw audio data collected by a source microphone within a preset time period, performs convolution processing on the raw audio data to obtain first intermediate audio data, performs equalization modeling processing on the first intermediate audio data to obtain second intermediate audio data, performs transient modeling processing on the second intermediate audio data to obtain third intermediate audio data, performs nonlinear distortion modeling processing on the third intermediate audio data to obtain fourth intermediate audio data, performs proximity effect simulation processing on the fourth intermediate audio data to obtain fifth intermediate audio data, and generates a target microphone model based on the fifth intermediate audio data. As can be seen, this application, through a multi-stage audio data processing flow, can accurately achieve microphone modeling, thereby generating a highly realistic target microphone model, meeting the precise simulation requirements for microphone timbre and characteristics in different scenarios, and further improving the quality and effect of audio processing.

[0035] In one embodiment, specifically in step S102, which involves convolutional processing of the original audio data to obtain the first intermediate audio data, the following steps are included: Acquire unit impulse response data corresponding to the target microphone, wherein the unit impulse response data is impulse response data collected by the target microphone within a preset time period to characterize its acoustic properties; The original audio data and the unit impulse response data are convolved to obtain the first intermediate audio data.

[0036] Specifically, the unit impulse response data of the target microphone is acquired in advance. This unit impulse response data is core data characterizing the acoustic properties of a linear time-invariant system (e.g., a microphone). For a microphone, this data comprehensively reflects its frequency response, phase characteristics, and transient behavior. The acquisition method involves inputting a known, approximately ideal impulse signal (such as a short pulse or sweep signal) into the target microphone in a controlled acoustic environment (e.g., an anechoic chamber) and recording the microphone's output. By appropriately processing the recorded output signal (e.g., deconvolution if the input is a sweep signal), impulse response data characterizing the target microphone's acoustic properties can be obtained. This impulse response data is acquired within a preset time period to ensure data consistency and integrity. Furthermore, it can be calculated and generated based on the target microphone's known acoustic parameters or design specifications using acoustic simulation software or mathematical models.

[0037] Furthermore, after acquiring the unit impulse response data, the original audio data and the unit impulse response data are convolved. This mainly combines the two signals to generate a third signal, which indicates how the shape of one signal is modified by the other. In audio processing, convolving the original audio signal with the microphone's unit impulse response data is equivalent to recording the original audio signal through the microphone. Specifically, this can be achieved directly using a time-domain convolution algorithm, which involves weighted summation of each sample point of the original audio data with the unit impulse response data; or, more efficiently, through frequency-domain convolution, which involves performing a Fast Fourier Transform (FFT) on the original audio data and the unit impulse response data separately, multiplying them in the frequency domain, and then performing an Inverse Fast Fourier Transform (IFFT) on the result back to the time domain to obtain the first intermediate audio data. The above technical solution can accurately capture the inherent linear acoustic characteristics of the target microphone, thereby simulating the acoustic performance of the target microphone. This ensures that the basic timbre and acoustic characteristics of the target microphone can be reproduced with high fidelity in the initial stage of microphone modeling, providing an accurate and reliable acoustic basis for subsequent data processing steps and significantly improving the realism and accuracy of the final generated target microphone model.

[0038] In one embodiment, specifically in step S103, which involves performing equalization modeling on the first intermediate audio data to obtain the second intermediate audio data, the following steps are included: Acquire audio data from the same source microphone and the target microphone simultaneously in the same scene; Fourier transforms are performed on the source audio data and the first intermediate audio data respectively to calculate the spectral amplitude difference curve; Equalization parameters are generated based on the number of filters selected by the user and the spectral amplitude difference curve. Based on the equalization parameters, the first intermediate audio data is subjected to equalization filtering to obtain the second intermediate audio data.

[0039] Specifically, by placing the source microphone and target microphone in the same acoustic environment, such as a professional recording studio or anechoic chamber, and simultaneously playing a test signal (such as powder noise, a swept-frequency signal, or human voice), a high degree of consistency in content and environment can be ensured between the two acquired audio data segments. This synchronous acquisition can be achieved using a multi-channel audio interface or precisely synchronized recording equipment. Then, Fourier transforms are performed on the source audio data and the first intermediate audio data respectively to calculate the spectral amplitude difference curve. The Fourier transform is a mathematical tool that converts a time-domain signal into a frequency-domain signal, revealing the energy distribution of the signal at different frequencies. By performing a Fast Fourier Transform (FFT) on the source audio data (representing the actual frequency response of the target microphone) and the first intermediate audio data (representing the convolutional response of the source microphone), their respective spectral amplitude information can be obtained. The audio data is processed by frame segmentation, that is, a window function (such as Hanning window) is applied to each frame, then FFT is performed, and the results of multiple frames are averaged to obtain the spectrum amplitude difference curve. The spectrum amplitude difference curve can be obtained by subtracting the spectrum amplitude of the first intermediate audio data (or the logarithm of the ratio of the two) from the spectrum amplitude of the target microphone. This curve intuitively represents the gain or attenuation required at different frequencies, thereby quantifying the difference in frequency response between the two.

[0040] Furthermore, based on the number of filters selected by the user and the spectral amplitude difference curve, equalization parameters are generated. The filter types include at least one of shelf filter, bell filter, lowpass filter, and highpass filter. The equalization parameters include center / cutoff frequency, Q value, gain, and slope. That is, the user can select different numbers of filters according to actual needs; for example, for a parametric equalizer, 3, 7, or more bands can be selected. The system can use optimization algorithms, such as least squares or genetic algorithms, to determine the center frequency, bandwidth (Q value), and gain of each filter based on the spectral amplitude difference curve. For example, if the difference curve shows insufficient gain in a certain frequency range, filter parameters that boost the gain in that frequency range will be generated to compensate for this gap. Then, based on the equalization parameters, corresponding digital filters (e.g., IIR or FIR filters) are constructed, and the first intermediate audio data is input into these filters. After filtering, the output audio data becomes the second intermediate audio data. Its frequency response characteristics are closer to those of the target microphone, thus achieving accurate simulation of the target microphone's timbre. This primarily involves finding a set of filter parameters so that the overall frequency response of these cascaded filters matches the difference curve as closely as possible. Mean squared error (MSE) is typically used as the loss function to measure the difference. Optimization algorithms employ advanced gradient optimization algorithms such as L-BFGS-B to quickly find the optimal parameter combination; genetic algorithms and particle swarm optimization can also be used for this type of problem. Through this approach, the second intermediate audio data can more realistically simulate the timbre characteristics of the target microphone in terms of frequency response, significantly improving the accuracy and simulation level of microphone modeling and solving the problem of model distortion caused by the lack of accurate references in traditional equalization processing.

[0041] In one embodiment, specifically in step S104, which involves performing transient modeling processing on the second intermediate audio data to obtain the third intermediate audio data, the following steps are included: The second intermediate audio data is processed by fast envelope processing and slow envelope processing respectively using a transient shaper to obtain fast envelope processing results and slow envelope processing results; Calculate the difference between the fast envelope processing result and the slow envelope processing result; Based on the subtraction value, it is determined whether the data status signal is in the onset state or the holding state; If the data status signal is in the start-up state, then obtain the start-up gain parameter; If the data status signal is in a sustained tone state, then obtain the sustained tone gain parameter; The target gain coefficient is determined based on the onset gain parameter or the sustain gain parameter; The modulated audio signal is determined based on the target gain coefficient and the second intermediate audio data; The modulated audio signal is processed using a compressor to obtain third intermediate audio data.

[0042] In one implementation, a transient shaper is used to perform fast and slow envelope processing on the second intermediate audio data to obtain fast and slow envelope processing results, respectively. The transient shaper simulates the unique way the target microphone responds to transient sounds. Fast envelope processing focuses on instantaneous changes in the signal, with a fast response time, capturing the onset of the sound; while slow envelope processing focuses on the average or long-term amplitude changes of the signal, with a slow response time, reflecting the sustain or decay of the sound. By performing these two envelope processing simultaneously, a multi-dimensional view of the dynamic characteristics of the audio signal can be obtained, thus more accurately identifying the transient phases of the sound. The subtraction value between the fast and slow envelope processing results is calculated. This subtraction value is the difference between the fast and slow envelope processing results. Based on the subtraction value, it is determined whether the data status signal is in an onset or sustained state. One or more thresholds can be set. When the subtraction value exceeds a certain positive threshold, it is determined to be in an onset state; when the subtraction value is below a certain negative threshold or within a small range, it is determined to be in a sustained state. For example, it can be determined whether the subtraction value is greater than or less than 0. When the subtraction value is greater than 0, it can be determined that the microphone is in an onset state; otherwise, it is in a sustained state. If the data status signal is in an onset state, the onset gain parameter is obtained; if the data status signal is in a sustained state, the sustained gain parameter is obtained. The onset gain parameter and the sustained gain parameter can be used to control the corresponding state intensity (the sustained gain parameter is set to 1 in the onset state, and the onset gain parameter is set to 1 in the sustained state). These parameters are then stored in a lookup table for subsequent querying.

[0043] Furthermore, based on the attack gain parameter or the sustain gain parameter, the target gain coefficient is determined. This step transforms the abstract gain parameter into a practically operable gain factor. If the microphone is in the attack state, the target gain coefficient equals the attack gain parameter multiplied by 1; if it is in the sustain state, the target gain coefficient equals the sustain gain parameter multiplied by 1. Then, based on the obtained target gain coefficient, it is multiplied by the amplitude of the second intermediate audio data at the same moment, completing the dynamic gain modulation of the second intermediate audio data to obtain the modulated audio signal. The modulated audio signal is then processed by a compressor to obtain the third intermediate audio data. This compressor is a dynamic range processor that adjusts its gain based on the input signal level, i.e., multiplying the target gain coefficient by the second intermediate audio data, and then using compression to further improve the transient response, resulting in the third intermediate audio data. This scheme overcomes the limitations of single gain adjustment in traditional transient processing, enabling the generated third intermediate audio data to more finely and accurately simulate the real response of the target microphone in terms of transient characteristics, thereby significantly improving the realism and accuracy of the microphone model.

[0044] In one embodiment, specifically in step S105, which involves performing nonlinear distortion modeling on the third intermediate audio data to obtain the fourth intermediate audio data, the following steps are included: Basic vacuum tube distortion data is generated using a distortion algorithm based on the tanh function, and the distortion saturation is adjusted using the drive parameter. The formula for the basic vacuum tube distortion data is as follows: Where x represents the third intermediate audio data, and drive represents the saturation control parameter; Harmonic generation processing was performed on the basic vacuum tube distortion data using a scheme where odd and even harmonics are individually adjustable, to obtain odd and even harmonic data after adjusting the wet / dry ratio. The odd harmonic data after the dry-to-wet ratio adjustment is superimposed with the even harmonic data to obtain the fourth intermediate audio data.

[0045] In one implementation, basic tube distortion data is generated using a distortion algorithm based on the tanh function, and the distortion saturation is adjusted using the drive parameter. The formula for this basic tube distortion data is as follows: Here, 'x' represents the third intermediate audio data, 'drive' represents the saturation control parameter, and the tanh function (hyperbolic tangent function), as an S-shaped curve, can well simulate this gradual saturation characteristic. When the input signal 'x' is multiplied by the 'drive' parameter and then mapped through the tanh function, the process of the signal being amplified in the vacuum tube circuit and gradually entering saturation can be simulated. As a saturation control parameter, the larger the value of the 'drive' parameter, the faster the signal enters the saturation region, and the stronger the distortion, thus allowing the user or system to adjust the intensity of the distortion effect as needed.

[0046] Furthermore, by utilizing a scheme where odd and even harmonics are individually adjustable, harmonic generation processing is performed on the basic tube distortion data to obtain odd and even harmonic data with adjusted wet / dry ratio. This is achieved by analyzing the signal spectrum through Fourier transform to identify and extract harmonics at specific frequencies, or by designing specific nonlinear functions. Then, the odd and even harmonics are subjected to gain, attenuation, or filtering processing respectively to finely adjust their proportions in the final distortion effect. Finally, the processed harmonic data (wet signal) is mixed with the original signal or unprocessed distortion data (dry signal). By adjusting the wet / dry ratio, odd and even harmonic data with adjusted wet / dry ratio are obtained. The odd and even harmonic data with adjusted wet / dry ratio are then superimposed, i.e., a linear summation operation is performed, to synthesize the final fourth intermediate audio data. This superposition process merges the previously finely adjusted odd and even harmonic components together to form an audio signal with specific timbre and distortion characteristics. The aforementioned technical solutions allow for precise control of the harmonic composition of the final distortion effect, thereby simulating the unique nonlinear distortion produced by microphones of different types or operating states. This significantly enhances the realism, adjustability, and applicability of the target microphone model. For example, it can simulate the distortion of a more aggressive rock microphone or the distortion of a retro microphone with a warmer, softer tone.

[0047] In one embodiment, specifically in step S106, which involves performing proximity effect simulation processing on the fourth intermediate audio data to obtain the fifth intermediate audio data, the following steps are included: Based on a preset proximity effect transfer function and open adjustable parameters, low-frequency gain processing is applied to the fourth intermediate audio data to obtain the fifth intermediate audio data. The open adjustable parameters include a cutoff frequency and a cutoff gain. The preset proximity effect transfer function characterizes the relationship between microphone directivity and proximity effect, and its formula is as follows: in, Represented as the cutoff frequency, This is expressed as the sampling frequency. This is represented as the cutoff gain.

[0048] In one implementation, the fourth intermediate audio data is processed with low-frequency gain by using a preset proximity effect transfer function and open adjustable parameters to obtain the fifth intermediate audio data. The open adjustable parameters include a cutoff frequency and a cutoff gain. The cutoff frequency defines the low-frequency boundary at which the proximity effect begins to have a significant impact; signals below this frequency will be more strongly affected by the proximity effect. The cutoff gain defines the amount of gain obtained by the low-frequency signal at or below the cutoff frequency, thereby controlling the intensity of the proximity effect. The preset proximity effect transfer function characterizes the relationship between microphone directivity and the proximity effect. This transfer function is usually in the form of a digital filter, and its coefficients are pre-designed and determined based on the microphone's physical characteristics and acoustic principles. The formula for the preset proximity effect transfer function is: in, Represented as the cutoff frequency, This is expressed as the sampling frequency. Represented as cutoff gain, through the above steps, further proximity effect simulation processing is introduced, which can achieve gain in the low-frequency part, and finally obtain the fifth intermediate audio data containing proximity effect simulation, thereby significantly improving the realism and applicability of the subsequently generated microphone model, so that the model can more accurately simulate the sound characteristics of the microphone when affected by proximity effect in actual use.

[0049] In one embodiment, that is, after step S106, i.e. after generating the target microphone model, the following steps are also included: The model algorithm of the target microphone model is determined and persistently saved, and the target microphone model is published to the model application platform.

[0050] In one implementation, after obtaining the target microphone model, the model algorithm is determined and persistently stored. The target microphone model can then be published to a model application platform via API interfaces, SDK packages, or containerized deployment (such as Docker). This allows other applications or users to easily call the target microphone model for microphone simulation or audio processing, thus achieving model sharing and reuse. Persistent storage can be achieved by serializing the model into a specific file format (e.g., ONNX, TensorFlowSavedModel, or PyTorchJIT) and storing it in a local file system, database, or cloud storage to ensure the model is not lost when the system is shut down. The model application platform can be a recording studio or other platforms. This technical solution allows for the systematic evaluation and optimization of model parameters, ensuring higher accuracy and stability of the generated target microphone model. It effectively avoids underfitting or overfitting, enabling the model to more accurately capture the acoustic characteristics of the target microphone, improving the realism and accuracy of the simulation, and providing users with a high-quality, deployable microphone simulation solution.

[0051] Please see Figure 2 , Figure 2 This is a schematic diagram of the microphone modeling device provided in an embodiment of the present invention. This microphone modeling device corresponds one-to-one with the microphone modeling method in the above embodiments. Please refer to [link / reference] for details. Figure 1 as well as Figure 1 The relevant descriptions in the corresponding embodiments are shown below. For ease of explanation, only the parts relevant to this embodiment are shown. See also... Figure 2 The microphone modeling device 20 includes: an acquisition module 21, a processing module 22, and a generation module 23.

[0052] The acquisition module 21 is used to acquire the raw audio data collected by the source microphone within a preset time period; Processing module 22 is used to perform convolution processing on the original audio data to obtain first intermediate audio data; perform equalization modeling processing on the first intermediate audio data to obtain second intermediate audio data; perform transient modeling processing on the second intermediate audio data to obtain third intermediate audio data; and perform nonlinear distortion modeling processing on the third intermediate audio data to obtain fourth intermediate audio data. The generation module 23 is used to perform proximity effect simulation processing on the fourth intermediate audio data to obtain the fifth intermediate audio data, and generate a target microphone model based on the fifth intermediate audio data.

[0053] Optionally, the above processing module 22 is specifically used for: Acquire unit impulse response data corresponding to the target microphone, wherein the unit impulse response data is impulse response data collected by the target microphone within a preset time period to characterize its acoustic properties; The original audio data and the unit impulse response data are convolved to obtain the first intermediate audio data.

[0054] Optionally, the processing module 22 is further configured to: Acquire audio data from the same source microphone and the target microphone simultaneously in the same scene; Fourier transforms are performed on the source audio data and the first intermediate audio data respectively to calculate the spectral amplitude difference curve; Equalization parameters are generated based on the number of filters selected by the user and the spectral amplitude difference curve. Based on the equalization parameters, the first intermediate audio data is subjected to equalization filtering to obtain the second intermediate audio data.

[0055] Optionally, the processing module 22 is further configured to: The second intermediate audio data is processed by fast envelope processing and slow envelope processing respectively using a transient shaper to obtain fast envelope processing results and slow envelope processing results; Calculate the difference between the fast envelope processing result and the slow envelope processing result; Based on the subtraction value, it is determined whether the data status signal is in the onset state or the holding state; If the data status signal is in the start-up state, then obtain the start-up gain parameter; If the data status signal is in a sustained tone state, then obtain the sustained tone gain parameter; The target gain coefficient is determined based on the onset gain parameter or the sustain gain parameter; The modulated audio signal is determined based on the target gain coefficient and the second intermediate audio data; The modulated audio signal is processed using a compressor to obtain third intermediate audio data.

[0056] Optionally, the processing module 22 is further configured to: Basic vacuum tube distortion data is generated using a distortion algorithm based on the tanh function, and the distortion saturation is adjusted using the drive parameter. The formula for the basic vacuum tube distortion data is as follows: Where x represents the third intermediate audio data, and drive represents the saturation control parameter; Harmonic generation processing was performed on the basic vacuum tube distortion data using a scheme where odd and even harmonics are individually adjustable, to obtain odd and even harmonic data after adjusting the wet / dry ratio. The odd harmonic data after the dry-to-wet ratio adjustment is superimposed with the even harmonic data to obtain the fourth intermediate audio data.

[0057] Optionally, the above-mentioned generation module 33 is specifically used for: Based on a preset proximity effect transfer function and open adjustable parameters, low-frequency gain processing is applied to the fourth intermediate audio data to obtain the fifth intermediate audio data. The open adjustable parameters include a cutoff frequency and a cutoff gain. The preset proximity effect transfer function characterizes the relationship between microphone directivity and proximity effect, and its formula is as follows: in, Represented as the cutoff frequency, This is expressed as the sampling frequency. This is represented as the cutoff gain.

[0058] Optionally, the generation module 33 described above is specifically used for: The model algorithm of the target microphone model is determined and persistently saved, and the target microphone model is published to the model application platform.

[0059] It should be noted that the information interaction and execution process between the above-mentioned units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0060] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 3 As shown, the computer device of this embodiment includes: at least one processor ( Figure 3 Only one is shown in the diagram), a memory, and a computer program stored in the memory and capable of running on at least one processor, which, when executing the computer program, implements the steps in any of the microphone modeling method embodiments described above.

[0061] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 3The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input systems.

[0062] In one embodiment, an aircraft is provided, which includes a power warning system and a computer device connected to the power warning system, enabling the computer device to perform various steps as described in any embodiment of the microphone modeling method disclosed in this invention, which will not be repeated here.

[0063] In one embodiment, a computer-readable storage medium is provided that, when the instructions in the computer-readable storage medium are executed by a processor in a computer device, enables the computer device to perform the steps of any embodiment of the microphone modeling method disclosed in this invention, which will not be repeated here. The computer-readable storage medium may be non-volatile or volatile.

[0064] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0065] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of the computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard. Furthermore, memory can include both internal storage units and external storage devices of the computer device. Memory is used to store the operating system, cooperative applications, bootloader, data, and other programs, such as program code of computer programs. Memory can also be used to temporarily store data that has been output or will be output.

[0066] Those skilled in the art will understand that implementing all or part of the processes in the above embodiments can be accomplished by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0067] Those familiar with the technical field will understand that, for ease of description and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0068] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A microphone modeling method, characterized in that, include: Acquire the raw audio data collected by the source microphone within a preset time period; The original audio data is convolved to obtain the first intermediate audio data; The first intermediate audio data is subjected to equalization modeling processing to obtain the second intermediate audio data; The second intermediate audio data is subjected to transient modeling processing to obtain the third intermediate audio data; The third intermediate audio data is subjected to nonlinear distortion modeling processing to obtain the fourth intermediate audio data; The fourth intermediate audio data is subjected to proximity effect simulation processing to obtain the fifth intermediate audio data, and a target microphone model is generated based on the fifth intermediate audio data.

2. The microphone modeling method as described in claim 1, characterized in that, The convolutional processing of the original audio data to obtain the first intermediate audio data includes: Acquire unit impulse response data corresponding to the target microphone, wherein the unit impulse response data is impulse response data collected by the target microphone within a preset time period to characterize its acoustic properties; The original audio data and the unit impulse response data are convolved to obtain the first intermediate audio data.

3. The microphone modeling method as described in claim 1, characterized in that, The process of performing equalization modeling on the first intermediate audio data to obtain the second intermediate audio data includes: Acquire audio data from the same source microphone and the target microphone simultaneously in the same scene; Fourier transforms are performed on the source audio data and the first intermediate audio data respectively to calculate the spectral amplitude difference curve; Equalization parameters are generated based on the number of filters selected by the user and the spectral amplitude difference curve. Based on the equalization parameters, the first intermediate audio data is subjected to equalization filtering to obtain the second intermediate audio data.

4. The microphone modeling method as described in claim 1, characterized in that, The transient modeling process performed on the second intermediate audio data to obtain the third intermediate audio data includes: The second intermediate audio data is processed by fast envelope processing and slow envelope processing respectively using a transient shaper to obtain fast envelope processing results and slow envelope processing results; Calculate the difference between the fast envelope processing result and the slow envelope processing result; Based on the subtraction value, it is determined whether the data status signal is in the onset state or the holding state; If the data status signal is in the start-up state, then obtain the start-up gain parameter; If the data status signal is in a sustained tone state, then obtain the sustained tone gain parameter; The target gain coefficient is determined based on the onset gain parameter or the sustain gain parameter; The modulated audio signal is determined based on the target gain coefficient and the second intermediate audio data; The modulated audio signal is processed using a compressor to obtain third intermediate audio data.

5. The microphone modeling method as described in claim 1, characterized in that, The process of performing nonlinear distortion modeling on the third intermediate audio data to obtain the fourth intermediate audio data includes: Basic vacuum tube distortion data is generated using a distortion algorithm based on the tanh function, and the distortion saturation is adjusted using the drive parameter. The formula for the basic vacuum tube distortion data is as follows: Where x represents the third intermediate audio data, and drive represents the saturation control parameter; Harmonic generation processing was performed on the basic vacuum tube distortion data using a scheme where odd and even harmonics are individually adjustable, to obtain odd and even harmonic data after adjusting the wet / dry ratio. The odd harmonic data after the dry-to-wet ratio adjustment is superimposed with the even harmonic data to obtain the fourth intermediate audio data.

6. The microphone modeling method as described in claim 1, characterized in that, The process of performing proximity effect simulation on the fourth intermediate audio data to obtain the fifth intermediate audio data includes: Based on a preset proximity effect transfer function and open adjustable parameters, low-frequency gain processing is applied to the fourth intermediate audio data to obtain the fifth intermediate audio data. The open adjustable parameters include a cutoff frequency and a cutoff gain. The preset proximity effect transfer function characterizes the relationship between microphone directivity and proximity effect, and its formula is as follows: in, Represented as the cutoff frequency, This is expressed as the sampling frequency. This is represented as the cutoff gain.

7. The microphone modeling method as described in claim 1, characterized in that, After generating the target microphone model, the process includes: The model algorithm of the target microphone model is determined and persistently saved, and the target microphone model is published to the model application platform.

8. A microphone modeling device, characterized in that, include: The acquisition module is used to acquire the raw audio data collected by the source microphone within a preset time period; The processing module is used to perform convolution processing on the original audio data to obtain first intermediate audio data; and to perform equalization modeling processing on the first intermediate audio data to obtain second intermediate audio data. The second intermediate audio data is subjected to transient modeling processing to obtain the third intermediate audio data; The third intermediate audio data is subjected to nonlinear distortion modeling processing to obtain the fourth intermediate audio data; The generation module is used to perform proximity effect simulation processing on the fourth intermediate audio data to obtain the fifth intermediate audio data, and generate a target microphone model based on the fifth intermediate audio data.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the microphone modeling method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the microphone modeling method as described in any one of claims 1 to 7.