Virtual play simulation method, apparatus, device, and medium
By acquiring complex frequency response data and adaptively selecting multiple virtual playback simulation algorithms, the problems of insufficient flexibility, accuracy and real-time performance in existing virtual simulation technologies are solved. This achieves high-fidelity reproduction of the acoustic characteristics of audio playback devices and full-scene adaptation, and the generated audio signals exhibit high flexibility and naturalness in different application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUAQIN TECH CO LTD
- Filing Date
- 2026-03-18
- Publication Date
- 2026-07-10
AI Technical Summary
Existing virtual simulation technology is insufficient in terms of flexibility, accuracy and real-time performance, making it difficult to meet diverse application needs. In particular, when processing amplitude-frequency response, ignoring phase characteristics leads to distortion of time-domain waveforms, and a single fixed solution cannot dynamically adjust delay, accuracy or computational complexity according to the application scenario.
By acquiring the complex frequency response data of the target audio playback device, including amplitude and phase frequency response data, a pool of multiple virtual playback simulation algorithms is constructed. The target algorithm is adaptively selected according to the preset application constraints. Combined with auditory perception optimization technology, high-fidelity restoration of audio signals and flexible adaptation to all scenarios are achieved.
It achieves complete characterization and accurate reproduction of the acoustic characteristics of the target audio playback device, improving the flexibility, accuracy and real-time performance of virtual simulation. The generated audio signal is more natural, softer and more immersive, and is suitable for different application scenarios.
Smart Images

Figure CN122363648A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio signal processing and audio playback device simulation technology, and in particular to a virtual playback simulation method, device, equipment and medium. Background Technology
[0002] The core objective of virtual playback simulation is to simulate the acoustic characteristics of real audio playback devices using digital signal processing technology. It is widely used in audio engineering, consumer electronics, car audio, virtual reality (VR), and remote collaboration. In professional audio production, recording and mixing engineers need to monitor audio effects in real time using virtual playback devices such as speakers to simulate the sound performance of different playback devices in a specific room or device, thereby optimizing the compatibility of audio content. In consumer electronics, such as smartphones, tablets, or wireless headphones, the size and power consumption of physical playback devices necessitate algorithmic compensation of the playback device's frequency response characteristics to improve the sound quality experience. In car audio systems, virtual playback technology can simulate the audio configurations of different car models, helping engineers optimize the sound field distribution during the design phase. Furthermore, in remote collaboration scenarios, team members can share audio processing effects through virtual playback technology, reducing reliance on physical devices. Therefore, achieving high-fidelity and highly natural virtual simulation while balancing latency, computational efficiency, and accuracy has become a core technical challenge that urgently needs to be addressed in these fields.
[0003] In related technologies, virtual simulation technology mainly relies on a single signal processing scheme, such as simple filter design based on frequency domain equalization or finite impulse response (FIR) convolution. Specifically, it first acquires acoustic characteristic data of the target playback device, such as the target loudspeaker, through measurement or simulation, mainly its frequency response in an anechoic chamber environment, i.e., sound pressure level (SPL) data, which reflects the loudspeaker's energy output characteristics at different frequencies. Subsequently, based on this SPL data, corresponding equalizer parameters are designed to construct a filter network for simulating the loudspeaker's characteristics. Finally, the input audio signal is fed into this filter network for processing, so that the amplitude-frequency characteristics of the output signal approximate those of the target loudspeaker, aiming to reproduce the tonal characteristics of the target loudspeaker at the headphone end. However, these schemes have significant shortcomings in terms of flexibility, accuracy, and real-time performance, making it difficult to meet diverse application requirements. Summary of the Invention
[0004] This application provides a virtual playback simulation method, apparatus, device, and medium to improve the problem that the solutions in the related technologies have significant deficiencies in terms of flexibility, accuracy, and real-time performance, making it difficult to meet diverse application needs.
[0005] Firstly, this application provides a virtual playback simulation method, including:
[0006] Acquire the complex frequency response data of the target audio playback device, which includes amplitude frequency response data and phase frequency response data;
[0007] Acquire the input audio signal;
[0008] Based on preset application constraints, a target algorithm is adaptively selected from multiple preset virtual playback simulation algorithms; among them, the multiple preset virtual playback simulation algorithms have different latency characteristics, accuracy levels and computational complexity.
[0009] The target algorithm is used to process the input audio signal based on complex frequency response data to generate an output audio signal that simulates the playback effect of the target audio playback device.
[0010] In one possible implementation, the preset application constraints include at least one of maximum allowable processing latency, available computing resources, phase accuracy requirements, and processing modes. Based on the preset application constraints, an algorithm is adaptively selected from multiple preset virtual playback simulation algorithms, including: obtaining at least one of maximum allowable processing latency, available computing resources, phase accuracy requirements, and processing modes as the current constraints; and matching an algorithm that is compatible with the current constraints from multiple preset virtual playback simulation algorithms as the target algorithm.
[0011] In one possible implementation, based on the current constraints, an algorithm that is suitable for the current constraints is selected as the target algorithm from a plurality of preset virtual playback simulation algorithms, including at least one of the following: when the processing mode is real-time processing mode and the maximum allowable processing latency is lower than a preset latency threshold, the algorithm with the lowest latency characteristic is selected as the target algorithm from a plurality of preset virtual playback simulation algorithms; when the processing mode is offline processing mode and the phase accuracy requirement is higher than a preset accuracy threshold, the algorithm with the highest accuracy level is selected as the target algorithm from a plurality of preset virtual playback simulation algorithms; when the available computing resources are lower than a preset resource threshold, the algorithm with the lowest computational complexity is selected as the target algorithm from a plurality of preset virtual playback simulation algorithms.
[0012] In one possible implementation, the multiple preset virtual playback simulation algorithms include at least three of the following algorithm sets: frequency domain multiplication algorithm, FIR filter convolution algorithm, cascaded infinite impulse response (IIR) dual second-order filter algorithm, overlapping summing block processing algorithm, overlapping retaining block processing algorithm, minimum phase reconstruction algorithm, impulse response convolution algorithm, parametric equalizer approximation algorithm, frequency warping FIR filter algorithm, partitioned convolution algorithm, frequency sampling algorithm, polyphase filter bank algorithm, weighted overlapping summing algorithm, constant Q transform processing algorithm, and zero-delay convolution algorithm; wherein, the zero-delay convolution algorithm refers to an algorithm that decomposes the complex frequency response data into minimum phase components and all-pass components for separate processing.
[0013] In one possible implementation, the algorithm employing a frequency resolution processing method that matches the characteristics of human auditory perception includes a frequency warping FIR filter algorithm or a constant Q transform processing algorithm. The frequency warping FIR filter algorithm determines a frequency warping factor based on the characteristics of human auditory perception and performs a warping transformation on the frequency axis based on the frequency warping factor to enhance the filter resolution in the low-frequency region. The constant Q transform processing algorithm analyzes the audio signal using logarithmically spaced frequency points with a constant quality factor, where the ratio of the center frequency to the bandwidth at each frequency point remains constant.
[0014] In one possible implementation, if the target algorithm is a partitioned convolution algorithm, the target algorithm is used to process the input audio signal based on complex frequency response data, including: dividing the target impulse response into multiple continuous segments, the target impulse response being obtained from the complex frequency response data through inverse transform; performing a fast Fourier transform on each segment to obtain the corresponding frequency domain segmentation coefficients; constructing a frequency-domain delay line (FDL) to store the frequency domain representations of multiple historical input audio signals; multiplying and accumulating the frequency domain representation of the current input audio signal with each frequency domain segmentation coefficient to obtain the frequency domain output result; and performing an inverse fast Fourier transform on the frequency domain output result to obtain the output audio signal.
[0015] In one possible implementation, if the target algorithm is a zero-delay convolution algorithm, the target algorithm is used to process the input audio signal based on the complex frequency response data, including: decomposing the complex frequency response data into minimum phase components and all-pass components; constructing a minimum phase impulse response based on the minimum phase components to achieve zero prering characteristics; designing a cascaded first-order all-pass filter based on the all-pass components to approximate the residual phase characteristics of the complex frequency response data; and cascading the minimum phase impulse response with the cascaded first-order all-pass filter to perform convolution processing on the input audio signal to obtain the output audio signal.
[0016] Secondly, this application provides a virtual playback simulation device, comprising:
[0017] The input module is used to acquire complex frequency response data of the target audio playback device, including amplitude frequency response data and phase frequency response data; and to acquire the input audio signal.
[0018] The algorithm selection module is used to adaptively select a target algorithm from multiple preset virtual playback simulation algorithms based on preset application constraints; wherein the multiple preset virtual playback simulation algorithms have different latency characteristics, accuracy levels and computational complexity.
[0019] The processing module is used to process the input audio signal based on the complex frequency response data using the target algorithm to generate an output audio signal that simulates the playback effect of the target audio playback device.
[0020] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0021] Memory is used to store instructions executed by the computer;
[0022] A processor for executing computer-executable instructions stored in memory to implement any of the methods in the first aspect.
[0023] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed, are used to implement the method of any one of the first aspects.
[0024] Fifthly, this application provides a computer program product, including a computer program that, when executed, implements the method of any one of the first aspects.
[0025] The virtual playback simulation method, apparatus, device, and medium provided in this application acquire complex frequency response data of a target audio playback device, including amplitude frequency response data and phase frequency response data; acquire an input audio signal; adaptively select a target algorithm from multiple preset virtual playback simulation algorithms according to preset application constraints; wherein the multiple preset virtual playback simulation algorithms have different delay characteristics, accuracy levels, and computational complexity; and use the target algorithm to process the input audio signal based on the complex frequency response data to generate an output audio signal that simulates the playback effect of the target audio playback device.
[0026] In this process, by acquiring complex frequency response data containing both amplitude and phase response data, the time-domain waveform distortion problem caused by existing technologies that only process amplitude response and ignore phase characteristics is effectively improved, achieving a complete characterization and accurate restoration of the acoustic characteristics of the target audio playback device. By adaptively selecting a target algorithm from multiple preset algorithms with distinct delay characteristics, accuracy levels, and computational complexity based on preset application constraints, the limitations of existing technologies with a single fixed solution are broken, achieving flexible adaptation across all scenarios from low-latency real-time monitoring to high-precision offline processing. Furthermore, by processing the input audio signal using the target algorithm based on the complex frequency response data, an output audio signal simulating the playback effect of the target audio playback device is generated, thereby comprehensively improving the flexibility, accuracy, and real-time performance of the virtual simulation. Attached Figure Description
[0027] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0028] Figure 1 A flowchart illustrating a virtual playback simulation method provided as an exemplary embodiment of this application;
[0029] Figure 2 A schematic diagram illustrating the target algorithm selection process provided for an exemplary embodiment of this application;
[0030] Figure 3 A schematic diagram of a virtual playback simulation device provided as an exemplary embodiment of this application;
[0031] Figure 4 A schematic diagram of the architecture of a virtual playback simulation device provided as an exemplary embodiment of this application;
[0032] Figure 5 A comparative schematic diagram of the input audio signal, complex frequency response data, and output audio signal provided for an exemplary embodiment of this application;
[0033] Figure 6 A schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application.
[0034] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0035] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0036] The terms “first,” “second,” etc., used in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, products, or apparatus.
[0037] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0038] In related technologies, virtual simulation technology mainly relies on a single signal processing scheme. These schemes typically only process the amplitude-frequency response while completely ignoring phase characteristics. Although simple amplitude-frequency equalization can adjust the timbre, it leads to distortion in the reproduction of time-domain waveforms, resulting in the loss of key information such as transient response and sound field localization, and a lack of realism in sound reproduction. Furthermore, these processing schemes use a single, fixed processing flow and cannot dynamically adjust latency, precision, or computational complexity according to the application scenario. For example, high-precision long-order FIR filters have excessive latency and are computationally intensive when processing ultra-long impulse responses, making them unsuitable for real-time monitoring or embedded mobile platforms. On the other hand, low-latency IIR equalizers are difficult to accurately reproduce the characteristics of complex audio playback devices. In addition, existing algorithm designs are often computationally inefficient, especially when processing ultra-long impulse responses such as room acoustic simulation. Traditional convolution operations consume huge amounts of resources and are difficult to run on resource-constrained devices.
[0039] To address the aforementioned issues, this application provides a virtual playback simulation scheme. By constructing a virtual simulation framework integrating multiple algorithms and introducing an adaptive selection mechanism and auditory perception optimization, it achieves high-fidelity reproduction of the acoustic characteristics of the target audio playback device and flexible adaptation to all scenarios. Specifically, firstly, complex frequency response data, including amplitude and phase responses, is acquired to fully characterize the acoustic characteristics of the target audio playback device, effectively improving the phase information loss problem caused by related technologies that only process amplitude responses. Based on this, multiple virtual playback simulation algorithms with distinct delay characteristics, accuracy levels, and computational complexity are pre-configured, forming a full-range algorithm pool covering near-zero latency to high-precision offline processing. Then, based on preset application constraints, the optimal matching target algorithm is adaptively selected from the algorithm pool to achieve dynamic adaptation under different application scenarios. Finally, the selected target algorithm is used to process the input audio signal based on the complex frequency response data to generate an output audio signal simulating the playback effect of the target audio playback device. Through the above processing mechanism, phase restoration, adaptive algorithm selection, and auditory perception optimization are organically integrated, thereby comprehensively improving the flexibility, accuracy, and real-time performance of virtual simulation.
[0040] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0041] Figure 1 This is a flowchart illustrating a virtual playback simulation method provided as an exemplary embodiment of this application. Figure 1 As shown, the virtual playback simulation method includes the following steps:
[0042] S101. Obtain the complex frequency response data of the target audio playback device. The complex frequency response data includes amplitude frequency response data and phase frequency response data.
[0043] Among them, complex frequency response data refers to the response characteristics of an audio playback device to an input signal in the frequency domain. It is represented in complex form and includes information in two dimensions: amplitude and phase. Amplitude frequency response data, also known as sound pressure level data, reflects the energy output characteristics of the audio playback device at different frequencies and determines the timbre balance of the sound. Phase frequency response data reflects the time delay characteristics of the audio playback device to different frequency components and determines the transient response, sound field localization, and spatial sense of the sound.
[0044] In one implementation, this data can be obtained by measuring the target audio playback device, such as the target loudspeaker, in a standard acoustic environment such as an anechoic chamber. Specifically, using equipment such as an audio analyzer and a high-precision microphone array, a test signal (such as a logarithmic sweep signal or a maximum length sequence signal) is input to the loudspeaker, and the loudspeaker's output response is simultaneously acquired. By performing a Fast Fourier Transform (FFT) analysis on the acquired signal, the amplitude and phase responses of the loudspeaker at various frequency points can be obtained.
[0045] In another approach, this data can be obtained by modeling and simulating the target loudspeaker using acoustic simulation software. For loudspeakers still in the design phase or in scenarios where physical measurement is difficult, the theoretical amplitude-frequency response and phase-frequency response data can be calculated based on the loudspeaker's physical parameters (such as diaphragm material, voice coil parameters, cabinet structure, etc.) and electroacoustic model using methods such as finite element analysis and boundary element analysis.
[0046] Optionally, the acquired raw measurement or simulation data usually needs to be preprocessed, including but not limited to removing outliers, smoothing, and interpolating to the required frequency resolution, in order to form complex frequency response data that can be directly used for subsequent processing.
[0047] It should be noted that the aforementioned target audio playback device, such as a target speaker, is merely an example. In practical applications, the target audio playback device includes, but is not limited to, speakers, headphones, over-ear headphones, in-ear headphones, high-fidelity headphones, monitoring headphones, smart glasses, smart helmets, smartphone receivers, built-in audio units in tablet computers, built-in speakers in laptops, built-in speakers in smart displays, stage monitors, in-ear monitoring systems, car audio systems, home theater audio systems, and any other electroacoustic transducers with audio playback functions. The specific type of target audio playback device is not limited here.
[0048] S102, Acquire the input audio signal.
[0049] For example, in one implementation, the input audio signal can come from a professional audio production environment, such as a multitrack audio file exported from a Digital Audio Workstation (DAW) or the bus output of a mixing project. These signals typically have high sampling accuracy and dynamic range.
[0050] In another implementation, the input audio signal can come from consumer electronic devices, such as voice signals captured in real time by a smartphone microphone, compressed audio streams decoded by a media player, or real-time audio rendered by a game engine.
[0051] In another implementation, the input audio signal can be a standard test signal, such as pink noise, white noise, sinusoidal sweep signal, pulse signal, etc., used to evaluate the performance of the virtual simulation system or to perform calibration.
[0052] Optionally, after acquiring the input audio signal, the input audio signal can also be preprocessed, such as format conversion (e.g., decoding the compressed format into Pulse Code Modulation (PCM) format), sampling rate matching (ensuring consistency with the sampling rate of subsequent processing), channel format conversion (e.g., downmixing multi-channel audio into stereo), and floating-point normalization (converting integer samples into floating-point numbers in the range of -1.0 to 1.0), etc.
[0053] S103. Based on preset application constraints, adaptively select a target algorithm from multiple preset virtual playback simulation algorithms; wherein, the multiple preset virtual playback simulation algorithms have different latency characteristics, accuracy levels and computational complexity.
[0054] In some embodiments, the preset application constraints include at least one of maximum allowable processing latency, available computing resources, phase accuracy requirements, and processing modes. Based on the preset application constraints, an algorithm is adaptively selected from multiple preset virtual playback simulation algorithms, including: obtaining at least one of maximum allowable processing latency, available computing resources, phase accuracy requirements, and processing modes as the current constraints; and matching an algorithm that is compatible with the current constraints from multiple preset virtual playback simulation algorithms as the target algorithm.
[0055] Among them, the preset application constraints refer to the objective description and limiting factors of the current application scenario, which are used to guide the selection of the target algorithm. These application constraints reflect the different requirements of the specific application scenario for audio processing, including at least one of the following: maximum allowable processing latency, available computing resources, phase accuracy requirements, and processing mode. Latency characteristics refer to the time delay required for the algorithm to generate output from receiving input, usually measured in milliseconds (ms) or sample count. Different algorithms have significant differences in latency characteristics due to their different implementation principles. Some can achieve near-zero latency, while others introduce larger latency for higher accuracy. Accuracy level refers to the algorithm's ability to reproduce the acoustic characteristics of the target audio playback device, including amplitude frequency response matching degree and phase frequency response fidelity. High-precision algorithms can accurately approximate complex frequency response curves and phase characteristics. Computational complexity refers to the computing resources required for the algorithm to run, including central processing unit (CPU) utilization, memory usage, power consumption, etc., which determines whether the algorithm can run in real time on a specific hardware platform.
[0056] Given that existing processing schemes often employ linear frequency resolution, neglecting the auditory perception characteristics of the human ear, the simulated sound, while objectively matching the data, suffers from a subjectively harsh and unnatural listening experience, lacking immersion. Therefore, in some embodiments, at least one of the multiple preset virtual playback simulation algorithms adopts a frequency resolution processing method that matches the auditory perception characteristics of the human ear. This frequency resolution processing method refers to using a non-linear frequency scale (such as the Bark scale, equivalent rectangular bandwidth scale, etc.) to analyze and process the audio signal, making the processing results more consistent with the auditory characteristics of the human ear and helping to improve the naturalness of the subjective listening experience.
[0057] For example, the system pre-constructs an algorithm pool containing various virtual playback simulation algorithms. These algorithms differ in latency characteristics, accuracy levels, and computational complexity, forming a full-range algorithm set covering near-zero latency to high-precision offline processing, and from low complexity to high-performance requirements. Furthermore, this algorithm pool includes at least one algorithm employing a frequency resolution processing method that matches the characteristics of human auditory perception, ensuring that it can provide processing effects that better conform to human auditory characteristics when needed. Accordingly, when performing adaptive selection, it is first necessary to determine the constraints of the current application scenario. These constraints can be obtained from multiple channels. For example, in one implementation, user-input parameters can be received through the user interface, such as preset options like "real-time monitoring mode," "high-precision offline processing mode," and "energy-saving mode" provided in professional audio software. In another implementation, the current operating environment can be automatically detected; for example, when it is identified that the system is running on a mobile device, the available computing resources are automatically marked as "limited," or when the audio signal is detected to be from real-time recording input, the processing mode is automatically set to "real-time processing."
[0058] Correspondingly, after obtaining the current constraints, an algorithm suitable for these constraints is matched from a pre-defined algorithm pool as the target algorithm. The matching process can employ various strategies, such as prioritizing different dimensions based on the nature of the constraints: when latency requirements are extremely strict, priority is given to the algorithm's latency characteristics; when accuracy requirements are a decisive factor, priority is given to the algorithm's accuracy level; when computational resources are a bottleneck, priority is given to the algorithm's computational complexity. For complex scenarios with multiple constraints that may conflict, a comprehensive scoring mechanism can be used to score each algorithm across different dimensions, calculate the total score based on the weights of the constraints, and thus select the optimal matching algorithm.
[0059] Through the aforementioned adaptive selection mechanism, the optimal virtual playback simulation algorithm can be dynamically selected according to different application requirements, achieving full-scenario coverage from embedded devices to high-performance workstations, and from real-time monitoring to offline mastering processing. This fundamentally improves the limited applicability of related technologies due to their singular processing methods. Simultaneously, the algorithm pool includes at least one algorithm that employs a frequency resolution processing method matching the characteristics of human auditory perception, providing support for scenarios requiring optimization of subjective listening experience, further enhancing the system's comprehensiveness and adaptability.
[0060] S104. Using a target algorithm, the input audio signal is processed based on complex frequency response data to generate an output audio signal that simulates the playback effect of the target audio playback device.
[0061] For example, in this step, based on the target algorithm adaptively selected in step S103 and combined with the complex frequency response data obtained in step S101, the input audio signal obtained in step S102 is subjected to corresponding signal processing operations. Specifically, the target algorithm is the optimal algorithm selected from the aforementioned algorithm pool based on application constraints, and the complex frequency response data contains the amplitude frequency response and phase frequency response information of the target audio playback device. During processing, the acoustic characteristics of the target audio playback device, represented by the complex frequency response data, are applied to the input audio signal through the target algorithm, so that the output audio signal approximates the playback effect of the target audio playback device in both frequency and time domain characteristics.
[0062] Optionally, after processing, the generated output audio signal is normalized to adapt its amplitude range to the input requirements of subsequent playback devices, and then converted to the target audio format as needed. This normalization process reduces signal clipping or excessively low amplitude caused by gain changes introduced during processing, helping to ensure the quality of the output signal. Furthermore, the normalized output audio signal can be played through a high-fidelity (HIFI) headphone player for subjective evaluation of the virtual simulation effect's sound quality. Accordingly, evaluators listen to the output audio signal through audio playback devices such as HIFI headphones to determine whether it faithfully reproduces the timbre, transient response, and spatial acoustic characteristics of the target audio playback device, thereby verifying the effect of the virtual simulation. This output audio signal can also be used for real-time monitoring, recording and storage, or subsequent audio production and analysis.
[0063] The virtual playback simulation method provided in this application effectively improves the time-domain waveform distortion problem caused by prior art that only processes the amplitude frequency response and ignores the phase characteristics by acquiring complex frequency response data containing amplitude frequency response data and phase frequency response data. This achieves a complete characterization and accurate restoration of the acoustic characteristics of the target audio playback device. By adaptively selecting a target algorithm from multiple preset algorithms with different delay characteristics, accuracy levels, and computational complexity according to preset application constraints, it breaks the limitations of the single fixed scheme in the prior art and achieves flexible adaptation to all scenarios from low-latency real-time monitoring to high-precision offline processing. In addition, by introducing at least one algorithm that adopts a frequency resolution processing method that matches the characteristics of human hearing perception, it significantly improves the harsh listening experience caused by the linear frequency resolution of the prior art, making the virtual simulated sound more natural, soft, and immersive. Furthermore, by processing the input audio signal based on the complex frequency response data using the target algorithm, an output audio signal simulating the playback effect of the target audio playback device is generated, thereby comprehensively improving the flexibility, accuracy, real-time performance, and perception optimization effect of the virtual simulation.
[0064] In some embodiments, multiple preset virtual playback simulation algorithms include at least three of the following algorithm sets: frequency domain multiplication algorithm, FIR filter convolution algorithm, cascaded IIR double second-order filter algorithm, overlapping summing block processing algorithm, overlapping retaining block processing algorithm, minimum phase reconstruction algorithm, impulse response convolution algorithm, parametric equalizer approximation algorithm, frequency warping FIR filter algorithm, partitioned convolution algorithm, frequency sampling algorithm, polyphase filter bank algorithm, weighted overlapping summing algorithm, constant Q transform processing algorithm, and zero-delay convolution algorithm; wherein, the zero-delay convolution algorithm refers to an algorithm that decomposes the complex frequency response data into minimum phase components and all-pass components for separate processing.
[0065] For example, an algorithm pool containing various virtual playback simulation algorithms is pre-built, which covers algorithm implementations for multiple technical paths such as frequency domain processing, time domain convolution, block processing, phase reconstruction, and perceptual optimization.
[0066] Specifically, the frequency domain multiplication algorithm is based on the Fast Fourier Transform (FFT). After converting the input audio signal to the frequency domain, it is directly multiplied with the complex frequency response characterizing the audio playback device, and then restored to the time domain through an inverse transform, achieving the highest accuracy in frequency domain matching. The FIR filter convolution algorithm constructs a finite impulse response (FIR) filter and performs time-domain convolution between the input audio signal and the impulse response, simultaneously controlling the amplitude-frequency and phase-frequency characteristics to achieve a precise linear phase response. The cascaded IIR dual second-order filter algorithm uses a cascaded second-order section structure, where each dual second-order section can achieve peak, low-profile, or high-profile filtering characteristics, offering advantages such as high computational efficiency and extremely low latency. The overlapping add-block processing algorithm and the overlapping retain-block processing algorithm both belong to block convolution methods, which... The input audio signal is segmented into fixed-length data blocks, and fast convolution processing is performed in the frequency domain. The continuous output signal is then reconstructed through overlapping addition or overlapping preservation, enabling efficient real-time convolution processing. The minimum phase reconstruction algorithm is used to reconstruct the corresponding minimum phase response from the amplitude-frequency response using the Hilbert transform of the logarithmic amplitude spectrum when phase data is missing, ensuring the causality and stability of the system. The impulse response convolution algorithm obtains the time-domain impulse response by performing an inverse fast Fourier transform on the complex frequency response data, and then performs direct time-domain convolution with the input audio signal. This algorithm can completely preserve the amplitude-frequency and phase-frequency characteristics of the target audio playback device, achieving accurate linear convolution. The parametric equalizer approximation algorithm uses multiple parametric equalizer sections to approximate the target audio signal. The frequency response of the audio playback device is evaluated, with each equalizer section allowing independent control of the center frequency, gain, and quality factor. The frequency warping FIR filter algorithm introduces a frequency warping factor, replacing the delay unit of the traditional FIR filter with a first-order all-pass filter to achieve a non-linear transformation of the frequency axis, enhancing filter resolution in the low-frequency region and better aligning with human hearing characteristics. The partitioned convolution algorithm is suitable for handling scenarios with extremely long impulse responses, including room reverberation. It divides the long impulse response into multiple segments for parallel processing, enabling real-time convolution with low latency. The frequency sampling algorithm directly samples the complex frequency response of the target audio playback device at the frequency points of the Fast Fourier Transform (FFT), and performs an inverse FFT on the sampling results to obtain the initial filter. The algorithm calculates the coefficients and uses a Kaiser window to suppress the Gibbs phenomenon. The polyphase filter bank algorithm combines long FIR filters with FFT through polyphase decomposition to achieve efficient filter bank analysis and synthesis. The weighted overlap addition algorithm uses matched analysis and synthesis windows to satisfy the constant overlap addition condition, achieving perfect reconstruction without the application of filtering. The constant Q transform processing algorithm uses logarithmically spaced frequency points, keeping the ratio of the center frequency to the bandwidth constant, so that the frequency resolution changes with the center frequency, which is more in line with the auditory perception characteristics of the human ear. The zero-delay convolution algorithm decomposes the complex frequency response into minimum phase components and all-pass components, processes them separately and then combines them to achieve accurate restoration of the original phase characteristics while maintaining low delay.
[0067] These algorithms differ in latency characteristics, accuracy levels, and computational complexity, collectively forming a full range of algorithms covering everything from near-zero latency to high-precision offline processing, and from low complexity to high-performance requirements.
[0068] In some embodiments, based on the current constraints, an algorithm that is suitable for the current constraints is matched from a plurality of preset virtual playback simulation algorithms as the target algorithm, including at least one of the following situations: when the processing mode is real-time processing mode and the maximum allowable processing latency is lower than a preset latency threshold, the algorithm with the lowest latency characteristic is selected from a plurality of preset virtual playback simulation algorithms as the target algorithm; when the processing mode is offline processing mode and the phase accuracy requirement is higher than a preset accuracy threshold, the algorithm with the highest accuracy level is selected from a plurality of preset virtual playback simulation algorithms as the target algorithm; when the available computing resources are lower than a preset resource threshold, the algorithm with the lowest computational complexity is selected from a plurality of preset virtual playback simulation algorithms as the target algorithm.
[0069] For example, based on different application constraints, a priority rule is used to match the optimal target algorithm from the algorithm pool, specifically including the following cases:
[0070] Scenario 1: Real-time processing with strict latency requirements. Specifically, when the processing mode is real-time and the maximum allowable processing latency is below a preset threshold (e.g., 10 milliseconds), latency characteristics are the primary consideration. In this scenario, the algorithm with the lowest latency characteristics is selected from the algorithm pool as the target algorithm. For example, in a scenario where a recording engineer monitors the effect of a virtual audio playback device in real time through headphones, a cascaded IIR dual second-order filter algorithm or a zero-latency convolution algorithm is selected. These algorithms can meet the strict latency requirements of real-time monitoring (usually from a few samples to a few milliseconds) while ensuring processing quality as much as possible, helping to reduce the "echo" or "out-of-sync" feeling caused by excessive latency.
[0071] Scenario 2: Offline Processing with High Precision Requirements. Specifically, when the processing mode is offline and the phase precision requirement is higher than a preset precision threshold, the precision level is the primary consideration. In this case, the algorithm with the highest precision level is selected from the algorithm pool as the target algorithm. For example, in mastering or professional audio production scenarios, engineers pursue the ultimate sound quality reproduction and can accept a longer processing time. The system will select a frequency domain multiplication algorithm, which achieves optimal frequency domain matching precision by performing frequency domain processing on the entire file, helping to accurately reproduce the amplitude and phase frequency characteristics of the target audio playback device.
[0072] Scenario 3: Limited Computational Resources. Specifically, when available computational resources are below a preset resource threshold (e.g., CPU utilization consistently exceeds 80%, memory usage approaches its limit), computational complexity becomes the primary consideration. In this scenario, the algorithm with the lowest computational complexity is selected from the algorithm pool as the target algorithm. For example, in embedded devices such as wireless headphones, smartphones, or car audio processors, considering the limited CPU clock speed, small memory capacity, and power consumption, a cascaded IIR dual second-order filter algorithm is typically chosen. This algorithm has low computational complexity, can run smoothly on resource-constrained devices, and provides acceptable sound quality.
[0073] It should be noted that the above three scenarios can be applied individually or in combination as needed. For example, Figure 2 A schematic diagram illustrating the target algorithm selection process provided for an exemplary embodiment of this application. For example... Figure 2 As shown, the selection process for the target algorithm includes the following steps:
[0074] S201. Obtain application constraints.
[0075] This involves obtaining the application constraints of the current application scenario, which include at least one of the following: processing mode, maximum allowable processing latency, available computing resources, and phase accuracy requirements.
[0076] S202. Based on the application constraints, determine whether the processing mode is real-time processing.
[0077] If it is in real-time processing mode, execute S203 to perform a delay criticality judgment;
[0078] If it is a non-real-time processing mode (i.e., offline processing mode), execute S208 to determine the accuracy requirement.
[0079] S203. Determine whether there are critical requirements for delay.
[0080] In real-time processing mode, it is further determined whether there are critical requirements for latency, that is, whether the maximum allowable processing latency is lower than the preset latency threshold.
[0081] If not, proceed to S204;
[0082] If so, execute S207.
[0083] S204. Determine if there is a requirement for a long pulse response.
[0084] If so, that is, when long impulse responses need to be processed (such as room acoustic simulation), execute S205;
[0085] If not, that is, when there is no need to process long impulse responses, select the algorithm with the lowest delay characteristics and execute S206.
[0086] S205. Select the partitioned convolution algorithm as the target algorithm.
[0087] S206. Select either the overlapping addition block processing algorithm or the overlapping retaining block processing algorithm as the target algorithm.
[0088] S207. Select the cascaded IIR double second-order filter algorithm as the target algorithm.
[0089] S208. Determine if there is a precision priority requirement.
[0090] If so, execute S212;
[0091] If not, proceed with S209.
[0092] S209. Determine if the phase information is available.
[0093] If so, that is, when phase information is available, execute S210;
[0094] If not, that is, when phase information is unavailable (only amplitude frequency data is available), execute S211.
[0095] S210. Select either the frequency warping FIR filter algorithm or the constant Q transform processing algorithm as the target algorithm.
[0096] S211. Select the minimum phase reconstruction algorithm as the target algorithm.
[0097] S212. Select the frequency domain multiplication algorithm as the target algorithm.
[0098] Through this adaptive selection process, the optimal target algorithm can be dynamically matched from a set of preset virtual playback simulation algorithms based on various application constraints such as processing mode, latency requirements, availability of phase information, and accuracy requirements. This enables full-scenario coverage from real-time monitoring to offline mastering, and from resource-constrained devices to high-performance workstations.
[0099] Accordingly, if there is a need to optimize subjective listening experience, an algorithm that matches the frequency resolution processing characteristics of human hearing is selected. In some embodiments, the algorithm that matches the frequency resolution processing characteristics of human hearing includes a frequency warping FIR filter algorithm or a constant Q transform processing algorithm. The frequency warping FIR filter algorithm determines the frequency warping factor based on the characteristics of human hearing perception, and performs a warping transformation on the frequency axis based on the frequency warping factor to enhance the filter resolution in the low-frequency region. The constant Q transform processing algorithm analyzes the audio signal using logarithmically spaced frequency points with a constant quality factor, and the ratio of the center frequency to the bandwidth of each frequency point remains constant.
[0100] For example, the frequency warping FIR filter algorithm first determines the frequency warping factor based on the characteristics of human auditory perception. The warp factor The value is typically between 0.3 and 0.9, calculated based on the sampling rate, to approximate the Bark frequency scale or equivalent rectangular bandwidth scale in psychoacoustics, thereby simulating the frequency selectivity characteristics of the human ear's basilar membrane. For example, for a sampling rate of 44.1 kHz, Typically, a value of 0.6 to 0.7 is used; for a 48kHz sampling rate, Use a value between 0.6 and 0.7; for a 96kHz sampling rate, Take a value of 0.7 to 0.8. This is used to determine the warpage factor. Subsequently, the algorithm employs a frequency warping technique to perform a nonlinear transformation on the frequency axis. This transformation is based on a variant of the bilinear transformation, converting the unit delay unit in a traditional FIR filter into a nonlinear transformation. Replace it with a warped delay unit composed of a first-order all-pass filter. The mathematical expression for the frequency warp transform is as follows:
[0101]
[0102] in, This refers to the frequency after warping. It is the normalized frequency, defined as the actual frequency. The ratio to the Nyquist frequency (half the sampling rate), i.e. , This refers to the Nyquist frequency, with values ranging from [-1,1] to [0,1]. Through the aforementioned nonlinear transformation, the frequency axis is remapped: in the low-frequency region ( (when close to 0) Compared to Being stretched means that the region achieves higher frequency resolution; in the high-frequency region ( (approaching 1) Compared to Compression implies a relative decrease in resolution. This transformation characteristic closely matches the physiological characteristics of the human auditory system, which is more sensitive to low frequencies and whose frequency resolution decreases as frequency increases. In the warped frequency domain, interpolation yields the sound pressure level gain and phase response of the target audio playback device corresponding to the original frequency. Due to the uneven distribution of frequency points in the warped frequency domain, the amplitude and phase response data of the target audio playback device need to be interpolated from the original frequency domain to each frequency point in the warped frequency domain to ensure that the filter design accurately matches the characteristics of the audio playback device. Next, FIR filter coefficients are designed on the warped frequency axis. Specifically, the initial FIR filter response is designed in the warped frequency domain based on the amplitude characteristics of the target audio playback device, and then the filter coefficients in the original frequency domain are obtained through inverse warping transformation. Since the warping transformation enhances the resolution in the low-frequency region, more precise control over low-frequency details can be achieved while maintaining the filter order, and fewer filter taps are required than in traditional linear FIR filters. Furthermore, the designed FIR filter undergoes phase adjustment to match the phase response characteristics of the target audio playback device. This can be achieved, for example, by multiplying the filter response in the frequency domain by a phase correction term or by cascading all-pass filters. Finally, the designed filter coefficients are used to filter the input audio signal, generating an output audio signal. Because the frequency resolution distribution better matches the characteristics of human hearing, the processed audio sounds more natural and realistic, providing a better listening experience.
[0103] Accordingly, the constant Q transform processing algorithm analyzes audio signals using logarithmically spaced frequency points with a constant quality factor. The quality factor Q is defined as the ratio of the center frequency to the bandwidth and remains constant throughout the analysis frequency range. First, the frequency point distribution for the constant Q transform is determined, with the frequency points set according to a logarithmic law. Its mathematical expression is: ,in, For the first The center frequency of each frequency point For the lowest analysis frequency, The index number of the frequency point ( =0, 1, 2, ...), This represents the number of frequency points per octave. Since the frequency points are logarithmically spaced, the ratio of adjacent frequency points remains constant. ,when Increase At that time, frequency Doubling (i.e., increasing by one octave). Based on the above frequency point settings, the quality factor Q of the constant Q transform is determined by the number of frequency points per octave. It is uniquely determined, and its mathematical expression is: For frequency points with logarithmic intervals, the interval between adjacent frequency points is the filter bandwidth. ,satisfy Therefore, the quality factor The frequency response remains constant throughout the analysis. This means that the low-frequency region has a narrower bandwidth (high frequency resolution) and the high-frequency region has a wider bandwidth (low frequency resolution), a characteristic consistent with the frequency selectivity of the human auditory system.
[0104] Next, the sound pressure level data of the target audio playback device is interpolated to each constant Q frequency point to obtain the gain value G_k corresponding to each frequency point. Based on the center frequency of each frequency point... Given the quality factor Q and gain G_k, design a corresponding bandpass filter for each frequency point. These filters can be implemented using FIR or IIR filters, with the passband center frequency of each filter being [value missing]. The bandwidth is determined by the Q value. During the analysis phase, the input audio signal is convolved through the bandpass filter corresponding to each frequency point to obtain the time-domain output signal for each frequency band. That is, for each frequency point k, the time-domain output signal is calculated. Where h_k[n] is the bandpass filter coefficient of the k-th frequency band. Then, the output signals of each frequency band are weighted and superimposed, and the weight is the gain value G_k corresponding to that frequency band, to obtain the weighted sum signal y_sum[n]=ΣG_k×y_k[n]. The weighted and superimposed signal is added to the original audio signal according to a certain ratio to obtain the final output signal y[n]=α×x[n]+β×y_sum[n], where α and β are mixing ratio coefficients, which can be adjusted as needed to control the processing intensity. Through the above constant Q transform processing algorithm, the bandwidth of the bandpass filter at each frequency point is adaptively adjusted with the center frequency, realizing variable time-frequency resolution: the low-frequency region obtains higher frequency resolution and lower time resolution, and the high-frequency region obtains lower frequency resolution and higher time resolution. Since the frequency resolution distribution of the transform domain is consistent with the hearing characteristics of the human ear, the processed signal is more natural in subjective listening and can better simulate the playback effect of the target audio playback device.
[0105] By using the two perception optimization algorithms mentioned above, while ensuring processing accuracy, the auditory effect of virtual simulation can be made more in line with the characteristics of human hearing, significantly improving the naturalness and immersion of the sound.
[0106] In some embodiments, if the target algorithm is a frequency domain multiplication algorithm, the target algorithm is used to process the input audio signal based on the complex frequency response data, including: performing a Fast Fourier Transform on the input audio signal to convert the time-domain signal x(t) to the frequency domain and obtain its spectrum X(f); simultaneously, processing the acquired complex frequency response data of the target audio playback device, interpolating the sound pressure level data and phase data to the same frequency point as the FFT, and constructing a complex frequency response H(f) = |H(f)| × exp(jΦ(f)); then, performing a complex multiplication of the spectrum X(f) of the input audio signal and the complex frequency response H(f) of the target audio playback device in the frequency domain to obtain the spectrum Y(f) = X(f) × H(f) of the output signal; further, performing an Inverse Fast Fourier Transform on Y(f) to restore the signal from the frequency domain to the time domain, obtaining the output audio signal y(t) = IFFT{Y(f)}, and taking the real part to obtain the output audio signal simulating the playback effect of the target audio playback device.
[0107] In some embodiments, if the target algorithm is an FIR filter convolution algorithm, the target algorithm is used to process the input audio signal based on the complex frequency response data, including: converting the sound pressure level data into a linear gain value, the conversion formula being: , That is, the gain value. The sound pressure level data is used; the frequency is normalized to the Nyquist frequency range (half the sampling rate); based on the processed amplitude-frequency response and phase-frequency response data, finite impulse response (FIR) filter coefficients are designed. Specifically, initial filter coefficients are designed in the frequency domain according to the target frequency response H(f), and then the initial coefficients are subjected to a fast Fourier transform (FFT) to adjust their phase response in the frequency domain to match the phase-frequency characteristics of the target audio playback device. The final FIR filter coefficients h[k] are then obtained through an inverse fast Fourier transform (IFF), where k = 0, 1, ..., M-1, and M is the filter order. Then, the input audio signal x[n] is convolved with the designed FIR filter coefficients h[k] in the time domain, and the convolution formula is y[n] = Σh[k] × x[nk] (k from 0 to M-1) to obtain the output audio signal y[n]. Through the above FIR filter convolution algorithm, the amplitude-frequency characteristics and phase-frequency characteristics of the target audio playback device can be accurately realized in the time domain, which is suitable for general processing scenarios that require simultaneous control of amplitude and phase.
[0108] In some embodiments, if the target algorithm is a cascaded IIR bi-second-order filter algorithm, the target algorithm is used to process the input audio signal based on complex frequency response data, including: generating logarithmically distributed center frequency points based on the complex frequency response data of the target audio playback device, these frequency points covering the main frequency response characteristic region of the target audio playback device; then, interpolating the sound pressure level data at the generated center frequency points to obtain the gain value corresponding to each center frequency point; simultaneously, determining the quality factor Q value of each frequency band based on the frequency response characteristics of the target audio playback device; and calculating the corresponding peak equalizer (PEQ) bi-second-order filter coefficients using the audio EQ cookbook formula based on the gain value, center frequency, and Q value of each center frequency point. The transfer function of each bi-second-order filter satisfies: , where the coefficient The center frequency, gain, and Q value are calculated using EQ cookbook formulas. Depending on the needs, some frequency bands can be handled by either low-profile or high-profile filters to better approximate the low- and high-frequency characteristics of the audio playback device's response. Then, all the designed bi-second-order filters are cascaded to form a complete IIR filter network, whose overall transfer function is the product of all cascaded stages: H_total(z) = ΠH_i(z). Finally, the input audio signal is sequentially filtered through the cascaded IIR bi-second-order filter network to obtain the output audio signal. This cascaded IIR bi-second-order filter algorithm can approximate the amplitude-frequency characteristics of the target audio playback device with low computational complexity and extremely low latency, making it suitable for embedded devices with limited computing resources and real-time monitoring scenarios with strict latency requirements.
[0109] In some embodiments, if the target algorithm is an overlapping block processing algorithm, the target algorithm is used to process the input audio signal based on complex frequency response data, including: dividing the input audio signal into multiple consecutive data blocks, each data block having a specified block size, with overlapping regions between adjacent data blocks, the overlap ratio being selectable between 20% and 80% according to actual needs; applying an analysis window function, such as a Hanning window, to each data block to reduce the spectral leakage effect that may be introduced by block processing, and padding the windowed data blocks with zeros to ensure that their length meets the requirements of the Fast Fourier Transform; then, performing a Fast Fourier Transform on each windowed and zero-pasted data block to convert it to frequency response data. In the frequency domain, the spectrum X_block(f) of each data block is obtained. Simultaneously, the complex frequency response data of the target audio playback device is processed to construct a complex frequency response H(f), which is then interpolated to the same frequency point as the FFT. In the frequency domain, the spectrum X_block(f) of each data block is multiplied by the complex frequency response H(f) of the target audio playback device to obtain the processed spectrum Y_block(f) = X_block(f) × H(f). Then, an inverse fast Fourier transform is performed on the processed spectrum Y_block(f) of each data block to restore it to the time domain, obtaining the time-domain processing result of each data block. Finally, the time-domain processing results of all data blocks are overlapped and added together according to the same overlap ratio as during block segmentation, i.e., the samples in the overlapping region are summed to reconstruct the continuous output audio signal y[n] = ΣIFFT{FFT{w[n] × x_block[n]} × H(f)}, where w[n] is the analysis window. The above-described overlapping and additive block processing algorithm can efficiently perform convolution operations in the frequency domain, while controlling processing latency through block processing. It is suitable for general audio processing scenarios that require a balance between processing efficiency and latency.
[0110] In some embodiments, if the target algorithm is an overlap-preserving block processing algorithm, the target algorithm is used to process the input audio signal based on complex frequency response data, including: determining the block size and filter length, assuming the filter length is M and the FFT length is N, where N is usually an integer power of 2 greater than M to meet the requirements of the Fast Fourier Transform; preprocessing the input audio signal by adding M-1 zero samples as leading zero padding at its beginning position to ensure that the initial block can be processed correctly; dividing the preprocessed audio signal into multiple overlapping data blocks, each data block having a length of N, with an overlapping region between adjacent data blocks, the overlap length being M-1 samples, that is, each new block contains the last M-1 samples of the previous block and new N-(M-1) samples. For each data block, a Fast Fourier Transform (FFT) is performed to convert it to the frequency domain, yielding the spectrum X_block(f). Simultaneously, the acquired complex frequency response data of the target audio playback device is processed to construct a complex frequency response H(f), which is then interpolated to the same frequency point as the FFT. In the frequency domain, the spectrum X_block(f) of each data block is multiplied by the complex frequency response H(f) of the target audio playback device, resulting in the processed spectrum Y_block(f) = X_block(f) × H(f). Then, an Inverse Fast Fourier Transform (IFFT) is performed on the processed spectrum Y_block(f) of each data block to restore it to the time domain, yielding the time-domain processing result of each data block. Since the frequency domain multiplication is actually a circular convolution rather than a non-linear convolution, the first M-1 samples in the time-domain result of each data block contain circular convolution artifacts and need to be discarded. The valid portion of each data block is retained, i.e., the samples from the M-1th sample to the end of the block, mathematically expressed as Valid. `output = IFFT{FFT{x_block} × H(f)} [M-1:N-1]`. Finally, the valid portions of all data blocks are concatenated sequentially to obtain the complete output audio signal. This overlapping block processing algorithm efficiently achieves linear convolution in the frequency domain without applying a window function to the input data, making it suitable for audio processing scenarios requiring precise convolution operations.
[0111] In some embodiments, if the target algorithm is a minimum phase reconstruction algorithm, the target algorithm is used to process the input audio signal based on the complex frequency response data, including: firstly, processing the amplitude frequency response data of the acquired target audio playback device, interpolating the sound pressure level data to the frequency point that matches the subsequent fast Fourier transform to obtain a continuous amplitude response |H(f)|; calculating the logarithmic amplitude spectrum log|H(f)| of the amplitude response, and performing a Hilbert transform on the logarithmic amplitude spectrum to obtain the corresponding minimum phase response Φ_min(f), whose mathematical expression is Φ_min(f)=-Hilbert{log|H(f)|}. Through this transformation, the minimum phase characteristic that uniquely corresponds to it is reconstructed from the only amplitude response. Next, based on the interpolated amplitude response |H(f)| and the reconstructed minimum phase response Φ_min(f), the minimum phase complex frequency response H_min(f) = |H(f)| × exp(jΦ_min(f)) is constructed; the minimum phase complex frequency response H_min(f) is then subjected to an inverse fast Fourier transform to obtain the minimum phase impulse response h_min[n]. Due to the minimum phase characteristic, this impulse response has the characteristics of concentrated energy and the fastest decay, and it is physically realizable. Optionally, a Hanning window is applied to the obtained minimum phase impulse response for optimization to reduce spectral leakage caused by the truncation effect and improve the stability and performance of the filter. Finally, the input audio signal x[n] is convolved with the optimized minimum phase impulse response h_min[n] in the time domain to obtain the output audio signal. The minimum phase reconstruction algorithm described above can reconstruct a physically realizable minimum phase filter with optimal transient characteristics, even with only the amplitude-frequency response data of the audio playback device. This is suitable for application scenarios where phase data is missing but system causality and stability still need to be guaranteed.
[0112] In some embodiments, if the target algorithm is an impulse response convolution algorithm, the target algorithm is used to process the input audio signal based on the complex frequency response data, including: processing the acquired complex frequency response data of the target audio playback device to construct a complete complex frequency response H(f), which includes amplitude frequency response data and phase frequency response data, fully characterizing the acoustic characteristics of the target audio playback device; performing an inverse fast Fourier transform on the constructed complex frequency response H(f) to obtain the corresponding time-domain impulse response h[n], with the mathematical expression h[n]=IFFT{H(f)}; performing a cyclic shift operation on the obtained impulse response to align the phase characteristics of the impulse response, ensuring its causality and the accuracy of time-domain alignment. Next, a Hanning window is applied to the aligned impulse response to reduce the spectral leakage effect that may be introduced by impulse response truncation, improving the performance and stability of the filter. Finally, the input audio signal x[n] and the processed impulse response h[n] are convolved in the time domain to obtain the output audio signal. In practical implementation, efficient fast Fourier transform convolution (such as the overlap-addition method or the overlap-preservation method) can be used to improve computational efficiency. Using the aforementioned impulse response convolution algorithm, a time-domain impulse response can be directly constructed from complex frequency response data, and the acoustic characteristics of the target audio playback device can be accurately realized through convolution operations. This is suitable for audio processing scenarios that require complete preservation of amplitude and phase frequency characteristics.
[0113] In some embodiments, if the target algorithm is a parametric equalizer approximation algorithm, the target algorithm is used to process the input audio signal based on complex frequency response data, including: first, determining a multi-band structure to approximate the response of the audio playback device, specifically using relevant standard frequency points, such as a 1 / 3 octave band 31-segment equalizer structure. These frequency points are logarithmically distributed and can effectively cover the entire audio frequency band; interpolating the obtained sound pressure level data of the target audio playback device onto these frequency points to obtain the gain value corresponding to the center frequency of each frequency band; simultaneously, determining the required phase correction amount for each frequency band based on the phase response data; next, designing a corresponding dual second-order filter for each effective frequency band. For the amplitude response, a peak equalizer dual second-order filter is used, where each peak equalizer is determined by its center frequency, gain, and quality factor Q value to adjust the amplitude characteristics of that frequency band; for the phase response, a corresponding all-pass filter is designed for phase correction, where the all-pass filter can adjust the phase characteristics while maintaining the amplitude response unchanged. All the designed second-order filters are organized into a second-order section (SOS) filter bank, where each second-order section corresponds to a processing unit for one frequency band. The overall transfer function of this filter bank is a cascade of the products of the second-order sections of each peak equalizer and the second-order sections of each all-pass filter, mathematically expressed as H_total(z) = ΠH_peaking(z) × ΠH_allpass(z), where H_total(z) is the overall transfer function, H_peaking(z) is the peak equalizer transfer function, and H_allpass(z) is the all-pass filter transfer function. Finally, the input audio signal is processed sequentially through this SOS filter bank, that is, through the cascaded peak equalizer and all-pass filter step by step, to obtain the output audio signal. Through the above parametric equalizer approximation algorithm, the amplitude-frequency and phase-frequency characteristics of the target audio playback device can be flexibly approximated in the manner of a parametric equalizer, which is suitable for audio processing scenarios that require independent control of parameters of each frequency band.
[0114] In some embodiments, if the target algorithm is a partitioned convolution algorithm, the target algorithm is used to process the input audio signal based on the complex frequency response data, including: dividing the target impulse response into multiple continuous segments, the target impulse response being obtained from the complex frequency response data through inverse transform; performing a fast Fourier transform on each segment to obtain the corresponding frequency domain segmentation coefficients; constructing a frequency domain delay line to store the frequency domain representations of multiple historical input audio signals; multiplying and accumulating the frequency domain representation of the current input audio signal with each frequency domain segmentation coefficient to obtain the frequency domain output result; and performing an inverse fast Fourier transform on the frequency domain output result to obtain the output audio signal.
[0115] For example, the acquired complex frequency response data of the target audio playback device is processed to construct a complete complex frequency response H(f). An inverse fast Fourier transform (IFFT) is then performed on this complex frequency response to obtain the corresponding long impulse response h[n]. This long impulse response may contain thousands or even tens of thousands of sampling points, such as the impulse response simulating the acoustic characteristics of a room. Then, the obtained long impulse response h[n] is divided into multiple continuous segments h_k[n], where k = 0, 1, ..., K-1, and K is the total number of segments. The length of each segment can be set according to actual needs, usually using an equal-length segmentation strategy. After zero-padding each segment, an IFFT is performed to obtain the corresponding frequency domain segmentation coefficients H_k(f). These coefficients can be pre-calculated and stored for real-time processing. Next, a frequency domain delay line is constructed. This frequency domain delay line is the core data structure in the partitioned convolution algorithm, used to store the frequency domain representations of multiple historical input audio signals. The depth of the frequency domain delay line corresponds to the number of segments K of the impulse response, and each storage unit corresponds to the frequency domain representation of a historical input block. In real-time processing, the input audio signal is segmented into input blocks matching the segment length. A Fast Fourier Transform (FFT) is performed on each input block to obtain its frequency domain representation X_current(f). X_current(f) is then multiplied and accumulated together with the frequency domain representations of each historical input block stored in the frequency domain delay line, and then multiplied and accumulated with the frequency domain segmentation coefficients H_k(f), i.e., Y(f) = ΣH_k(f) × X_delayed_k(f), where X_delayed_k(f) is the frequency domain representation of the k-th historical input block stored in the frequency domain delay line. An Inverse Fast Fourier Transform (IFT) is performed on the frequency domain output Y(f) obtained from the multiplication and accumulation to obtain the time domain output block corresponding to the current input block. These output blocks are then concatenated according to an appropriate overlap pattern to reconstruct the continuous time-domain output audio signal. This partitioned convolution algorithm enables real-time convolution processing of ultra-long impulse responses while maintaining low latency, making it suitable for applications requiring simulation of room acoustics, reverberation effects, and other scenarios involving long impulse responses.
[0116] In some embodiments, if the target algorithm is a frequency sampling algorithm, the target algorithm is used to process the input audio signal based on the complex frequency response data, including: directly sampling the complex frequency response of the target audio playback device at the frequency points of the Fast Fourier Transform (FFT) to construct a complete complex frequency response H(f). These frequency points are uniformly distributed in the range from 0 to the Nyquist frequency, matching the frequency resolution of the subsequent FFT; adding a phase offset term e to the complex frequency response. (-j2πfτ)τ is the group delay parameter used for signal alignment. This phase offset term is used to adjust the group delay characteristics of the filter, ensuring that the designed FIR filter is causal and achieves time-domain alignment. Next, an inverse fast Fourier transform is performed on the complex frequency response after adding the phase offset to obtain the initial FIR filter coefficients h'[n] = IFFT{H(f) × e (-j2πfτ) To suppress the Gibbs phenomenon (i.e., time-domain ripple caused by frequency domain abrupt changes) that may be introduced by direct sampling of frequency points, an analysis window w[n] is applied to the initial filter coefficients for optimization. This analysis window is, for example, a Kaiser window. The parameter β of the Kaiser window can be set as needed, for example, β=8, to achieve a good balance between sidelobe suppression and main lobe width. The filter coefficients after windowing are h[n]=h'[n]×w[n]. Furthermore, the FIR filter coefficients after windowing are normalized to ensure that the overall gain of the filter meets expectations, so as to reduce unnecessary amplitude changes introduced during signal processing. Finally, the input audio signal x[n] and the designed FIR filter coefficients h[n] are convolved in the time domain to obtain the output audio signal. In practical implementation, efficient Fast Fourier Transform (FFT) convolution can be used to improve computational efficiency. Using the aforementioned frequency sampling algorithm, the response of the target audio playback device can be sampled directly in the frequency domain. Furthermore, a high-performance FIR filter can be designed through windowing optimization, making it suitable for audio processing scenarios requiring precise frequency response control and tolerating a certain degree of pre-ringing.
[0117] In some embodiments, if the target algorithm is a polyphase filter bank algorithm, the target algorithm is used to process the input audio signal based on complex frequency response data, including: dividing the entire audio frequency band into multiple equally spaced frequency sub-bands according to processing requirements, each sub-band corresponding to a center frequency, and these center frequencies being evenly distributed in the range from 0 to the Nyquist frequency; processing the obtained sound pressure level data of the target audio playback device, interpolating to obtain the gain value G_k corresponding to the center frequency point of each frequency sub-band, where k is the sub-band index, and these gain values reflect the amplitude response characteristics of the target audio playback device in each frequency band. Next, a polyphase filter bank is designed, including an analysis filter bank and a synthesis filter bank. The analysis filter bank is used to decompose the input audio signal into various frequency sub-bands, and the synthesis filter bank is used to reconstruct the processed sub-band signals into a complete audio signal. In the analysis phase, the input audio signal is divided into data blocks that match the filter bank design. An analysis window function w[n] is applied to each data block, and then a Discrete Fourier Transform (DFT) is performed to obtain the frequency domain representation of each sub-band, X_k[m] = DFT{x[n] × w[n]}, where k is the sub-band index and m is the time frame index. In the gain adjustment phase, the corresponding gain value G_k is applied to the frequency domain representation X_k[m] of each sub-band to obtain the processed sub-band signal Y_k[m] = X_k[m] × G_k. This step enables independent control of the amplitude of each frequency band to match the amplitude-frequency characteristics of the target audio playback device. In the synthesis stage, the processed sub-band signals Y_k[m] undergo an inverse discrete Fourier transform (IDFT) to restore them to the time domain. Then, a windowed overlap-addition method is used to synthesize the time-domain signals of all sub-bands: a synthesis window function is applied to the time-domain output blocks of each sub-band, and the signals are overlapped and added according to the same overlap ratio as during block division, reconstructing a continuous output audio signal. Through this multiphase filter bank algorithm, independent gain adjustment of each sub-band can be performed in the frequency domain, achieving flexible frequency response control, suitable for audio applications requiring fine-grained frequency domain processing.
[0118] In some embodiments, if the target algorithm is a weighted overlap addition algorithm, the target algorithm is used to process the input audio signal based on complex frequency response data, including: first, setting the size of the data block and the overlap factor. The block size determines the length of data processed each time, and the overlap factor determines the overlap ratio between adjacent data blocks. It is usually set to 50% overlap to ensure that the constant overlap addition condition is met; then, the input audio signal is divided into multiple consecutive data blocks, and the length of each data block is consistent with the set block size; an analysis window function w_a[n] is applied to each data block. The analysis window adopts the square root Hanning window. The square root Hanning window is characterized by its square value satisfying the shape of the Hanning window, so that perfect reconstruction can be achieved when overlapping and adding. Accordingly, a Fast Fourier Transform (FFT) is performed on each windowed data block to convert it to the frequency domain, yielding the spectrum X_block(f) of each data block. Simultaneously, the complex frequency response data of the target audio playback device is processed to construct a complex frequency response H(f), which is then interpolated to the same frequency point as the FFT. In the frequency domain, the spectrum X_block(f) of each data block is multiplied by the complex frequency response H(f) of the target audio playback device to obtain the processed spectrum Y_block(f) = X_block(f) × H(f). An Inverse Fast Fourier Transform (IFFT) is performed on the processed spectrum Y_block(f) of each data block to restore it to the time domain, yielding the time-domain processing result of each data block. A synthesis window function w_s[n] is applied to each time-domain result. This synthesis window also uses a square root Hanning window, matching the analysis window, i.e., w_a = w_s = square root Hanning window. Furthermore, all data blocks after applying the synthesis window are overlapped and added together according to the same overlap ratio as during block segmentation, i.e., the samples in the overlapping region are summed. After overlap and addition, each output sample is normalized and divided by the sum of the products of the analysis window and the synthesis window, Σ(w_a×w_s). Since both the analysis window and the synthesis window use square root Hanning windows, they satisfy the constant overlap-add (COLA) constraint condition Σw²=constant, thus achieving perfect signal reconstruction without filtering. Finally, through the above normalization process, the complete output audio signal y[n]=Σw_s[n]×IFFT{FFT{w_a[n]×x_block[n]}×H(f)} / Σ(w_a×w_s). Through the above weighted overlap-add algorithm, convolution operations can be efficiently implemented in the frequency domain. At the same time, the matching analysis window and synthesis window design ensures that the COLA condition is met, achieving perfect reconstruction without filtering. This is suitable for audio application scenarios that require high-precision frequency domain processing.
[0119] In some embodiments, if the target algorithm is a zero-delay convolution algorithm, the target algorithm is used to process the input audio signal based on the complex frequency response data, including: decomposing the complex frequency response data into minimum phase components and all-pass components; constructing a minimum phase impulse response based on the minimum phase components to achieve zero prering characteristics; designing a cascaded first-order all-pass filter based on the all-pass components to approximate the residual phase characteristics of the complex frequency response data; and cascading the minimum phase impulse response with the cascaded first-order all-pass filter to perform convolution processing on the input audio signal to obtain the output audio signal.
[0120] For example, the acquired complex frequency response data H(f) of the target audio playback device is decomposed into a minimum phase component H_minphase(f) and an all-pass component H_allpass(f), satisfying H(f) = H_minphase(f) × H_allpass(f). The minimum phase component is responsible for the amplitude-frequency characteristics and can be obtained from the amplitude-frequency response using minimum phase reconstruction techniques. The all-pass component is responsible for the excess phase characteristics, i.e., the remaining portion after subtracting the minimum phase from the original phase. Then, based on the decomposed minimum phase component H_minphase(f), a minimum phase impulse response h_minphase[n] is constructed using inverse fast Fourier transform. Due to the characteristics of the minimum phase system, this impulse response has concentrated energy and can achieve zero prering, meaning the signal will not exhibit oscillatory distortion before the input excitation arrives. This is crucial for maintaining the transient response of the audio signal. Next, the difference between the target phase Φ_target(f) and the minimum phase Φ_minphase(f) is calculated to obtain the excess phase characteristic Φ_excess(f) = Φ_target(f) - Φ_minphase(f). Based on this excess phase characteristic, a cascaded first-order all-pass filter is designed to approximate the residual phase characteristic of the complex frequency response data. The transfer function of each first-order all-pass filter is H_allpass(z) = (a + z⁻¹) / (1 + az⁻¹), where the coefficient a is determined by the phase difference to be compensated. By cascading multiple first-order all-pass filters, arbitrarily complex excess phase characteristics can be accurately approximated while keeping the amplitude response unchanged (the amplitude-frequency response of the all-pass filter is always 1). Finally, the constructed minimum phase impulse response h_minphase[n] is cascaded with the cascaded first-order all-pass filter H_allpass(z) to form a complete zero-delay convolution processing chain. The input audio signal x[n] is processed sequentially through this cascaded system: first, it is convolved with the minimum phase impulse response to accurately restore the amplitude-frequency characteristics; then, it is passed through a cascaded all-pass filter to accurately approximate the excess phase characteristics, finally obtaining the output audio signal y[n]. This zero-delay convolution algorithm achieves accurate restoration of the original phase characteristics while maintaining near-zero delay. It reduces the pre-ringing distortion that may occur with traditional linear-phase FIR filters, while ensuring the signal's temporal waveform and spatial sense. It is suitable for applications sensitive to delay and requiring high phase accuracy, such as real-time monitoring and virtual reality audio.
[0121] In summary, this application has at least the following advantages:
[0122] First, it abandons the single and fixed processing solutions of related technologies, and integrates a variety of virtual playback simulation algorithms with complementary characteristics, such as frequency domain multiplication, FIR / IIR filtering, partitioned convolution, zero-delay convolution, and frequency warping. This forms a full-range algorithm pool covering embedded systems to high-performance workstations, and from real-time monitoring to offline mastering processing. It can flexibly meet the needs of diverse application scenarios such as professional audio production, consumer electronics, car audio, and virtual reality.
[0123] Second, through an adaptive algorithm selection mechanism, the system can automatically select the optimal target algorithm from the algorithm pool based on application constraints such as maximum allowable processing latency, available computing resources, and phase accuracy requirements. For example, in real-time monitoring scenarios with strict latency requirements, it automatically selects a cascaded IIR dual second-order filter algorithm with near-zero latency; in mastering scenarios that pursue ultimate accuracy, it automatically selects the frequency domain multiplication algorithm with the highest accuracy. This adaptive mechanism resolves the contradiction in existing technologies where a single solution cannot simultaneously address latency, accuracy, and computational efficiency.
[0124] Third, by acquiring complex frequency response data including amplitude and phase responses and employing various phase preservation or reconstruction methods, the time-domain waveform distortion problem caused by related technologies that only process the amplitude response and ignore phase characteristics is effectively improved. Even in the case of missing phase data, a physically realizable filter with optimal transient response can be obtained through minimum phase reconstruction technology, significantly improving the realism of virtual simulation and the accuracy of spatial positioning.
[0125] Fourth, by employing a frequency warping FIR filter algorithm and a constant Q transform processing algorithm, and using a frequency resolution processing method that matches the characteristics of human auditory perception, higher frequency resolution is achieved in the low-frequency region, while the resolution in the high-frequency region is relatively reduced, which highly matches the frequency selectivity characteristics of the human auditory system. Compared with the harsh listening experience caused by linear frequency resolution in related technologies, the audio processed by the solution in this application is more natural, smooth, and immersive in subjective listening.
[0126] V. The various algorithms provided in this application have significant differences in computational complexity, ranging from a low-complexity cascaded IIR double second-order filter algorithm suitable for resource-constrained embedded devices to a high-precision partitioned convolution algorithm suitable for high-performance workstations, forming a complete gradient of computational complexity. This enables flexible adaptation to different hardware platforms, from wireless headphones and smartphones to professional audio workstations, achieving a balance between energy efficiency and performance.
[0127] VI. Virtual simulation technology significantly reduces reliance on physical equipment such as physical audio playback devices and standard listening rooms, greatly lowering R&D, testing, and collaboration costs in fields such as professional audio production, consumer electronics audio tuning, and automotive audio design. Simultaneously, the virtualization solution supports remote workflows, enabling team members to share unified audio processing effects across geographical locations, significantly improving engineering efficiency and collaboration convenience.
[0128] Figure 3 A schematic diagram of a virtual playback simulation device provided as an exemplary embodiment of this application. (See diagram below.) Figure 3 As shown, the virtual playback simulation device 30 includes an input module 31, an algorithm selection module 32, and a processing module 33, wherein:
[0129] Input module 31 is used to acquire complex frequency response data of the target audio playback device, including amplitude frequency response data and phase frequency response data; and to acquire input audio signals;
[0130] The algorithm selection module 32 is used to adaptively select a target algorithm from multiple preset virtual playback simulation algorithms according to preset application constraints; wherein, the multiple preset virtual playback simulation algorithms have different latency characteristics, accuracy levels and computational complexity.
[0131] The processing module 33 is used to process the input audio signal based on the complex frequency response data using the target algorithm to generate an output audio signal that simulates the playback effect of the target audio playback device.
[0132] In one possible implementation, the preset application constraints include at least one of the following: maximum allowable processing delay, available computing resources, phase accuracy requirements, and processing mode. The algorithm selection module 32 can be specifically used to: obtain at least one of the following as the current constraints: maximum allowable processing delay, available computing resources, phase accuracy requirements, and processing mode; and based on the current constraints, match an algorithm that is compatible with the current constraints from a plurality of preset virtual playback simulation algorithms as the target algorithm.
[0133] In one possible implementation, based on the current constraints, an algorithm that is suitable for the current constraints is selected as the target algorithm from a plurality of preset virtual playback simulation algorithms, including at least one of the following: when the processing mode is real-time processing mode and the maximum allowable processing latency is lower than a preset latency threshold, the algorithm with the lowest latency characteristic is selected as the target algorithm from a plurality of preset virtual playback simulation algorithms; when the processing mode is offline processing mode and the phase accuracy requirement is higher than a preset accuracy threshold, the algorithm with the highest accuracy level is selected as the target algorithm from a plurality of preset virtual playback simulation algorithms; when the available computing resources are lower than a preset resource threshold, the algorithm with the lowest computational complexity is selected as the target algorithm from a plurality of preset virtual playback simulation algorithms.
[0134] In one possible implementation, the multiple preset virtual playback simulation algorithms include at least three of the following algorithm sets: frequency domain multiplication algorithm, FIR filter convolution algorithm, cascaded IIR double second-order filter algorithm, overlapping summing block processing algorithm, overlapping retaining block processing algorithm, minimum phase reconstruction algorithm, impulse response convolution algorithm, parametric equalizer approximation algorithm, frequency warping FIR filter algorithm, partitioned convolution algorithm, frequency sampling algorithm, polyphase filter bank algorithm, weighted overlapping summing algorithm, constant Q transform processing algorithm, and zero-delay convolution algorithm; wherein, the zero-delay convolution algorithm refers to an algorithm that decomposes the complex frequency response data into minimum phase components and all-pass components for separate processing.
[0135] In one possible implementation, the frequency warping FIR filter algorithm determines the frequency warping factor based on the characteristics of human auditory perception, and performs a warping transformation on the frequency axis based on the frequency warping factor to enhance the filter resolution in the low-frequency region.
[0136] In one possible implementation, the constant Q transform processing algorithm analyzes the audio signal using logarithmically spaced frequency points with a constant quality factor, where the ratio of the center frequency to the bandwidth at each frequency point remains constant.
[0137] In one possible implementation, the processing module 33 may be specifically used to: divide the target impulse response into multiple continuous segments, the target impulse response being obtained from the complex frequency response data through inverse transform; perform a fast Fourier transform on each segment to obtain the corresponding frequency domain segmentation coefficients; construct a frequency domain delay line to store the frequency domain representations of multiple historical input audio signals; multiply and accumulate the frequency domain representation of the current input audio signal with each frequency domain segmentation coefficient to obtain the frequency domain output result; and perform an inverse fast Fourier transform on the frequency domain output result to obtain the output audio signal.
[0138] In one possible implementation, the processing module 33 can also be used to: decompose the complex frequency response data into minimum phase components and all-pass components; construct a minimum phase impulse response based on the minimum phase components to achieve zero prering characteristics; design a cascaded first-order all-pass filter based on the all-pass components to approximate the residual phase characteristics of the complex frequency response data; cascade the minimum phase impulse response with the cascaded first-order all-pass filter to perform convolution processing on the input audio signal to obtain the output audio signal.
[0139] In some embodiments, the virtual playback simulation device further includes a preprocessing module and an output module. For example, Figure 4 This is a schematic diagram of the architecture of a virtual playback simulation device provided for an exemplary embodiment of this application. Figure 4As shown, the device includes an input module, a preprocessing module, an algorithm selection module, a processing module, an output module, and an audio playback device. The input module receives digital audio signals of various formats and audio playback device characterization data. The digital audio signals can come from various sources, such as PCM format audio files, streaming media data, or real-time recording input. The audio playback device characterization data includes the sound pressure level (SPL) data and phase response data of the target audio playback device, which together constitute the complex frequency response data of the audio playback device. The preprocessing module, connected to the input module, preprocesses the input digital audio signals and audio playback device characterization data. Specifically, it normalizes the digital audio signals to floating-point representation for easier subsequent signal processing; it dewinds the phase data of the audio playback device to eliminate phase jumps; and it interpolates the SPL and phase data to the frequency bins required for subsequent processing to ensure they match the frequency resolution of the processing module, thus constructing a complete transfer function. The algorithm selection module, connected to the preprocessing module, adaptively selects a target algorithm from multiple preset virtual playback simulation algorithms based on preset application constraints. These constraints include at least one of the following: maximum allowable processing delay, available computing resources, phase accuracy requirements, and processing mode. The algorithm selection module acquires... After these constraints are met, the optimal target algorithm is matched from the algorithm pool. The processing module is connected to the algorithm selection module, implementing all preset virtual playback simulation algorithms through a unified input / output interface. These algorithms include frequency domain multiplication algorithm, FIR filter convolution algorithm, cascaded IIR double second-order filter algorithm, overlapping adder block processing algorithm, overlapping retainer block processing algorithm, minimum phase reconstruction algorithm, impulse response convolution algorithm, parametric equalizer approximation algorithm, frequency warping FIR filter algorithm, partitioned convolution algorithm, frequency sampling algorithm, polyphase filter bank algorithm, weighted overlapping adder algorithm, constant Q transform processing algorithm, and zero-delay convolution algorithm. Accordingly, the processing module processes the input audio signal based on the preprocessed complex frequency response data according to the target algorithm determined by the algorithm selection module, generating an output audio signal that simulates the playback effect of the target audio playback device. The output module is connected to the processing module and is used to perform level normalization processing on the processed output audio signal to ensure that its amplitude range is compatible with the input requirements of the subsequent playback device, and converts it to the target audio format for output as needed. Furthermore, subjective evaluators judge sound quality by playing sound through audio playback devices such as Hi-Fi headphones or Hi-Fi speakers.
[0140] For example, Figure 5 A comparative schematic diagram of the input audio signal, complex frequency response data, and output audio signal provided for an exemplary embodiment of this application. (See attached diagram.) Figure 5 As shown, the schematic diagram includes an input waveform, a loudspeaker acoustic characteristic diagram, and an output waveform. Figure 3The system consists of two parts. The input waveform diagram shows the amplitude of the original input audio signal changing over time in the time domain, with the horizontal axis representing time (s) and the vertical axis representing amplitude. This waveform diagram reflects the original time-domain characteristics of the input audio signal, such as the specific waveform shape of speech, music, or test signals. The loudspeaker acoustic characteristic diagram consists of two sub-diagrams: the loudspeaker sound pressure level response diagram and the loudspeaker phase response diagram. In the loudspeaker sound pressure level response diagram, the horizontal axis represents frequency (Hz) and the vertical axis represents sound pressure level (dB). This curve shows the energy output characteristics of the target audio playback device, such as the target loudspeaker, at different frequencies. For example, the diagram shows that the target loudspeaker has a high sound pressure level in the low-to-mid frequency range, which gradually decreases in the high-frequency range, reflecting the loudspeaker's tonal balance characteristics. In the loudspeaker phase response diagram, the horizontal axis represents frequency (Hz) and the vertical axis represents phase (degrees or radians). This curve shows the phase delay characteristics of the target loudspeaker at different frequencies. For example, the diagram shows that the phase changes non-linearly with frequency, reflecting the time delay distribution of the loudspeaker for each frequency component of the signal, which has a significant impact on the transient response and spatial perception of the sound. The aforementioned sound pressure level response and phase response together constitute the complex frequency response data of the target loudspeaker, fully characterizing its acoustic properties. The output waveform diagram shows the amplitude of the output audio signal processed by the target algorithm as a function of time in the time domain, with the horizontal axis representing time (s) and the vertical axis representing amplitude. By employing the target algorithm to process the input audio signal and the complex frequency response data of the target loudspeaker, the output waveform retains the basic outline of the original signal while its waveform shape changes accordingly, reflecting the modulation effect of the target loudspeaker's acoustic characteristics on the input signal.
[0141] By comparing the input and output waveforms, the effect of the virtual simulation processing can be observed: the overall amplitude envelope, transient response, and waveform details of the output waveform all exhibit changes corresponding to the acoustic characteristics of the target audio playback device. For example, in the time-domain waveform portion corresponding to the higher sound pressure level response of the target loudspeaker, the amplitude is correspondingly enhanced; in the frequency region where the phase response changes, the time alignment of the waveform is also adjusted accordingly. This comparative diagram intuitively demonstrates the core processing procedure of the virtual playback simulation method of this application: based on the complex frequency response data of the target loudspeaker (including amplitude frequency response and phase frequency response), the input audio signal is precisely processed to generate an output audio signal that acoustically approximates the playback effect of the target loudspeaker. It should be noted that... Figure 5 The target algorithm shown is a frequency domain multiplication algorithm, which is only an example. In practical applications, the target algorithm is adaptively selected from multiple preset virtual playback simulation algorithms according to the application constraints.
[0142] The virtual playback simulation device provided in this application embodiment can execute the technical solution shown in the above-described virtual playback simulation method embodiment. Its implementation principle and beneficial effects are similar, and will not be described again here.
[0143] It should be noted that the division of the various modules in the above device is merely a logical functional division. In actual implementation, they can be fully or partially integrated into a single physical entity, or they can be physically separated. Furthermore, these modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented by processing element calls to software, while others are implemented in hardware. For example, a processing module can be a separate processing element, or it can be integrated into a chip within the device. Alternatively, it can be stored as program code in the device's memory, and its functions can be called and executed by a processing element. The implementation of other modules is similar. Moreover, these modules can be fully or partially integrated together, or they can be implemented independently. The processing element here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.
[0144] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more Digital Signal Processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together as a System-On-a-Chip (SOC).
[0145] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Video Discs, DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).
[0146] Figure 6 A schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. For example... Figure 6 As shown, the electronic device 60 in this embodiment includes:
[0147] At least one processor 61; and a memory 62 communicatively connected to the at least one processor;
[0148] The memory 62 stores instructions that can be executed by at least one processor 61 to cause the electronic device to perform the method as described in any of the above embodiments.
[0149] Alternatively, the memory 62 can be either standalone or integrated with the processor 61.
[0150] The memory 62 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device.
[0151] The processor 61 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. Specifically, when implementing the virtual playback simulation method described in the foregoing method embodiments, the electronic device may be, for example, an electronic device with processing capabilities such as a server.
[0152] Optionally, the electronic device may also include a communication interface 63. In specific implementations, if the communication interface 63, memory 62, and processor 61 are implemented independently, they can be interconnected via a bus to complete communication. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc., but this does not imply that there is only one bus or one type of bus.
[0153] Optionally, in a specific implementation, if the communication interface 63, memory 62, and processor 61 are integrated on a single chip, then the communication interface 63, memory 62, and processor 61 can communicate through an internal interface.
[0154] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the foregoing embodiments, and will not be repeated here.
[0155] This application also provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed, they are used to implement the method steps as described in the above method embodiments. The specific implementation methods and technical effects are similar and will not be repeated here.
[0156] The aforementioned computer-readable storage media can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0157] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an application-specific integrated circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in a virtual playback simulation device.
[0158] This application also provides a computer program product, including a computer program that, when executed, implements the method steps as described in the above method embodiments. The specific implementation and technical effects are similar and will not be repeated here.
[0159] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0160] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A virtual playback simulation method, characterized in that, include: Acquire complex frequency response data of the target audio playback device, wherein the complex frequency response data includes amplitude frequency response data and phase frequency response data; Acquire the input audio signal; Based on preset application constraints, a target algorithm is adaptively selected from multiple preset virtual playback simulation algorithms; wherein, the multiple preset virtual playback simulation algorithms have different latency characteristics, accuracy levels and computational complexity. Using the target algorithm, the input audio signal is processed based on the complex frequency response data to generate an output audio signal that simulates the playback effect of the target audio playback device.
2. The virtual playback simulation method according to claim 1, characterized in that, The preset application constraints include at least one of the following: maximum allowable processing latency, available computing resources, phase accuracy requirements, and processing mode. The adaptive selection of a target algorithm from multiple preset virtual playback simulation algorithms based on the preset application constraints includes: At least one of the maximum allowable processing latency, the available computing resources, the phase accuracy requirement, and the processing mode is obtained as the current constraint. Based on the current constraints, an algorithm that is compatible with the current constraints is selected from the plurality of preset virtual playback simulation algorithms as the target algorithm.
3. The virtual playback simulation method according to claim 2, characterized in that, The step of matching an algorithm that is compatible with the current constraints from among the plurality of preset virtual playback simulation algorithms as the target algorithm includes at least one of the following: When the processing mode is real-time processing mode and the maximum allowable processing latency is lower than a preset latency threshold, the algorithm with the lowest latency characteristic is selected from the plurality of preset virtual playback simulation algorithms as the target algorithm; When the processing mode is offline processing mode and the phase accuracy requirement is higher than the preset accuracy threshold, the algorithm with the highest accuracy level is selected from the plurality of preset virtual playback simulation algorithms as the target algorithm; When the available computing resources are lower than a preset resource threshold, the algorithm with the lowest computational complexity is selected from the plurality of preset virtual playback simulation algorithms as the target algorithm.
4. The virtual playback simulation method according to any one of claims 1 to 3, characterized in that, The plurality of preset virtual playback simulation algorithms include at least three of the following algorithm sets: The algorithms include frequency domain multiplication, FIR filter convolution, cascaded IIR dual second-order filter, overlapping summing block processing, overlapping retaining block processing, minimum phase reconstruction, impulse response convolution, parametric equalizer approximation, frequency warping FIR filter, partitioned convolution, frequency sampling, multiphase filter bank, weighted overlapping summing, constant Q transform processing, and zero-delay convolution; wherein, the zero-delay convolution algorithm refers to an algorithm that decomposes the complex frequency response data into minimum phase components and all-pass components for separate processing.
5. The virtual playback simulation method according to claim 4, characterized in that, The frequency warping FIR filter algorithm determines the frequency warping factor based on the characteristics of human auditory perception, and performs a warping transformation on the frequency axis based on the frequency warping factor to enhance the filter resolution in the low-frequency region.
6. The virtual playback simulation method according to claim 4, characterized in that, The constant Q transform processing algorithm analyzes the audio signal using logarithmically spaced frequency points with a constant quality factor, and the ratio of the center frequency to the bandwidth at each frequency point remains constant.
7. The virtual playback simulation method according to claim 4, characterized in that, If the target algorithm is the partitioned convolution algorithm, the step of using the target algorithm to process the input audio signal based on the complex frequency response data includes: dividing the target impulse response into multiple continuous segments, wherein the target impulse response is obtained by inverse transform of the complex frequency response data; performing a fast Fourier transform on each segment to obtain the corresponding frequency domain segmentation coefficients; constructing a frequency domain delay line to store the frequency domain representations of multiple historical input audio signals; multiplying and accumulating the frequency domain representation of the current input audio signal with each of the frequency domain segmentation coefficients to obtain the frequency domain output result; and performing an inverse fast Fourier transform on the frequency domain output result to obtain the output audio signal. If the target algorithm is the zero-delay convolution algorithm, the step of using the target algorithm to process the input audio signal based on the complex frequency response data includes: decomposing the complex frequency response data into minimum phase components and all-pass components; constructing a minimum phase impulse response based on the minimum phase components to achieve zero pre-ringing characteristics; designing a cascaded first-order all-pass filter based on the all-pass components to approximate the residual phase characteristics of the complex frequency response data; cascading the minimum phase impulse response with the cascaded first-order all-pass filter to perform convolution processing on the input audio signal to obtain the output audio signal.
8. A virtual playback simulation device, characterized in that, include: The input module is used to acquire the complex frequency response data of the target audio playback device, wherein the complex frequency response data includes amplitude frequency response data and phase frequency response data; And to acquire the input audio signal; The algorithm selection module is used to adaptively select a target algorithm from multiple preset virtual playback simulation algorithms according to preset application constraints; wherein the multiple preset virtual playback simulation algorithms have different latency characteristics, accuracy levels and computational complexity. The processing module is used to process the input audio signal based on the complex frequency response data using the target algorithm to generate an output audio signal that simulates the playback effect of the target audio playback device.
9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory is used to store computer-executed instructions; The processor is configured to execute the computer execution instructions to implement the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed, are used to implement the method as described in any one of claims 1-7.