Audio reproduction method and system for on-site sound field data collection and digital signal processing

By collecting data in a real acoustic space to generate a dedicated sound field parameter model and performing digital signal processing, the problem of insufficient sound field reproduction accuracy in existing technologies has been solved. This enables high-fidelity audio devices to reproduce professional acoustic characteristics in a home environment, enhancing spatial sense and sound image localization.

CN122637792APending Publication Date: 2026-08-25CHERY AUTOMOBILE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610699465.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In existing technologies, high-fidelity audio devices struggle to reproduce the acoustic characteristics of professional concert halls, recording studios, or cinemas in home or personal environments, resulting in a significant gap between the user experience and that of professional venues. The existing acoustic models differ from the real acoustic spaces, leading to insufficient accuracy in sound field reproduction. The simulation processing link introduces noise and distortion, and the processing parameters lack objective measured data support.

Method used

By collecting sound field data in a real acoustic space, a unique sound field parameter model is generated, which includes phase response characteristics and amplitude response characteristics. Digital signal processing is used for convolution processing to output an audio signal with enhanced spatial sense, avoiding analog conversion and ensuring parameter accuracy.

Benefits of technology

It achieves highly accurate reproduction of the target acoustic space, enhances reverberation characteristics and sound image localization, and significantly improves the accuracy and completeness of sound field reproduction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122637792A_ABST
    Figure CN122637792A_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of audio processing, and particularly relates to an audio reproduction method and system for live sound field data collection and digital signal processing. The method comprises collecting sound field data of a real acoustic space and generating a dedicated sound field parameter model; receiving an original digital audio signal and a user instruction, obtaining a dedicated sound field parameter model corresponding to the user instruction as a target dedicated sound field parameter model; performing convolution processing on the original digital audio signal and the target dedicated sound field parameter model to obtain a processed audio signal with enhanced spatial sense; and outputting the processed audio signal after multi-channel digital amplification, thereby realizing audio reproduction of the acoustic response characteristics of the dedicated sound field. The present application generates a dedicated sound field parameter model based on impulse response data collected in a real acoustic space, and the reproduced reverberation characteristics, spatial sense and sound image positioning are highly consistent with the target space, and the sound field restoration accuracy is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of audio processing technology, and in particular relates to an audio reproduction method and system for on-site sound field data acquisition and digital signal processing. Background Technology

[0002] The goal of high-fidelity audio equipment is to reproduce sound as realistically as possible. However, ordinary audio playback, especially in home or personal environments, is limited by factors such as speaker units, listening environment, and signal processing links, making it difficult to reproduce the unique acoustic characteristics of professional concert halls, recording studios, or cinemas (such as rich reflections, reverberation characteristics, and precise sound image localization). This lack of "presence" or "spatial feel" results in a significant gap between the user experience and that of professional settings.

[0003] Existing technologies include some techniques that use digital signal processing (DSP) to simulate different sound field environments, such as adding simple reverberation effects or using general room acoustic models. However, the acoustic models used by these techniques are mostly general models based on theory or generated with a small number of parameters, which cannot accurately correspond to any real, high-quality acoustic space, and the reproduced sound appears "artificial" or unnatural. Summary of the Invention

[0004] To overcome the shortcomings of the prior art, this invention provides an audio reproduction method and system for on-site sound field data acquisition and digital signal processing. It generates a unique sound field parameter model based on impulse response data collected on-site in a real acoustic space, rather than using a general theoretical model. The reproduced reverberation characteristics, spatial sense, and sound image localization are highly consistent with the target space, and the accuracy of sound field reproduction is significantly improved. This solves the problem that the acoustic model used in the sound field simulation of the prior art is a general model, which differs from the real acoustic space, resulting in insufficient accuracy of sound field reproduction.

[0005] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions: The first aspect of this invention provides an audio reproduction method for on-site sound field data acquisition and digital signal processing.

[0006] An audio reproduction method based on on-site sound field data acquisition and digital signal processing includes the following steps: Acquire sound field data of a real acoustic space and generate a dedicated sound field parameter model, wherein the dedicated sound field parameter model includes acoustic response characteristic parameters of the real acoustic space; Receive raw digital audio signals and user commands, and obtain the exclusive sound field parameter model corresponding to the user commands as the target exclusive sound field parameter model; The original digital audio signal is convolved with the target-specific sound field parameter model to obtain a processed audio signal with enhanced spatial sense. The processed audio signal is then digitally amplified through multiple channels and output to achieve audio reproduction of the specific acoustic response characteristics of the sound field.

[0007] In some embodiments, acquiring sound field data from a real acoustic space specifically includes: A microphone array is placed at the listening position in a real acoustic space; Play test signals into a real acoustic space; The pulse response data formed by the test signal after reflection and reverberation in the real acoustic space is recorded by a microphone array, and the pulse response data is used as the sound field data.

[0008] In some embodiments, the acoustic response characteristics include phase response characteristics and amplitude response characteristics, and the acoustic response characteristic parameters include phase response parameters and amplitude response parameters; The phase response parameters reflect the phase delay characteristics of each frequency component when sound propagates in the real acoustic space, including the phase of direct sound, the phase of early reflected sound, and the phase of reverberation field. The amplitude response parameters are objective acoustic indicators that reflect the energy attenuation characteristics of each frequency component, including early decay time, reverberation time, sound intensity, and clarity.

[0009] In some embodiments, the specific process of generating a dedicated sound field parameter model includes: The acquired sound field data is preprocessed, including removing DC bias, time alignment, and removing background noise interference. Phase response analysis is performed on the preprocessed sound field data to extract the phase delay characteristics of each frequency component and construct the phase response parameter matrix; simultaneously, amplitude response analysis is performed to calculate the energy attenuation characteristics of each frequency component and extract objective acoustic indicators. By summarizing the phase response parameter matrix and objective acoustic indicators, objective analytical data can be obtained. Based on objective data analysis, professional audio engineers make fine adjustments to the early reflection delay time, reverberation attenuation curve, amplitude and phase response parameters of each frequency band; After the adjustments are complete, package all parameters to generate a custom sound field parameter model file.

[0010] In some embodiments, it also includes: The dedicated sound field parameter model is stored in the memory of the user terminal device. The dedicated sound field parameter model includes a concert hall sound field parameter model, a recording studio sound field parameter model and a cinema sound field parameter model. According to user instructions, the dedicated sound field parameter model for the corresponding sound field mode is retrieved from memory to obtain the target dedicated sound field parameter model.

[0011] In some embodiments, the original digital audio signal is a multi-channel digital audio signal; The original digital audio signal is convolved with the target-specific sound field parameter model, specifically including: Each channel of the original digital audio signal is convolved with the target-specific sound field parameter model to obtain a processed audio signal with enhanced spatial sense.

[0012] In some embodiments, after generating the dedicated sound field parameter model file, the method further includes: The generated model file is validated by loading the model file onto the test device and playing the reference audio. A professional audio engineer then conducts a subjective listening evaluation to confirm that the model's performance matches the acoustic characteristics of the target acoustic space.

[0013] The second aspect of the present invention provides an audio reproduction system for on-site sound field data acquisition and digital signal processing.

[0014] An audio reproduction system for on-site sound field data acquisition and digital signal processing includes: The on-site parameter generation module is configured to: collect sound field data of a real acoustic space and generate a dedicated sound field parameter model, wherein the dedicated sound field parameter model includes acoustic response characteristic parameters of the real acoustic space; The parameter matching module is configured to: receive the original digital audio signal and user command, and obtain the exclusive sound field parameter model corresponding to the user command as the target exclusive sound field parameter model; The convolution processing module is configured to: perform convolution processing on the original digital audio signal and the target-specific sound field parameter model to obtain a processed audio signal with enhanced spatial sense; The audio playback module is configured to: amplify the processed audio signal through multiple channels and then output it to achieve audio reproduction of the specific acoustic response characteristics of the sound field.

[0015] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of the audio reproduction method for on-site sound field data acquisition and digital signal processing as described in the first aspect of the present invention.

[0016] A fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the audio reproduction method for on-site sound field data acquisition and digital signal processing as described in the first aspect of the present invention.

[0017] The above one or more technical solutions have the following beneficial effects: This invention provides an audio reproduction method and system for on-site sound field data acquisition and digital signal processing. It generates a dedicated sound field parameter model based on impulse response data collected on-site in a real acoustic space, rather than using a general theoretical model. The processing parameters have reliable data support, and the accuracy of the parameters and the degree of reproduction of the target acoustic space are improved compared with subjective tuning schemes. The reproduced reverberation characteristics, spatial sense and sound image positioning are highly consistent with the target space, and the accuracy of sound field reproduction is significantly improved.

[0018] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0019] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0020] Figure 1 This is a flowchart illustrating the steps of Example 1.

[0021] Figure 2 This is a flowchart illustrating the audio reproduction method of Example 1.

[0022] Figure 3 This is a schematic diagram of the convolution processing principle of the phase filtering processing unit in Embodiment 1.

[0023] Figure 4 This is a schematic diagram of the overall architecture of the audio reproduction system in Embodiment 2.

[0024] Figure 5 This is a schematic diagram of the user-side audio processing subsystem in Embodiment 2. Detailed Implementation

[0025] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0026] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0027] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0028] Example 1 With the continuous development of audio playback technology, high-fidelity audio devices have been widely used in home theaters, car audio systems, and personal listening scenarios. Users' demands for audio playback quality are increasing, focusing not only on frequency response and dynamic range but also on subjective listening experiences such as spatiality and presence. However, in ordinary home or personal environments, limited by the acoustic conditions of the listening space and the processing capabilities of the equipment, the sound field effect presented by audio playback differs significantly from that of professional acoustic venues such as concert halls, recording studios, and cinemas.

[0029] To bridge the aforementioned gap, existing technologies have employed digital signal processing (DSP) techniques to simulate sound field effects in audio signals. These solutions typically pre-configure several acoustic parameters in the audio processing equipment. By applying reverberation, delay, and equalization to the input audio signal, they impart a sense of space to the output audio. The simulation of reverberation is the core processing method in these solutions. The processor processes the audio signal based on preset parameters such as reverberation time and early reflection delay to simulate the propagation of sound within a closed space. However, the acoustic parameters used in these solutions are usually derived from theoretical modeling or statistical summarization of limited measurement data. The constructed acoustic models are general-purpose models, differing from the actual acoustic characteristics of any specific real acoustic space. This results in insufficient accuracy in sound field reproduction in the processed audio output, making it difficult to realistically reproduce the sound field characteristics of a specific professional acoustic space.

[0030] Another type of existing technology introduces an analog processing stage into the audio processing chain, using analog audio circuits to colorize and enhance the listening experience. However, in this type of solution, the nonlinear characteristics and noise of the analog circuit components are superimposed on the audio signal, affecting the integrity of the original digital audio signal. This is particularly noticeable in multi-stage processing chains, where the noise floor and distortion are lower compared to all-digital processing chains.

[0031] Furthermore, the processing parameters of some existing sound field simulation schemes rely on audio engineers to adjust and set them based on subjective listening perception, lacking the support of measured acoustic data obtained from objective and systematic measurements of a specific target acoustic space. As a result, the accuracy of the parameters and the degree of reproduction of the target space are somewhat limited.

[0032] Therefore, there is a need for a technical solution that can construct an acoustic model based on objective measurement data of real acoustic space and process audio signals through digital signal processing to accurately reproduce the sound field characteristics of the target acoustic space at the user end.

[0033] This embodiment discloses an audio reproduction method for on-site sound field data acquisition and digital signal processing, which solves the problems in the prior art where the acoustic model used in sound field simulation is a general model and differs from the real acoustic space, resulting in insufficient accuracy of sound field reconstruction; noise and distortion introduced by the simulation processing link leading to a decrease in the integrity of the original audio signal; and the lack of objective measured data to support the processing parameters, resulting in limited accuracy.

[0034] like Figure 1 As shown, the audio reproduction method based on on-site sound field data acquisition and digital signal processing includes the following steps: Acquire sound field data from real acoustic spaces and generate a dedicated sound field parameter model, which includes acoustic response characteristic parameters of the real acoustic space; Receive raw digital audio signals and user commands, and obtain the exclusive sound field parameter model corresponding to the user commands as the target exclusive sound field parameter model; The original digital audio signal is convolved with the target-specific sound field parameter model to obtain a processed audio signal with enhanced spatial sense. The processed audio signal is digitally amplified through multiple channels before being output, enabling audio reproduction that captures the acoustic response characteristics of the specific sound field.

[0035] Overall, this embodiment includes two stages: sound field data acquisition and modeling, and user-end audio processing. (1) The sound field data acquisition and modeling stage includes the following steps: Deploy measuring equipment in a real acoustic space to collect sound field data of the real acoustic space; A unique sound field parameter model is generated based on the sound field data. The unique sound field parameter model contains the acoustic response characteristic parameters of the real acoustic space. Store the proprietary sound field parameter model in the memory of the user's device.

[0036] Among them, acoustic response characteristics include phase response characteristics and amplitude response characteristics, and acoustic response characteristic parameters include phase response parameters and amplitude response parameters; Phase response parameters reflect the phase delay characteristics of each frequency component when sound propagates in a real acoustic space, including the phase of direct sound, the phase of early reflected sound, and the phase of reverberation field. Amplitude response parameters are objective acoustic indicators that reflect the energy decay characteristics of each frequency component, including early decay time, reverberation time, sound intensity, and clarity.

[0037] (2) The user-side audio processing stage includes the following steps: Receive raw digital audio signals through a digital audio interface; Load the dedicated sound field parameter model corresponding to the user's instructions from memory according to the user's instructions; The phase filtering processing unit in the digital signal processor convolves the original digital audio signal with a dedicated sound field parameter model generated based on impulse response data collected in real acoustic space, and outputs a processed audio signal with enhanced spatial sense. The processed audio signal is input into a multi-channel pure digital power amplifier for multi-channel digital amplification before being output.

[0038] In this embodiment, the steps for collecting sound field data of a real acoustic space include: arranging a microphone array at the listening position in the real acoustic space, playing a test signal into the real acoustic space, recording the impulse response data formed after the test signal is reflected and reverberated in the real acoustic space by the microphone array, and using the impulse response data as sound field data.

[0039] The dedicated sound field parameter model incorporates the phase response characteristics of a real acoustic space; the specific process for generating the dedicated sound field parameter model includes: The acquired sound field data is preprocessed, including removing DC bias, time alignment, and removing background noise interference. Phase response analysis is performed on the preprocessed sound field data to extract the phase delay characteristics of each frequency component and construct the phase response parameter matrix; simultaneously, amplitude response analysis is performed to calculate the energy attenuation characteristics of each frequency component and extract objective acoustic indicators. By summarizing the phase response parameter matrix and objective acoustic indicators, objective analytical data can be obtained. Based on objective data analysis, professional audio engineers make fine adjustments to the early reflection delay time, reverberation attenuation curve, amplitude and phase response parameters of each frequency band; After the adjustments are complete, package all parameters to generate a custom sound field parameter model file.

[0040] The memory stores multiple dedicated sound field parameter models corresponding to the sound field of a concert hall, a recording studio, and a cinema. The step of loading the dedicated sound field parameter model corresponding to the user's instruction includes: calling the dedicated sound field parameter model corresponding to the sound field mode selected by the user from the memory.

[0041] The original digital audio signal is a multi-channel digital audio signal; the steps of convolution processing through the phase filtering processing unit include: convolving each channel of the multi-channel digital audio signal with a dedicated sound field parameter model; the number of output channels of the multi-channel pure digital power amplifier is not less than the number of channels of the multi-channel digital audio signal.

[0042] (a) Concert Hall Sound Field Mode.

[0043] This embodiment takes the sound field mode of a concert hall as a specific application scenario, and describes in detail the complete implementation process of the audio reproduction method based on on-site sound field data acquisition and digital signal processing, covering two stages: the sound field data acquisition and modeling stage and the user-end audio processing stage.

[0044] like Figure 2 As shown, the audio reproduction method includes two main stages: the sound field data acquisition and modeling stage (S1 stage) and the user-end audio processing stage (S2 stage). The two stages are independent of each other in time. After the S1 stage is completed in a professional venue, its results are solidified in the form of digital files and transmitted to the user-end device. The S2 stage runs in real time on the user-end device.

[0045] (1) In the sound field data acquisition and modeling stage, it is necessary to first set up measurement equipment in the real acoustic space to collect the sound field data of the real acoustic space.

[0046] This embodiment selects a famous concert hall as the target real acoustic space. The concert hall has a building volume of approximately 18,000 cubic meters and an audience seating of approximately 2,500. Its natural reverberation time is approximately 1.8 to 2.2 seconds when the hall is full and approximately 2.2 to 2.5 seconds when the hall is empty, exhibiting typical acoustic characteristics of a large concert hall.

[0047] Microphone arrays were placed in the typical listening positions of the concert hall. The typical listening positions were selected in the area 20 to 25 meters horizontally from the center of the stage and 1.2 to 1.5 meters above the ground, corresponding to the golden listening area in the back row of the main hall.

[0048] The microphone array consists of eight high-precision condenser microphones arranged in a three-dimensional array. The condenser microphones have a frequency response range of 20Hz to 20kHz, a frequency response flatness not exceeding ±1dB within this range, an equivalent noise level not exceeding 10dB (A-weighted), and a maximum sound pressure level not less than 140dB SPL. The spacing between adjacent microphones in the array is 0.3 meters to 0.5 meters. Four microphones are arranged horizontally, and two layers are arranged vertically, forming a three-dimensional acquisition array covering a horizontal azimuth angle of 360 degrees and a vertical pitch angle, to comprehensively capture the sound field distribution characteristics of the target acoustic space in three-dimensional space.

[0049] Test signals were played to the concert hall. These test signals included two types: The first type is a logarithmic swept sinusoidal signal covering the frequency range of 20Hz to 20kHz, with a sweep duration of 10 to 30 seconds and a logarithmically uniform sweep rate, which can obtain high signal-to-noise ratio frequency response measurement results across the entire frequency band. The second category is broadband pulse signals, with a pulse width of no more than 1 millisecond and a peak sound pressure level of 100 dB SPL to 110 dB SPL.

[0050] The test signals were played through a professional monitoring-grade speaker system, which was placed in the center of the stage and played at a sound pressure level of 85dB SPL to 95dB SPL (at the reference measurement point) to ensure a sufficient signal-to-noise ratio (not less than 60dB) across the entire frequency range.

[0051] The test signal was recorded by a microphone array, which recorded the impulse response data formed after reflection and reverberation of the walls, ceiling, floor, seats, reflectors and other interfaces of the concert hall. The recording sampling rate was 96kHz, the quantization bit depth was 32 bits, and the recording time for each measurement was no less than 5 seconds, so as to fully capture the reverberation tail until it decayed to below the noise floor.

[0052] The impulse response data contains the phase and amplitude response characteristics of the actual acoustic space of the concert hall at each measurement location, serving as the raw data foundation for generating a dedicated sound field parameter model. Measurements were repeated at least three times at each microphone location, and the average value was taken to reduce occasional noise interference. The impulse response data was used as the sound field data.

[0053] Subsequently, a custom sound field parameter model is generated based on the sound field data. Specifically, the acquired impulse response data is transmitted digitally to a computer workstation via an audio interface card. The audio interface card supports sampling rates from 44.1kHz to 192kHz and quantization bits from 24 to 32 bits, ensuring no precision loss during data transmission. Professional acoustic analysis software runs on the computer workstation to systematically analyze and adjust the phase response and amplitude response parameters in the sound field data.

[0054] Phase response parameters reflect the phase delay characteristics of each frequency component when sound propagates in the concert hall space, including three components: direct sound phase, early reflection phase, and reverberation field phase; amplitude response parameters reflect the energy attenuation characteristics of each frequency component, including objective acoustic indicators such as early decay time (EDT), reverberation time RT60, sound intensity G value, and intelligibility C80.

[0055] Based on objective measurement data, and combined with their professional understanding of concert hall audio and mixing skills, professional audio engineers made fine adjustments to the following parameters: The early reflection delay time was adjusted to the range of 15 to 30 milliseconds to restore the unique spatial immersion of the concert hall; The reverberation decay profile was calibrated to the target RT60 value (1.8 seconds to 2.5 seconds); The reverberation time in the low-frequency range (20Hz to 200Hz) is extended by 10% to 20% compared to the mid-frequency range to reproduce the fullness of the low frequencies in a concert hall. The reverberation time in the high-frequency range (8kHz to 20kHz) is appropriately shortened compared to the mid-frequency range to reproduce the air absorption effect.

[0056] After the above analysis and adjustments, a dedicated sound field parameter model file for concert halls is generated. This dedicated sound field parameter model includes the acoustic response characteristics of a real acoustic space, especially the phase response characteristics of a real acoustic space. This is the key difference between this invention and a general acoustic model. The generated dedicated sound field parameter model file is stored in digital format, with a file size ranging from 2MB to 50MB, depending on the model's time resolution and frequency resolution.

[0057] Finally, the dedicated sound field parameter model is stored in the user's device memory, completing the sound field data acquisition and modeling stage. The memory also stores dedicated sound field parameter models corresponding to both recording studio and cinema sound fields, allowing users to select any sound field mode as needed.

[0058] (2) During the user-side audio processing stage, the raw digital audio signal is received through the digital audio interface.

[0059] The digital audio interface supports at least one of HDMI ARC (Audio Return Channel) and USB digital interfaces. The raw digital audio signal is input to the digital audio interface in digital form via these interfaces, maintaining its digital format throughout the entire process without any analog conversion, thus fundamentally avoiding the noise and distortion introduced by analog signal transmission. In this embodiment, the raw digital audio signal is a 2-channel stereo digital audio signal with a sampling rate of 44.1kHz to 192kHz, a quantization bit depth of 16 bits to 32 bits, and supports PCM (Pulse Code Modulation) lossless digital audio format.

[0060] According to user instructions, the device loads the dedicated sound field parameter model corresponding to the user's instructions from the memory. Specifically, when the user selects "Concert Hall Mode" on the device's operating interface, after receiving the user's instructions, the digital signal processor (DSP) retrieves the dedicated sound field parameter model corresponding to the concert hall sound field mode from the memory and loads the dedicated sound field parameter model into the phase filtering processing unit embedded in the DSP.

[0061] The loading process includes: reading the dedicated sound field parameter model file from the memory, parsing the phase response parameter and amplitude response parameter data in the file, and writing the parameter data into the working register of the phase filter processing unit. The entire loading process takes no more than 500 milliseconds and has no significant impact on the user experience.

[0062] like Figure 3 As shown, the phase filtering processing unit in the digital signal processor convolves the original digital audio signal with a dedicated sound field parameter model generated based on impulse response data collected in real acoustic space, and outputs a processed audio signal with enhanced spatial sense. The phase filtering processing unit performs phase convolution processing on the original digital audio signal according to the phase response characteristics in the dedicated sound field parameter model. The convolution processing is performed according to the following formula (1): (1) in, Represents the original digital audio signal. This represents the target acoustic space impulse response function represented by the dedicated sound field parameter model. This indicates the processed audio signal with enhanced spatial sense output after convolution. This represents the convolution operator. It is the integral variable.

[0063] The physical meaning of the above convolution operation is that the original audio signal is passed through a linear time-invariant system characterized by the impulse response of the target acoustic space, so that the output signal contains all the acoustic features of the target space, including the time structure and phase characteristics of direct sound, early reflections and reverberation tails.

[0064] Convolution processing is accelerated in the frequency domain using Fast Fourier Transform (FFT). The specific steps are as follows: the input signal and impulse response function are transformed to the frequency domain by FFT, pointwise complex multiplication is performed in the frequency domain, and then transformed back to the time domain by Inverse Fast Fourier Transform (IFFT) to obtain the convolution result.

[0065] The convolution algorithm accelerated by FFT reduces computational complexity and significantly improves real-time processing capabilities. The end-to-end latency of the entire convolution process does not exceed 5ms, meeting the latency requirements of real-time audio processing.

[0066] After convolution processing, the phase and amplitude features of the target acoustic space are embedded in the processed audio signal, and the sense of space, sense of immersion, and sound image localization are significantly enhanced.

[0067] Finally, the processed audio signal is input into a multi-channel pure digital power amplifier for multi-channel digital amplification before output. The multi-channel pure digital power amplifier has 8 output channels, directly receiving digital audio signals from the digital signal processor. After digital power amplification within the amplifier, the signals drive the corresponding speaker units. There is no analog signal conversion throughout the process, effectively ensuring the integrity of the original audio signal. Each output channel of the multi-channel pure digital power amplifier has a rated output power of 50W to 200W, a signal-to-noise ratio of no less than 100dB, a total harmonic distortion (THD) of no more than 0.01%, a dynamic range of no less than 110dB, a frequency response range of 20Hz to 20kHz, and a frequency response flatness of no more than ±0.5dB.

[0068] (ii) Recording Studio Sound Field Mode.

[0069] This embodiment takes the recording studio sound field mode as an example to focus on describing the channel-by-channel convolution processing of multi-channel digital audio signals.

[0070] In the sound field data acquisition and modeling phase, this embodiment selects a professional recording studio as the target real acoustic space. Recording studios and concert halls differ significantly in acoustic characteristics: the volume of a recording studio is typically 100 to 500 cubic meters, much smaller than a concert hall; the reverberation time of a recording studio is extremely short, with an RT60 typically between 0.2 and 0.5 seconds; recording studios utilize a large amount of sound-absorbing materials, resulting in weak early reflected sound energy and short reflected sound delay times (typically 2 to 8 milliseconds); the frequency response of a recording studio is extremely flat to ensure a neutral listening environment. These acoustic characteristics determine that the sound field parameter model of a recording studio differs fundamentally from that of a concert hall in terms of phase and amplitude response parameters, necessitating specialized data acquisition and modeling tailored to the acoustic characteristics of the recording studio.

[0071] In a typical listening position in a recording studio, the recording engineer's monitoring position, a microphone array is set up 1.5 to 2.5 meters away from the main monitor speakers. The microphone array consists of 6 high-precision condenser microphones arranged in a near-field three-dimensional array, with the distance between adjacent microphones being 0.2 to 0.3 meters. The smaller spacing is used to adapt to the sound field characteristics of the near-field monitoring environment in the recording studio.

[0072] Test signals were played into the recording studio. These signals included a logarithmic sweep signal with a frequency range of 20Hz to 20kHz, a sweep duration of 5 to 10 seconds, and sinusoidal steady-state signals at various frequencies, used to analyze the phase response characteristics at specific frequency points. Impulse response data was recorded using a microphone array at a sampling rate of 96kHz and a quantization bit depth of 32 bits. Each measurement recording lasted 2 to 3 seconds to fully capture the brief reverberation tails of the recording studio.

[0073] The phase response and amplitude response parameters in the sound field data are analyzed and adjusted, with a focus on the near-field phase characteristics and extremely short reverberation characteristics of the recording studio. A dedicated sound field parameter model for the recording studio is generated, which includes the phase response and amplitude response characteristics of the actual acoustic space of the recording studio. This model is then stored in the memory of the user's device.

[0074] During the user-side audio processing stage, the raw digital audio signal is input to the digital audio interface in digital signal form via the USB digital interface. The USB digital interface conforms to the USB Audio Class 2.0 specification, supports a maximum sampling rate of 192kHz, a quantization bit depth of 32 bits, and a maximum transmission bandwidth of no less than 480Mbps (USB 2.0 High Speed), which can meet the lossless transmission requirements of multi-channel high-resolution digital audio signals. In this embodiment, the raw digital audio signal is a 5.1-channel multi-channel digital audio signal, including six channels: left channel (L), right channel (R), center channel (C), left surround channel (Ls), right surround channel (Rs), and low-frequency effects channel (LFE), with a sampling rate of 48kHz and a quantization bit depth of 24 bits.

[0075] When the user selects "Recording Studio Mode" on the device's interface, the digital signal processor retrieves the corresponding studio sound field parameter model from memory and loads it into the phase filter processing unit. For example... Figure 3 As shown, the phase filtering processing unit performs convolution processing on each channel of the 5.1-channel multi-channel digital audio signal with the dedicated sound field parameter model of the recording studio sound field mode. Specifically, for each of the six channels, the phase filtering processing unit independently performs the convolution operation shown in the above formula (1): The left channel (L) signal is convolved with the impulse response function corresponding to the left front direction in the dedicated sound field parameter model; The right channel (R) signal is convolved with the impulse response function corresponding to the right front direction; The center channel (C) signal is convolved with the impulse response function corresponding to the front direction; The left surround channel (Ls) signal is convolved with the impulse response function corresponding to the left rear direction; The right surround channel (Rs) signal is convolved with the corresponding impulse response function in the right rear direction; The low-frequency effects channel (LFE) signal is convolved with the corresponding low-frequency omnidirectional impulse response function.

[0076] The convolution processing of each channel is independent, and the impulse response function of each channel is derived from the measured data of the microphone array in the corresponding spatial direction, which ensures the accuracy of the sound image localization of each channel.

[0077] The phase filtering processing unit adopts a parallel processing architecture, capable of performing real-time convolution operations on all six channels simultaneously. The convolution processing of each channel is completely synchronized in time, with no processing delay differences between channels. The computational cost of single-channel convolution processing does not exceed 2 × 10⁻⁶ per second. 7 Sub-floating-point operations, with the total computational load for parallel processing of 6 audio channels, do not exceed 1.2 × 10⁻⁶ per second. 8 The total processing latency is no more than 5ms, meeting the performance requirements of real-time multi-channel audio processing. The processed audio signals from each channel are then combined and sent to the corresponding input channels of a multi-channel pure digital power amplifier.

[0078] The multi-channel pure digital power amplifier has 8 to 10 output channels, which is no less than the number of channels in a multi-channel digital audio signal (6 channels), ensuring that the processed audio signal of each channel can be independently output to the corresponding speaker unit. In this embodiment, the multi-channel pure digital power amplifier has 8 output channels, of which 6 channels correspond to the 6 channels of a 5.1 channel system, and the other 2 channels can be used for expansion configurations (such as overhead speakers or subwoofers). The rated output power of each output channel is 50W to 200W, the signal-to-noise ratio is no less than 100dB, and the total harmonic distortion does not exceed 0.01%.

[0079] (iii) Cinema sound field mode.

[0080] This embodiment takes the 7.1 channel cinema sound field mode as an example to focus on describing the complete implementation of the multi-channel processing and all-digital signal processing link.

[0081] In the sound field data acquisition and modeling phase, this embodiment selects a professional Hollywood mixing studio as the target real acoustic space. The volume of the professional mixing studio is 800 to 1500 cubic meters, the reverberation time RT60 is 0.8 to 1.2 seconds, and the early reflection delay is 8 to 20 milliseconds, which meets the SMPTE (Society of Motion Picture and Television Engineers) acoustic standards for professional mixing rooms.

[0082] Microphone arrays were placed in the typical mixing engineer listening positions in the cinema. The array design fully considered the channel distribution characteristics of a 7.1 channel system. Measurement microphones were arranged in seven directions on the horizontal plane: left (L), right (R), center (C), left surround (Ls), right surround (Rs), left rear surround (Lbs), and right rear surround (Rbs). An additional low-frequency effect (LFE) measurement microphone was placed in the low-frequency omnidirectional position, for a total of eight measurement positions, corresponding one-to-one with the channel distribution of the 7.1 channel system. The collected impulse response data was analyzed and adjusted by professional audio engineers to generate a dedicated sound field parameter model that included the phase response characteristics of the cinema's actual acoustic space, and stored in the memory of the user's terminal device.

[0083] During the user-side audio processing stage, the original digital audio signal is a 7.1-channel multi-channel digital audio signal, including eight channels: left channel (L), right channel (R), center channel (C), left surround channel (Ls), right surround channel (Rs), left rear surround channel (Lbs), right rear surround channel (Rbs), and low-frequency effects channel (LFE). The sampling rate is 48kHz and the quantization bit depth is 24 bits.

[0084] The raw digital audio signal is input to the digital audio interface in digital form via the HDMI ARC (Audio Return Channel) interface. The HDMI ARC interface conforms to the HDMI 2.1 specification, supports a maximum digital audio transmission bandwidth of no less than 36Mbps, supports lossless multi-channel audio formats such as Dolby TrueHD and DTS-HD Master Audio, and supports lossless transmission of multi-channel high-resolution digital audio signals with a maximum of 7.1 channels, a sampling rate of 192kHz, and a quantization bit depth of 24 bits, fully meeting the transmission requirements of the aforementioned 7.1-channel multi-channel digital audio signals.

[0085] The HDMI ARC interface connects to the digital signal processor via digital signals, transmitting the received multi-channel digital audio signals to the input of the digital signal processor in digital form, without any analog conversion.

[0086] When a user watches a movie, they select "Cinema Mode" on the device's interface. The digital signal processor then retrieves the corresponding cinema sound field parameter model from memory and loads it into the phase filter processing unit. For example... Figure 3As shown, the phase filtering processing unit performs convolution processing on each channel of the 7.1-channel multi-channel digital audio signal with the dedicated sound field parameter model of the cinema sound field mode. Specifically, for each of the eight channels, the phase filtering processing unit independently performs the convolution operation shown in the above formula (1), convolving the original digital audio signal of each channel with the impulse response function corresponding to the spatial position of that channel in the dedicated sound field parameter model: The left channel (L) signal is convolved with the impulse response function in the dedicated sound field parameter model corresponding to the left front 30 degrees direction; The right channel (R) signal is convolved with the impulse response function corresponding to the right front 30 degrees direction; The center channel (C) signal is convolved with the impulse response function corresponding to the 0-degree direction directly in front; The left surround channel (Ls) signal is convolved with the impulse response function corresponding to the left side in the 90-110 degree direction; The right surround channel (Rs) signal is convolved with the impulse response function corresponding to the right side in the 90-110 degree direction; The left rear surround channel (Lbs) signal is convolved with the impulse response function in the corresponding left rear 135-degree to 150-degree direction; The right rear surround channel (Rbs) signal is convolved with the impulse response function corresponding to the right rear 135-degree to 150-degree direction; The low-frequency effects channel (LFE) signal is convolved with the corresponding low-frequency omnidirectional impulse response function.

[0087] Through the above channel-by-channel convolution processing, the output signal of each channel contains the reflected sound and reverberation characteristics of the target cinema acoustic space in the corresponding direction, thereby restoring a complete three-dimensional sound field spatial sense in multi-channel playback, including the sense of front sound image localization, the sense of side envelopment, the sense of rear surround, and the sense of vertical depth.

[0088] The phase filtering processing unit adopts a parallel processing architecture, capable of performing real-time convolution operations on eight channels simultaneously, with each channel's convolution processing being completely synchronized in time. The computational cost of single-channel convolution processing does not exceed 2 × 10⁻⁶ per second. 7 Sub-floating-point operations, with the total computational load for parallel processing of 8 channels, do not exceed 1.6 × 10⁻⁶ per second. 8 The total processing latency for each floating-point operation is no more than 5ms.

[0089] The processed multi-channel audio signal is input to a multi-channel pure digital power amplifier for multi-channel digital amplification before output. The multi-channel pure digital power amplifier has 12 output channels, which is no less than the number of channels in the multi-channel digital audio signal (8 channels). Each output channel has a rated output power of 50W to 200W, a signal-to-noise ratio of no less than 100dB, a total harmonic distortion of no more than 0.01%, and a dynamic range of no less than 110dB. Alternatively, the multi-channel pure digital power amplifier can have 8 to 16 output channels, which can be flexibly selected according to the user's speaker system configuration. In this embodiment, it is configured with 12 channels, of which 8 channels correspond to the 8 channels of a 7.1 channel system, and the other 4 channels can be used to expand the configuration of overhead speakers to achieve playback of 3D audio formats such as Dolby Atmos.

[0090] The entire digital signal processing link is formed by digital signal connections between the digital audio interface and the digital signal processor, and between the digital signal processor and the multi-channel pure digital power amplifier. In this link, the digital signal transmission protocol between the digital audio interface and the digital signal processor uses the I²S (Inter-IC Sound) bus protocol. The I²S bus protocol supports sampling rates from 44.1kHz to 192kHz and quantization bits from 16 bits to 32 bits, meeting the lossless transmission requirements of multi-channel high-resolution digital audio signals. The digital signal transmission protocol between the digital signal processor and the multi-channel pure digital power amplifier uses the TDM (Time Division Multiplexing) multiplexing digital audio transmission protocol. The TDM protocol transmits multiple channels of digital audio data simultaneously on a single digital audio bus using time division multiplexing, with a digital signal transmission bandwidth of no less than 100Mbps, meeting the lossless transmission requirements of 16-channel, 192kHz sampling rate, and 32-bit quantization bit multi-channel high-resolution digital audio signals.

[0091] The all-digital signal processing link maintains full digitization from digital audio interface input to multi-channel pure digital power amplifier output, effectively avoiding noise and distortion introduced by traditional analog processing links. The integrity of the original audio signal is effectively guaranteed, and the signal-to-noise ratio is improved by no less than 20dB compared to the mixed digital-analog processing link.

[0092] The table below lists typical acoustic parameters for three sound field modes according to embodiments of the present invention, to facilitate comparative analysis of the characteristic differences among the various sound field modes:

[0093] As shown in the table above, the three sound field modes exhibit significant differences in reverberation time RT60 and early reflection delay, reflecting the unique acoustic characteristics of each of the three types of professional acoustic spaces: concert halls, recording studios, and cinemas. The memory stores dedicated sound field parameter models corresponding to each of the three sound field modes, allowing users to flexibly select the appropriate sound field mode based on the type of content being listened to (symphony, studio music, film, etc.), enabling on-demand reproduction of various professional acoustic spaces.

[0094] The method in this embodiment has the following technical advantages: (1) Based on impulse response data collected on-site in real acoustic space, a unique sound field parameter model is generated instead of a general theoretical model. The reproduced reverberation characteristics, spatial sense and sound image localization are highly consistent with the target space, and the accuracy of sound field reproduction is significantly improved. (2) The entire digital signal processing link is adopted, and the digital audio interface input to the multi-channel pure digital power amplifier output is kept digital throughout. Combined with the phase filtering processing unit, the original digital audio signal is subjected to phase convolution processing, which effectively avoids the noise and distortion introduced by the analog processing link, and the integrity of the original audio signal is effectively guaranteed. (3) By storing multiple dedicated sound field parameter models corresponding to the sound field of a concert hall, the sound field of a recording studio, and the sound field of a cinema in the memory, the same device can flexibly switch the sound field mode according to the user's instructions, realize the on-demand reproduction of various professional acoustic spaces, and significantly improve the user experience. (4) The dedicated sound field parameter model is constructed based on objective measured acoustic data. The processing parameters have reliable data support, and the accuracy of the parameters and the degree of restoration of the target acoustic space are improved compared with the subjective tuning scheme.

[0095] Example 2 This embodiment discloses an audio reproduction system for on-site sound field data acquisition and digital signal processing.

[0096] An audio reproduction system for on-site sound field data acquisition and digital signal processing includes: The on-site parameter generation module is configured to: collect sound field data of the real acoustic space and generate a dedicated sound field parameter model, which includes the acoustic response characteristic parameters of the real acoustic space; The parameter matching module is configured to receive the raw digital audio signal and user command, and obtain the exclusive sound field parameter model corresponding to the user command as the target exclusive sound field parameter model. The convolution processing module is configured to convolve the original digital audio signal with the target-specific sound field parameter model to obtain a processed audio signal with enhanced spatial sense. The audio playback module is configured to output the processed audio signal after multi-channel digital amplification, thereby achieving audio reproduction based on the acoustic response characteristics of the specific sound field.

[0097] like Figure 4 As shown, the audio reproduction system mentioned in this embodiment includes two main parts: a sound field data acquisition subsystem and a user-end audio processing subsystem. These two subsystems cooperate functionally but are physically independent. The sound field data acquisition subsystem is responsible for acquiring and modeling sound field data in a professional acoustic space. After generating a dedicated sound field parameter model, it pre-loads the model file into the memory of the user-end audio processing subsystem via network transmission or copying to a storage medium. The user-end audio processing subsystem is integrated into the user's audio playback device and is responsible for loading the sound field parameter model in real time and processing and outputting the input digital audio signals.

[0098] like Figure 4 As shown, the sound field data acquisition subsystem comprises two core components: a microphone array and a computer workstation. The microphone array is positioned at typical listening positions in a real acoustic space to record the impulse response data formed after the test signal is reflected and reverberated in the real acoustic space.

[0099] The microphone array consists of 4 to 16 high-precision condenser microphones. The technical specifications of the microphones include: a frequency response range of 20Hz to 20kHz, a frequency response flatness of no more than ±1dB; an equivalent noise level of no more than 10dB (A-weighted); a maximum sound pressure level of no less than 140dB SPL; and directional types including omnidirectional, cardioid, or supercardioid, which can be selected according to the acquisition requirements.

[0100] The spatial arrangement of the microphone array is determined according to the characteristics of the target acoustic space and the number of channels required. The distance between adjacent microphones is 0.2 meters to 0.5 meters. The array as a whole covers a three-dimensional space with a horizontal azimuth angle of 360 degrees and a vertical pitch angle, so as to fully capture the three-dimensional sound field distribution characteristics of the target acoustic space.

[0101] The computer workstation communicates with the microphone array to receive impulse response data and generate a dedicated sound field parameter model. The microphone array and the computer workstation are connected digitally via an audio interface card. The audio interface card supports sampling rates from 44.1kHz to 192kHz, quantization bits from 24 to 32 bits, and has at least as many input channels as the microphone array, ensuring high resolution and low noise in the acquired data.

[0102] The computer workstation configuration includes: a processor with a clock speed of at least 3GHz, at least 16GB of memory, at least 2TB of storage, and the ability to run professional acoustic analysis software. The complete process for the computer workstation to generate a custom sound field parameter model is as follows: First, the acquired raw impulse response data is preprocessed, including removing DC bias, time-aligning the data of each microphone channel, and removing background noise interference. Then, phase response analysis is performed on the preprocessed data to extract the phase delay characteristics of each frequency component and construct the phase response parameter matrix; Simultaneously, amplitude response analysis is performed to calculate the energy attenuation characteristics of each frequency component and extract objective acoustic indicators such as RT60, EDT, and C80. Based on the objective analysis data above, professional audio engineers, combined with their professional mixing experience, finely adjusted parameters such as early reflection delay time, reverberation decay curve, amplitude and phase response of each frequency band. After the adjustments are completed, all parameters will be packaged to generate a custom sound field parameter model file. The file format is a customized digital format, and the file size is 2MB to 50MB. Finally, the generated model file is verified by loading the model file onto the test device and playing the reference audio. A professional audio engineer then conducts a subjective listening evaluation to confirm that the model effect matches the acoustic characteristics of the target acoustic space before the model file can be published to the user's device.

[0103] like Figure 5 As shown, the user-end audio processing subsystem includes four core modules: digital audio interface, digital signal processor, memory, and multi-channel pure digital power amplifier. The modules are connected by digital signals to form a fully digital signal processing link.

[0104] The digital audio interface is used to receive raw digital audio signals from external audio source devices. The digital audio interface includes at least one digital interface. In this embodiment, the digital audio interface includes three types of digital interfaces: an HDMI ARC interface, a USB digital interface, and an optical TOSLINK interface, to accommodate the connection needs of different audio source devices.

[0105] The HDMI ARC interface complies with the HDMI 2.1 specification, supports a maximum transmission bandwidth of 48Gbps, and supports 7.1-channel, 192kHz sampling rate, and 24-bit quantization for multi-channel lossless digital audio transmission. The USB digital interface conforms to the USB Audio Class 2.0 specification, supports a maximum sampling rate of 192kHz, a quantization bit depth of 32 bits, and a maximum transmission bandwidth of 480Mbps; The fiber optic TOSLINK interface conforms to the IEC 60958 standard, supports a maximum sampling rate of 96kHz and a quantization bit depth of 24 bits, and is suitable for 2-channel high-resolution digital audio transmission.

[0106] The digital audio interface is connected to the digital signal processor (DSP) to transmit the received raw digital audio signal in digital form to the DSP's input terminal. The connection method adopts the I²S bus protocol, and the digital format is maintained throughout the transmission process.

[0107] The digital signal processor is connected to the digital audio interface. The digital signal processor has an embedded phase filter processing unit. The digital signal processor is configured to: load a dedicated sound field parameter model corresponding to the user instruction from the memory, and convolve the original digital audio signal received by the digital audio interface with the dedicated sound field parameter model through the phase filter processing unit to output a processed audio signal with enhanced spatial sense.

[0108] The technical specifications of the digital signal processor include: a main frequency of not less than 500MHz, an internal arithmetic precision of 32-bit floating point, an on-chip random access memory (RAM) capacity of not less than 256MB to meet the computation and caching requirements of real-time multi-channel convolution processing; an on-chip program memory (Flash) capacity of not less than 64MB for storing the firmware program of the digital signal processor; and a number of digital audio interfaces not less than the sum of the number of input channels and the number of output channels, supporting both I²S and TDM digital audio bus protocols.

[0109] The phase filtering processing unit, as a dedicated functional module within the digital signal processor, uses a hardware-accelerated FFT operation unit to achieve efficient frequency domain convolution processing. It is configured to perform phase convolution processing on the original digital audio signal based on the phase response characteristics in the dedicated sound field parameter model.

[0110] The physical significance of phase convolution processing is: By accurately reproducing the phase response characteristics of the target acoustic space, the phase relationship of each frequency component of the original audio signal after convolution processing is completely consistent with the phase relationship when actually listening in the target acoustic space. This reconstructs the spatial positioning of the sound (the perception of the location of the sound image in three-dimensional space), the surround sound (the sense of being surrounded from the sides and rear), and the depth sound (the perception of the distance of the sound), thus achieving a high-fidelity reproduction of the immersive acoustic experience of the target acoustic space.

[0111] The digital signal processor is also configured to execute the following control logic: Receive user commands from the user interface, which include target sound field mode information (concert hall sound field, recording studio sound field, or cinema sound field). Based on the target sound field mode information, a read request is sent to the memory to select the exclusive sound field parameter model file corresponding to the target sound field mode from the memory; The read dedicated sound field parameter model file is parsed into a phase response parameter matrix and an amplitude response parameter matrix, and the above parameter matrices are written into the working register of the phase filtering processing unit; Once the parameters have been loaded, the phase filtering processing unit is started to perform real-time convolution processing on the input digital audio signal. The processing latency is monitored in real time during the process to ensure that the end-to-end processing latency does not exceed 5ms.

[0112] The execution time of the above control logic is no more than 500 milliseconds. The waiting time for users when switching sound field modes is extremely short and does not affect the user experience.

[0113] The memory is connected to the digital signal processor and is used to store dedicated sound field parameter models. These models are generated from impulse response data collected in real acoustic spaces and include the acoustic response characteristics of the actual acoustic space, especially its phase response characteristics. The memory stores multiple dedicated sound field parameter models corresponding to concert hall, recording studio, and cinema sound fields, respectively. The memory capacity is no less than 4GB, and the read rate is no less than 100MB / s to support fast loading of multiple sound field parameter model files. The memory uses non-volatile storage media (such as NAND Flash or eMMC), ensuring data retention even after power failure and guaranteeing long-term reliable storage of the sound field parameter model files.

[0114] The multi-channel pure digital power amplifier is connected to a digital signal processor (DSP) to receive processed audio signals and amplify them through multiple channels before outputting them. The multi-channel pure digital power amplifier has 8 to 16 output channels; in this embodiment, it has 12 channels. Each channel has a rated output power of 50W to 200W, a signal-to-noise ratio of no less than 100dB, a total harmonic distortion of no more than 0.01%, a dynamic range of no less than 110dB, a frequency response range of 20Hz to 20kHz, and a frequency response flatness of no more than ±0.5dB.

[0115] The multi-channel pure digital power amplifier directly receives digital audio signals from the digital signal processor (DSP), performs digital power amplification internally, and drives the corresponding speaker units. The entire process eliminates the need for a digital-to-analog conversion (DAC) stage, effectively ensuring high signal fidelity. Digital signal transmission between the multi-channel pure digital power amplifier and the DSP uses the TDM protocol, with a transmission bandwidth of no less than 100Mbps, supporting lossless transmission of multi-channel digital audio signals up to 16 channels, a 192kHz sampling rate, and 32-bit quantization.

[0116] See Figure 5The entire digital signal processing link is formed by digital signal connections between the digital audio interface and the digital signal processor, and between the digital signal processor and the multi-channel pure digital power amplifier. This link maintains full digitality from the digital audio interface input to the multi-channel pure digital power amplifier output, effectively avoiding the noise and distortion introduced by traditional analog processing, and ensuring the integrity of the original audio signal.

[0117] When the original digital audio signal is a multi-channel digital audio signal, the phase filtering processing unit is configured to convolve each channel of the multi-channel digital audio signal with a dedicated sound field parameter model. The number of output channels of the multi-channel pure digital power amplifier is no less than the number of channels of the multi-channel digital audio signal, so as to ensure that the processed audio signal of each channel can be independently output to the corresponding speaker unit, and fully restore the spatial characteristics of the multi-channel sound field.

[0118] The dedicated sound field parameter model contains the phase response characteristics of the real acoustic space. The phase filtering processing unit is configured to perform phase convolution processing on the original digital audio signal based on the phase response characteristics in the dedicated sound field parameter model. By accurately reproducing the phase response characteristics of the target acoustic space, the spatial positioning, surround sound and depth of the sound are reconstructed, and a high-fidelity reproduction of the immersive acoustic experience of the target acoustic space is achieved.

[0119] Example 3 The purpose of this embodiment is to provide a computer-readable storage medium.

[0120] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the audio reproduction method for on-site sound field data acquisition and digital signal processing as described in Embodiment 1 of this disclosure.

[0121] Example 4 The purpose of this embodiment is to provide an electronic device.

[0122] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the audio reproduction method for on-site sound field data acquisition and digital signal processing as described in Embodiment 1 of this disclosure.

[0123] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0124] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0125] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. An audio reproduction method for on-site sound field data acquisition and digital signal processing, characterized in that, Includes the following steps: Acquire sound field data of a real acoustic space and generate a dedicated sound field parameter model, wherein the dedicated sound field parameter model includes acoustic response characteristic parameters of the real acoustic space; Receive raw digital audio signals and user commands, and obtain the exclusive sound field parameter model corresponding to the user commands as the target exclusive sound field parameter model; The original digital audio signal is convolved with the target-specific sound field parameter model to obtain a processed audio signal with enhanced spatial sense. The processed audio signal is then digitally amplified through multiple channels and output to achieve audio reproduction of the specific acoustic response characteristics of the sound field.

2. The audio reproduction method for on-site sound field data acquisition and digital signal processing as described in claim 1, characterized in that, Collecting sound field data from real acoustic spaces, specifically including: A microphone array is placed at the listening position in a real acoustic space; Play test signals into a real acoustic space; The pulse response data formed by the test signal after reflection and reverberation in the real acoustic space is recorded by a microphone array, and the pulse response data is used as the sound field data.

3. The audio reproduction method for on-site sound field data acquisition and digital signal processing as described in claim 1, characterized in that, The acoustic response characteristics include phase response characteristics and amplitude response characteristics, and the acoustic response characteristic parameters include phase response parameters and amplitude response parameters; The phase response parameters reflect the phase delay characteristics of each frequency component when sound propagates in the real acoustic space, including the phase of direct sound, the phase of early reflected sound, and the phase of reverberation field. The amplitude response parameters are objective acoustic indicators that reflect the energy attenuation characteristics of each frequency component, including early decay time, reverberation time, sound intensity, and clarity.

4. The audio reproduction method for on-site sound field data acquisition and digital signal processing as described in claim 3, characterized in that, The specific process of generating a custom sound field parameter model includes: The acquired sound field data is preprocessed, including removing DC bias, time alignment, and removing background noise interference. Phase response analysis is performed on the preprocessed sound field data to extract the phase delay characteristics of each frequency component and construct the phase response parameter matrix; simultaneously, amplitude response analysis is performed to calculate the energy attenuation characteristics of each frequency component and extract objective acoustic indicators. By summarizing the phase response parameter matrix and objective acoustic indicators, objective analytical data can be obtained. Based on objective data analysis, professional audio engineers make fine adjustments to the early reflection delay time, reverberation decay curve, amplitude and phase response parameters of each frequency band; After the adjustments are complete, package all parameters to generate a custom sound field parameter model file.

5. The audio reproduction method for on-site sound field data acquisition and digital signal processing as described in claim 4, characterized in that, Also includes: The dedicated sound field parameter model is stored in the memory of the user terminal device. The dedicated sound field parameter model includes a concert hall sound field parameter model, a recording studio sound field parameter model and a cinema sound field parameter model. According to user instructions, the dedicated sound field parameter model for the corresponding sound field mode is retrieved from memory to obtain the target dedicated sound field parameter model.

6. The audio reproduction method for on-site sound field data acquisition and digital signal processing as described in claim 1, characterized in that, The original digital audio signal is a multi-channel digital audio signal; The original digital audio signal is convolved with the target-specific sound field parameter model, specifically including: Each channel of the original digital audio signal is convolved with the target-specific sound field parameter model to obtain a processed audio signal with enhanced spatial sense.

7. The audio reproduction method for on-site sound field data acquisition and digital signal processing as described in claim 5, characterized in that, After generating the custom sound field parameter model file, the following is also included: The generated model file is verified by loading the model file onto the test device and playing the reference audio. A professional audio engineer conducts a subjective listening evaluation to confirm that the model effect conforms to the acoustic characteristics of the target acoustic space. or, Also includes: It receives raw digital audio signals through a digital audio interface, maintaining the digital format throughout the process without any analog conversion. The dedicated sound field parameter model is retrieved from memory by the digital signal processor and loaded into the phase filtering processing unit embedded in the digital signal processor. The original digital audio signal is convolved with a dedicated sound field parameter model by the phase filtering processing unit in the digital signal processor, and the processed audio signal with enhanced spatial sense is output. The processed audio signal is input into a multi-channel pure digital power amplifier for multi-channel digital amplification and then output. After digital power amplification, it drives the corresponding speaker unit. There is no analog signal conversion throughout the process.

8. An audio reproduction system for on-site sound field data acquisition and digital signal processing, characterized in that, include: The on-site parameter generation module is configured to: collect sound field data of a real acoustic space and generate a dedicated sound field parameter model, wherein the dedicated sound field parameter model includes acoustic response characteristic parameters of the real acoustic space; The parameter matching module is configured to: receive the original digital audio signal and user command, and obtain the exclusive sound field parameter model corresponding to the user command as the target exclusive sound field parameter model; The convolution processing module is configured to: perform convolution processing on the original digital audio signal and the target-specific sound field parameter model to obtain a processed audio signal with enhanced spatial sense; The audio playback module is configured to: amplify the processed audio signal through multiple channels and then output it to achieve audio reproduction of the specific acoustic response characteristics of the sound field.

9. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the audio reproduction method for on-site sound field data acquisition and digital signal processing as described in any one of claims 1-7.

10. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the audio reproduction method for on-site sound field data acquisition and digital signal processing as described in any one of claims 1-7.