An immersive adaptive rendering method, processor, and system for audio files
By combining sound source localization and distance-aware adaptive labeling in the immersive adaptive rendering method for audio files, and taking into account physical space design, the problem of unrealistic audio reproduction in existing technologies is solved, and realistic spatial sound effects are achieved under different speaker configurations and playback capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2026-03-13
AI Technical Summary
Existing immersive audio processing methods cannot meet the needs of immersive production and playback in small and medium-sized venues. HOA technology may result in the loss of high-frequency cues when reconstructing 3D sound fields, and VBAP technology produces jumps when rendering moving sound sources. It is impossible to achieve realistic spatial sound effects in listening environments with different speaker configurations and playback capabilities.
The sound source localization adaptive labeling method and the distance perception adaptive labeling method are used to render static sound sources. Combined with the actual physical space design, the adaptive rendering algorithm considers factors such as horizontal plane, vertical axis plane, frequency and azimuth angle. The motion perception adaptive labeling method is used to render dynamic sound sources. The motion trajectory of the moving sound source and the motion speed are designed to increase the sense of presence.
It achieves spatial reproduction of audio in real physical space, and is suitable for traditional sound transmission systems and new digital audio systems with surround-screen speaker arrays, thus enhancing the user's listening experience.
Smart Images

Figure CN115604644B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio file rendering technology, and in particular to an immersive adaptive rendering method, processor and system for audio files. Background Technology
[0002] In recent years, with the continuous development of high-definition video, from 2K to 4K and even 8K, and with the development of virtual reality (VR) and augmented reality (AR), people's auditory requirements for audio have also increased. People are no longer satisfied with the stereo, 5.1, and 7.1 sound effects that have been popular for many years, and are beginning to pursue more immersive and realistic 3D sound effects or immersive audio experiences. Currently, immersive audio processing is mainly based on technologies such as channel-based audio (CBA), object-based audio (OBA), and Ambisonics scene-based audio (SBA), encompassing audio production, encoding / decoding, packaging, and rendering techniques.
[0003] Specifically, Ambisonics scene audio uses spherical harmonic functions to record the sound field and drive the speakers, with strict requirements for speaker placement, and can reconstruct the original sound field with high quality at the center of the speakers. When rendering moving sound sources, HOA (Higher Order Ambisonics) creates a smoother and more fluid listening experience.
[0004] Furthermore, Vector Based Amplitude Panning (VBAP) is based on the sine law in three-dimensional space. It uses three adjacent speakers in space to form a three-dimensional sound vector, without affecting the interaural time difference (ITD) in low frequencies or the spectral cues in high frequencies, resulting in more accurate sound localization in three-dimensional space. Due to its simplicity, VBAP has become the most commonly used multi-channel three-dimensional audio processing technique.
[0005] However, existing immersive audio processing methods cannot meet the needs of immersive production and playback in small and medium-sized venues. Furthermore, HOA uses an intermediate format to reconstruct a 3D sound field, which may result in a lack of high-frequency cues due to the limited order used, thus affecting the accuracy of the listener's positioning. On the other hand, VBAP produces jumps when rendering moving sound sources, resulting in disjointed spatial sound effects.
[0006] While advanced 3D audio systems, such as the Atmos™ system, are largely designed and deployed for cinematic applications, consumer-grade systems are being developed to bring cinematic, adaptive audio experiences to home and office environments. These environments are significantly constrained by venue size, acoustic characteristics, system power, and speaker configurations compared to cinemas. Current professional-grade spatial audio systems therefore need to be adapted to render advanced object audio content to listening environments characterized by varying speaker configurations and playback capabilities. To this end, certain virtualization techniques have been developed to extend the capabilities of traditional stereo or surround sound speaker arrays, thereby reconstructing spatial sound cues using sophisticated rendering algorithms and techniques such as content-dependent rendering algorithms, reflected sound transmission, etc. Such rendering techniques have led to the development of DSP-based renderers and circuitry optimized for rendering different types of adaptive audio content, such as Object Audio Metadata Content (OAMD) beds and ISF (Intermediate Spatial Format) objects.
[0007] The introduction of digital cinema and the development of true 3D (“3D”) or virtual 3D content have created new sound standards, such as the merging of multiple audio channels to allow for greater creativity among content creators and a more immersive and realistic auditory experience for viewers. As a means of distributing spatial audio, extending beyond traditional speaker feeds and channel-based audio is crucial, and there has been considerable interest in model-based audio descriptions, which allow listeners to choose their desired playback configuration, thereby rendering audio specifically for their chosen configuration. Spatial representation of sound utilizes audio objects, which are audio signals with associated parameterized source descriptions having apparent source location (e.g., 3D coordinates), apparent source width, and other parameters. Further developments include the development of next-generation spatial audio (also known as “adaptive audio”) formats that incorporate audio objects and traditional channel-based speaker feeds, along with the location metadata of the audio objects. In a spatial audio decoder, channels are either directly transmitted to their associated speakers or downmixed into existing speaker arrays, and the audio objects are rendered by the decoder in a flexible (adaptive) manner. The parametric source description associated with each object (such as its positional trajectory in 3D space) along with the number and location of the speakers connected to the decoder are taken as input. The renderer then uses certain algorithms (such as translation laws) to distribute the audio associated with each object across the attached set of speakers. The creative spatial intent of each object is thus optimally represented on the specific speaker configuration present in the listening room.
[0008] In existing technologies, spatial audio only considers the number and location of the speakers used after decoding the audio file, and it lacks consideration of the actual physical space, thus failing to reproduce the true spatial nature of the audio. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to provide an immersive adaptive rendering method, processor and system for audio files that not only considers the number and location of the speakers, but also incorporates the actual physical space design, thereby restoring the real spatial reproduction of the audio.
[0010] The technical solution adopted in this invention is an immersive adaptive rendering method for audio files, which includes the following steps:
[0011] S1. Set the length of the engineering physical space to L, the width to W, and the height to H. Set the position coordinates of the speakers placed in the engineering physical space to (xi,yi,zi), where i represents the total number of speakers, i = 1, 2, 3...N;
[0012] S2. Decode the audio file to be processed, preprocess the decoded audio file, and obtain the preprocessed audio file;
[0013] S3. The static sound sources in the preprocessed audio file obtained in step S2 are marked and rendered using the sound source localization adaptive marking method or the distance-aware adaptive marking method.
[0014] The specific process of using the sound source localization adaptive marking method for marking and rendering includes:
[0015] S301. Locate the static sound source and set its coordinates as (x, y, z) and sound pressure level as SPL0.
[0016] S302. Calculate the spatial distance between the static sound source and each speaker arranged in the engineering physical space.
[0017] S303. Sort the calculated spatial distances from smallest to largest, and select the three speakers with the smallest spatial distances to play the static sound source.
[0018] The specific process of using the distance-aware adaptive labeling method for label rendering includes:
[0019] S310. Locate the static sound source, set its coordinates as (x, y, z), and the sound pressure level as SPL0. Calculate the spatial distance between the static sound source and each speaker in the engineering physical space. Sort the calculated spatial distances from smallest to largest and select the three speakers with the smallest spatial distance.
[0020] S311. Set the length of the listening area in the engineering physical space to L1, the width to W1, and the height to H1. Set the coordinates of the listening point of the listener in the engineering physical space to (x0, y0, z0), where x0 = (1 / 2) * L1, y0 = (1 / 2) * W1, and z0 = (1 / 2) * H1.
[0021] S312. Obtain the sound pressure level of the static sound source signal in the audio file, and decompose the static sound source signal into signals in the horizontal direction and signals in the vertical direction.
[0022] S313. Set a virtual static sound source in the engineering physical space, and set the theoretical rendering distance Dis between the virtual static sound source and the listening point. At the same time, calculate the sound source orientation angle of the virtual static sound source relative to the listener, that is: angle=Actan((y-y0) / (x-x0)).
[0023] S314. When the set theoretical rendering distance Dis < 2.5m, based on the influence of the sound source azimuth angle of the virtual static sound source on the perception of rendering distance, the sound source azimuth angle of the horizontal signal obtained by decomposition in S312 is processed to obtain the sound source azimuth angle after angle processing. The expression of the processed sound source azimuth angle is: Angle = -0.236*lg(angle) + 2.75.
[0024] S315. The horizontal signal processed by angle in S314 and the vertical signal obtained by decomposition in step S312 are filtered by an octave band IIR filter to obtain a filtered static sound source signal. The filtered static sound source signal is processed by a frequency distance sensing algorithm and then processed to obtain sound signals of different frequency bands. All the sound signals of different frequency bands are linearly added together to synthesize a new static sound source signal.
[0025] S316. Use the three speakers selected in step S310 to distribute and replay the new static sound source signal obtained in S315.
[0026] S4. The dynamic sound sources in the preprocessed audio file obtained in step S2 are marked and rendered using the motion-aware adaptive marking method.
[0027] The beneficial effects of this invention are as follows: The aforementioned immersive adaptive rendering method for audio files, during the static sound source marking and rendering process, considers not only the sound pressure level factor of the actual distance of the static sound source in real physical space, but also the differences in distance perception rendering caused by factors such as the horizontal plane, vertical axis plane, frequency, and azimuth angle. Furthermore, the aforementioned immersive adaptive rendering method for audio files also employs a motion-aware adaptive marking method to mark and render dynamic sound sources. This method considers not only the number and location of the speakers but also incorporates the actual physical space design, thereby restoring the true spatial reproduction of the audio file. The aforementioned immersive adaptive rendering method for audio files is applicable not only to traditional sound transmission systems but also to novel digital sound reproduction systems with surround-screen speaker arrays, demonstrating strong applicability.
[0028] Preferably, in step S4, the specific process of using the motion-aware adaptive labeling method to label and render the dynamic sound sources in the preprocessed audio file includes:
[0029] S401. Design the motion sensing trajectory of the moving sound source, set the number of sound channels on the motion sensing trajectory to N, the total duration of motion sensing to T1, the moving sound source signal to S, and the sampling rate to 44.1kHz.
[0030] S402. Set the fade-in and fade-out time of sound playback on each audio channel to t, and set the fade-in and fade-out method of sound playback on each audio channel.
[0031] S403. Calculate the playback duration of the moving sound source on each audio channel:
[0032] S404. Based on the playback duration T of the moving sound source on each audio channel, perform signal cutting processing on the moving sound source signal on each corresponding channel to obtain a new signal S' after cutting processing on each channel, S' = S(N:N(fs*T-1));
[0033] S405. The signal S' after clipping on each channel is faded in and out for time t, and then mixed to obtain the final multi-channel motion sensing signal.
[0034] In the above-mentioned adaptive rendering motion algorithm for dynamic sound sources, the process of marking and rendering dynamic sound sources not only considers the design of the sound source motion trajectory, but also the selection of motion speed, which increases the sense of presence of the sound effects of moving sound sources and realistically restores the spatial reproduction of audio.
[0035] Preferably, in step S2, the preprocessing of the decoded audio file specifically includes equalization processing, volume processing, delay processing, and mixing processing.
[0036] Preferably, in step S312, the specific process of obtaining the sound pressure level of the static sound source signal in the audio file is as follows: select the sound power of a 75 dB white noise signal as the reference sound power, calculate the sound power of the static sound source signal in the audio file, compare the calculated sound power of the static sound source signal with the reference sound power, obtain the comparison difference between the two, and process the sound pressure level of the static sound source signal according to the comparison difference.
[0037] Preferably, in step S315, after processing the obtained filtered static sound source signal using a frequency distance sensing algorithm, the horizontal signal sound pressure level of the filtered static sound source signal is: Dis1 represents the horizontal straight-line distance of the static sound source in the audio file; the sound pressure level of the filtered static sound source signal in the vertical direction is: Here, Dis2 represents the vertical distance of a static sound source in the audio file.
[0038] A processor includes an audio adaptive rendering module, which includes operation instructions for executing the aforementioned immersive adaptive rendering method for audio files. Using this processor, audio files can be directly adaptively rendered, and the adaptively rendered audio files can be realistically reproduced in physical space. It is powerful and highly applicable.
[0039] A system comprising the aforementioned processor, speakers arranged in a physical space, and an LED immersive surround screen array channel, which can realistically reproduce audio files, is powerful, and greatly enhances the user experience. Attached Figure Description
[0040] Figure 1 This is a flowchart of an immersive adaptive rendering method for audio files according to the present invention;
[0041] Figure 2 This is a schematic diagram of the engineering physical space when using the distance-aware adaptive marking method for marking and rendering in this invention;
[0042] Figure 3 This is a comparison diagram of the results of distance compression perception in the horizontal and vertical axes during the distance-aware adaptive labeling and rendering process in the invention.
[0043] Figure 4This is a general distribution diagram of the sensing data of the motion image angle range of the moving sound source in the invention.
[0044] Figure 5 This is a distribution diagram showing the differences in the perceived data of the motion image angle range of the moving sound source in the invention.
[0045] As shown in the figure: 1. Speaker; 2. Virtual static sound source; 3. Listening point for the listener. Detailed Implementation
[0046] The utility model will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can implement it based on the description. The scope of protection of this utility model is not limited to the specific embodiments.
[0047] Embodiments of the present invention provide an immersive adaptive rendering method for audio files, such as... Figure 1 As shown, the method includes the following steps:
[0048] S1. Set the length of the engineering physical space to L, the width to W, and the height to H. Set the position coordinates of the speakers placed in the engineering physical space to (xi,yi,zi), where i represents the total number of speakers, i = 1, 2, 3...N;
[0049] S2. Decode the audio file to be processed, preprocess the decoded audio file, and obtain the preprocessed audio file;
[0050] S3. The static sound sources in the preprocessed audio file obtained in step S2 are marked and rendered using the sound source localization adaptive marking method or the distance-aware adaptive marking method.
[0051] The specific process of using the sound source localization adaptive marking method for marking and rendering includes:
[0052] S301. Locate the static sound source and set its coordinates as (x, y, z) and sound pressure level as SPL0.
[0053] S302. Calculate the spatial distance between the static sound source and each speaker arranged in the engineering physical space.
[0054] S303. Sort the calculated spatial distances from smallest to largest, and select the three speakers with the smallest spatial distances to play the static sound source.
[0055] The specific process of using the distance-aware adaptive labeling method for label rendering includes:
[0056] S310. Locate the static sound source, set its coordinates as (x, y, z), and the sound pressure level as SPL0. Calculate the spatial distance between the static sound source and each speaker in the engineering physical space. Sort the calculated spatial distances from smallest to largest and select the three speakers with the smallest spatial distance.
[0057] S311. Set the length of the listening area in the engineering physical space to L1, the width to W1, and the height to H1. Set the coordinates of the listening point of the listener in the engineering physical space to (x0, y0, z0), where x0 = (1 / 2) * L1, y0 = (1 / 2) * W1, and z0 = (1 / 2) * H1.
[0058] S312. Obtain the sound pressure level of the static sound source signal in the audio file, and decompose the static sound source signal into signals in the horizontal direction and signals in the vertical direction;
[0059] S313, such as Figure 2 As shown, a virtual static sound source is set in the engineering physical space, and the theoretical rendering distance Dis between the virtual static sound source and the listening point is set. At the same time, the sound source orientation angle of the virtual static sound source relative to the listener is calculated, that is: angle=Actan((y-y0) / (x-x0)), which is convenient for the calculation needs of distance perception in the later stage.
[0060] S314. When the set theoretical rendering distance Dis < 2.5m, based on the influence of the sound source azimuth angle of the virtual static sound source on the perception of rendering distance, the sound source azimuth angle of the horizontal signal obtained by decomposition in S312 is processed to obtain the processed sound source azimuth angle. The expression of the processed sound source azimuth angle is: Angle = -0.236*lg(angle) + 2.75, which increases the accuracy of the perception of theoretical rendering distance.
[0061] S315. The horizontal signal processed by angle in S314 and the vertical signal obtained by decomposition in step S312 are filtered by an octave band IIR filter. The frequency range of the octave band IIR filter is 250Hz to 8000Hz. The filtered static sound source signal is obtained. The filtered static sound source signal is processed by a frequency distance sensing algorithm (different frequencies of the audio signal have different weight ratios for the distance sensing algorithm, and different distance sensing algorithms are required for different frequencies). Then, signal processing is performed to obtain sound signals of different frequency bands. All the sound signals of different frequency bands are linearly added together to synthesize a new static sound source signal.
[0062] S316. Using the three speakers selected in step S310, the new static sound source signal obtained in S315 is distributed and replayed to achieve the virtual rendering distance Dis theoretically set in step S313. Dis1 represents the straight-line distance of the static sound source rendered in the horizontal direction, and Dis2 represents the straight-line distance of the static sound source rendered in the vertical direction.
[0063] S4. The dynamic sound sources in the preprocessed audio file obtained in step S2 are marked and rendered using the motion-aware adaptive marking method.
[0064] The beneficial effects of this invention are as follows: The aforementioned immersive adaptive rendering method for audio files, during the static sound source marking and rendering process, considers not only the sound pressure level factor of the actual distance of the static sound source in real physical space, but also the differences in distance perception rendering caused by factors such as the horizontal plane, vertical axis plane, frequency, and azimuth angle. Furthermore, the aforementioned immersive adaptive rendering method for audio files also employs a motion-aware adaptive marking method to mark and render dynamic sound sources. This method considers not only the number and location of the speakers but also incorporates the actual physical space design, thereby restoring the true spatial reproduction of the audio file. The aforementioned immersive adaptive rendering method for audio files is applicable not only to traditional sound transmission systems but also to novel digital sound reproduction systems with surround-screen speaker arrays, demonstrating strong applicability.
[0065] Preferably, in step S4, the specific process of using the motion-aware adaptive labeling method to label and render the dynamic sound sources in the preprocessed audio file includes:
[0066] S401. Design the motion sensing trajectory of the moving sound source, set the number of sound channels on the motion sensing trajectory to N, the total duration of motion sensing to T1, the moving sound source signal to S, and the sampling rate to 44.1kHz.
[0067] S402. Set the fade-in and fade-out time of sound playback on each audio channel to t, and set the fade-in and fade-out method of sound playback on each audio channel.
[0068] S403. Calculate the playback duration of the moving sound source on each audio channel:
[0069] S404. Based on the playback duration T of the moving sound source on each audio channel, perform signal clipping on the moving sound source signal on each corresponding channel to obtain the clipped signal S' on each channel, S' = S(N:N(fs*T-1));
[0070] S405. The signal S on each channel, after being cut, is faded in and out for time t, and then mixed to obtain the final multi-channel motion sensing signal.
[0071] In the above-mentioned adaptive rendering motion algorithm for dynamic sound sources, the process of marking and rendering dynamic sound sources not only considers the design of the sound source motion trajectory, but also the selection of motion speed, which increases the sense of presence of the sound effects of moving sound sources and realistically restores the spatial reproduction of audio.
[0072] Preferably, in step S2, the preprocessing of the decoded audio file specifically includes equalization processing, volume processing, delay processing, and mixing processing.
[0073] Preferably, in step S312, the specific process of obtaining the sound pressure level of the static sound source signal in the audio file is as follows: select the sound power of a 75 dB white noise signal as the reference sound power, calculate the sound power of the static sound source signal in the audio file, compare the calculated sound power of the static sound source signal with the reference sound power, obtain the comparison difference between the two, and obtain the sound pressure level of the static sound source signal based on the comparison difference.
[0074] Preferably, in step S315, after processing the obtained filtered static sound source signal using a frequency distance sensing algorithm, the horizontal signal sound pressure level of the filtered static sound source signal is: Dis1 represents the horizontal straight-line distance of the static sound source in the audio file; the sound pressure level of the filtered static sound source signal in the vertical direction is: Here, Dis2 represents the vertical distance of a static sound source in the audio file.
[0075] A processor includes an audio adaptive rendering module, which includes operation instructions for executing the aforementioned immersive adaptive rendering method for audio files. Using this processor, audio files can be directly adaptively rendered, and the adaptively rendered audio files can be realistically reproduced in physical space. It is powerful and highly applicable.
[0076] A system includes a processor as described above, speakers arranged in a physical space, and an LED immersive ring screen array channel. Using this system, audio files processed by the processor are passed through the LED immersive ring screen array channel and finally output by the speakers. This system can realistically reproduce audio files, has powerful functions, and greatly improves the user experience.
[0077] In the research process of an immersive adaptive rendering method for audio files provided in the embodiments of the present invention, relevant experimental results were obtained:
[0078] I. Differences between horizontal and vertical planes in distance-aware rendering:
[0079] Subjective distance judgments along the horizontal and vertical axes may yield different psychological results. Similarly, the compression rate of auditory distance perception along the vertical axis is plotted as a compression function graph in supine and lateral positions: a comparison of the compression rate of subjective distance perception along the four axes under pure-tone stimulation, and a comparison of the compression rate of subjective distance perception along the four axes under narrow-band stimulation, see [link to relevant documentation]. Figure 3 As shown, the actual fixed distance of the sound source is 3.5m.
[0080] Depend on Figure 3 It can be seen that when the stimulus source is a pure tone signal, the change trend of the subjective distance perception compression rate in the horizontal axis and the vertical axis is consistent. That is, as the frequency increases, the subjective distance perception compression rate decreases in a logarithmic function. At the same time, the compression rate decreases from 0.7 to about 0.4. Therefore, the pure tone signal has a similar response mechanism in terms of psychological perception in the horizontal axis and the vertical axis.
[0081] When the stimulus source is a narrowband signal, the trends in the subjective distance perception compression rates along the four axes remain the same as those for pure-tone stimuli. However, a significant difference exists: the subjective distance perception compression rate along the horizontal axis is significantly lower than that along the vertical axis. Therefore, for narrowband stimuli, the psychological compression of subjective distance perception along the vertical axis is greater. The horizontal and vertical axes maintain a stable gap in subjective distance perception compression rates, indirectly reflecting that the psychological response mechanisms for perceived distance along the horizontal and vertical axes are similar, only differing in the degree of compression.
[0082] II. Differences exist between different spatial regions of a moving sound source:
[0083] The total number of data points obtained in the experiment was 280 (stimulus signals) * 15 (subjects) = 4200. All data were subjected to a Cronbach's α coefficient reliability test, and the results showed that the overall experimental data had high internal consistency (Cronbach's Alpha = 0.906, N = 15). Since the range of motion sound image angle changes is 90° in real-world environments, preliminary statistics of the overall data were performed (see below). Figure 4 The total number of motion image angles perceived is greater than 90°: 1308; the total number of angles equal to 90° is 699; and the total number of angles less than 90° is 2193.
[0084] The overall data were categorized, with 560 data points per spatial region, 600 data points per stimulus source, and 840 data points per rotational speed. Under identical experimental conditions (same stimulus source, spatial region, and rotational speed), for 16 participants, invalid data points (a total of 18 invalid data points) were removed using a normal distribution with one standard deviation. The mean of the remaining data points was then calculated, yielding 280 mean data points. Preliminary ANOVA statistical analysis was performed on each factor (stimulus source, rotational speed, and spatial region) with a significance level set at 0.05, to determine the significance of the main effects.
[0085] As shown in Table 2, the data results indicate that the main effects of spatial region and rotational speed are significant, while the main effect of the stimulus source is not significant.
[0086] Table 2 One-way ANOVA for each factor
[0087]
[0088] The main effects of different factors and the interactions between factors are independent of each other. Therefore, an interaction analysis was conducted on the three factors of spatial region, rotation speed and stimulus source type. The results are shown in Table 3. The data show that there is a significant interaction between spatial region and rotation speed. At the same time, there are significant interactions between spatial region and stimulus source, and between rotation speed and stimulus source.
[0089] Table 3. Analysis of interactions between factors
[0090]
[0091] III. Perception Analysis of the Angle Range of Moving Sound Sources:
[0092] Statistical results of valid data for the entire group of participants (see) Figure 5 It can be seen that 31% of the motion sound image angle ranges were perceived as being in an expanded state, with a mean of 125° and a standard deviation of 27°; 52% were perceived as being in a compressed state, with a mean of 54° and a standard deviation of 22°; and only 17% of the motion sound image angle ranges could be correctly judged by the human ear.
[0093] Therefore, the human ear's perception of the range of motion angles of a sound source does not match the actual range of motion; the perceived range of motion sound image angles is compressed to a greater extent. Different factors have varying degrees of influence on the state and value of the perceived range of motion sound image angles. Therefore, single-factor analysis and interaction analysis are performed for each influencing factor. In all the following figures, the red dividing line represents the actual angular range of sound source movement, i.e., 90°.
Claims
1. A method of immersive, adaptive rendering of an audio file, characterized by: The method comprises the following steps: S1, setting the length, width and height of the engineering physical space as L, W and H respectively, and setting the position coordinates of the sound equipment arranged in the engineering physical space as (xi, yi, zi), wherein i represents the total number of the sound equipment, i = 1, 2, 3, …, n; S2, decoding the audio file to be processed, and pre-processing the decoded audio file to obtain a pre-processed audio file; S3, using a distance perception adaptive marking method to mark and render the static sound source in the pre-processed audio file obtained in step S2; The specific process of using the distance perception adaptive marking method to mark and render comprises: S310, positioning the static sound source, setting its coordinates as (x, y, z) and sound pressure level as SPL0, respectively calculating the spatial distance between the static sound source and each sound equipment arranged in the engineering physical space, sorting the calculated spatial distances from small to large, and selecting three sound equipments with the smallest spatial distance; S311, setting the length, width and height of the listening area in the engineering physical space as L1, W1 and H1 respectively, and setting the listening point coordinates of the listening personnel in the engineering physical space as (x0, y0, z0), wherein x0 = (1 / 2)*L1, y0 = (1 / 2)*W1, and z0 = (1 / 2)*H1; S312, obtaining the sound pressure level of the static sound source signal in the audio file, and decomposing the static sound source signal into a signal in the horizontal direction and a signal in the vertical direction; S313, setting a virtual static sound source in the engineering physical space, setting the theoretical rendering distance Dis formed between the virtual static sound source and the listening point, and calculating the sound source azimuth angle of the virtual static sound source relative to the listening personnel, i.e. angle = Actan((y-y0) / (x-x0)); S314, when the set theoretical rendering distance Dis is less than 2.5 m, according to the influence of the sound source azimuth angle of the virtual static sound source on the rendering distance perception, performing angle processing on the sound source azimuth angle of the signal in the horizontal direction obtained by decomposing in step S312, to obtain an angle-processed sound source azimuth angle, and the expression of the processed sound source azimuth angle is Angle = -0.236*lg(angle) + 2.75; S315, filtering the signal in the horizontal direction processed in step S314 and the signal in the vertical direction obtained by decomposing in step S312 through an octave IIR filter respectively to obtain a filtered static sound source signal, performing a frequency distance perception algorithm processing on the obtained filtered static sound source signal to obtain sound signals of different frequency bands, and linearly adding all the sound signals of different frequency bands to synthesize a new static sound source signal; S316, using the three sound equipments selected in step S310 to distribute and play back the new static sound source signal obtained in step S315; S4, using a motion perception adaptive marking method to mark and render the dynamic sound source in the pre-processed audio file obtained in step S2.
2. The method of claim 1, wherein: In step S4, the specific process of motion-aware adaptive tagging for the dynamic sound source in the pre-processed audio file obtained in step S2 includes: S401, design a motion-aware trajectory route of the motion sound source, set the number of sound channels on the motion-aware trajectory route as N, the total time length of motion-aware as T1, the motion sound source signal as S, and the sampling rate fs as 44.1 kHz; S402, set the fade-in and fade-out time of sound playback on each sound channel as t, and set the fade-in and fade-out mode of sound playback on each sound channel; S403, the playback duration of the moving sound source on each sound channel is calculated as: S404. Based on the playback duration T of the moving sound source on each audio channel, perform signal clipping processing on the moving sound source signal on each corresponding channel to obtain the new signal S after clipping processing on each channel. ’ S ’ =S(N:N(fs*T-1)); S405, the signal S after the shearing process on each channel is obtained ’ The fade-in and fade-out process is performed for time t, and then the mixing process is performed to obtain the final multi-channel motion perception signal.
3. The method of claim 1 or claim 2, wherein: In step S2, the pre-processing of the decoded audio file specifically includes: equalization processing, volume processing, delay processing and mixing processing on the decoded audio file.
4. The method of claim 1 or claim 2, wherein: In step S312, the specific process of obtaining the sound pressure level of the static sound source signal in the audio file is: selecting the sound power of a 75 decibel white noise signal as a reference sound power, calculating the sound power of the static sound source signal in the audio file, comparing the calculated sound power of the static sound source signal with the reference sound power to obtain the comparison difference, and obtaining the sound pressure level of the static sound source signal according to the comparison difference.
5. The method of claim 4, wherein: In step S315, after the obtained filtered static sound source signal is processed by the frequency distance perception algorithm, the signal sound pressure level of the filtered static sound source signal in the horizontal direction is: Dis1 represents the straight-line distance of the static sound source rendered in the horizontal direction in the audio file; the signal sound pressure level of the filtered static sound source signal in the vertical direction is: Wherein, Dis2 represents the straight-line distance of the static sound source rendered in the vertical direction in the audio file.
6. A processor for immersive adaptive rendering of an audio file, characterized in that: The processor includes an audio adaptive rendering module, and the audio adaptive rendering module includes operation instructions for executing the immersive adaptive rendering method of an audio file according to any one of claims 1 to 5.
7. A system for immersive adaptive rendering of an audio file, characterized in that: It comprises a sound in a physical space, an LED immersive ring screen array channel arranged in a physical space, and a processor according to claim 6.
Citation Information
Patent Citations
Adaptive sound field control
CN103733648A
Method, device and equipment for restoring sound field space and tracking attitude
CN114173256A