A method and system for generating a three-dimensional virtual surround sound

By decomposing and processing audio signals, constructing multiple channels, and utilizing crosstalk cancellation technology, the problems of narrow virtual surround sound field and insufficient surround sound richness are solved, achieving an immersive experience of three-dimensional sound effects.

CN116847269BActive Publication Date: 2026-06-02SHENZHEN XINZHONGXIN TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN XINZHONGXIN TECH CO LTD
Filing Date
2023-03-07
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing virtual surround sound technology suffers from insufficient sound field width and lack of surround sound richness, leading to auditory fatigue and muddy sound, especially in two-channel stereo products where it is difficult to achieve a good virtual sound field.

Method used

The system decomposes audio signals into high-frequency and low-frequency signals, constructs multiple channels through an infinite N-level echo module and filter processing, and combines crosstalk cancellation technology and the Laue effect to generate virtual surround sound with three-dimensional sound effects.

Benefits of technology

It achieves a virtual surround sound effect with a wide sound field and rich surround sound, enhances the sense of space, distance and diffusion, eliminates crosstalk interference, and provides an immersive listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116847269B_ABST
    Figure CN116847269B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional virtual surround sound generation method and generation system, and relates to the field of surround sound. The application adopts an excellent structure of a howling elimination technology to eliminate howling, enhance spatial positioning and speech clarity. Reflection sound is constructed by infinite N-level echo simulation to construct a surround sound channel, so that the generated surround sound has more distant background sound, thereby making the virtual surround sound have strong spatial sense, distance sense, diffusivity and surrounding sense, rather than just positioning sense. A post sound channel is constructed by inversely superimposing the delayed signal on the direct signal, so that the spatial sense of the surround sound is obviously enhanced, and the sound seems to come from all directions. The original sound bass is preserved by separation technology, and the virtual sound in each direction is freely controlled, thereby controlling the sound intensity, tone brightness and tone softness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of surround sound, and in particular to a method and system for generating three-dimensional virtual surround sound. Background Technology

[0002] Audio equipment is an essential part of home entertainment. As people increasingly demand realistic video playback and a sense of presence, their requirements for audio playback quality are also rising. Simple sound effects alone are no longer sufficient; surround sound is one of the main directions of current audio-visual equipment development. Traditional surround sound systems use five or more speakers placed throughout the room to achieve a surround effect, resulting in complex wiring. Wireless connections for these speakers are limited by network transmission, leading to poor synchronization and significant sound quality loss. Therefore, listeners desire to enjoy surround sound without the inconvenience of adding additional rear speakers.

[0003] Therefore, some integrated speakers use side-mounted horns on both sides, forming a reflection angle with the left and right walls. This utilizes the walls of the room to reflect sound, achieving a virtual surround sound effect. However, this type of reflective virtual surround sound device, which uses reflective materials for direct reflection, not only requires additional speakers, increasing cost and product size, but also places high demands on product placement due to its direct reflection principle. This is unsuitable for today's fast-paced lifestyle where people strongly value aesthetics and convenience. Furthermore, unlike headphones where the two speakers are independent in the auditory space (isolated by the head), the auditory environment created by audio equipment located in a three-dimensional space can cause crosstalk.

[0004] The size and portability requirements of mobile communication systems make it impossible to achieve 3D surround sound effects through multi-channel audio. Virtual surround sound technology is one of the important technologies for the future development of surround sound audio. Existing virtual sound generation systems only utilize two speakers to provide surround sound effects such as 5.1 channel systems. WO 99 / 49574 (PCT / AU / 00002, entitled "Audio Signal Processing Method and Apparatus") discloses an implementation method for virtual sound systems, such as... Figure 1 and Figure 2As shown. However, this technology also has many defects and problems. Virtual surround sound uses front speakers and software algorithms to expand the sound field, and most front speakers are relatively close together, making the sound sound muddy and the sound field not wide enough, making it difficult to achieve a good virtual sound field. In addition, most virtual surround sound technologies cause "ear pressure"—creating a feeling of pressure on the human auditory nervous system. Therefore, establishing a wider virtual sound field and richer and more detailed three-dimensional spatial sound effects, adaptable to any two-channel stereo speaker product, is an urgent problem to be solved. Therefore, there is an urgent need in this field for a technical solution that can provide a wide sound field and rich surround sound in virtual surround sound. Summary of the Invention

[0005] The purpose of this invention is to provide a virtual surround sound generation scheme with a wide sound field and rich ambient sound, so as to solve the problems of auditory fatigue and muddy sound and narrow surround sound that are easily caused by general virtual surround sound.

[0006] To achieve the above objectives, the present invention provides the following solution:

[0007] A method for generating three-dimensional virtual surround sound includes:

[0008] Decompose the audio signal into high-frequency and low-frequency signals;

[0009] The center channel signal is obtained by summing and averaging the main left channel signal and the main right channel signal of the high-frequency signal.

[0010] The main left channel signal and the main right channel signal of the high-frequency signal are respectively passed through an infinite N-level echo module to construct the left surround channel and the right surround channel.

[0011] The main left channel signal and the main right channel signal of the high-frequency signal are delayed and then inverted and superimposed on the original signal to construct the rear left channel and the rear right channel;

[0012] The center channel signal is filtered by a first filter to obtain a virtual center channel signal for the left ear; the center channel signal is filtered by a second filter to obtain a virtual center channel signal for the right ear.

[0013] The main left channel signal is filtered by the third filter to obtain the positioning signal in the direction of the main left channel. The signal after the main right channel signal is filtered by the fourth filter is then superimposed to obtain the signal after the main left channel signal is filtered to obtain the signal after the main left channel signal is filtered to obtain the signal after the main left channel signal is filtered to obtain the signal after the main right channel signal is filtered to obtain the positioning signal in the direction of the main right channel. The signal after the main right channel signal is filtered by the fifth filter to obtain the positioning signal in the direction of the main right channel. The signal after the main left channel signal is superimposed to obtain the signal after the main right channel signal is filtered by the sixth filter to obtain the signal after the main right channel signal is filtered to obtain the signal after the main right channel signal is filtered.

[0014] The left surround channel signal is filtered by the seventh filter to obtain the positioning signal in the direction of the left surround channel, and then the signal after the right surround channel signal is filtered by the eighth filter is superimposed to obtain the signal three for eliminating crosstalk in the left surround channel; the right surround channel signal is filtered by the ninth filter to obtain the positioning signal in the direction of the right surround channel, and then the signal after the left surround channel signal is filtered by the tenth filter is superimposed to obtain the signal four for eliminating crosstalk in the direction of the right surround channel.

[0015] The rear left channel signal is filtered by the eleventh filter to obtain the positioning signal in the rear left channel direction, and then superimposed with the signal obtained by the rear right channel signal filtered by the twelfth filter to obtain the signal five for eliminating crosstalk in the rear left channel direction; the rear right channel signal is filtered by the thirteenth filter to obtain the positioning signal in the rear right channel direction, and then superimposed with the signal obtained by the rear left channel signal filtered by the fourteenth filter to obtain the signal six for eliminating crosstalk in the rear right channel direction.

[0016] The signals three, four, five, and six are mixed in a prescribed ratio;

[0017] The left channel signal with three-dimensional sound effects is obtained by weighted summation of the main left channel signal, the left surround channel signal, the rear left channel signal, signal one, signal three, and signal five; the right channel signal with three-dimensional sound effects is obtained by weighted summation of the main right channel signal, the right surround channel signal, the rear right channel signal, signal two, signal four, and signal six.

[0018] The low-frequency signal, after being delayed, is combined with the left channel signal and the right channel signal with three-dimensional sound effects in a certain proportion to obtain a virtual surround signal.

[0019] Preferably, the main left channel signal and the main right channel signal of the high-frequency signal are delayed and then inverted and superimposed onto the original signal to construct the rear left channel and the rear right channel, which is used to enhance the spatial sense of surround sound.

[0020] Preferably, the signals three, four, five and six are mixed in a prescribed ratio to obtain a change in the sound image range of 120° to 240°.

[0021] Preferably, the first to fourteenth filters are all different.

[0022] Preferably, the number of the infinite N-level echo modules is two.

[0023] Preferably, by passing the main left channel signal and the main right channel signal of the high-frequency signal through two infinite N-level echo modules respectively, a left surround channel and a right surround channel are constructed, so that each left surround channel and right surround channel is composed of two echo modules with different time delays and gains, resulting in differentiated left and right surround sound.

[0024] The present invention also provides a system for generating three-dimensional virtual surround sound, comprising:

[0025] The decomposition module is used to decompose audio signals into high-frequency and low-frequency signals;

[0026] The center channel signal generation module is used to sum and average the main left channel signal and the main right channel signal of the high-frequency signal to obtain the center channel signal.

[0027] An infinite N-level echo module is used to pass the main left channel signal and the main right channel signal of the high-frequency signal through the infinite N-level echo module to construct a left surround channel and a right surround channel.

[0028] The inverse superposition module is used to delay and then invert the main left channel signal and the main right channel signal of the high-frequency signal and superimpose them onto the original signal to construct the rear left channel and the rear right channel;

[0029] The filtering module is used to filter signals;

[0030] Left and right channel signal generation modules with three-dimensional sound effects: used to generate left and right channel signals with three-dimensional sound effects;

[0031] The virtual surround signal generation module is used to combine the low-frequency signal obtained by delay processing with the left channel signal with three-dimensional sound effects and the right channel signal with three-dimensional sound effects in a certain proportion to obtain a virtual surround signal.

[0032] Preferably, the filtering module includes fourteen filters.

[0033] Preferably, all of the filters are different.

[0034] Preferably, the number of the infinite N-level echo modules is two.

[0035] This invention employs superior crosstalk cancellation technology to eliminate crosstalk, enhance spatial positioning, and improve speech clarity. It constructs surround sound channels by simulating reflected sound through infinite N-level echoes, resulting in more distant background sound in the generated surround sound. This gives the virtual surround sound a stronger sense of space, distance, diffusion, and immersion, rather than just a sense of positioning. Utilizing the "Lloyd's effect," the delayed signal is inverted and superimposed onto the direct signal to construct rear channels, significantly enhancing the spatial sense of the surround sound. Separation technology preserves the original bass frequencies, and the virtual sound in each direction can be freely adjusted, thereby controlling sound intensity, timbre brightness, and timbre smoothness. Attached Figure Description

[0036] Figure 1 This is a conceptual diagram of traditional 5.1 channel virtual surround sound in existing technologies.

[0037] Figure 2 This is a schematic diagram of the traditional 5.1 channel virtual surround sound implementation architecture in existing technologies.

[0038] Figure 3 This is a schematic diagram of the virtual sound image location based on sound provided in an embodiment of the present invention.

[0039] Figure 4 This is a schematic diagram of the process framework for generating 7.1 three-dimensional virtual surround sound based on dual-channel stereo sound, provided in an embodiment of the present invention.

[0040] Figure 5 This is a schematic diagram of crosstalk provided in an embodiment of the present invention.

[0041] Figure 6 Figure a is a schematic diagram of the superior crosstalk cancellation structure provided in the embodiment of the present invention, and Figure b is a schematic diagram of the traditional crosstalk cancellation structure.

[0042] Figure 7 This is a schematic diagram of the N-level echo implementation architecture provided in an embodiment of the present invention.

[0043] Figure 8 The flowchart illustrates the processing of virtual surround sound channels SL and SR as provided in this embodiment of the invention. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0045] Example 1:

[0046] The purpose of this embodiment is to provide a method for providing a wide sound field and rich surround sound in virtual surround sound, in order to solve the problems of auditory fatigue and muddy sound and narrow surround sound that are easily caused by general virtual surround sound. It enables the creation of a 7.1 channel three-dimensional surround sound field with only two speakers, giving the listener a 360° sound field illusion, a strong sense of presence, immersion, and power, and an immersive experience.

[0047] To achieve the above objectives, this embodiment adopts the following technical solution:

[0048] (1) The left and right channel signals are decomposed by high-pass HP and low-pass LP respectively to obtain two signals HPF and LPF corresponding to the left channel and two signals HPF and LPF corresponding to the right channel.

[0049] (2) The left and right channel signals of HPF are summed and averaged to obtain the center FC signal. The left channel of HPF signal is FL and the right channel of HPF signal is FR.

[0050] (3) Pass the left and right channel signals of HPF through two infinite N-level echo modules respectively to construct surround channels SL and SR;

[0051] (4) The left and right channel signals of HPF are delayed and then inverted and superimposed on the original signal to construct the rear channels RL and RR;

[0052] (5) The FC signal is filtered by HRTF_L_Fc to obtain the virtual center signal FC_L′ of the left ear, and filtered by HRTF_R_Fc to obtain the virtual center signal FC_R′ of the right ear;

[0053] (6) The FL signal is filtered by HRTF_L_FL to obtain the positioning signal in the FL direction, and then the FR signal filtered by HRTF_R_FL is superimposed to obtain the crosstalk-eliminating signal FL′ in the FL direction; the FR signal is filtered by HRTF_R_FR to obtain the positioning signal in the FR direction, and then the FL signal filtered by HRTF_L_FR is superimposed to obtain the crosstalk-eliminating signal FR′ in the FR direction.

[0054] (7) The SL signal is filtered by HRTF_L_SL to obtain the positioning signal in the SL direction, and then the SR signal filtered by HRTF_R_SL is superimposed to obtain the crosstalk-eliminating signal SL′ in the SL direction; the SR signal is filtered by HRTF_R_SR to obtain the positioning signal in the SR direction, and then the SL signal filtered by HRTF_L_SR is superimposed to obtain the crosstalk-eliminating signal SR′ in the SR direction.

[0055] (8) The RL signal is filtered by HRTF_L_RL to obtain the positioning signal in the RL direction, and then the RR signal filtered by HRTF_R_RL is superimposed to obtain the crosstalk cancellation signal RL′ in the RL direction; the RR signal is filtered by HRTF_R_RR ​​to obtain the positioning signal in the RR direction, and then the RL signal filtered by HRTF_L_RR is superimposed to obtain the crosstalk cancellation signal RR′ in the RR direction.

[0056] (9) SL′, SR′ and RL′, RR′ are mixed in a certain proportion to achieve spectrum compensation and adjustment of virtual sound image angle;

[0057] (10) The left channel signal with three-dimensional sound effect is obtained by weighted summation of all the left channel direction-related signals FC_L′, FL′, SL′, RL′ and the right channel signal with three-dimensional sound effect is obtained by weighted summation of all the right channel direction-related signals FC_R′, FR′, SR′, RR′.

[0058] (11) The LPF signal in (1) above is delayed and then combined with the three-dimensional sound effect signal after (10) in proportion to obtain the final virtual surround signal;

[0059] Example 2:

[0060] The technical solution of the invention will be described in detail below with reference to the accompanying drawings:

[0061] (1) As Figure 4 As shown, the process of achieving 7.1 three-dimensional virtual surround sound based on a dual-channel stereo system in this invention involves decomposing the signal into two paths. One path is processed by high-pass filtering and then by relevant virtual transformation to obtain virtual signals corresponding to the seven directions. The virtual sound image positions are as follows: Figure 3 As shown, the signal is then processed through HRTF filtering and crosstalk cancellation to form the final three-dimensional audio signal. This signal is then combined with another original low-frequency signal that has undergone low-pass filtering and delay processing, and finally reproduced through a speaker. The final virtual surround sound is represented as follows:

[0062] L_out=k1*Delay(L_LPF)+k2*(a*FC_L+b*FL+(c*SL+d*SR)+(e*RL+f*RR))

[0063] R_out=k1*Delay(R_LPF)+k2*(a*FC_R+b*FR+(c*SR+d*SL)+(e*RR+f*RL))

[0064] L_out and R_out represent the final synthesized left and right channel signals, respectively. L_LPF and R_LPF represent the left and right channel signals of the LPF obtained after low-pass LP processing of the original audio signal, respectively. FC_L′ and FC_R′ represent the left and right ear reception signals of FC′ obtained after transforming the original audio signal into a center signal and filtering it with HRTF, respectively, after high-pass HP processing. FL′ and FR′ represent the left and right front and right front sound signals of the original audio signal after high-pass HP processing and virtual transformation, respectively. SL′ SR′ and RL′ represent the left and right surround sound signals of the original sound signal after virtual transformation, which are located in the left and right surround directions, respectively. RL′ and RR′ represent the left and right rear sound signals of the original sound signal after virtual transformation, which are located in the left and right rear directions, respectively. k1, k2, a, b, c, d, e, and f represent the gain of the sound signal in each direction. Free adjustment of the gain value in each direction can change the proportion of sound signals in different directions, thereby adjusting the intensity, brightness, and smoothness of the sound to meet the needs of users and different products.

[0065] (2) Figure 7As shown, in order to obtain a virtual surround sound with a wide sound field and rich ambient sound, this invention constructs surround sound channels by simulating reflected sound through infinite N-level echoes. Each surround sound channel is composed of two echo modules with different time delays and gains, resulting in differentiated left and right surround sound. This makes the constructed surround sound feel like a distant sound spreading and expanding, and also has obvious differences and a sense of distance in spatial hearing. By using the "Liaoning effect", the delayed signal is inverted and superimposed on the direct signal to construct the rear channels, which significantly enhances the spatial sense of the surround sound, as if it comes from all directions.

[0066] The acoustic signals constructed in each direction are realized as follows:

[0067] FC = (HP(L) + HP(R)) / 2;

[0068] FL = HP(L);

[0069] FR = HP(R);

[0070] SL=HP(L)+g1*Echos(delay1(HP(L)))+g2*Echos(delay2(HP(L)));

[0071] SR=HP(R)+g3*Echos(delay3(HP(R)))+g4*Echos(delay4(HP(R)));

[0072] RL=HP(L)+Invert(delay5(HP(L)));

[0073] RR=HP(R)+Invert(delay6(HP(R)));

[0074] like Figure 5 As shown, crosstalk between the signals emitted by the left and right speakers and the left and right ears will make the sound muddy. Figure 6 The traditional crosstalk cancellation method shown in (a) is implemented as follows:

[0075] FL=HRTF_L_FL(FL)+HRTF_L_FR(FR);

[0076] FR=HRTF_R_FR(FR)+HRTF_R_FL(FL);

[0077] SL=HRTF_L_SL(SL)+HRTF_L_SR(SR);

[0078] SR=HRTF_R_SR(SR)+HRTF_R_SL(SL);

[0079] RL=HRTF_L_RL(RL)+HRTF_L_RR(RR);

[0080] RR=HRTF_R_RR(RR)+HRTF_R_RL(RL);

[0081] Crosstalk cancellation technology employing superior structures, such as Figure 6 The crosstalk cancellation method shown in (b) is implemented as follows:

[0082] FL=HRTF_L_FL(FL)+HRTF_R_FL(FR);

[0083] FR=HRTF_R_FR(FR)+HRTF_L_FR(FL);

[0084] SL=HRTF_L_SL(SL)+HRTF_R_SL(SR);

[0085] SR=HRTF_R_SR(SR)+HRTF_L_SR(SL);

[0086] RL=HRTF_L_RL(RL)+HRTF_R_RL(RR);

[0087] RR=HRTF_R_RR(RR)+HRTF_L_RR(RL);

[0088] Therefore, it can be seen that crosstalk cancellation methods with superior structures are more in line with real auditory scenarios, and crosstalk cancellation technology with superior structures can enhance spatial positioning and speech clarity.

[0089] (3) Figure 8 As shown, the surround sound signal, after being mixed in a certain proportion, is achieved as follows:

[0090] LS=LS+g1*LS1+g2*LS2+g3*RS1+g4*RS2

[0091] RS=RS+k1*LS1+k2*LS2+k3*RS1+k4*RS2

[0092] The virtual sound image angle sinθ = (LS - RS′) / (LS + RS′) * sin120°

[0093] =(LS / RS-1) / (LS / RS+1)*sin120°;

[0094] It can be seen that by adjusting the relevant mixing ratio coefficient and SL / SR, the sound image range of 120° to 240° can be changed.

[0095] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0096] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for generating three-dimensional virtual surround sound, characterized in that, include: Decompose the audio signal into high-frequency and low-frequency signals; The center channel signal is obtained by summing and averaging the main left channel signal and the main right channel signal of the high-frequency signal. The main left channel signal and the main right channel signal of the high-frequency signal are respectively passed through an infinite N-level echo module to construct the left surround channel and the right surround channel. The main left channel signal and the main right channel signal of the high-frequency signal are delayed and then inverted and superimposed on the original signal to construct the rear left channel and the rear right channel; The center channel signal is filtered by a first filter to obtain a virtual center signal for the left ear; the center channel signal is filtered by a second filter to obtain a virtual center signal for the right ear. The main left channel signal is filtered by the third filter to obtain the positioning signal in the direction of the main left channel. Then, the signal after the main right channel signal is filtered by the fourth filter is superimposed to obtain the signal after the main left channel signal is filtered to obtain the signal after the main left channel signal is filtered to obtain the signal after the main left channel signal is filtered by the sixth filter to obtain the signal after the main right channel signal is filtered to obtain the signal after the main right channel signal is filtered to obtain the signal after the main right channel signal is filtered to obtain the signal after the main left channel signal is filtered to obtain the signal after the main right ... The left surround channel signal is filtered by the seventh filter to obtain the positioning signal in the direction of the left surround channel, and then the signal after the right surround channel signal is filtered by the eighth filter is superimposed to obtain the signal three for eliminating crosstalk in the left surround channel; the right surround channel signal is filtered by the ninth filter to obtain the positioning signal in the direction of the right surround channel, and then the signal after the left surround channel signal is filtered by the tenth filter is superimposed to obtain the signal four for eliminating crosstalk in the direction of the right surround channel. The rear left channel signal is filtered by the eleventh filter to obtain the positioning signal in the rear left channel direction, and then superimposed with the signal obtained by the rear right channel signal filtered by the twelfth filter to obtain the signal five for eliminating crosstalk in the rear left channel direction; the rear right channel signal is filtered by the thirteenth filter to obtain the positioning signal in the rear right channel direction, and then superimposed with the signal obtained by the rear left channel signal filtered by the fourteenth filter to obtain the signal six for eliminating crosstalk in the rear right channel direction. The signals three, four, five, and six are mixed in a prescribed ratio; The left channel signal with three-dimensional sound effects is obtained by weighted summation of the main left channel signal, the left surround channel signal, the rear left channel signal, signal one, signal three, and signal five; the right channel signal with three-dimensional sound effects is obtained by weighted summation of the main right channel signal, the right surround channel signal, the rear right channel signal, signal two, signal four, and signal six. The low-frequency signal, after being delayed, is combined with the left channel signal and the right channel signal with three-dimensional sound effects in a certain proportion to obtain a virtual surround signal.

2. The method for generating three-dimensional virtual surround sound according to claim 1, characterized in that, The high-frequency signal is delayed and then inverted and superimposed onto the original signal to construct the rear left and rear right channels, which are used to enhance the spatial sense of surround sound.

3. The method for generating three-dimensional virtual surround sound according to claim 1, characterized in that, The signals three, four, five, and six are mixed in a prescribed ratio to obtain a change in the sound image range of 120° to 240°.

4. The method for generating three-dimensional virtual surround sound according to claim 1, characterized in that, The first through fourteenth filters are all different.

5. The method for generating three-dimensional virtual surround sound according to claim 1, characterized in that, The number of infinite N-level echo modules is two.

6. The method for generating three-dimensional virtual surround sound according to claim 5, characterized in that, By passing the main left channel signal and the main right channel signal of the high-frequency signal through two infinite N-level echo modules respectively, a left surround channel and a right surround channel are constructed, so that each left surround channel and right surround channel is composed of two echo modules with different time delays and gains, resulting in differentiated left and right surround sound.

7. The generation system of the method for generating three-dimensional virtual surround sound according to any one of claims 1-6, characterized in that, include: The decomposition module is used to decompose audio signals into high-frequency and low-frequency signals; The center channel signal generation module is used to sum and average the main left channel signal and the main right channel signal of the high-frequency signal to obtain the center channel signal. An infinite N-level echo module is used to pass the main left channel signal and the main right channel signal of the high-frequency signal through the infinite N-level echo module to construct a left surround channel and a right surround channel. The inverse superposition module is used to delay and then invert the main left channel signal and the main right channel signal of the high-frequency signal and superimpose them onto the original signal to construct the rear left channel and the rear right channel; The filtering module is used to filter signals; Left and right channel signal generation modules with three-dimensional sound effects: used to generate left and right channel signals with three-dimensional sound effects; The virtual surround signal generation module is used to combine the low-frequency signal obtained by delay processing with the left channel signal with three-dimensional sound effects and the right channel signal with three-dimensional sound effects in a certain proportion to obtain a virtual surround signal.

8. The three-dimensional virtual surround sound generation system according to claim 7, characterized in that, The filtering module includes fourteen filters.

9. The three-dimensional virtual surround sound generation system according to claim 7, characterized in that, The filters are all different.

10. The three-dimensional virtual surround sound generation system according to claim 7, characterized in that, The number of infinite N-level echo modules is two.